Synchronization in a Multi-Chip System
By determining and synchronizing maximum loop latencies and adjusting buffer sizes, the solution ensures deterministic data transmission in multi-chip systems, enhancing reliability and reducing buffer requirements.
Patent Information
- Application Number
- JP2023175966
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-08-16
- Filing Date
- 2023-10-11
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-08-14
AI Technical Summary
In multi-chip systems, data communication between chips is non-deterministic due to varying latency times, which affects the predictability and reliability of operations, particularly in applications like neural networks where precise timing is crucial.
The solution involves determining the maximum loop latency between chip pairs, synchronizing local counters based on this latency, and adjusting buffer sizes to ensure deterministic data transmission by incorporating a fixed delay in data transfer operations.
This approach minimizes variations in data arrival times, enables deterministic data communication, reduces the need for large receive buffers, and improves the reliability and efficiency of operations in multi-chip systems, especially in neural network applications.
Smart Images

Figure 0007698016000001 
Figure 0007698016000002 
Figure 0007698016000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to synchronization and data transfer in a multi-chip system.
Background Art
[0002] An electronic device may be composed of a plurality of different chips, and the plurality of different chips need to communicate data among themselves for the electronic device to operate. Data communication between chips may be non-deterministic. For example, the latency from the transmission time in one chip to the reception time in another chip in data communication between chips tends to vary. That is, the time it takes for data to move from one chip to another is not constant and is influenced by many different factors in the transmission time.
Summary of the Invention
Means for Solving the Problems
[0003] Generally, innovative aspects of the subject matter described herein include, for each pair of chips among a plurality of chips of a semiconductor device, an action of determining a corresponding loop latency of round-trip data transmission between the pair of chips around a transmission path through the plurality of chips, an action of identifying the maximum loop latency from among the loop latencies, an action of determining a full-path latency until data transmission transmitted from a chip among the plurality of chips is transmitted around the path and returns to the chip, an action of comparing half of the maximum loop latency with one-Nth of the full-path latency, where N is the number of chips in the transmission path of the chips, and an action of storing a larger value as the inter-chip latency of the semiconductor device, where the inter-chip latency represents an operating characteristic of the semiconductor device, which can be embodied in a method for evaluating inter-chip latency characteristics including.
[0004] In a second general aspect, the innovative features of the subject matter described in this application can be embodied in a method for evaluating inter-chip latency characteristics that includes actions to determine the corresponding loop latency of round-trip data transmission between pairs of adjacent chips for each pair of adjacent chips in a series ring arrangement of semiconductor devices. The actions include an action to identify the maximum loop latency from among the loop latencies. The actions include an action to determine the ring latency until data transmission sent from a chip among the plurality of chips returns to the chip after traversing the series ring arrangement. The actions include an action to compare half of the maximum loop latency to one Nth of the ring latency, where N is the number of chips within the plurality of chips, and an action to store the larger value as the inter-chip latency of the semiconductor device, where the inter-chip latency represents the operating characteristics of the semiconductor device. Other implementations of this aspect include corresponding systems, devices, and computer programs configured to execute the actions of the method encoded on a computer storage device.
[0005] These and other implementations can each optionally include one or more of the following features.
[0006] In some implementations, the actions for determining the loop latency of round-trip data transmission between a pair of chips include transmitting first timestamped data from the first chip of the pair of chips to the second chip of the pair of chips, an action for determining a first relative one-way latency between the pair of chips based on the first timestamped data, an action for transmitting second timestamped data from the second chip to the first chip, an action for determining a second relative one-way latency between the pair of chips based on the second timestamped data, and an action for determining the loop latency of round-trip data transmission between the pair of chips based on the first relative one-way latency and the second relative one-way latency. In some implementations, the first timestamped data indicates the local counter time of the first chip when the first timestamped data was transmitted. In some implementations, the action for determining the first relative one-way latency between the pair of chips includes an action of calculating the difference between the time indicated in the timestamped data and the local counter time of the second chip when the second chip received the first timestamped data. In some implementations, the action for determining the loop latency of round-trip data transmission between the pair of chips includes an action of calculating the difference between the first relative one-way latency and the second relative one-way latency.
[0007] In some implementations, one or more of the plurality of chips are application-specific integrated circuit (ASIC) chips configured to perform neural network operations.
[0008] In a third general aspect, the innovative features of the subject matter described herein can be embodied in an inter-chip timing synchronization method that includes, for each pair of chips within a plurality of chips of a semiconductor device, an action of determining a first one-way latency of transmission from a first chip within the pair to a second chip within the pair of chips, and an action of determining a second one-way latency of transmission from the second chip within the pair to the first chip within the pair of chips. The actions include an action of receiving, in a semiconductor device driver, the first one-way latency and the second one-way latency for each pair of chips. The actions include an action of determining, by the semiconductor device driver, a loop latency between each pair of chips from the respective first one-way latency and second one-way latency for each pair of chips. The actions include an action of adjusting, by the semiconductor device driver, a local counter of a second chip within at least one pair of chips based on a characteristic inter-chip latency of the semiconductor device and a first one-way latency of at least one pair of chips. Other implementations of this aspect include corresponding systems, devices, and computer programs configured to execute the actions of the method encoded on a computer storage device.
[0009] These and other implementations can each optionally include one or more of the following features.
[0010] In some implementations, the actions include an action of determining, by the semiconductor device driver, that each loop latency is less than or equal to a characteristic inter-chip latency of the semiconductor device.
[0011] In some implementations, the action of adjusting the local counter of the second chip within at least one pair of chips includes an action of increasing the value of the local counter by an adjustment value. In some implementations, the adjustment value is equal to the characteristic inter-chip latency of the semiconductor device plus the first one-way latency of transmission from the first chip within the pair to the second chip within the pair.
[0012] In some implementations, the action of determining the loop latency between each pair of chips includes, for each pair of chips, an action of calculating the difference between a first relative one-way latency associated with the pair of chips and a second relative one-way latency associated with the pair of chips.
[0013] In some implementations, the action of determining the first one-way latency of transmission from the first chip in a pair to the second chip in the pair of chips includes an action of transmitting first timestamped data from the first chip to the second chip, and an action of determining a first relative one-way latency between the pair of chips based on the first timestamped data. In some implementations, the first timestamped data indicates the local counter time of the first chip when the first timestamped data was transmitted. In some implementations, the action of determining the first relative one-way latency between the pair of chips includes an action of calculating the difference between the time indicated in the timestamped data and the local counter time of the second chip when the second chip received the first timestamped data.
[0014] In some implementations, one or more of the plurality of chips are application-specific integrated circuits (ASICs) configured to perform neural network operations.
[0015] In a fourth general aspect, an innovative aspect of the subject matter described herein can be embodied in a method for transmitting data between chips that, at a first time, includes the action of transmitting data from a first chip to a second adjacent chip in a serial ring arrangement of semiconductor devices. The action includes the action of storing the data in a buffer in the second chip. The action includes the action of releasing the data from the buffer at a second time, where the interval between the first time and the second time is based on a characteristic inter-chip latency of the serial ring arrangement of the chips. The action includes the action of transmitting the data from the second chip to a third chip, where the third chip is adjacent to the second chip in the serial ring arrangement of the chips. Other implementations of this aspect include corresponding systems, devices, and computer programs configured to execute the actions of the method encoded on a computer storage device.
[0016] These and other implementations can optionally include one or more of the following features, respectively.
[0017] In some implementations, the characteristic inter-chip latency represents the maximum predicted one-way data transmission latency between two chips within the serial ring arrangement of the chips.
[0018] In some implementations, the second time is a pre-scheduled time of an operating schedule for the second chip.
[0019] In some implementations, the action includes the action of passing data from the buffer of the second chip along an internal bypass path to a communication interface of the second chip coupled to the third chip.
[0020] In some implementations, one or more of the first, second, and third chips are application-specific integrated circuit (ASIC) chips configured to perform neural network operations.
[0021] Various implementations provide one or more of the following advantages. For example, in some implementations, the processes described herein minimize variations in the potential data arrival times of inter-chip communications. Reducing variations in data communication may enable the use of smaller receive data buffers in the chips of the system. In some implementations, the processes described herein make the transmission operations between chips deterministic. For example, an implementation may enable a program compiler to use a constant (e.g., deterministic) latency when calculating the local counter time at which a receiving chip accesses data from an input buffer transmitted from an adjacent chip at a particular time.
[0022] Details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below. Other features and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0023]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 3C
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 7
DETAILED DESCRIPTION OF THE INVENTION
[0024] Generally, the present disclosure relates to inter-chip time synchronization and data transmission in a multi-chip system. More specifically, the present disclosure provides a chip operation process that improves the predictability of data transmission around a serial ring topology of chips, in some examples, between chips. The present disclosure provides an exemplary process for synchronizing local counters of chips within a system and performing data transmission in a manner that takes into account the inherent variable data arrival times of inter-chip data transmission to make the data reception time more deterministic, and in some cases, completely deterministic.
[0025] Referring first to inter-chip time synchronization, time synchronization includes two aspects. The first aspect is the characterization of the inter-chip latency of data transmission between each pair of chips on a processing system. This process provides the operating characteristics of the board (e.g., maximum inter-chip latency) that function as constants for synchronizing local chip counters each time the board is powered on. The second aspect is synchronizing the local chip counters when the board is powered on (e.g., "synchronization at power-on").
[0026] More specifically, the characterization process must be completed for each board redesign. For example, the maximum inter-chip latency is generally a physical characteristic that depends on the layout of the chips on the board. The characterization process includes steps of measuring the "round-trip" loop latency of transmissions between pairs of chips on the board that are involved in direct communication with each other (e.g., pairs of adjacent chips). Further, in an implementation that includes chips connected in a serial ring arrangement, the characterization process can also include steps of measuring the round-trip transmission latency of the entire ring. The data collected from these measurements can be used to determine the maximum inter-chip latency that occurs between any two chips.
[0027] Synchronization at startup is performed to synchronize the local counters of the chips each time the board is powered on, reset, or both. Each chip is clocked by a local clock that is synchronized with the local clocks of other chips (e.g., the clocks of each chip have the same frequency and phase), but the chips operate using local counters to clock individual chip operations, and when the board is powered on or a chip exits reset, the individual counters generally have different count values. Therefore, synchronization at startup is used to approximately synchronize the local count values of the chips.
[0028] The synchronization process at startup includes a step of measuring the one-way latency of transmission between pairs of chips on the board. The board driver determines a local counter adjustment for one chip within each pair based on the maximum inter-chip latency characterized for the board and one of the one-way latencies between chips within the pair. For example, the driver can adjust the local counter of one of the chips within the pair by increasing the counter value by the sum of the maximum inter-chip latency and one of the one-way latencies between chips. In some implementations, the startup process includes a step of adjusting the round-trip latency between one or between chip pairs, for example, by adjusting the FIFO buffer of one of the chips.
[0029] In some implementations, the semiconductor chip can be an application specific integrated circuit (ASIC) designed to perform machine learning operations. An ASIC is an integrated circuit (IC) customized for a specific application. For example, the ASIC can be designed to perform the operations of a machine learning model, including operations such as recognizing objects in an image as part of, for example, a deep neural network, machine translation, speech recognition, or other machine learning algorithms. For example, when used as an accelerator for a neural network, the ASIC can receive an input to the neural network and compute the inference of the neural network for the input. The data input to a neural network layer, such as either an input to the neural network or the output of another layer of the neural network, may be referred to as an activation input. The inference can be computed according to each set of weight inputs associated with the layer of the neural network. For example, some or all of the layers can receive a set of activation inputs and process the activation inputs according to a set of weight inputs for the layer to generate an output. Further, the neural network operations can be executed by the system of the ASIC according to an explicit operation schedule. In that way, deterministic and synchronized data transfer between ASIC chips can improve the reliability of neural network operations and simplify the debugging operations.
[0030] FIG. 1 is a schematic diagram showing an exemplary multi-chip system 100. The multi-chip system 100 can be a network of integrated circuits configured to perform machine learning operations. For example, the multi-chip system 100 can be configured to implement a neural network architecture. The multi-chip system includes a plurality of semiconductor chips 102. The chips 102 can be general-purpose integrated circuit chips or application-specific integrated circuit chips. For example, one or more of the chips 102 can be an ASIC, a field-programmable gate array (FPGA), a graphics processing unit (GPU), or any other suitable integrated circuit chip. A clock 106 is coupled to each of the chips 102 to provide a synchronous timing signal. For example, the clock 106 can include a crystal oscillator that provides a common timing signal (e.g., a 1 GHz clock signal) to each of the chips 102.
[0031] The system 100 also includes a system driver 104. The system driver 104 can be an external computing system such as, for example, a laptop computer, a desktop computer, or a server system. The system driver 104 can be used to execute or manage the chip synchronization process described herein or a portion thereof. For example, the system driver 104 can be configured to program the chips, manage the startup operation of the system 100, debug the chips, or combinations thereof. The system driver can be coupled to the chips 102 via a communication link. The system driver 104 can be coupled to the chips 102 via a configuration status register (e.g., a low-speed interface for programming and debugging the chips).
[0032] In the illustrated example, the multi-chip system 100 includes eight ASIC chips 102 and one FPGA chip 102 arranged in a serial ring topology. More specifically, each chip 102 communicates with two adjacent chips, one on each side, such that data is communicated from the chip around the ring to the adjacent chip. The chips 102 and their data communication links form a closed loop. Further, the multi-chip system 100 includes two data paths, a clockwise path 108 and a counterclockwise path 110, between each pair of chips.
[0033] In some implementations, each ASIC chip (P0 - P7) may be configured to implement a layer of a neural network. Input activation data may be received by the FPGA chip 104 and transmitted to P0. P0 may be configured to implement, for example, the input layer of a neural network. P0 performs calculations on the activation data to generate layer output data that is transmitted to P1. P1 may be configured to implement the first hidden layer of the neural network, perform calculations on the output from P0, and then transmit its output to the next neural network layer implemented by P2. The process continues around the ring through each of the ASICs 102 and, by extension, may be processed by each layer of the neural network. Such a process may depend on the precise timing of data transfers between adjacent chips (and throughout the ring) for the neural network to operate reliably and accurately. Thus, synchronization of data transmission between each ASIC may be important to ensure proper operation coordination between the chips.
[0034] The internal operation of a single chip within a synchronous system is synchronous and deterministic, meaning there are no differences in the timing of such internal operations. However, for inter-chip operations such as data transmission, even in a synchronous system, there are inherent and non-deterministic variations in the timing of operations. One cause of the timing variability is the characteristics of the physical link between two adjacent chips, which can introduce variations of, for example, about 0 to 3 clock cycles in the latency of data transmission between adjacent chips. A second and larger cause of the timing variability is the lack of synchronization between the internal chip operation and the forward error correction scheme implemented by the multi-chip system. In the forward error correction scheme, error correction data is added to the data transmission between chips, but the added error correction data is not necessarily synchronized with the data transmission. The introduction of asynchronous data into the data transmission can introduce variations of, for example, up to 16 clock cycles in the latency of data transmission between adjacent chips.
[0035] When data is transmitted from one chip to another non - adjacent chip (e.g., from P0 to P7), the variations in the latency for each inter - chip transmission (e.g., from P0 to P1, P1 to P2, etc.) accumulate into the cumulative delay at the destination chip (P7). Taking only the variations due to forward error correction as an example, the latency for a single inter - chip (e.g., from P0 to P1) transmission has a variation of ±16 clock cycles. However, some operations may require data transmission from one chip 102 to another non - adjacent chip 102, e.g., data transmission from chip P0 to chip P3, or even data transmission around a ring from the first chip P0 to the last chip P7. As described in more detail below, to transmit data from one chip to another non - adjacent chip (e.g., from P0 to P7), the data can be transmitted through each of the intervening chips (e.g., through chips P1 - P6) using a bypass operation. However, the latency variations between chips accumulate over 8 chips, and the total variation in latency around the ring approaches ±128 clock cycles. The processes described below improve the predictability of data transmission between chips and, in some examples, enable inter - chip data transmission to be performed in a deterministic manner.
[0036] FIG. 2 shows a flowchart of an exemplary process 200 for characterizing the maximum latency in the multi - chip system 100. Process 200 is described with reference to FIGS. 1, 2, 3A - 3B. In some implementations, process 200 or a part thereof is executed or controlled by the system driver 104. In some examples, the process or a part thereof is executed by the individual chips 102 of the multi - chip system 100. The characterization process 200 is used to determine the characteristic inter - chip latency of the multi - chip system design, e.g., the maximum inter - chip latency (L max )). For example, process 200 can be executed for an initial chip placement and / or for a new system topology.
[0037] The first step of process 200 includes determining the loop latency between each pair of chips within multi-chip system 100 (step 202). For example, as shown in FIG. 1, the illustrated multi-chip system 100 has ten individual inter-chip communication loops (112, 114) with independently measurable latencies. There are nine loops 112 between adjacent chips 102 and one loop 114 around the entire ring. In a multi-chip system, since there is no available common time reference, absolute latency values may only be measurable in these loops 112, 114. That is, each of chips 112, 114 is driven by a common clock 106, but the local counters on each chip 112, 114 are not necessarily synchronized to the same count value. In other words, the "local time" on each chip 112, 114 may be different. As will be described in more detail below, measuring loop latency rather than individual one-way latencies between chips can be used to account for differences between local counters on each chip.
[0038] Since these loop latencies are simply the sum of the latencies in each direction, the nine single-chip loops 112 that first proceed clockwise and then counterclockwise have the same latency as the first clockwise loop. Similarly, the complete system counterclockwise loop 114 has the same latency as the sum of all nine single-chip loops 112 minus the latency of the clockwise system loop 114. Measuring the difference in latency in different directions around a loop between two chips does not provide more information since these differences can be derived from the nine small loops 112 and the single system loop 114.
[0039] Figures 3A - 3C are a series of block diagrams showing loop latency measurements between adjacent chips 102. Figures 3A - 3C show simplified block diagrams of two adjacent chips 102, chip A and chip B. Each chip 102 includes a controller 304 that controls the local operation of the chip, a local counter 306, and a communication interface 308. For clarity of explanation, the communication interface 308 is represented as from a transmitter interface (Tx) to a receiver interface (Rx). The communication interface 308 includes a first-in first-out (FIFO) buffer.
[0040] To measure the loop latency, each chip initializes its local counter 306, for example, by activating the chip 102. The local counter 306 of each chip represents its local time, as discussed above. In some implementations, the chips 102 execute their individual operations (e.g., calculations, reading data from an input buffer, and transmitting data to other chips) at pre-scheduled counter times. The counter 306 does not need to be synchronized with the process 200 in any way. For example, in the example shown in Figure 3A, the local counter 306 of chip A is initialized at time 0, and the local counter 306 of chip B is initialized at time 150, and thus the local counters of chip A and chip B are out of sync by 150 clock cycles. The startup synchronization process discussed below is used to synchronize the local counters 306 within the chips 102. It should be noted that the counter times used in Figures 3A - 3C (and Figures 6A and 6B) are simplified for purposes of explanation.
[0041] Referring to FIGS. 3B and 3C, to measure the round-trip latency between chip A and chip B, chips A and C first perform a series of timestamped data transmissions from chip A to chip B and then from chip B to chip A. For example, first chip A transmits timestamped data 309 to chip B to measure the relative one-way latency of transmission in the first direction, e.g., from chip A to chip B on clockwise data path 108. Chip A transmits data 309 including a timestamp having the local counter time of chip A (e.g., 10) when the data 309 was transmitted to chip B. For clarity of explanation, FIG. 3B shows only one data transmission being transmitted to chip B. In reality, for example, chip A can transmit a series of data transmissions 309 at various points in a 512-cycle physical coding sublayer (PCS) period, each of which is timestamped with the local counter time of chip A at the time of transmission. Chip B receives the data 309 and records its own local counter time (e.g., 180). The difference between the local time of chip A (e.g., 10) when the data 309 was transmitted and the local time of chip B (e.g., 180) when the data 309 was received is equal to the relative one-way relative latency from chip A to chip B. For example, the relative one-way latency as shown in FIG. 3B is 170 clock cycles.
[0042] As shown in FIG. 3C, chip B executes the same process to measure, for example, the relative one-way latency in the second direction of transmission from chip B to chip A in the counterclockwise data path 110. Chip B transmits data 310 to chip A that includes a timestamp having the local counter time of chip B (e.g., 200) when the data 310 was transmitted. For clarity of explanation, FIG. 3C shows only one data transmission sent to chip A. In reality, for example, chip B can send a series of data transmissions 309 at various points in a 512-cycle PCS period, each of which is timestamped with the local counter time of chip B at the time of transmission. Chip A receives the data 310 and records its own local counter time (e.g., 60). The difference between the local time of chip B (e.g., 200) when the data 310 was sent and the local time of chip A (e.g., 60) when the data 310 was received is equal to the relative one-way latency from chip B to chip A. For example, the relative one-way latency shown in FIG. 3C is -140 clock cycles. It should be noted that the relative one-way latency can be negative due to the difference in local counters between two adjacent chips 102.
[0043] When a series of data transmissions are executed, each chip 102 (e.g., chips A and B) calculates a relative one-way latency in one direction based on the timestamp value included in the data (e.g., data 309 and data 310) and its own local counter time when the data is received. Each chip 102 can then identify the maximum relative one-way latency it has measured and transmit the maximum relative one-way latency to the system driver 104 for the calculation of each maximum loop latency. In some implementations, each chip 102 transmits the timestamp data from each transmission in a series of transmissions to the system driver 104, along with its associated local counter value at the time each transmission is received. The system driver 104 then calculates the relative one-way latency in each direction for each pair of chips, identifies the maximum one-way latency in each direction, and calculates each maximum loop latency.
[0044] Since the local counter on each chip 102 becomes an unknown state, the relative one-way latency value is meaningless by itself. However, when two relative one-way latencies (e.g., the relative one-way latency from chip A to chip B and the relative one-way latency from B back to A) between a given pair of chips 102 are summed, the difference in local counters is canceled, leaving only the absolute latency around the loop between chip A and chip B. For example, the calculation of the loop latency can be represented by the following equations. max(R b -S a )=L ab +C ba 、 max(R a -S b )=L ba -C ba 、および L inter-chip_loop_max =max(R a -S b )+max(R b -S a )=L ab +C ba +L ba-C ba =L ab +L ba R a 、R b represents the local counter time at which time-stamped data was received on chip A or chip B, respectively (e.g., in this example, R a is 60, and R b is 180). S a 、S b represents the counter time when data was transmitted by chip A or chip B, respectively (e.g., in this example, S a is 10, and S b is 200). C ba is the difference in counter time between the local counter time of chip B and the local counter time of chip A, and C ba =C b -C a is (this cannot be directly observed) (e.g., in this example, C ba is 150). L ab is the maximum jitter absolute waiting time from chip A to chip B (this cannot be directly observed). L ba is the maximum jitter absolute waiting time from chip B to chip A (this cannot be directly observed). max(R b -S a ) represents the maximum relative one-way waiting time from chip A to chip B. max(R b -S a ) is the difference between the local counter time of chip B when data was received from chip A and the local counter time of chip A when the data was transmitted. This also corresponds to the sum of the actual waiting time (L ab ) in the direction from chip A to chip B and the difference (C ba ) between the counter of chip B and the counter of chip A. max(R a -S b ) is the maximum relative one-way waiting time from chip B to chip A. max(R a -S b) is the difference between the local counter time of chip A when data is received from chip B and the local counter time of chip B when the data is transmitted. This is the actual latency (L in the direction from chip B to chip A ba ) and also corresponds to subtracting the difference (C ba ) between the counter of chip B and the counter of chip A. This relationship can also be rewritten as max(R a - S b ) = L ba + C ab and C ab is the result of subtracting the counter value from chip B from the counter value of chip A, for example, it is the inverse of C ba . Briefly speaking, the offset between the local counters on the two chips seems like the "additive" latency of transmission in one direction and the "subtractive" latency of transmission in the opposite direction. L inter-chip_loop_max represents the maximum loop latency of a given loop 112 between the two chips.
[0045] After performing some measurements of a single chip for adjacent loops 112, the system driver 104 identifies the maximum loop latency among all chip pairs (step 204). For example, the system driver 104 can compare the maximum measured loop latencies from the transmission loops 112 between each chip pair to identify the maximum chip - to - chip loop latency (L loop_max ).
[0046] One of the chips 102 or the system driver 104 determines the ring latency of data transmission around the entire ring 114 (step 206). For example, similar techniques as described with respect to FIGS. 3A - 3C are used to measure and calculate the latency around the full ring, except that time - stamped data is transmitted around the full ring and received at the same chip 102 that transmitted the data. The maximum transmission time measured around the full - ring loop 114 is the maximum full - ring latency (L ring_max ). Therefore, the difference in local counters is not a problem.
[0047] System driver 104 determines the characteristic inter-chip latency (L max ) of the multi-chip system 100 (step 208). For example, the system driver 104 can compare half of the maximum inter-chip loop latency and one-Nth of the maximum full-ring latency in the system 100 to estimate the maximum one-way latency in the system 100, where N is the total number of chips 102 within the multi-chip system 100. The larger of these two values is the characteristic inter-chip latency (L max ) of the multi-chip system 100. The system driver 104 can store the characteristic inter-chip latency for future operations. For example, the characteristic inter-chip latency can be a constant used in other operations such as startup synchronization and data transmission, as discussed below. In some implementations, the characteristic inter-chip latency is also used by the compiler to generate an operation schedule for each chip 102 to execute a specific software application, for example, a specific machine learning algorithm. For example, the characteristic inter-chip latency represents the longest time it takes for data to be transferred from one chip to an adjacent chip. The compiler can use the characteristic inter-chip latency to schedule the receiving chip to read data from the input FIFO buffer after the adjacent chip has transmitted the data and ensure that all data arrives by the scheduled read time.
[0048] In some implementations, L max can be increased by a design factor to account for any possible variations that may not have been measured during the characteristic evaluation process. For example, the measured L max may not account for the maximum possible variations in data transmission between adjacent chips. Therefore, in some implementations, L max can be increased to ensure that the actual inter-chip latency experienced by the multi-chip system 100 does not exceed the value of L max .
[0049] FIG. 4 is a flowchart of an exemplary process 400 for synchronizing the local counters of the multi-chip system 100. Process 400 is described with reference to FIGS. 1, 3A-3B, and 4. In some implementations, the process or a portion thereof is executed or controlled by the system driver 104. In some examples, process 400 or a portion thereof is executed by individual chips 102 of the multi-chip system 100. The synchronization process 400 is used to synchronize the local counters 306 of the chips 102 within the multi-chip system 100. Process 400 may be executed when the system 100 is powered on and is thus referred to as a “power-on synchronization” process. However, process 400 may similarly be executed at other times, for example, when the multi-chip system is reset.
[0050] For each chip pair, a first relative one-way latency for data transmission from the first chip (e.g., chip A) within the pair to the second chip (e.g., chip B) within the pair is determined (step 402a), and a second relative one-way latency for data transmission from the second chip (e.g., chip B) within the pair to the first chip (e.g., chip A) within the pair is determined (step 402b). For example, the relative one-way latency in the clockwise data path 108 between two chips may be determined, and then the relative one-way latency in the counterclockwise data path 108 between the two chips may be determined. The first and second relative one-way latencies may be measured using the techniques described above with reference to FIGS. 3A-3C. The chip 102 sends the measured relative one-way latency back to the system driver 104. In some implementations, the system driver 104 controls the individual chips 102 to perform the measurement of the relative one-way latency. In some implementations, the individual chips 102 include software (e.g., firmware) that controls the individual chips 102 to perform the measurement of the relative one-way latency when the system is powered on or reset.
[0051] System driver 104 determines the loop latency between each pair of chips (step 404). For example, system driver 104 can determine the loop latency between pairs of chips based on the respective one-way latencies measured between pairs of chips. For example, system driver 104 can use the formula L loop =(R a -S b )+(R b -S a ) to calculate the loop latency between a given pair of chips. System driver 104 can repeat the calculation for each loop 112 between each pair of chips within multi-chip system 100.
[0052] Optionally, system driver 104 verifies that each loop latency is less than or equal to the characteristic inter-chip latency (L max ) of the multi-chip system (step 406). For example, system driver 104 can compare the loop latency calculated for each pair of chips to the stored value of the characteristic inter-chip latency. In some implementations, if any of the calculated loop latencies is greater than the characteristic inter-chip latency, system driver 104 may re-execute the loop latency measurement. For example, system driver 104 can cause steps 402 and 404 to be re-executed. In some implementations, system driver 104 can generate an error signal if any of the calculated loop latencies is greater than the characteristic inter-chip latency.
[0053] System driver 104 uses the characteristic inter-chip latency (L max) Synchronize chip 102 by adjusting the local counter of one or more chips based on (step 408). For example, referring to FIG. 1, one chip 102 of the multi-chip system 100 can be selected as the reference chip. For example, the counter value of the reference chip will function as a basis for adjusting the respective local counters 306 of the other chips 102 within the multi-chip system 100 to synchronize chip 102. In this example, the FPGA chip is used as the reference chip. The system driver 104 starts from the reference chip and adjusts the local counter in a pairwise manner. The system driver 104 adjusts the local counter time of one chip within each pair of adjacent chips based on L max and one of the measured one-way latencies between the chips. For example, starting from the FPGA and P0, the system driver 104 adjusts the local counter time of P0 based on L max and the measured one-way latency of data transmission from the FPGA to P0 (e.g., the measured one-way latency of data transmission along the clockwise data path 108). After the local counter within P0 is adjusted, the system driver 104 adjusts the local counter of chip P1. For example, the system driver 104 adjusts the local counter time of P1 based on L max and the measured one-way latency of data transmission from P0 to P1. The system driver 104 repeats this process to adjust the local counters of each chip 102 around the ring until all chips are synchronized. However, the local counter of the FPGA (e.g., the reference chip) is not adjusted.
[0054] More specifically, using the examples shown in FIGS. 3A and 3B, the system driver 104 adjusts the local counter time of one chip within a pair of chips based on L max and one of the measured relative one-way latencies between the two chips. The system driver 104 adjusts the local counter time of one chip within a pair of chips based on L maxFrom the measured relative one-way latency from one chip in the pair to the chip whose local counter is adjusted, the local counter 306 of the chip can be adjusted by increasing the counter value by the difference. That is, the new counter value (T new ) can be determined by T new =T old +L max -(R b -S a ), where T old is the original counter value, and (R b -S a ) represents the relative one-way latency measured by the chip whose counter is adjusted. For example, in Figure 3B, the relative one-way latency from chip A to chip B was measured to be 170. If L max is 30, the adjustment to the counter of chip B is L max -(R b -S a ), that is, 30 - 170 = -140. Therefore, the system driver 104 increases the local counter of chip B by -140 counts (for example, decreases the local counter by 140). Using the simplest case as an example (for example, the counter time shown in Figure 3A), the system driver 104 will adjust the local counter of chip B from 150 to 10. The adjusted value of the local counter of chip B is not the same as the value of the local counter of chip A (for example, 0), but the two chips can be considered synchronized for the purposes of this disclosure. For example, the synchronization process does not necessarily need to make the local counters of the two chips equal, but synchronizes the data transmission latency between each pair of chips in the multi-chip system so that the maximum relative one-way latency between each pair of chips is L max or less.
[0055] In some implementations, the FIFO buffer in the Rx communication interface 308 of the chip 102 can also be adjusted. For example, the system driver 102 ensures that all of the chip loops 112 are within [2L max -3, 2Lmax Increase or decrease the receive buffer size (e.g., add or remove latency in 4 ns increments) until the loop latency within the range of max [-3, NL max is achieved, where N is the number of chips in the loop. Generally, latency only needs to be added, e.g., all 2-chip loops are within their limits, but if the full system clockwise loop 114 requires more latency, some latency may need to be removed from some of the counterclockwise pointing data paths 110. In that case, the system driver 104 removes some latency in some of the counterclockwise data paths 110 (e.g., by decreasing one or more of the receive buffers of the chips coupled to the counterclockwise data path 110) and adds the same amount of latency to the clockwise links (e.g., by increasing an appropriate data receive buffer), thereby maintaining the latency in each 2-chip loop 112 while adding latency on the clockwise system loop 114. In some implementations, latency can be adjusted by increasing or decreasing the appropriate transmitter FIFO buffer size, rather than or in addition to adjusting the receiver-side FIFO buffer.
[0056] Referring to FIG. 1, even after the chips 102 are synchronized, there may be a need to address one or more other remaining variations between the chips and in data transmission. For example, latency variations caused by forward error correction (FEC) operations may need to be addressed. This variation has a greater impact on transmission from one chip to another non-adjacent chip than on transmission between adjacent chips. For example, to transmit data from one chip to another non-adjacent chip (e.g., around a ring from the first chip P0 to the last chip P7), the data can be transmitted through each of the intervening chips (e.g., through chips P1 to P6) using a bypass operation. The latency of each data path between adjacent chips has a constant component and a variable component. The exact value of the variable component may be difficult or impossible to determine. When transmitting data from one chip (P0) to another non-adjacent chip (P7), for example, in a bypass operation, the cumulative latency variation is the sum of the latency variations of each link through which the data is transmitted. At the destination chip (e.g., P7), the cumulative latency variation can be large, resulting in a significant amount of variation in the arrival time of the transmitted data at the destination chip (e.g., P7). In a synchronization system, this variability in arrival time may require significant buffering of the data at the destination. To eliminate the arrival time variation at the destination chip (e.g., P7) resulting from the cumulative latency variation, a delay can be imposed on the data transmission from each chip 102. The introduction of a delay in each chip 102 cancels out the effect of the latency variation, and the arrival time of the data at the destination chip becomes deterministic and compatible with the synchronization system.
[0057] In some implementations, the delay is incorporated into the operation of each chip by a program compiler. For example, the program compiler generates program instructions as explicitly scheduled operations for each chip in order to L maxis used. As will be described in more detail below with reference to FIGS. 5, 6A, and 6B, each chip can be pre-scheduled to retransmit data received from an adjacent chip at a time that is the maximum inter-chip latency (e.g., L max ). For example, the operation of chip P0 can be pre-scheduled to transmit data to chip P1 at local counter time t. Chip P1 can be pre-scheduled to retransmit the data to chip P2 at local counter time t+L max .
[0058] One cause of timing variability is the characteristics of the physical link between two adjacent chips (e.g., PCS jitter), which can introduce variability in the latency of data transmission between adjacent chips 102. This cause of variability is addressed by the system characteristic evaluation and synchronization processes (200 and 400) described above. However, a second cause of timing variability is the lack of synchronization between the internal chip operations and the forward error correction scheme implemented by the multi-chip system 100. In the forward error correction scheme, error correction data is added to the data transmission between chips 102, but the added error correction data is not necessarily synchronized with the data transmission. The introduction of asynchronous data into the data transmission can introduce variability in the latency of data transmission between adjacent chips, e.g., up to 16 clock cycles.
[0059] When data is transmitted from one chip 102 to another non - adjacent chip 102 (e.g., from P0 to P7), the variation in the latency for each inter - chip transmission (e.g., from P0 to P1, from P1 to P2, etc.) accumulates in the cumulative delay at the destination chip (P7). Taking only the variation due to forward error correction as an example, the latency for a single inter - chip (e.g., from P0 to P1) transmission has a variation of ±16 clock cycles. However, some operations may require data transmission from one chip 102 to another non - adjacent chip 102, for example, data transmission from chip P0 to chip P3, or even data transmission around the ring from the first chip P0 to the last chip P7. As described in more detail below, to transmit data from one chip to another non - adjacent chip (e.g., from P0 to P7), the data can be transmitted through each of the intervening chips (e.g., through chips P1 - P6) using a bypass operation. However, the latency variations between chips accumulate over 8 chips, and the total variation in latency around the ring approaches ±128 clock cycles. To make this large variability in arrival time at the destination chip compatible with the synchronization system, a significant amount of data buffering can be implemented at the receiver interface, for example, the Rx communication interface 308, by increasing the size of the receive FIFO buffer at each chip 102.
[0060] However, additional buffering can be avoided by preventing the accumulation of latency variations throughout the multi - chip data transmission process. To achieve this, a small delay can be introduced into the data transmission operation at each chip 102 such that the latency for data transmission between each pair of adjacent chips 102 is fixed rather than variable. Specifically, the maximum inter - chip latency (L max) is determined as discussed above. During data transmission, when data is received at chip 102 in a bypass operation, instead of being transmitted immediately to the next chip, the data is stored in a receive buffer such as a FIFO buffer. The data is released from the buffer only after a maximum inter-chip latency (e.g., L max ) has elapsed since data transmission was initiated at the previous chip 102. When controlling the timing of each bypass operation in the data transmission process, the exact time of the entire data transmission process then becomes a known value, which means there is no variation in the perceived arrival time of the data at the destination chip 102.
[0061] FIG. 5 is a flowchart of an exemplary process 500 for performing data transmission between chips within the multi-chip system 100. Process 500 will be described with reference to FIGS. 1, 5, and 6A-6B. FIGS. 6A and 6B show a series of block diagrams illustrating the data transmission process 500. The block diagrams are similar to those of FIGS. 3A-3C except that the internal bypass data paths 602 and 604 are labeled. For example, chip 102 can include bypass data paths in two directions that allow chip 102 to directly route data to the next chip in the ring topology.
[0062] Process 500 or a portion thereof is executed by the individual chips 102 of the multi-chip system 100. The data transmission process 500 is used to reduce the variability of the data arrival time at the destination chips 102 within the multi-chip system 100 in order to make the data communication between chips 102 more deterministic. Further, the data transmission process 500 can reduce the required data input buffer size at each chip 102. Process 500 also enables a series of operations executed by each chip 102 within the multi-chip system 100 to be pre-scheduled and executed at pre-scheduled local counter times.
[0063] As shown in FIGS. 6A and 6B, data 606 is transmitted from a first chip (e.g., chip A) to a second chip (e.g., chip B) (step 502). For example, data 606 is bypass data targeted at another chip 102 within system 100 and not bypass data targeted at chip B. Chip B receives data 606 and stores data 606 in a buffer (step 504). For example, chip B stores data 606 in a FIFO buffer. Chip B stores the data until a maximum inter-chip latency has elapsed since chip A transmitted the data. In the example shown, chip B receives data 606 at local counter time 32, and the maximum inter-chip latency is assumed to be L max = 30 counter cycles. Thus, chip B does not transmit data 606 to the next chip (e.g., chip C) within system 100 until, for example, local counter time 40 (e.g., 40 = local counter time of chip A (10) when data 606 was transmitted to chip B plus the maximum inter-chip latency (30)). This represents an exemplary delay time of 8 counter cycles.
[0064] After the maximum inter-chip latency (e.g., L max ) has elapsed since the first chip (e.g., chip A) transmitted data 606, the second chip (e.g., chip B) releases the stored data from the buffer (step 506) and transmits the released data 608 (FIG. 6B) to a third chip (e.g., chip C) (step 508). For example, after 30 cycles of chip B's counter have elapsed since chip A transmitted data 606 to chip B, chip B releases data 606 from its FIFO buffer, passes the data along an internal bypass path 602, and can transmit the data (shown as 608 in FIG. 6B) to chip C.
[0065] In some implementations, the chip operations are explicitly scheduled at a given counter value. Thus, for example, the latency for storing bypass data in a given chip buffer is taken into account in the scheduled operations. For example, referring to the example described above, the scheduled operation instruction for chip A instructs to transmit data 606 to chip B at the local counter time of 10 for chip A. The scheduled operation instruction for chip B instructs chip B to release data 606 from its input buffer and re-transmit the data to chip C at the local counter time of 40 for chip B. Thus, chip B does not need to internally calculate the latency for re-transmitting data 606.
[0066] FIG. 7 is a schematic diagram showing an example of a dedicated logic chip (e.g., ASIC 700) that can be used as one of the chips 102 within the system 100 of FIG. 1, for example, as ASIC chips P0 - P7. The ASIC 700 includes a plurality of tiles 702, and one or more of the tiles 702 include dedicated circuits configured to perform operations such as multiplication operations and addition operations. In particular, each tile 702 can include a computational array of cells (similar to the computational unit 24 of FIG. 1), and each cell is configured to perform mathematical operations (e.g., as shown in FIG. 4, see the exemplary tile 200 described herein). In some implementations, the tiles 702 are arranged in a grid pattern, and the tiles 702 are arranged along a first direction 701 (e.g., rows) and along a second direction 703. For example, in the example shown in FIG. 7, the tiles 702 are divided into four different sections (710a, 710b, 710c, 710d), and each section includes 288 tiles arranged in a grid of 18 tiles vertically by 16 tiles horizontally. In some implementations, the ASIC 700 shown in FIG. 7 can be understood to include a single systolic array of cells subdivided / arranged into separate tiles, and each tile includes a subset / subarray of cells, local memory, and bus lines (see, e.g., FIG. 4).
[0067] ASIC 700 also includes a vector processing unit 704. The vector processing unit 704 includes circuitry configured to receive an output from tile 702 and calculate a vector calculation output value based on the output received from tile 702. For example, in some implementations, the vector processing unit 704 includes circuitry (e.g., a multiplication circuit, an adder circuit, a shifter, and / or a memory) configured to perform a cumulative operation on the output received from tile 702. Alternatively, or in addition, the vector processing unit 704 includes circuitry configured to apply a non-linear function to the output of tile 702. Alternatively, or in addition, the vector processing unit 704 generates a normalized value, a pooled value, or both. The vector calculation output of the vector processing unit may be stored within one or more tiles. For example, the vector calculation output may be stored within a memory uniquely associated with tile 702. Alternatively, or in addition, the vector calculation output of the vector processing unit 704 may be transferred as an output of the calculation to circuitry external to ASIC 700.
[0068] In some implementations, the vector processing unit 704 is segmented such that each segment includes circuitry configured to receive an output from a corresponding collection of tiles 702 and calculate a vector calculation output based on the output received. For example, in the example shown in FIG. 7, the vector processing unit 704 includes two rows extending along a first dimension 701, each row including 32 segments arranged in 32 columns. Each segment 706 includes circuitry (e.g., a multiplication circuit, an adder circuit, a shifter, and / or a memory) configured to perform a vector calculation as described herein based on the output (e.g., the accumulated sum) from the corresponding column of tile 702. The vector processing unit 704 may be positioned at the center of the grid of tiles 702 as shown in FIG. 7. Other positioning of the vector processing unit 704 is possible.
[0069] ASIC 700 also includes a communication interface 708 (e.g., interfaces 7010A, 7010B). The communication interface 708 includes one or more sets of serializer / deserializer (SerDes) interfaces and a general-purpose input / output (GPIO) interface. The SerDes interface is configured to receive input data regarding ASIC 700 and output data from ASIC 700 to an external circuit. For example, the SerDes interface can be configured to transmit and receive data at a rate of 32 Gbps, 56 Gbps, or any suitable data rate via a set of SerDes interfaces included within the communication interface 708. For example, ASIC 700 can execute a startup program when powered on. The GPIO interface can be used to load instructions (e.g., an operation schedule) into ASIC 700 and communicate with the system driver 400 to execute a startup synchronization process (e.g., process 400).
[0070] ASIC 700 further includes a plurality of controllable bus lines (see, e.g., FIG. 4) configured to transfer data among the communication interface 708, the vector processing unit 704, and the plurality of tiles 702. The controllable bus lines include, for example, wires extending along both a first dimension 701 (e.g., rows) and a second dimension (e.g., columns) of the grid. A first subset of the controllable bus lines extending along the first dimension 701 can be configured to transfer data in a first direction (e.g., to the right in FIG. 7). A second subset of the controllable bus lines extending along the first dimension 701 can be configured to transfer data in a second direction (e.g., to the left in FIG. 7). A first subset of the controllable bus lines extending along the second dimension 703 can be configured to transfer data in a third direction (e.g., upward in FIG. 7). A second subset of the controllable bus lines extending along the second dimension 703 can be configured to transfer data in a fourth direction (e.g., downward in FIG. 7).
[0071] Each controllable bus line includes a plurality of transmission elements, such as flip-flops, that are used to transfer data along the line according to a clock signal. Transferring data via a controllable bus line can include shifting data from a first transmission element of the controllable bus line to a second adjacent transmission element of the controllable bus line in each clock cycle. In some implementations, data is transferred via a controllable bus line at the rising edge or falling edge of the clock cycle. For example, in the first clock cycle, data present on a first transmission element (e.g., a flip-flop) of a controllable bus line can be transferred to a second transmission element (e.g., a flip-flop) of the controllable bus line in the second clock cycle. In some implementations, the transmission elements can be periodically spaced at a fixed distance from each other. For example, in some cases, each controllable bus line includes a plurality of transmission elements, and each transmission element is disposed within or proximate to a corresponding tile 702.
[0072] To minimize the latency associated with the internal operation of the ASIC chip 700, the tile 702 and the vector processing unit 704 can be arranged to reduce the distance that data travels between various components. In certain implementations, both the tile 702 and the communication interface 708 can be segmented into a plurality of sections, and both the tile section and the communication interface section are arranged such that the maximum distance that data travels between the tile and the communication interface is reduced. For example, in some implementations, a first group of tiles 702 can be disposed within a first section on a first side of the communication interface 708, and a second group of tiles 702 can be disposed within a second section on a second side of the communication interface. As a result, the distance to the tile furthest from the communication interface can be halved compared to a configuration where all of the tiles 702 are disposed within a single section on one side of the communication interface.
[0073] Alternatively, the tiles may be arranged in a different number of sections, such as four sections. For example, in the example shown in FIG. 7, a plurality of tiles 702 of the ASIC 700 are arranged in a plurality of sections 710 (710a, 710b, 710c, 710d). Each section 710 includes the same number of tiles 702 arranged in a grid pattern (e.g., each section 710 can include 256 tiles arranged in 16 rows and 16 columns). The communication interface 708 is also divided into a plurality of sections, and a first communication interface 7010A and a second communication interface 7010B are arranged on either side of the section 710 of the tile 702. The first communication interface 7010A can be coupled to the two tile sections 710a, 710c on the left side of the ASIC chip 700 via controllable bus lines. The second communication interface 7010B can be coupled to the two tile sections 710b, 710d on the right side of the ASIC chip 700 via controllable bus lines. As a result, the maximum distance (and thus the latency associated with data propagation) for data to move to and / or from the communication interface 708 can be halved compared to an arrangement where only a single communication interface is available. Other coupling arrangements of the tiles 702 and the communication interface 708 are also possible to reduce data latency. The coupling arrangement of the tiles 702 and the communication interface 708 can be programmed by providing control signals to the transmission elements of the controllable bus lines and multiplexers.
[0074] In some implementations, one or more tiles 702 are configured to initiate read and write operations with respect to controllable bus lines and / or other tiles within the ASIC 700 (referred to herein as "control tiles"). The remaining tiles within the ASIC 700 can be configured to perform calculations based on input data (e.g., to compute layer inferences). In some implementations, the control tiles include the same components and configuration as other tiles within the ASIC 700. The control tiles can be added as additional tiles, additional rows, or additional columns of the ASIC 700. For example, in the case of a symmetric grid of tiles 702 where each tile 702 is configured to perform calculations on input data, one or more additional rows of control tiles can be included to handle read and write operations for the tiles 702 that perform calculations on input data. For example, each section 710 can include 18 rows of tiles, and the last two rows of tiles can include control tiles. Providing separate control tiles, in some implementations, increases the amount of memory available in the other tiles used to perform calculations. However, a dedicated separate tile for providing control as described herein is not necessary and, in some cases, no separate control tiles are provided. Rather, each tile can store instructions for initiating read and write operations for that tile within its local memory.
[0075] Further, each section 710 shown in FIG. 7 includes tiles arranged in 18 rows by 16 columns, although the number and arrangement of tiles 702 within a section can vary. For example, in some cases, section 710 can include an equal number of rows and columns.
[0076] Further, although FIG. 7 shows it as being divided into four sections, tile 702 can be divided into other different groups. For example, in some implementations, tile 702 is grouped into two different sections, such as a first section above vector processing unit 704 (closer to the top of the page shown in FIG. 7) and a second section below vector processing unit 704 (closer to the bottom of the page shown in FIG. 7). In such an arrangement, each section can include, for example, 596 tiles arranged in a grid that is 18 tiles tall (along direction 703) by 32 tiles wide (along direction 701). The sections can include other total numbers of tiles and can be arranged in arrays of different sizes. In some cases, the division between sections is demarcated by the hardware functionality of ASIC 700. For example, as shown in FIG. 7, sections 710a, 710b can be separated from sections 710c, 710d by vector processing unit 704.
[0077] Embodiments and functional operations of the subject matter described in this specification can be implemented in digital electronic circuitry, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of a computer program encoded on a tangible non-transitory program carrier for execution by, or to control the operation of, a data processing apparatus. Alternatively, or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, generated to carry information for transmission to an appropriate receiver apparatus for execution by a data processing apparatus. A computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0078] The term "data processing apparatus" includes, by way of example, all kinds of apparatus, devices, and machines for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus can include dedicated logic circuitry, such as, for example, an FPGA (Field Programmable Gate Array) or an ASIC. The apparatus can also include, in addition to hardware, code for creating an execution environment for the computer programs in question, such as processor firmware, a protocol stack, a database management system, an operating system, or code constituting one or a combination of more of them.
[0079] The processes and logical flows described herein can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by dedicated logic circuitry, such as, for example, an FPGA, an ASIC, or a GPGPU (General-Purpose Graphics Processing Unit), and the apparatus can also be implemented as dedicated logic circuitry, such as, for example, an FPGA, an ASIC, or a GPGPU (General-Purpose Graphics Processing Unit).
[0080] This specification includes details of many specific implementations, but these should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. The specific features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable partial combination. Further, features may be described above as acting in a particular combination and may even initially be claimed as such, but one or more features from the claimed combination may in some cases be excised from the combination, and the claimed combination may be directed to a partial combination or a variation of a partial combination.
[0081] Similarly, operations are depicted in the drawings in a particular order, but this should not be understood as requiring that such operations be performed in the particular order or sequence shown, or that all illustrated operations be performed, to achieve the desired result. Multitasking and parallel processing may be advantageous in certain circumstances. Further, the separation of various modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the program components and systems described may generally be integrated into a single software product or packaged into multiple software products.
[0082] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, although the bus lines are described as "controllable", not all bus lines need to have the same level of control. For example, various degrees of controllability may exist, and some bus lines may be controllable only if they are limited with respect to the number of tiles that can supply data to or transmit data from those bus lines. In another example, some bus lines may be dedicated to supplying data along a single direction such as north, east, west, or south, as described herein. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. As an example, the processes depicted in the accompanying drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing may be advantageous.
Explanation of Signs
[0083] 24 computing units 100 multi-chip system, system 102 semiconductor chip, chip, ASIC chip, FPGA chip, ASIC, destination chip 104 system driver, FPGA chip 106 clock 108 clockwise path, clockwise data path 110 counterclockwise path, counterclockwise pointing data path, counterclockwise data path 112 inter-chip communication loop, loop, chip loop, transmission loop, 2-chip loop 114 inter-chip communication loop, loop, ring, full ring loop, full system loop, clockwise system loop 200 tile 304 controller 306 local counter 308 communication interface, Rx communication interface 309 Data with Timestamp, Data, Data Transmission 310 Data 602 Internal Bypass Data Path, Internal Bypass Path 604 Internal Bypass Data Path 606 Data 608 Data 700 ASIC, ASIC Chip 701 First Direction, First Dimension 702 Tile 703 Second Direction, Second Dimension 704 Vector Processing Unit 706 Segment 708 Communication Interface 710 Section 710a Section 710b Section 710c Section 710d Section 7010A Interface, First Communication Interface 7010B Interface, Second Communication Interface
Claims
1. A method for evaluating the inter-chip waiting time characteristics between integrated circuit chips, comprising: For each pair of integrated circuit chips within the topology of the integrated circuit chips, Determining the inter-chip loop waiting time for each pair of the integrated circuit chips; Identifying the maximum inter-chip loop waiting time from among the inter-chip loop waiting times determined for each pair of the integrated circuit chips; Calculating the full-path waiting time for the topology of the integrated circuit chips; Generating an inter-chip waiting time representing the operating characteristics of the topology based on the maximum inter-chip loop waiting time and the full-path waiting time; And each step is executed by a system driver of a multi-chip system and a controller of each integrated circuit chip, the method for evaluating the inter-chip waiting time characteristics between integrated circuit chips.
2. The step of determining the inter-chip waiting time for each pair of the integrated circuit chips comprises: Transmitting first timestamped data from a first chip of the pair of integrated circuit chips to a second chip of the pair of integrated circuit chips; Determining a first relative one-way waiting time between the pair of the integrated circuit chips based on the first timestamped data; Transmitting second timestamped data from the second chip to the first chip; Determining a second relative one-way waiting time between the pair of the integrated circuit chips based on the second timestamped data; Determining the inter-chip loop waiting time for the pair of the integrated circuit chips based on the first relative one-way waiting time and the second relative one-way waiting time; The method according to claim 1, comprising.
3. The method according to claim 2, wherein the first timestamped data indicates the local counter time of the first chip when the first timestamped data is transmitted.
4. The step of determining the first relative one-way waiting time between the pair of the integrated circuit chips comprises calculating a difference between the time indicated in the timestamped data and the local counter time of the second chip when the second chip receives the first timestamped data. The method according to claim 2, comprising.
5. The method according to claim 2, further comprising determining a loop latency of round-trip data transmission between the pairs of the integrated circuit chips by calculating a difference between the first relative one-way latency and the second relative one-way latency. **Claim 6** The method according to claim 1, wherein two or more of the integrated circuit chips are application-specific integrated circuit (ASIC) chips configured to perform neural network operations. **Claim 7** The method according to claim 6, wherein one of the integrated circuit chips is a field programmable gate array chip. **Claim 8** The method according to claim 6, wherein the ASIC chips are each configured to implement a corresponding layer of a neural network. **Claim 9** The method according to claim 8, wherein a first ASIC chip is configured to implement an input layer of the neural network, and a second ASIC chip is configured to implement a second layer of the neural network based on an output from the first ASIC chip. **Claim 10** The method according to claim 1, wherein the integrated circuit chips form a closed loop. **Claim 11** The method according to claim 1, wherein each of the integrated circuit chips is configured to perform an internal operation that is synchronous and deterministic. **Claim 12** A processing device, a plurality of integrated circuit chips, a non-transitory machine-readable storage device for storing instructions, and a system comprising: when the instructions are executed by the processing device, for each pair of integrated circuit chips in the topology of the integrated circuit chips, determining an inter-chip loop latency for each pair of the integrated circuit chips; identifying a maximum inter-chip loop latency from among the inter-chip loop latencies determined for each pair of the integrated circuit chips; calculating a full-path latency for the topology of the integrated circuit chips; generating an inter-chip latency representing an operating characteristic of the topology based on the maximum inter-chip loop latency and the full-path latency; causing execution of operations including a system. **Claim 13** The system according to claim 12, wherein two or more of the integrated circuit chips are application-specific integrated circuit (ASIC) chips configured to perform neural network operations. **Claim 14** The system according to claim 13, wherein one of the integrated circuit chips is a field programmable gate array chip.
15. The system according to claim 13, wherein each of the ASIC chips is configured to implement a corresponding layer of the neural network.
16. The system according to claim 15, wherein a first ASIC chip is configured to implement an input layer of the neural network, and a second ASIC chip is configured to implement a second layer of the neural network based on an output from the first ASIC chip.
17. The system according to claim 12, wherein the integrated circuit chips form a closed loop.
18. The system according to claim 12, wherein each of the integrated circuit chips is configured to perform an internal operation that is synchronous and deterministic.
19. The system according to claim 12, comprising a clock, the clock being coupled to each of the integrated circuit chips, the clock providing a synchronization timing signal to the integrated circuit chips.
20. The system according to claim 19, wherein each of the integrated circuit chips includes a corresponding local clock or counter.
Citation Information
Patent Citations
Method for measuring performance of network by network node
JP2000156721A
Ring-shape synchronous network system
JP2011199420A
Network traffic mapping and performance analysis
JP2016512416A
Terminal, communication system, and operation sharing method
WO2013073125A1