Error checking for systolic array computations.
The integration of checksum circuits within the computational unit of a systolic array allows for real-time error detection during matrix multiplication, addressing the challenges of pre- and post-processing, and enhancing the efficiency and accuracy of systolic array computations.
Patent Information
- Application Number
- JP2023548914
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-24
- Filing Date
- 2022-07-13
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-07-13
AI Technical Summary
Existing systolic array computations face challenges in efficiently detecting errors during matrix multiplication, particularly in the absence of pre-processing to generate checksums for input matrices and post-processing to compare checksums in output matrices.
A computational unit with a systolic array and integrated checksum circuits that calculate and compare checksums in real-time, allowing for error detection without pre-processing the input matrices or post-processing the output matrices, and enabling rapid testing of hardware faults.
This approach enables rapid and efficient error detection in systolic array computations, allowing for the identification of soft and hard errors without slowing down the operation of the systolic array, and facilitating better energy efficiency and accuracy in critical tasks.
Smart Images

Figure 0007682283000001 
Figure 0007682283000002 
Figure 0007682283000003
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of U.S. Patent Application No. 17 / 410,558, filed August 24, 2021, which claims the benefit of the filing date of U.S. Provisional Application No. 63 / 222,549, entitled "Error Checking For Systolic Array Computation," filed July 16, 2021, the disclosures of which are incorporated herein by reference. [Background technology]
[0002] background A systolic array is an array of processing elements, such as processors, microprocessors, or special purpose circuits, configured to process some data. Adjacent processing elements of a systolic array may be connected via one or more interconnects, such as wires or other physical connections on a printed circuit board.
[0003] Algorithm-based fault tolerance (ABFT) refers to a scheme or technique for detecting and correcting errors during the execution of different types of arithmetic or logic algorithms such as matrix multiplication, Fourier transform, etc. In the case of matrix multiplication, for example, A×B=C for input matrices A, B, and output matrix C, ABFT for matrix multiplication involves generating a checksum row of A and a checksum column of matrix B. Each element of the checksum row of A is the result of a linear operation performed on the elements of the respective column of matrix A. For example, each element of the column of matrix A is added to generate a checksum value in the checksum row of matrix A. Similarly, each checksum value in column B is the result of a linear operation performed on the elements of the respective row of matrix B. Summary of the Invention [Problem to be solved by the invention]
[0004] After multiplying matrices A and B (with corresponding checksum rows / columns), the output matrix includes sub-matrices that represent the product of multiplying matrices A and B, as well as checksum rows and checksum columns. As part of the ABFT of the matrix multiplication, the checksum values of the checksum rows and columns of matrix C are compared to the results of performing the same linear operation on matrices A and B, but now performed on the sub-matrices of C. If the result of performing the linear operation on a row or column of a sub-matrix of C does not match the corresponding checksum value in C, then the mismatch indicates that an error occurred during the matrix multiplication. [Means for solving the problem]
[0005] Quick Overview Aspects of the present disclosure are directed to a computational unit configured to implement a systolic array and detect errors while processing data on the systolic array. A checksum circuit in communication with the systolic array is configured to calculate checksums and perform error detection while the systolic array is processing the input data. Instead of pre-generating the checksums on the input matrix, the input matrix can be directly fed to the systolic array via the checksum circuit. On the output side, according to one aspect, the checksum circuit can generate a checksum of an output to the systolic array when the systolic array finishes the corresponding operation to generate the output or while the output is being streamed from the systolic array. According to another aspect, on the output side, the checksum circuit can generate and compare a checksum in an output matrix generated by the systolic array with the checksum.
[0006] Aspects of the present disclosure provide many technical advantages. Error detection through checksums of the operands and outputs of the systolic array can be performed without pre-processing operations to generate checksums of the inputs or post-processing operations to compare the correctness of the checksums. Systolic arrays on a chip can be rapidly tested according to a variety of random or pseudo-random test patterns to determine whether hardware faults exist in the array or in individual processing elements of the array. Since no known inputs or outputs are required, and no pseudo-random test signatures are required, the variety of test patterns required for testing can be generated and deployed. Additionally, error checking of the operations to generate the output matrix can be performed without slowing down the operation of the systolic array and without pre-processing the input matrix.
[0007] At run time, a processor implemented according to aspects of the present disclosure can detect soft or hard errors, e.g., errors not caused by hardware defects, and errors caused by hardware defects. Rapid identification of errors can be particularly important when the processor is part of an accelerator that performs critical, time-sensitive tasks. For example, soft error detection can be particularly important when the processor is processing inputs for machine learning models trained to perform tasks related to banking, autonomous vehicle navigation / control, airplane or spacecraft navigation, etc.
[0008] In some examples, the computational units described herein may be configured to detect timing violation errors, which may be detected and addressed to provide better energy efficiency to the systolic array. Circuits for performing error detection are implemented to reduce the likelihood of timing errors and allow accurate error detection to continue even when the systolic array is operating at a supply voltage below a predetermined critical voltage. Aspects of the present disclosure also provide for detection of errors in several types of operations, including matrix multiplication, which may be modified to improve the accuracy and reliability of the processor. Different computational units may be tuned to different voltage and / or frequency levels to tune the units for better performance, for example, during inference of a machine learning model performed by the computational units. Detecting errors may improve the overall accuracy of the computation.
[0009] Aspects of the present disclosure can be implemented with minimal overhead and without affecting the performance of the systolic array of the computational unit. No pre-processing of the inputs is required, further improving the efficiency of error checking of the computational unit over other approaches in which the inputs are pre-processed. Furthermore, as described herein, the input to the checksum circuitry to perform error detection does not delay the input to the systolic array for processing.
[0010] "Relaxed" fault tolerance can also be applied to further reduce overhead, especially software interaction with the computational units, and only applied when an error is actually detected. In some instances, relaxed fault tolerance may be beneficial when the primary use case is to detect the presence of an error without needing to identify the exact cause of the error.
[0011] With error detection applied to the computation unit as described herein, an error correction mechanism, e.g., an ABFT-based error correction process for matrix multiplication, can be implemented to recover from hard errors detected during the execution of the computation unit.
[0012] One aspect of the disclosure is directed to a computational unit including a two-dimensional systolic array of processing elements. The systolic array is configured to receive first input elements from a first input matrix along a first direction of the systolic array and second input elements from a second input matrix along a second direction of the systolic array. The computational unit further includes the two-dimensional systolic array, a first checksum circuit configured to generate one or more groups of first checksums from the first input elements while the systolic array is receiving the first input elements, and a second checksum circuit configured to generate one or more groups of second checksums while the systolic array is receiving the second input elements. The systolic array is further configured to generate an output matrix from the first input matrix, the second input matrix, the one or more groups of first checksums, and the one or more groups of second checksums. The computational unit further includes an output checksum circuit configured to receive the output matrix and determine from the output matrix the occurrence of one or more errors in the generation of the output matrix.
[0013] The aforementioned aspects can include one or more of the following features, either alone or in any combination. In some examples, the aforementioned aspects include all of the following features together.
[0014] The output matrix includes a data sub-matrix, an output checksum row, and an output checksum column. The data sub-matrix includes values generated by the systolic array using the first input elements and the second input elements. To determine whether one or more errors occurred in generating the output matrix, the output checksum circuit is configured to generate a row checksum from at least one row of the data sub-matrix, compare the row checksum to a checksum in the output checksum column, and determine the occurrence of an error in generating the output matrix from a comparison of the row checksum to the checksum in the output checksum column.
[0015] To determine the occurrence of one or more errors in the generation of the output matrix, the output checksum circuitry is further configured to generate a column checksum from at least one column of the data sub-matrix, compare the column checksum to a checksum of the output checksum row, and determine the occurrence of an error in the generation of the output matrix from the comparison of the column checksum to the checksum of the output checksum row.
[0016] To determine the occurrence of one or more errors in the generation of the output matrix, the output checksum circuitry is further configured to generate a row checksum from at least one row of the data sub-matrix, compare the row checksum to a checksum in the output checksum column, and determine the occurrence of an error in the generation of the output matrix from a comparison of the row checksum to the checksum in the output checksum column.
[0017] To compare the row checksum to the checksum in the output checksum column, the output checksum circuit is further configured to determine whether an absolute difference between the row checksum and the checksum in the output checksum column is within a predetermined threshold.
[0018] The computing unit further includes one or more checksum processing elements configured to receive the checksum from one or both of the first and second checksum circuits.
[0019] Either or both of the first and second checksum circuits are configured to transmit the first or second checksum to the systolic array for processing based on the control signal.
[0020] The timing of the control signals to one or both of the first and second checksum circuits is based on the number of time steps to load the first or second input value across the systolic array.
[0021] The computation unit is further configured to, in response to determining the occurrence of the one or more errors in the generation of the output matrix, transmit an indication of the occurrence of the one or more errors to one or more devices coupled to the computation unit.
[0022] The systolic array is further configured to receive a regulated voltage from the one or more devices after transmitting an indication of the occurrence of the one or more errors, the regulated voltage being greater than a critical voltage of the computational unit.
[0023] The systolic array is further configured to receive a first voltage that is less than a critical voltage of the computational unit until receiving an adjusted voltage in response to transmitting the indication.
[0024] One or both of the first and second checksum circuits are configured to receive a second voltage that is greater than the critical voltage.
[0025] One or both of the first and second checksum circuits include a two-input, two-stage pipelined adder circuit configured to delay generation of one or both of the first and second checksums.
[0026] One or both of the first and second checksum circuits include one or more two-cycle adding circuits and a plurality of registers, the plurality of registers including one or more first registers configured to transmit and receive data according to a first clock frequency, one or more second registers configured to transmit and receive data according to a second clock frequency, and one or more third registers configured to transmit and receive data according to a third clock frequency, wherein the first, second, and third clock frequencies are all different frequencies.
[0027] In some examples, the systolic array is fixed weight. In other examples, the systolic array is a fixed output systolic array, and the first and second checksum circuits are both connected to a plurality of checksum processing elements, and the checksum processing elements are configured to receive the checksums generated from one or both of the first and second checksum circuits.
[0028] A number of checksum processing elements are arranged along the periphery of the systolic array. The computation unit is configured to generate the checksum only from the first checksum circuit or to generate the checksum only from the second checksum circuit.
[0029] One aspect of the disclosure is directed to a data processing system, the data processing system including one or more processors, one or more memory devices, and a computation unit. The computation unit includes a two-dimensional systolic array of processing elements. The systolic array is configured to receive first input elements from a first input matrix along a first direction of the systolic array and to receive second input elements from a second input matrix along a second direction of the systolic array. The computation unit further includes a first checksum circuit configured to generate one or more groups of first checksums from the first input elements while the systolic array is receiving the first input elements, and a second checksum circuit configured to generate one or more groups of second checksums while the systolic array is receiving the second input elements. The systolic array is further configured to generate an output matrix from the first input matrix, the second input matrix, the one or more groups of first checksums, and the one or more groups of second checksums. The computation unit further includes an output checksum circuit configured to receive the output matrix and determine from the output matrix the occurrence of one or more errors in the generation of the output matrix.
[0030] The aforementioned aspects can include one or more of the following features, either alone or in combination. In some examples, the aforementioned aspects can include all of the following features together.
[0031] The data processing system is configured to apply a voltage to the systolic array, where the applied voltage is below a critical voltage of the systolic array, and is configured to receive an indication of one or more errors from the computational unit and, in response, to increase the applied voltage above the critical voltage of the systolic array.
[0032] The data processing system is further configured to apply a voltage below the critical voltage to one or more of the first checksum circuit, the second checksum circuit, and the output checksum circuit of the systolic array, and in response to receiving the indication, to continue applying a voltage below the critical voltage to one or more of the first checksum circuit, the second checksum circuit, and the output checksum circuit.
[0033] The data processing system is configured to send a control signal to one or both of the first and second checksum circuits, the timing of the sending being based on a number of time steps for loading the first or second input value across the systolic array.
[0034] One aspect of the present disclosure is directed to one or more non-transitory computer-readable storage media encoded with computer instructions to be executed by a computing unit. The computing unit includes a two-dimensional systolic array of processing elements, a first checksum circuit, and a second checksum circuit. An output checksum circuit causes the computing unit to perform an operation. The operation includes receiving, by the systolic array, a first input element from a first input matrix along a first direction of the systolic array, receiving, by the systolic array, a second input element from a second input matrix along a second direction of the systolic array, generating, by the first checksum circuit, one or more groups of first checksums from the first input element while the systolic array is receiving the first input element, generating, by the second checksum circuit, one or more groups of second checksums from the second input element while the systolic array is receiving the second input element, generating, by the systolic array, an output matrix from the first input matrix, the second input matrix, one or more groups of first checksums, and one or more groups of second checksums, receiving, by the output checksum circuit, the output matrix, and determining, by the output checksum circuit, the occurrence of one or more errors in the generation of the output matrix from the output matrix.
Brief Description of the Drawings
[0035] [Figure 1] A block diagram of a computing unit according to an aspect of the present disclosure. [Diagram 2] An exemplary input matrix processed by a computing unit to generate respective output matrices with corresponding checksums according to an aspect of the present disclosure is shown. [Diagram 3] A block diagram of a vertical checksum circuit and checksum processing elements. [Figure 4] A block diagram of a horizontal checksum circuit. The horizontal checksum circuit 115 includes several addition circuits, registers, and multiplexers. [Diagram 5]1 is a block diagram of an output checksum circuit configured to generate a checksum from the rows and columns of data sub-matrices of an output matrix C. [Figure 6] 1 is an exemplary computation unit having a vertical checksum circuit and an output checksum circuit, but no horizontal checksum circuit. [Figure 7] 1 is an exemplary computational unit having an output fixed systolic array. [Figure 8A] 1 is an exemplary vertical checksum circuit configured to generate checksums from rows of an input matrix of length 8. [Figure 8B] 1 is an exemplary vertical checksum circuit having a two-input, two-stage pipelined adder circuit. [Figure 8C] FIG. 11 is a circuit diagram of another example vertical checksum circuit with a two-cycle non-pipelined adder circuit. [Figure 9A] 1 is a flowchart of an example process for performing matrix multiplication with error detection on a computation unit in accordance with an aspect of the present disclosure. [Figure 9B] 1 is a flowchart of an exemplary process for detecting the occurrence of an error during processing of a computing unit, according to an aspect of the present disclosure. [Figure 10A] FIG. 1 is a block diagram of a data processing system implementing an exemplary computing unit. [Figure 10B] 1 is a flowchart of an exemplary process for adjusting a supply voltage to a computing unit according to an aspect of the present disclosure. [Figure 11] 10 is a block diagram of an example environment for implementing a data processing system including a computing unit 1000. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0036] Detailed Description 1 is a block diagram of a computational unit 100 according to an embodiment of the present disclosure. Computational unit 100 includes a systolic array 110 of processing elements 110A-P, a horizontal checksum circuit 115, a vertical checksum circuit 120, an output checksum circuit 125, and a checksum processing element 130.
[0037] The systolic array 110 may be implemented in a variety of different ways, and may generally include one or more data buses or interconnects connecting different processing elements to adjacent processing elements. For example, the systolic array may include a data bus along each row and / or column of the systolic array 110, whose processing elements 110A-P may be configured in a square or rectangular arrangement (although the processing elements 110A-P may also be configured in other geometric arrangements, such as a hexagonal arrangement, in some examples). In some examples, the systolic array 110 may include one or more data buses along processing elements that are diagonally above or below each other in the array. In some examples, the systolic array 110 may implement two-phase non-overlapping clocks and latches to take advantage of time borrowing.
[0038] The data bus allows data to flow from processing element to processing element, or from external input to processing element, or from processing to external output. External inputs / outputs can include, for example, other devices communicatively coupled to the computing unit 100, or memory devices that are part of the computing unit 100 or located outside the computing unit 100. In some examples, the systolic array 110 can be configured to receive or generate inputs or outputs outside the systolic array 110 only via the processing elements around the systolic array 110. In other examples, the systolic array 110 can be configured to receive or generate external inputs or outputs from any of the processing elements 110A-P.
[0039] A processing element of the systolic array 110 may be a processor, microprocessor, or any dedicated circuit configured to receive one or more inputs, process the inputs, and generate one or more outputs. A processing element may be, for example, a data processing unit (DPU) configured to perform a particular type of operation. In some examples, a processing element may be configured to store memory temporarily, for example in a register, or using one or more circuits such as latches or flip / flop circuits.
[0040] The processing elements 110A-P may be arranged according to several topologies. For example, the processing elements 110A-P may be arranged as an array, a mesh, or a cube. Each processing element may have one or more channels that connect to adjacent processing elements. The one or more channels may be, for example, wires or other physical connections on the silicon wafer on which the systolic array 110 is implemented. Depending on the number of channels of each processing element, the systolic array 110 may have one or more dimensions. For purposes of explanation, the systolic array 110 is assumed to be a two-dimensional mesh, although it is understood that in other implementations, the systolic array 110 may be configured according to other topologies without loss of generality.
[0041] Although systolic array 110 is shown having 16 processing elements arranged as a 4×4 mesh, it will be appreciated that in various examples the number of processing elements in systolic array 110 may vary, such as, for example, 65,536 processing elements arranged as a 256×256 mesh, or 16,384 processing elements arranged as a 128×128 mesh.
[0042] Data may travel across each data bus or interconnect of the systolic array 110 independently of one another. In other words, different processing elements may receive and transmit different data to adjacent processing elements at different times, e.g., measured in time steps. A time step may refer to one or more clock cycles of the computational unit 100, or a larger system communicatively coupled to the computational unit 100. For ease of description in this specification, a time step in this specification refers to a single clock cycle unless otherwise specified. For each time step, each processing element 110A-P may perform a different operation. For example, in a single time step, some processing elements may be idle due to receiving new input, transmitting generated output, processing received input, or lack of data to receive or transmit.
[0043] The functionality of the processing elements 110A-P may depend on the functionality of the systolic array 110, e.g., the operations that the systolic array 110 is configured to perform. For example, the systolic array 110 may be configured to perform a matrix multiplication on two input matrices, e.g., a weight matrix of a neural network layer and an input matrix to that layer. In another example, the systolic array 110 may be configured to perform a convolution operation as part of the received input of a convolutional neural network ("ConvNet"). In a matrix multiplication example, each processing unit 110A0P may be configured to perform a multiply and accumulate ("MAC") operation. A processing element configured to perform a MAC operation may be referred to as a MAC unit. Aspects of the present disclosure are directed to error checking and correction during matrix multiplication performed by the systolic array of a MAC unit, and the following examples assume that the processing elements 110A-P are configured for MAC operations unless otherwise noted.
[0044] As an example, a processing element configured as a MAC unit may process data as follows: At each time step, the processing element may multiply a matrix element loaded into the unit by a received input matrix element and add the product of the multiplication to a running partial sum maintained by the MAC unit. The MAC unit may forward the input matrix element in a first direction (e.g., to the right) along a data bus to an adjacent MAC unit and pass the updated partial sum in a second direction (e.g., down) to the MAC unit along the same or a different data bus.
[0045] The systolic array 110 can be configured to perform operations at various levels, such as pipelined sub-operations, different size inputs, debug, etc. For example, one operation on the systolic array 110 may be performed by an operation that is performed individually by a group of processing elements, e.g., a subarray of 2×2 or 4×4 processing elements. In other examples, the systolic array 110 can be configured to perform more complex operations, such as a discrete Fourier transform, or to define different control flows that govern how data is input and output to the systolic array 110. The systolic array 110 can also perform operations, such as matrix multiplication, on input blocks that are part of a larger matrix that the systolic array 110 is tasked with multiplying. In other examples, the different processing units of the systolic array 110 can be implemented with different types or combinations of circuits, e.g., adders, multipliers, accumulators, circuits configured to compare two inputs, etc.
[0046] Although the examples in this specification focus on algorithm-based fault tolerance for matrix multiplication on systolic array 110, it is understood that aspects of the disclosure are applicable to error detection on systolic arrays configured to perform other types of operations under any of a variety of conditions, including, for example, the operations and conditions described herein.
[0047] Systolic array 110 is configured to perform matrix multiplication on input matrices having different data types for each element. For example, systolic array 110 can perform matrix multiplication on binary matrices having 1s and 0s as elements; integer matrices having positive or negative integers as elements; floating point matrices with floating point values as elements, etc. The elements can be of different levels of precision; for example, 8-bit integers, or floating point values with different ranges and levels of precision, etc.
[0048] As described in more detail herein, the vertical checksum circuit 120 and the horizontal checksum circuit 115 are configured to generate checksums of the incoming input matrices 140A, B. References to "vertical" and "horizontal" checksum circuits are not intended to limit the direction of inputs received by the systolic array 110. However, this is merely provided for ease of explanation in conjunction with the direction of the input matrices for the example circuit shown in the figures. In some examples, the vertical checksum circuit 120 described herein can be oriented to provide inputs to the systolic array 110 along the columns of the array instead of the columns, and vice versa for the horizontal checksum circuit 115 and rows instead of columns of the array. The checksum circuits 115, 120 can be positioned in any orientation relative to one another to provide inputs to the systolic array 110.
[0049] The checksum processing element 130 includes an additional processing element for storing the checksums generated for each column of the input matrix A 140A. As each column is processed by the vertical checksum circuit 120, the resulting checksums are pushed out to the checksum processing element 130, which may be communicatively coupled to the systolic array 110 and may be included during the multiplication of the input matrices 140A-B that are augmented with the respective checksum columns or rows by the checksum circuits 115, 120. As described in more detail herein, the orientation and location of the checksum processing element 130 may also vary relative to the orientation and location of the vertical checksum circuit 120.
[0050] Output checksum circuit 125 is configured to generate a checksum of a resulting output matrix that is the product of systolic array 110 multiplying input matrix A 140A (and its checksum) by input matrix 140B (and its checksum). The output matrix includes sub-matrices that are the products of input matrix A multiplied by input matrix B, and checksum columns and checksum rows generated by output checksum circuit 125. Output checksum circuit 125 is configured to compare the checksum columns and rows to the sums of the respective columns or rows of the sub-matrices and detect errors indicated by a discrepancy between the calculated sum compared to the checksum value.
[0051] FIG. 2 illustrates example input matrices 140A,B processed by computation unit 100 to generate respective output matrices 140C with corresponding checksums in accordance with an embodiment of the present disclosure.
[0052] In this example, input matrix A 140A ("matrix A") is multiplied by input matrix B 140B ("matrix B"), which can be represented as A×B. In FIG. 2, the input matrices are 4×4, matching the dimensions of the systolic array 110. In other examples, the dimensions of the input matrices may be different. For example, matrix A is n×4 and matrix B is 4×m, where n, m are positive integers. As described herein, the systolic array 110 may be of any size, so the internal dimensions of matrices A and B (four in this example) may also vary based on the size of the systolic array 110.
[0053] Based on the exemplary orientation and location of the checksum circuits 115, 120, input matrix A 140A ("input matrix A") and input matrix B 140B ("input matrix B") are first manipulated to ensure that the order in which the elements of the matrices are received by the systolic array 110 allows for correct multiplication according to the rules of matrix multiplication. A sequencer configured as part of a data processing system implementing the computational unit 100 can be implemented to perform some or all of the operations currently described in connection with the manipulation of matrices A, B and adjusting the timing at which the elements of the input matrices are provided to the systolic array 110. The sequencer and other components of an exemplary data processing system are described herein with reference to FIG. 10A.
[0054] The input matrix A is transposed (A T ), each of whose columns is shifted by one time step, e.g., one clock cycle. A time step may define one or more clock cycles or periods, which may be managed by a timing circuit, e.g., a timing circuit described herein with reference to FIG. 10A. T Each column of A passes through a respective processing element of the vertical checksum circuit 120. As described herein with reference to FIG. T The vertical checksum circuit 120 generates a checksum for each column of A, which corresponds to a row of the original input matrix A. T After processing the received elements, the received elements are passed to the systolic array 110 for matrix multiplication. The generated checksums are also pushed out to the checksum processing element 130.
[0055] The computation unit 100 can be configured to receive input matrices of different block sizes. The computation unit 100 can receive multiple matrices that are sub-matrices of a larger input. For example, the computation unit 100 can be 128×128 elements and can be configured to perform block matrix multiplication, where the original input matrix is large, e.g., 4,096×4,096 elements. The “A” input matrix can be 128×128 elements, and the “B” input matrix to the computation unit 100 can be 1024×128 elements.
[0056] The input matrix B is also pre-processed before its elements are received by the horizontal checksum circuit 115. The input matrix B is the inverse (B r ), the positions of the elements in each row are swapped. For example, the first element becomes the last element, the second element becomes the second-to-last element, etc. r The horizontal checksum circuit 115 also shifts by one time step for each row. As described herein with reference to FIG. r The horizontal checksum circuit 115 generates a checksum for each row of the inverse matrix B r After processing the received elements of , the received elements are passed to a systolic array 110 to perform matrix multiplication.
[0057] In some examples, elements of matrix A or elements of matrix B are loaded into the systolic array 110, e.g., one element loaded into each processing element of the systolic array 110, and then elements of the other matrix are streamed into the systolic array 110. For example, input matrix A may be a weight matrix corresponding to the weights of a layer of a neural network, while input matrix B may include input values for multiplying the weights of the neural network layer. The elements of matrix A may be loaded into the systolic array 110, e.g., by latches or weight loads, and elements of matrix B may be streamed into the systolic array 110 for processing. The selection of streaming or loading elements of the input matrices may depend on the dimensions of the input matrices, e.g., a wider or taller matrix may be streamed into the systolic array, while a matrix that fits into the systolic array 110 may be loaded. The output matrix C 140C ("matrix C") may be streamed out of the systolic array 110, or in some examples, may be pushed out entirely from the systolic array 110 only after the multiplications are completed.
[0058] After a number of time steps have passed, e.g., based on the configuration and number of processing elements of the systolic array to perform the matrix multiplication, the output checksum circuit 125 begins to receive the elements of matrix C. Because the input matrices A and B are misaligned, the elements of matrix C are also misaligned. However, in some examples, the systolic array 110 can be configured as an output fixed systolic array, where the elements of output matrix C remain in the systolic array 110 until the multiplication is complete, and are only then popped out via the output checksum circuit 125. As will be described in more detail with reference to FIG. 5, the output checksum circuit 125 generates an output matrix checksum and compares those values to the checksum from the output matrix C. If the generated checksum does not match the checksum of the output matrix C, the output checksum circuit 125 can raise a flag or other indication that an error has been detected.
[0059] Output matrix C has two parts, a data sub-matrix and a checksum part. The checksum part of output matrix C is indicated by the surrounding elements of output matrix C, shown by the dashed outline of the checksum row 202 of output matrix C and the dotted outline of the column checksum 204. The data sub-matrix 140D of the checksum matrix is indicated by the cross-satch elements. The data sub-matrix 140D corresponds to the product of input matrices A and B without the addition of a checksum. The overall output matrix C (elements 140C) including the checksum row 202, column checksum 204 and data sub-matrix 140D corresponds to the product of matrices A and B with the respective checksum rows or columns generated by the horizontal and vertical checksum circuits 115, 120.
[0060] Note that the checksums of the rows and columns 202, 204 share a common element, indicated by the element in the upper right corner of the output matrix C. The checksums of the rows and columns 202, 204 of the output matrix C may be discarded after the output checksum circuit 125 performs the checksum generation and comparison.
[0061] Overall, the introduction of horizontal, vertical, and output checksum circuits incurs low or no additional cost in terms of increasing the time steps, i.e., the additional time steps required to pass inputs to systolic array 110 in the case of horizontal / vertical checksum circuits, or to pass output elements from systolic array 110 to their respective destinations from output checksum circuit 125. The additional logic required also scales linearly for, for example, two-dimensional systolic arrays, where only the length of the array dimension increases, and not the number of processing elements in the array.
[0062] 3 is a block diagram of vertical checksum circuit 120 and checksum processing element 130. Checksum processing element 130 includes checksum processing elements 130A, B, C, and D, which may be configured similarly to the processing elements of systolic array 110. In some examples, the error checking functionality provided by vertical checksum circuit 120 may be disabled, for example, based on user input. In some cases, checksum processing element 130 may be used as part of performing matrix multiplication on systolic array 110 even when error checking is disabled.
[0063] The vertical checksum circuit 120 includes several adder circuits 305A, B, and C ("adder circuits 305"); as well as registers 310A, B, C, and D ("registers 310"). The registers 310 may be implemented according to any of a variety of techniques for storing data in a circuit, for example as a respective group of flip / flop circuits. The adder circuits 305 may be implemented according to any of a variety of techniques for constructing a circuit configured to receive two inputs and generate a resulting sum of the two inputs. The number of adder circuits 305 varies depending on the length of the row inputs of the input matrix. The examples provided herein focus on a 4×4 input matrix, and thus the vertical checksum circuit 120 includes three adder circuits and four registers.
[0064] The vertical checksum circuit 120 receives an input matrix A and performs one time step (A T ) The first row of entries is the input matrix A T (shown in FIG. 3 as input matrix 140AT). The first row of input matrix 140AT corresponds to the first column of input matrix A. Input matrix 140AT includes elements a1, a2, a3, a4, a5, a6, a7, a8, a9, a10, a11, a12, a13, a14, a15, and a16.
[0065] In general, vertical checksum circuit 120 accumulates partial sums of the checksum in registers 310A-C and stores the calculated checksum in register 310D before pushing the checksum out to checksum processing element 130. Because the inputs are staggered, in a first time step, vertical checksum circuit 120 receives input a1. Input a1 is stored in register 310A and is also passed to systolic array 110 for processing, i.e., matrix multiplication by the matrix multiplication algorithm on the systolic array.
[0066] In the second time step, inputs a1 and a2 are received by adder circuit 305A. The partial sum of a1+a2 is stored in register 310B. Input a2 is also passed to systolic array 110. Although not shown in FIG. 3, also in the second time step, the next input above a1, a5, is stored in register 310A and also passed to systolic array 110.
[0067] In a third time step, adder circuit 305B adds the partial sum stored in register 310B with input a3. The updated partial sum produced by adder circuit 305B is stored in register 310C. Input a3 is also passed to systolic array 110. Also, the next input in the same column as a1 (input a9) is passed to register 310A and the previous contents of register 310A are added to the next input in the same column as a2, a6, and stored in register 310B.
[0068] In the fourth time step, adder circuit 305C adds the partial sums stored in register 310C with input a4. Because adder circuit 305C is the last adder, it adds the first column of A (i.e., A T The checksum is stored in register 310D. Input a4 is also passed to systolic array 110, and the successive entries of input matrix A (a7) are passed to systolic array 110 and summed to update the respective partial sums across registers 310A-C.
[0069] At the fifth time step, the checksum stored in register 310D is pushed down to the last available processing element of checksum processing element 130. For example, the checksum generated for the first column of input matrix A (inputs a1-a4) is pushed down to processing element 130D. Also at the fifth time step, input a8 is pushed out and the checksum for the second column of input matrix A can be calculated and stored in register 310D. At a subsequent time step, the checksum for the second column (inputs a5-a8) is pushed down to processing element 130C, then the third checksum for the third column (inputs a9-a12) of processing element 130B, and the checksum for the fourth column (inputs a13-a16) are pushed down to checksum processing element 130A.
[0070] After the checksums for input matrix A are generated, the systolic array 110 can use the checksums stored in checksum processing element 130 to generate an output matrix C having corresponding rows and columns of the checksums, for example, by multiplying each checksum with a corresponding row of input matrix B.
[0071] In the previous example, the transpose of matrix A T Although the vertical checksum circuit 120 has been described as generating a column checksum of the input matrix A by summing the elements of each row of A, it will be appreciated that the vertical checksum circuit 120 may be configured to perform any linear operation to generate the checksum, such as adding and multiplying a constant value or coefficient, or multiplying the checksum row by a predetermined vector.
[0072] 4 is a block diagram of horizontal checksum circuit 115. Horizontal checksum circuit 115 includes a number of summing circuits 405A-D ("summing circuits 405"), registers 410A-D ("registers 410A-D"), and multiplexers 415A-D ("multiplexers 415"). Summing circuits 405 and registers 410 may be implemented as described herein with reference to summing circuits 305 and registers 310 shown in FIG. 3. Multiplexers 415 may be two-to-one multiplexers configured to pass either the value of the corresponding register or the corresponding input value to systolic array 110. The multiplexers may be replaced with any variant of decision-making circuitry, such as a switch controlled by one or more control signals to disable and enable certain portions of horizontal checksum circuit 115, or one or more circuits configured to select between multiple inputs and gate unselected inputs.
[0073] The horizontal checksum circuit 115 receives the input matrix B as shown in FIG. 3 and converts it into an inverse and staggered version (B r ) For ease of explanation, the inputs are displayed in ascending order. r The first column of inputs received by (i.e., the last column of input matrix B) includes elements b1, b2, b3, b4, b5, b6, b7, b8, b9, b10, b11, b12, b13, b14, b15, and b16.
[0074] In general, the horizontal checksum circuit 115 accumulates partial sums of the checksums in registers 410A-D and determines when to push the contents of registers 410A-D to the systolic array 110 as a completed checksum based on a control signal. For example, a control signal, e.g., an is_checksum flag, can be propagated from multiplexer 415A to multiplexer 415D based on the size of an inner dimension of input matrix A. As described herein, input matrix B can be much longer (e.g., n×4) than input matrix A (e.g., 4×m). Thus, a checksum is generated over a segment of input matrix B and periodically sent to the systolic array 110. The control signal allows the horizontal checksum circuit 115 to determine when the data passed to the systolic array 110 is the normal input from matrix B or the generated checksum. Based on the control signal, the horizontal checksum circuit 115 can pass the checksum and reset the accumulation of partial sums for a new set of inputs from matrix B. The control signal may be provided automatically by a sequencer communicatively coupled to the computation unit 100 and / or may be provided programmatically by software including instructions executed by the computation unit 100 to perform matrix multiplication on the input matrix A and the input matrix B.
[0075] For example, multiplexer 415A is configured to pass inputs from input matrix B to systolic array 110 until a control signal is received. Upon receiving the control signal, multiplexer 415A instead passes the contents of register 410A to systolic array 110 for processing. The timing of the control signal to multiplexer 415A should coincide with the completion of the corresponding row checksum of input matrix A. For example, the control signal may be passed to multiplexer 415A after an appropriate number of time steps have elapsed to receive and accumulate each input of a row of input matrix B. An example breakdown of the operations performed per time step for horizontal checksum circuit 115 follows.
[0076] Because the inputs are staggered, in the first time step, horizontal checksum circuit 115 receives input b13. Register 410A is initially empty, so input b4 is added to the contents of register 410A, which is zero. Multiplexer 415A checks the control signal and, if there is no control signal, passes input b4 to systolic array 110.
[0077] In the second time step, the horizontal checksum circuit 115 receives input b8. Since register 410B is initially empty, input b8 is added to the contents of register 410B, which is zero. Multiplexer 415B checks the control signal and, if the control signal is not present, passes input b8 to the systolic array 110. At the same time, the horizontal checksum circuit 115 also receives the next input b3, which follows b4.
[0078] In a third time step, horizontal checksum circuit 115 receives input b12. Input b12 is added to the contents of register 410C. Horizontal checksum circuit 115 also receives successive inputs from the top row of input matrix B and adds each received input to a respective partial sum stored in registers 410A-C. Horizontal checksum circuit 115 also receives inputs b7 and b2.
[0079] In the fourth time step, horizontal checksum circuit 115 receives input b16, which is added to the contents of register 410D and passed to the systolic array by multiplexer 415D. Horizontal checksum circuit 115 also receives inputs b11, b6, and b1.
[0080] Also during the fourth time step, multiplexer 415A receives a control signal indicating that the contents of register 410A are the checksum of the first row of matrix B. At this time step, register 410A contains the sum b4+b3+b2+b1. Multiplexer 415A pushes the checksum out of register 410A and into the systolic array 110 for inclusion in the matrix multiplication operation. Register 410A is then reset to zero to begin calculating the checksum for the next set of inputs. The control signal is pushed down to the next multiplexer, multiplexer 415B.
[0081] In the fifth time step, multiplexer 415B pushes the contents of register 410B out to systolic array 110. By the fifth time step, register 410B may contain the sum b8+b7+b6+b5. Register 410B is cleared and the control signal is pushed down to the next multiplexer, multiplexer 415C.
[0082] In the sixth time step, multiplexer 415C pushes the contents of register 410C out to systolic array 110. By the sixth time step, register 410C may contain the sum b12+b11+b10+b9. Register 410C is cleared and the control signal is pushed down to the next multiplexer, multiplexer 415D.
[0083] At the seventh time step, multiplexer 415D pushes the contents of register 410D out to systolic array 110. By the seventh time step, register 410D may contain the sum b16+b15+b14+b13. Register 410D is cleared. At a subsequent time step, horizontal control circuit 115 may generate new respective checksums for the new group of inputs from matrix B in each of its rows.
[0084] In some examples, the vertical checksum circuit 120 is configured with multiplexers that manage whether the input from matrix A or the generated checksum is passed to the systolic array 110. For example, the vertical checksum circuit 120 may receive an input matrix that is much longer than what can fit into the systolic array 110, and its contents are instead streamed to the systolic array 110. In these examples, the checksum processing element 130 may not be used, or may be used instead by the horizontal checksum circuit 115. For example, the checksum processing element 130 may be used by the horizontal checksum circuit 115 when it receives an input matrix that is not streamed to the systolic array 110, but is instead loaded directly into the processing elements 110A-P. In these examples, the checksum processing element 130 may be arranged as an additional row to the systolic array 110. In other examples, both the horizontal and vertical checksum circuits 115, 120 may be configured with multiplexers as described herein to handle streaming inputs. In some implementations, the particular orientation and location of the checksum circuits 115, 120 and the checksum processing element 130 may vary based on the direction in which the systolic array 110 is configured to receive inputs.
[0085] While the previous example described generating checksums for rows of input matrix B by summing the elements of each row, it is understood that vertical checksum circuit 120 can be configured to perform any linear operation to generate the checksum, such as adding and multiplying a constant value or coefficient, or multiplying the checksum row by a predetermined vector. An additional consideration for horizontal checksum circuit 115 is generating a checksum that is a valid input for systolic array 110. While some data types, such as floating point, may have sufficient precision to represent very large sums of individual elements of the input matrix, for example, systolic array 110 may be configured for low precision integer values for matrix multiplication, for example matrices with 8-bit quantized values. The checksum calculated by horizontal checksum circuit 115 may require more precision than the input to systolic array 110.
[0086] To address this potential problem without modifying the hardware or configuration of the systolic array, in some examples, the horizontal checksum circuit 115 may be configured to generate the checksum modulo some value, for example the maximum value supported by the systolic array 110. In other examples, Galois field arithmetic may be applied to make the precision required to represent the checksum the same as the precision required to represent the elements of the matrix being multiplied by the systolic array 110. In these examples, the horizontal checksum circuit 115 may be configured with additional circuitry configured to perform Galois field arithmetic.
[0087] Although the example operations by the horizontal and vertical checksum circuits are described as being performed over successive time steps, it will be understood that several delays or idle time steps may be inserted, as necessary, to synchronize the operation of the horizontal and vertical checksum circuits with the operations performed by the processing elements of the systolic array 110.
[0088] 5 is a block diagram of output checksum circuit 140C. Output checksum circuit 140C is configured to generate a checksum from the rows and columns of data sub-matrix 140D of output matrix C. As described with reference to FIG. 2, output matrix C includes checksum row 202, checksum column 204, and data sub-matrix 140D. The latter corresponds to the product of matrices A and B, excluding the checksum row and checksum column generated by the horizontal and vertical checksum circuits. Data sub-matrix 140D has output elements d1, d2, d3, d4, d5, d6, d7, d8, d9, d10, d11, d12, d13, d14, d15, and d16. Checksum row 202 has checksums c1, c2, c3, c4, and c5. Checksum column 204 has checksums c5, c6, c7, c8, and c9.
[0089] Output checksum circuit 140C in some examples may be a combination of horizontal and vertical checksum circuits 115, 120. Registers 500A-D and summing circuits 502A-C may be configured similarly to the summing circuits and registers of horizontal checksum circuit 115.
[0090] Because input matrices A and B are shifted before their elements are pushed out to systolic array 110, the input elements of output matrix C are also shifted as they are pushed out of systolic array 110. Output checksum circuit 140C generates a checksum for each row of output matrix C and compares it to the corresponding checksum in checksum column 204. For example, outputs d13, d14, d15, and d16 pass through registers 500A-D and summing circuits 502A-D as described herein with reference to FIG. 3 and the row inputs of input matrix 140AT. When 500D receives the inputs, i.e., the checksums generated from outputs d13, d14, d15, and d16, comparison circuit 505 compares the checksum in register 500D to the first checksum pushed out of checksum column 204, i.e., checksum c9.
[0091] If the checksum c9 matches the contents of the register 500D, the error check is successful. If the checksum c9 does not match the contents of the register 500D, the output checksum circuit 125 generates an error. In response to the generation of an error, the computation unit 100 may perform one or more of a variety of actions, as described herein with reference to FIG. 2. The output checksum circuit 125 may continue to calculate each row checksum of the data sub-matrix 140D using the checksum corresponding to the checksum column 204. For example, the comparison circuit 505 compares the sum d9+d10+d11+d12 with the checksum c8, the sum d5+d6+d7+d8 with the checksum c7, the sum d1+d2+d3+d4 with the checksum c6, and the sum c1+c2+c3+c4 with the checksum c5.
[0092] A comparison circuit implemented as part of the output checksum circuit 125 can either directly compare the checksums or calculate the absolute value of the difference between the compared checksums. The comparison circuit can then determine whether the absolute value of the difference between the checksums is within a predetermined threshold. In some examples, the threshold can be programmatically defined, for example based on a use case that has a higher or lower tolerance for error. In other examples, the threshold can be specified at design time of the computation unit 100. In some applications, for example, the execution of some machine learning models, absolute precision is not required and some degree of error in the computed floating-point values can be tolerated, for example within a predetermined threshold.
[0093] Although output checksum circuit 125 generates a checksum for each row of data sub-matrix 140D, output checksum circuit 125 may also generate a checksum for each column of output matrix C and compare each generated checksum to a respective checksum in checksum row 202. Output checksum circuit 125 may include summing circuits 504A-E, demultiplexers 503A-E, registers 507A-E, registers 510A-E, and comparison circuits 515A-E. Although output checksum circuit 125 is shown as including a demultiplexer, it will be understood that demultiplexers 503A-E may be replaced with any of a variety of different decision making circuits, such as, for example, a control signal for disabling or enabling certain portions of output checksum circuit 125, or a switch configured to direct an input to one of a number of outputs based, for example, on a control signal.
[0094] For example, in a first time step, output d13 is passed to register 500A and demultiplexer 503A (to compute a first row checksum of output matrix C, as described herein). Demultiplexers 503A-E are configured to transmit the checksums of checksum row 202 to respective registers 510A-E, while transmitting the remaining portions of the columns of output matrix C to respective summation circuits 504A-E and registers 507A-E.
[0095] In the second time step, d13 is pushed to summing circuit 504A, which adds output d13 to the contents of register 507A, which is initially zero. Also during the second time step, output d14 is pushed to summing circuit 502A and demultiplexer 503B. Output checksum circuit 125 can continue to receive and process the output elements of output matrix C until it reaches the checksum of checksum row 202. At the same time, output checksum circuit 125 can receive a control signal, for example a control signal used for horizontal checksum circuit 115, and based on the presence of the control signal, pushes the checksum value of checksum row 202 to the corresponding register 510A-E.
[0096] When the checksums corresponding to both contents of register 510A have been loaded, e.g., checksum c1 of register 510A, checksum c2 of register 510B, checksum c3 of register 510C, checksum c4 of register 510D, and checksum c5 of register 510E, then each comparison circuit compares the stored checksum with the checksum generated and stored in register 507A-E. For example, comparison circuit 515A compares the sum d1+d5+d9+d13 with checksum c1. Comparison circuit 515B compares the sum d2+d6+d10+d14 with checksum c2. Comparison circuit 515C compares the sum d3+d7+d11+d15 with checksum c3. Comparison circuit 515D compares the sum d4+d8+d12+d16 with checksum c4. Comparison circuit 515E compares the sum c6+c7+c8+c9 with the checksum c5.
[0097] In some examples, the output checksum circuit 125 can be configured to generate one or more measurements or indicators corresponding to the number and severity of different errors detected. For example, the output checksum circuit 125 can calculate the absolute value of the difference between the compared checksums, which can be transmitted to a coupled external processor in addition to the flag raised. The absolute value of the difference between the compared checksums can be a measure of the severity of the errors detected by the output checksum circuit 125. For example, a larger value can indicate a more severe error than a smaller value.
[0098] Similar to the comparisons performed by comparison circuit 505, if any of comparison circuits 515A-E identify any mismatch (or a mismatch above a predetermined threshold), then output checksum circuit 125 may flag an error that occurred during the matrix multiplication of input matrices A and B.
[0099] Figure 6 is an exemplary computational unit 600 having a vertical checksum circuit 620 and an output checksum circuit 625, but no horizontal checksum circuit. Also shown in Figure 6 are input matrices 640A, 640B, an output matrix 640C with checksum column 602, and a checksum processing element 630. The output checksum circuit 625 can be configured to generate a checksum for each row of the output matrix C and compare it to the checksums generated by the vertical checksum circuit.
[0100] Omitting the horizontal checksum circuit still allows for error detection, further reducing the complexity and inputs to the computation unit 600, at least since there is no need to manage the timing of control signals, e.g., through a program executable by the computation unit 600, as described herein with reference to the horizontal checksum circuit 115. In some cases, detecting an error without specifically indicating where the error occurred, which may be obtained by comparing both the checksum row and the checksum column, may still be of value if the corrective action taken does not require the exact location of the error. For example, the exact location of the error is not required if the default corrective action is to repeat the previous matrix multiplication or to cold restart the computation unit.
[0101] Furthermore, by omitting the horizontal checksum circuitry, the logic for generating the checksums can be completely separated from the systolic array. Referring back to FIG. 4, the checksums generated by the horizontal checksum circuitry 115 have been pushed out to the systolic array 110. Thus, the input specifications of the systolic array 110 must be adhered to if the systolic array 110 is configured to receive inputs only up to a certain level of precision. Because the computation unit 600 omits the horizontal checksum circuitry as in FIG. 4, the remaining checksums are stored in the checksum processing element 630, which can be separated from the systolic array 610. In some examples, the checksums generated by the vertical checksum circuitry 620 can have a different measurement of precision and / or be of an entirely different type than the elements of the input matrix.
[0102] The omission of horizontal checksum circuitry may allow for more flexibility between different types of linear operations for checksum generation that may be implemented. For example, a floating-point checksum may be used for 8-bit integer elements. A 16-bit integer may be used for 8-bit integer elements. Galois field arithmetic or modulo addition or multiplication operations may generally be used for the checksum calculation.
[0103] In some implementations, the vertical checksum circuit 620 can be extended to generate multiple checksum columns based on different linear operations for generating checksums from the input elements of the rows of the input matrix A. For all linear codes defined over a finite field, if there is a corresponding linear real code with a similar error detection and correction function, when generating checksums using Galois field operations, additional error detection processes such as N check symbol Reed-Solomon codes can be applied. Otherwise, the corresponding real code can be used.
[0104] In some examples, when the input matrix A is streamed and the input matrix B is latched or preloaded into the systolic array, the computing units of the systolic array can instead implement a properly configured horizontal checksum circuit, i.e., a checksum processing element, configured to generate checksum columns as described herein with reference to FIG. 3, i.e., calculate column checksums and compare them with the checksums generated by the horizontal checksum circuit. In other words, different implementations of the computing units can have different directions and positions of the checksum circuits compared to how the systolic array receives inputs, without affecting the overall error detection function described herein.
[0105] The computational units described herein can be configured to operate using none, one, or both of the horizontal and vertical checksum circuits. The computational units can use neither the horizontal nor the vertical checksum circuits, for example, because error detection has been programmatically disabled. The same computational units can also be configured to operate using only one of the horizontal and vertical checksum circuits to perform error detection, for example, as described herein with reference to FIG. 6. The same computational units can also be configured to operate using both the horizontal and vertical checksum circuits to perform error detection, for example, as described herein with reference to FIG. 1-5.
[0106] 7 is an exemplary computational unit 700 with an output fixed systolic array 710. In the output fixed systolic array, each processing element holds a respective output element of output matrix C 740C, and input matrices A, B, 740A, B can be streamed into the systolic array 710. Once all input elements of matrices A and B have been processed, the output matrix C is pushed out through an output checksum circuit 725. The computational unit 700 also includes a vertical checksum circuit 720 and a horizontal checksum circuit 722. circuit 715. Both checksum circuits 715, 720 are coupled to a checksum processing element 730 which includes a row 730A of checksum processing elements and a column 730B of checksum processing elements.
[0107] The vertical checksum circuit 720 is configured similarly to the vertical checksum circuit 120, and includes an input matrix A 740A (input matrix A and matrix A T Column 730B of checksum processing elements may correspond to checksum processing elements 130, which are configured to receive and store the generated checksums for later comparison by output checksum circuit 725.
[0108] The horizontal checksum circuit 715 may also be implemented similarly to the vertical checksum circuit 120, but with a corresponding row 730A of checksum processing elements, both of which are rotated 90 degrees relative to the corresponding circuit 720 and column 730B of processing elements. However, the horizontal checksum circuit 715 may also be configured to generate a checksum for a corresponding column of input matrix B 740B (or a corresponding row, if input matrix B 740B is transposed).
[0109] The output checksum circuit 725 may be configured to generate and compare checksums as described above with reference to the output checksum circuit 125 and Figure 5. In this exemplary arrangement of horizontal and vertical checksum circuits, the checksum row 702 is pushed out first, before the values of the data sub-matrix 740D of the output matrix 740C. The output checksum circuit 725 may be configured to push out each first output element of each column of the output matrix 740C to a corresponding register of a comparison circuit, followed by accumulating the remaining column output elements in respective registers coupled to respective summation circuits.
[0110] In some examples, instead of adding separate partial sums for comparison to the received checksum of output matrix C740C, output checksum circuit 725 can be configured to receive a checksum from a checksum row and subtract from that checksum each output element of output matrix C740C that exceeds the checksum. If after subtracting all output elements, the result is not zero or close to zero within a predetermined threshold, output checksum 725 can transmit an indication that an error has been detected. Depending on the outcome of the error detection by comparison of the checksums of the row 704 of checksum processing elements (e.g., by subtracting from the output elements of the corresponding row, or by comparing each checksum to the corresponding cumulative checksum as described herein), output checksum circuit 740 can also identify where the error occurred, e.g., the intersection of the column and row with the erroneous checksum.
[0111] 7, the computation unit 700 fully isolates the data path of the systolic array, e.g., with error checking logic implemented by checksum circuits 715, 720, and 725, in receiving inputs and performing matrix multiplication. This allows, e.g., augmenting the variety of different linear operations of ABFT for matrix multiplication without worrying about matching data types and precisions (e.g., as described with reference to omitting the horizontal checksum circuitry of computation unit 600 of FIG. 6).
[0112] This is the reverse of the operation of the output checksum circuit 125 described with reference to Figure 5, where a control signal was used to indicate when the output checksum circuit 125 is receiving the checksum for the output matrix 140C. The output checksum circuit 725 in some examples may also use a control signal, for example, when a large input matrix 740A,B is being streamed to the systolic array 710. In these examples, a control signal may be sent to the output checksum circuit 125 at the beginning of each group of output elements for checksum generation and comparison, rather than at the end.
[0113] The systolic array can operate according to different voltage levels. The critical supply voltage is a value indicating a voltage sufficient for the correct operation of the systolic array, for example in the face of various environmental or process-related variations that may affect the performance of the circuit. A voltage supplied to the systolic array at a level lower than the critical supply voltage may be more energy efficient, for example because less energy is required to operate the systolic array, but there is a risk of errors, for example timing errors, in the face of the aforementioned variations. Timing errors can quickly cascade into more serious errors if not corrected or addressed, especially in computational units using systolic arrays where proper timing of inputs and outputs between processing elements is important. A data processing system implementing a computational unit as described herein can supply a reduced voltage to the computational unit and increase the voltage as necessary in response to receiving an error detection flag from the computational unit. Since errors are relatively rare compared to normal operation of a computational unit using a systolic array, especially when multiple computational units are executed in parallel, running the systolic array at a supply voltage below the critical supply voltage level can usually outperform occasional errors and corrective measures.
[0114] The computational unit can detect errors associated with low voltages according to the same mechanisms described herein with reference to Figures 1-6. While other approaches require additional logic in multiple different latch, shadow latch, or flip / flop circuits, the horizontal, vertical, and output checksum circuits are independent of the systolic array and do not need to delay the execution of matrix multiplications or other operations on the array. Furthermore, the predetermined thresholds used by the comparison circuits of the output checksum circuit to compare generated and received checksums can be adjusted to tolerate smaller errors caused by a drop in the voltage supplied to the systolic array.
[0115] The control logic may for example be implemented by a data processing system implementing the computational unit and may also adjust the supplied voltage and / or frequency, for example the clock frequency at which the data processing system operates, depending on the observed error rate in order to further improve the energy efficiency of the computational unit.
[0116] Timing errors may occur not only in the processing of the systolic array, but also in the execution of error detection logic provided by the horizontal checksum circuitry, the vertical checksum circuitry, and / or the output checksum circuitry. Timing errors in these circuits may be addressed with a variety of approaches. In some examples, higher voltages may be provided to one or more of the horizontal checksum circuitry, the vertical checksum circuitry, and the output checksum circuitry. In other examples, specialized transistors or other components may be implemented in the vertical, horizontal, and / or output checksum circuitry to improve the speed at which error checking operations described herein, such as checksum generation and checksum comparison, are performed.
[0117] In yet another example, as described herein with reference to Figures 8A-C, the horizontal checksum circuitry can be modified to allow for a wider margin in timing since different operations are performed for each time step. The increased margin can mitigate the possibility of timing errors occurring in the error detection logic itself, thereby reducing the chance of inaccurate error detection or missing undetected errors by operating the computational units at a lower supply voltage.
[0118] 8A is an exemplary vertical checksum circuit 800A. Vertical checksum circuit 800A is configured to generate checksums from rows of an input matrix of length 8. Vertical checksum circuit 800A can receive inputs a0-a7 and includes registers 805A-H and summing circuits 804A-H. The generated checksum stored in register 805H can be pushed out to checksum processing element 830. Inputs a0-a7 are also pushed out to processing elements 803A-H of the systolic array.
[0119] FIG. 8B is an exemplary vertical checksum circuit 800B having a two-input, two-stage pipelined adder circuit. Circuit 800B is a modified version of circuit 800A, replacing adder circuits 804A-H with two-input, two-stage pipelined adder circuits 820A, B, and C ("adder circuits 820"). Circuit 800B can be configured to delay the generation of the checksum. Adder circuits 820A, B, C have a latency of two time steps, e.g., two clock cycles, with a throughput of one sum or addition operation performed per clock cycle. Each adder circuit 820 can include a two-to-one adder circuit including intermediate adder circuits A and B, and a register between the two intermediate adder circuits for storing intermediate sums.
[0120] Vertical checksum circuit 800B splits the processing of input elements a0-a7 into two pipeline stages 806A and 806B. The odd-numbered input elements (e.g., a1, a3, a5, and a7) are accumulated across summation circuit 820A of stage 806A, while the even-numbered input elements (e.g., a0, a2, a4, and a6) are accumulated across summation circuit 820B of stage 806B. The partial sums produced via stages 806A,B are added together in summation circuit 806C, and the result is passed to checksum processing element 830.
[0121] The additional latency provided by the adder circuit 820 increases the timing margin for the vertical checksum circuit 800B to operate correctly even when the vertical checksum circuit 800B is operating below the critical voltage level, reducing the possibility of miscalculation or miscalculation occurring as a result of timing errors caused by operating at a lower level of supply voltage. Compared to the vertical checksum circuit 800A, the vertical checksum circuit 800B requires three additional clock cycles to generate a checksum, but the additional latency in the vertical checksum circuit 800B does not affect the throughput of input elements a0-a7 to the systolic array connected to the vertical checksum circuit 800B, and does not affect the processing latency. For example, the number of clock cycles for the systolic array to process the input matrix is not increased.
[0122] In other examples of vertical checksum circuit 800B, vertical checksum circuit 800B may include a different number of stages and / or a different number of adder circuits or other processing circuits configured to generate a checksum according to a specified linear operation over a different number of clock cycles.
[0123] 8C is a circuit diagram of another exemplary vertical checksum circuit 800C having a two-cycle non-pipelined adder circuit 850. The vertical checksum circuit 800C includes multiple two-cycle adder circuits 850, as well as flip / flop circuits 860, 865, and 870. The two-cycle adder circuits are configured to receive two inputs and generate a sum of these two inputs over two clock cycles. As indicated by legend 880, the flip / flop circuit 860 is colored and solid in the circuit diagram of the vertical checksum circuit 800C, the flip / flop circuit 865 has a vertical thatch mark, and the flip / flop circuit 870 has a diagonal thatch mark. It is understood that in some examples, other types of circuits configured to store data can be used in place of the flip / flop circuits 860, 865, and 870. The flip / flop circuits 860, 865, and 870 operate at different clock frequencies.
[0124] The flip / flop circuit 860 operates at a clock frequency φ 0 The clock frequency can be set by a timing circuit connected to the computation unit implementing vertical checksum circuit 800C, as described herein with reference to Figures 10 and 11. Flip / flop circuits 865 and 870 operate at different respective clock frequencies φ 1 and φ 2 It can operate at a clock frequency of φ 1 and φ 2 may have different phases, as shown in chart 890. Vertical checksum circuit 800C also provides increased timing margins compared to circuit 800A, mitigating the likelihood of timing errors affecting the error checking functionality of a computation unit implementing vertical checksum circuit 800C, as described herein with reference to vertical checksum circuit 800B.
[0125] In other examples of vertical checksum circuit 800C, vertical checksum circuit 800B may include different flip / flop circuits operating at different frequencies, and / or different numbers of summing circuits or other processing circuits configured to generate checksums according to a specified linear operation over different numbers of clock cycles.
[0126] FIG. 9A is a flowchart of an example process for performing matrix multiplication with error detection on a computation unit according to an aspect of the disclosure.
[0127] According to block 905A, the systolic array of the computational unit receives a first input element from a first input matrix along a first direction of the systolic array. For example, the first input element may be an input element of input matrix A received from a top periphery of the systolic array, for example, as described herein with reference to FIG. 3. It is understood that the systolic array may be configured to receive inputs from any direction with corresponding checksum circuitry implemented consistent with aspects of the present disclosure.
[0128] The systolic array receives second input elements from a second input matrix along a second direction of the systolic array according to block 910A. The second direction may be horizontal, e.g., left to right, while the first direction may be vertical, e.g., top to bottom, relative to a fixed direction of the systolic array. The second input elements may be, for example, input elements from input matrix B, as described herein with reference to FIG. 4.
[0129] According to block 915A, a first checksum circuit of the computation unit generates one or more groups of first checksums from the first input elements while the systolic array receives the first input elements. The first checksum circuit may be a vertical checksum circuit, e.g., 1 The first group of checksums may be generated and later fed to a checksum processing element, e.g., vertical checksum circuit 120 of FIG. 1The first and second checksum circuits may each include a column of checksums stored in the checksum processing element 130. In some examples, the one or more groups of checksums may refer to multiple columns or rows of checksums generated according to different linear operations. In some examples, the computation unit includes one or more checksum processing elements configured to receive the checksums from one or both of the first and second checksum circuits.
[0130] According to block 920A, the second checksum circuit of the computation unit generates one or more groups of second checksums from the second input elements while the systolic array receives the second input elements. The second checksum circuit may be a horizontal checksum circuit, e.g., the horizontal checksum circuit 115 of FIG. 1. The second checksum circuit may generate the checksums, e.g., as described herein with reference to the horizontal checksum circuit 115 and FIG. 4. In some examples, when the input matrix to the second checksum circuit is streamed to the systolic array, the second checksum circuit may push a control signal across the circuit to indicate when the checksum can be pushed out to the systolic array and generate a new checksum. The timing of the control signal to one or both of the first and second checksum circuits may be based on the number of time steps to load the first or second input values across the systolic array. The time step may be one or more clock cycles.
[0131] In some examples, the systolic array is a fixed output systolic array, and the first and second checksum circuits are both connected to a plurality of checksum processing elements, and the checksum processing elements are configured to receive the checksums generated from one or both of the first and second checksum circuits. The plurality of checksum processing elements may be arranged around the periphery of the systolic array, for example, as described herein with reference to FIG.
[0132] The systolic array generates an output matrix from the first input matrix, the second input matrix, the one or more groups of first checksums, and the one or more groups of second checksums, in accordance with block 925 A. The output matrix may be, for example, an output matrix C, as described herein with reference to FIG.
[0133] According to block 930A, an output checksum circuit receives the output matrix. As described herein with reference to FIG. 5, the output checksum circuit can receive the output matrix as it is pushed out of the systolic array. In some examples, the computation unit generates checksums only from the first checksum unit or only from the second checksum unit.
[0134] The output checksum circuit determines whether an error was detected in the calculation of the output matrix, per diamond 935A. If an error was detected ("YES"), the output checksum circuit transmits an indication of the occurrence of an error (or more than an error, if applicable), per block 940A. Otherwise ("NO"), process 900A can continue with new input elements for processing and performing error detection.
[0135] The systolic array may receive a first voltage below the critical voltage until receiving a response to the indication of error detection, which may be predetermined as described herein with reference to FIGS. 8A-C and the preceding discussion. In response to transmitting the indication or identifying one or more errors, the systolic array may begin receiving a second voltage above the critical voltage. The systolic array may receive the second voltage automatically, or the second voltage may be applied to the systolic array by another device. In some examples, the systolic array automatically returns to receiving the first voltage, for example, after a period of time has elapsed without detecting an error.
[0136] One or more of the first, second, and output checksum circuits may continue to receive a voltage that is less than the critical voltage of the systolic array. The first, second, and output checksum circuits may be implemented according to aspects of the present disclosure to increase timing margins for operation of the circuits and mitigate timing errors due to voltage sagging. For example, one or both of the first and second checksum circuits may include a two-input, two-stage pipelined adder circuit as described herein with reference to FIG. 800B. As another example, one or both of the first and second checksum circuits may include one or more two-cycle adder circuits and a plurality of registers, each of which may include one or more flip / flop circuits that operate at different clock frequencies. For example, some registers may operate at a first frequency, some at a second frequency, and other registers at a third frequency. The second and third frequencies may be half the first frequency, out of phase.
[0137] 9B is a flow chart of an example process 900B for detecting the occurrence of an error during processing of a computational unit according to an embodiment of the present disclosure. For example, according to diamond 935A in FIG. 9A, process 900B may be performed as part of determining whether an error is detected.
[0138] The output checksum circuit generates a column checksum from at least one row of the data sub-matrix, according to block 910B. The data sub-matrix may be an output from processing both input matrices without a corresponding checksum, e.g., data sub-matrix 140D, as described with reference to FIG. 2. For example, the output checksum circuit may accumulate output elements in a received row of the output matrix and store the checksum in a register, e.g., register 500D, described herein and shown in FIG. 5. Also, according to block 910B, the output checksum circuit generates a column checksum from at least one column of the data sub-matrix. The column checksums may be generated independently of the row checksums and may be generated in parallel for each column of the output matrix. For example, the output checksum circuit may generate column checksums and store them in respective registers of the circuit, such as registers 510A-E, as described and shown with reference to FIG. 5. In some examples, both column and row checksums are generated by the output checksum circuit. In other examples, only column and row checksums are generated.
[0139] The output checksum circuit compares the row checksum to a checksum of an output checksum column of the output matrix, according to block 920B. The output checksum column may correspond to checksum column 204, for example. The comparison by the output checksum circuit may be performed by one or more comparison circuits, for example comparison circuit 505, as described and shown with reference to FIG. 5. Also according to block 920B, the output checksum circuit compares the column checksum to a checksum of an output checksum row of the output matrix. The output checksum row may be, for example, checksum row 202 of output matrix C, as shown in FIG. 2. The one or more output checksum rows may be checked in parallel by the output checksum circuit, for example, using comparison circuits 515A-E, as described and shown in FIG. 5.
[0140] Comparisons between row / column checksums and output checksum column / row checksums can be done in parallel or sequentially, for example. In some examples, only comparisons using row checksums are performed with the checksums of output checksum columns, and vice versa for column checksums and output checksum rows.
[0141] The output checksum circuit determines the occurrence of an error in the generation of the output matrix from a comparison of the row checksums to the checksums of the output checksum columns, in accordance with block 930B. The output checksum circuit may, for example, determine whether the compared checksums match, or whether they match within a predetermined threshold. If not, the output checksum circuit may transmit an indication that an error has occurred, for example, as described herein with reference to block 940A and FIG. 9A. Also in accordance with block 930B, the output checksum circuit determines the occurrence of an error in the generation of the output matrix from a comparison of the column checksums to the checksums of the output checksum rows.
[0142] The determination from the comparison of the row / column checksums with the corresponding checksums in the output checksum column / row can occur simultaneously, in parallel, or sequentially, as examples. In some examples, the output checksum circuit can send an indication after one error is detected, or after a threshold number of errors are detected. In some examples, the indication can include a measure of the severity of the error, such as the absolute difference between the compared checksums. In other examples, the indication can include information regarding the source of the error, such as the intersection of errors detected for the checksum row checksum and the checksum column checksum.
[0143] 10A is a block diagram of a data processing system 1001 implementing an exemplary computing unit 1000. Computing unit 1000 may be any of a variety of different computing units, such as computing units 100 described herein with reference to FIGS. 1-5. Computing unit 1000 may implement any of a variety of combinations of horizontal, vertical, and output checksum circuits as described throughout this specification.
[0144] The data processing system may include a host interface 1005, a sequencer circuit 1010, one or more processors 1015, a memory 1020, and a timing circuit 1025. The data processing system 1001 may be implemented in one or more devices across one or more physical locations, as described herein with reference to FIG. 11. In some examples, the components of the data processing system 1001 described may be implemented on one or more chips that may interface with a host device according to any of a variety of data buses or other physical interconnect interfaces. In some examples, the data processing system 1001 may be implemented on one or more devices on a network, for example, one or more servers of a cloud platform.
[0145] The processor 1015 and memory 1020 may be any of a variety of different types of processors and memories, as described herein with reference to Figure 11. In some examples, the processor 1015 receives instructions executable by the computing unit 1000 to process data. For example, the instructions may be part of a computer program written to perform operations using the computing unit 1000.
[0146] The sequencer circuit 1010 can convert received instructions into one or more signals understood by the computation unit 1000, which causes the computation unit 1000 to perform any of a variety of preconfigured operations. These operations can include, for example, loading data from memory 1020 into a systolic array (not shown) of the computation unit 1000, moving the data to one or more processing elements of the systolic array, processing the data by one or more processing elements, and pushing the data out of the systolic array. The sequencer circuit 1010 can also be configured to generate one or more control signals to control when checksums are pushed out to the computation unit 1000, for example, as described herein with reference to the horizontal checksum circuit 115 and FIG. 4.
[0147] The host interface 1005 may be configured to receive data from outside the data processing system, e.g., from a processor or another device, and to transmit data generated by the computation unit 1000, e.g., products of matrix multiplication, to one or more devices or processors.
[0148] The timing circuit 1025 can be configured to control the timing of the computation unit, e.g., its clock frequency or clock rate. The time steps described herein with respect to the operation of the checksum circuit can be measured in terms of clock cycles managed by the timing circuit 1025.
[0149] The data processing system 1001 may also be connected to a power source 1030. The power source 1030 may be a battery or other form of power available on the host device implementing the data processing system, or may be a source external to the host device and connected to the host device and data processing system 1001 via some wireless or physical connection, e.g., a wire. The power source 1030 may provide a voltage to the computational unit 1000 that may be managed by the processor 1015, e.g., by regulation higher or lower.
[0150] FIG. 10B is a flowchart of an example process for adjusting a supply voltage to a computing unit according to an aspect of the disclosure.
[0151] According to block 1050, a data processing system implementing the compute unit may apply a voltage to a systolic array of the compute unit. The voltage provided may be less than a critical voltage of the systolic array, which may be predetermined based on use case operation of the compute unit, environmental factors, and / or architectural features of the systolic array, etc.
[0152] According to block 1060, the data processing system receives an indication of one or more errors from the computing unit and, in response, increases the applied voltage above a critical voltage of the systolic array. The computing unit may, for example, send an indication of the detected error as part of executing processes 900A-B as described in Figures 9A-B. The data processing system may increase the voltage supplied to the systolic array to reduce the risk of further errors occurring as a result of a timing violation.
[0153] In response to receiving the indication, the data processing system may continue to apply a voltage below the critical voltage of the systolic array to one or more of the first checksum circuit, the second checksum circuit, and the output checksum circuit of the computational unit, per block 1070. As previously discussed, the horizontal, vertical, and output checksum circuits may be configured according to one or more of a variety of approaches to increase timing margins for operations performed by the circuits (e.g., generating checksums, storing checksums, comparing checksums) to reduce the risk of timing errors.
[0154] In some examples, the data processing system may reduce the voltage to the systolic array after a period or condition is met, such as after a period of time has passed without receiving additional indications of an error from the computational units. In some examples, the data processing system may take various actions in response to receiving an indication of an error, such as rolling back operations performed by the computational units to a previous checkpoint or causing the computational units to re-perform one or more operations.
[0155] FIG. 11 is a block diagram of an exemplary environment 1100 for implementing a data processing system 1001 including a computing unit 1000. The system 1001 can be implemented in one or more devices having one or more processors at one or more locations, such as a server computing device 1105. The user computing device 1112 and the server computing device 1105 can be communicatively coupled to one or more storage devices 1130 via a network 1160. The storage device 1130 can be a combination of volatile and non-volatile memory and can be in the same or different physical location as the computing devices 1112, 1105. For example, the storage device 1130 can include any type of non-transitory computer-readable medium that can store information, such as hard drives, solid-state drives, tape drives, optical storage, memory cards, ROM, RAM, DVDs, CD-ROMs, writable, and read-only memory.
[0156] The server computing device 1105 may include one or more processors 1113 and memory 1114. The memory 1114 may store information accessible by the processor 1113, including instructions 1121 executable by the processor 1113. The memory 1114 may also include data 1123 that may be retrieved, manipulated, or stored by the processor 1113. The memory 1114 may be a type of non-transitory computer-readable medium that may store information accessible by the processor 1113, such as volatile and non-volatile memory. The processor 1113 may include one or more central processing units (CPUs), graphics processing units (GPUs), field programmable gate arrays (FPGAs), and / or application specific integrated circuits (ASICs), such as tensor processing units (TPUs).
[0157] The instructions 1121 may include one or more instructions that, when executed by the processor 1113, cause the one or more processors to perform actions defined by the instructions. The instructions 1121 may be stored in object code format for direct processing by the processor 1113, or in other formats including an interpretable script or a collection of independent source code modules that are interpreted on demand or pre-compiled. The instructions 1121 may include instructions for implementing a system 400 consistent with aspects of the present disclosure. The system 400 may be executed using the processor 1113 and / or using other processors located remotely from the server computing device 1105.
[0158] The data 1123 may be retrieved, stored, or modified by the processor 1113 according to the instructions 1121. The data 1123 may be stored in computer registers, in a relational or non-relational database, as a table with multiple different fields and records, or as a JSON, YAML, Proto, or XML document. The data 1123 may also be formatted in a computer readable format, such as, but not limited to, binary values, ASCII or Unicode. Additionally, the data 1123 may include sufficient information to identify relevant information, such as numerical values, descriptive text, unique codes, pointers, references to data stored in other memory, including other network locations, or information used by a function to calculate the relevant data.
[0159] The user computing device 1112 may also be configured similarly to the server computing device 1105, with one or more processors 1116, memory 1117, instructions 1118, and data 1119. The user computing device 1112 may also include a user output 1126, and a user input 1124. The user input 1124 may include any suitable mechanism or technology for receiving input from a user, such as a keyboard, a mouse, a mechanical actuator, a soft actuator, a touch screen, a microphone, and a sensor.
[0160] The server computing device 1105 may be configured to transmit data to the user computing device 1112, which may be configured to display at least a portion of the received data on a display implemented as part of the user output 1126. The user output 1126 may also be used to display an interface between the user computing device 1112 and the server computing device 1105. The user output 1126 may alternatively or additionally include one or more speakers, transducers or other audio output, a haptic interface, or other haptic feedback that provides non-visual and non-auditory information to a platform user of the user computing device 1112.
[0161] 11 illustrates the processors 1113, 1116 and memories 1114, 1117 as being within the computing devices 1105, 1112, but the components described herein, including the processors 1113, 1116 and memories 1114, 1117, may include multiple processors and memories that may operate in different physical locations rather than within the same computing device. For example, some of the instructions 1121, 1118 and data 1123, 1119 may be stored on a removable SD card, while others may be stored within a read-only computer chip. Some or all of the instructions and data may be stored in a location physically separate from the processors 1113, 1116, but still accessible by the processors 1113, 1116. Similarly, the processors 1113, 1116 may include a collection of processors that may perform simultaneous and / or sequential operations. The computing devices 1105, 1112 may each include one or more internal clocks that provide timing information, which may be used to time operations and programs executed by the computing devices 1105, 1112.
[0162] The server computing device 1105 can be configured to receive requests to process data from the user computing device 1112. For example, the environment 1100 can be part of a computing platform configured to provide various services to users via various user interfaces and / or APIs that expose platform services. One or more services can be a machine learning framework or set of tools for generating a neural network or other machine learning model according to a specified task and training data. The user computing device 1112 can send and receive data specifying operations to be performed by the computation unit 1000.
[0163] The devices 1112, 1105 can communicate directly and indirectly through the network 1160. The devices 1105, 1112 can set up listening sockets that can accept initiating connections to send and receive information. The network 1160 itself can include a variety of configurations and protocols, including the Internet, the World Wide Web, intranets, virtual private networks, wide area networks, local networks, and private networks using one or more corporate proprietary communication protocols. The network 1160 can support a variety of short and long range connections. The short and long range connections can be made at various bandwidths, such as 2.402 GHz to 2.480 GHz (typically associated with the Bluetooth® standard), 2.4 GHz and 11 GHz (typically associated with the Wi-Fi® communication protocol), or using various communication standards, such as the LTE® standard for wireless broadband communication. The network 1160 can additionally or alternatively support wired connections between the devices 1112, 1105, including various types of Ethernet® connections.
[0164] Although FIG. 11 illustrates a single server computing device 1105, user computing device 1112, and data processing system 1001, it is understood that aspects of the disclosure can be implemented according to a variety of different configurations and quantities of computing devices, including in sequential or parallel processing paradigms, or on a distributed network of multiple devices. In some implementations, aspects of the disclosure can be performed on a single device, and any combination thereof. In some examples, one or more devices implement one or more data processing systems, each data processing system including one or more computing units according to aspects of the disclosure. In some examples, a single device can implement multiple computing units, each of the multiple computing units configured to communicate with at least one other computing unit to perform distributed data processing tasks, e.g., in sequential or parallel processing.
[0165] Aspects of the present disclosure can be implemented as one or more computer programs in a digital circuit, a computer-readable storage medium, or as a combination of one or more of the foregoing. The computer-readable storage medium can be non-transitory, such as executable by a cloud computing platform, as one or more instructions stored on a tangible storage device.
[0166] In this specification, the phrase "configured" is used in different contexts relating to a computer system, hardware, or part of a computer program, engine, or module. When a system is said to be configured to perform one or more operations, this means that the system has appropriate software, firmware, and / or hardware installed thereon that causes the system to perform one or more operations during operation. When some hardware is said to be configured to perform one or more operations, this means that the hardware includes one or more circuits that receive inputs during operation and generate outputs corresponding to the one or more operations according to the inputs. When a computer program, engine, or module is said to be configured to perform one or more operations, this means that the computer program includes one or more program instructions and, when executed by one or more computers, causes the one or more computers to perform one or more operations.
[0167] Although the operations illustrated in the figures and recited in the claims are shown in a particular order, it is understood that the operations may be performed in an order different from that shown, and that some operations may be omitted, performed multiple times, and / or performed in parallel with other operations. Furthermore, the separation of different system components configured to perform different operations should not be understood as requiring the components to be separated. The components, modules, programs, and engines described may be integrated together as a single system or may be part of multiple systems.
[0168] Unless otherwise stated, the aforementioned alternatives are not mutually exclusive, but may be implemented in various combinations to achieve inherent advantages. These and other variations and combinations of the above-mentioned features may be utilized without departing from the subject matter defined by the claims, and therefore the foregoing description of the embodiments should be construed as illustrative, rather than limiting, the subject matter defined by the claims. Furthermore, the provision of embodiments described herein, and clauses such as "for example," "including," and the like, should not be construed as limiting the subject matter of the claims to any particular embodiment. Rather, the embodiments are intended to describe only one of many possible implementations. Furthermore, the same reference numbers in different drawings may identify the same or similar elements.
Claims
1. A computing unit, comprising: a two-dimensional systolic array of processing elements configured to receive first input elements from a first input matrix along a first direction of the two-dimensional systolic array and to receive second input elements from a second input matrix along a second direction of the two-dimensional systolic array, the computation unit further comprising: a first checksum circuit configured to generate one or more groups of first checksums from the first input elements while the two-dimensional systolic array receives the first input elements; a second checksum circuit configured to generate one or more groups of second checksums while the two-dimensional systolic array is receiving the second input elements; The two-dimensional systolic array is further configured to generate an output matrix from the first input matrix, the second input matrix, the one or more groups of the first checksums, and the one or more groups of the second checksums, and the computation unit further comprises: a computation unit comprising an output checksum circuit configured to receive the output matrix and determine from the output matrix the occurrence of one or more errors in the generation of the output matrix;
2. the output matrix includes a data sub-matrix, an output checksum row, and an output checksum column, the data sub-matrix including values generated by the two-dimensional systolic array using the first input elements and the second input elements; To determine the occurrence of one or more errors in the generation of the output matrix, the output checksum circuitry comprises: generating a row checksum from at least one row of the data sub-matrix; comparing the row checksum with the checksum in the output checksum string; The computation unit of claim 1 , configured to determine the occurrence of an error in the generation of the output matrix from the comparison of the row checksum and the checksum of the output checksum column.
3. To determine the occurrence of one or more errors in the generation of the output matrix, the output checksum circuitry comprises: generating a column checksum from at least one column of the data sub-matrix; comparing the column checksums to the output checksum row checksums; The computation unit of claim 2 , further configured to determine an occurrence of an error in the generation of the output matrix from the comparison of the column checksums and the checksums of the output checksum rows.
4. To determine the occurrence of one or more errors in the generation of the output matrix, the output checksum circuitry comprises: generating a row checksum from at least one row of the data sub-matrix; comparing the row checksum with the checksum in the output checksum string; The computation unit of claim 3 , further configured to determine an occurrence of an error in the generation of the output matrix from the comparison of the row checksum and the checksum of the output checksum column.
5. 3. The computing unit of claim 2, wherein to compare the row checksum to the checksum in the output checksum column, the output checksum circuitry is further configured to determine whether an absolute difference between the row checksum and the checksum in the output checksum column is within a predetermined threshold.
6. 2. The computing unit of claim 1, further comprising one or more checksum processing elements configured to receive a checksum from one or both of the first and second checksum circuits.
7. 2. The computational unit of claim 1, wherein one or both of the first and second checksum circuits are configured to transmit first or second checksums to the two-dimensional systolic array for processing based on a control signal.
8. 8. The computational unit of claim 7, wherein timing of the control signals to one or both of the first and second checksum circuits is based on a number of time steps for loading first or second input values across the two-dimensional systolic array.
9. 2. The computing unit of claim 1, wherein the computing unit is further configured, in response to the determination of an occurrence of one or more errors in the generation of the output matrix, to transmit an indication of the occurrence of the one or more errors to one or more devices coupled to the computing unit.
10. The two-dimensional systolic array further comprises:
10. The computational unit of claim 9, configured to receive a regulated voltage from the one or more devices after an indication of the occurrence of the one or more errors is transmitted, the regulated voltage being greater than a predetermined critical voltage of the two-dimensional systolic array.
11. 11. The computing unit of claim 10, wherein the two-dimensional systolic array is further configured to receive a first voltage that is less than the predetermined critical voltage of the computing unit until receiving the adjusted voltage in response to transmitting the indication.
12. The computing unit of claim 11 , wherein one or both of the first and second checksum circuits are configured to receive a second voltage that is greater than the predetermined critical voltage.
13. 11. The computational unit of claim 10, wherein one or both of the first and second checksum circuits comprise a two-input, two-stage pipelined adder circuit configured to delay the generation of one or both of the first and second checksums.
14. One or both of the first and second checksum circuits one or more two-cycle add circuits; 11. The computing unit of claim 10, comprising: a plurality of registers, the plurality of registers comprising: one or more first registers configured to transmit and receive data according to a first clock frequency, one or more second registers configured to transmit and receive data according to a second clock frequency, and one or more third registers configured to transmit and receive data according to a third clock frequency, the first, second, and third clock frequencies all being different frequencies.
15. 2. The computational unit of claim 1, wherein the two-dimensional systolic array is a fixed-output systolic array, and the first and second checksum circuits are both connected to a plurality of checksum processing elements, the checksum processing elements configured to receive generated checksums from one or both of the first and second checksum circuits.
16. The computational unit of claim 15 , wherein the plurality of checksum processing elements are arranged along a periphery of the two-dimensional systolic array.
17. The computing unit of claim 1 , wherein the computing unit is configured to generate a checksum only from the first checksum circuit or to generate a checksum only from the second checksum circuit.
18. 1. A data processing system comprising: The data processing system includes one or more processors, one or more memory devices, and a computing unit according to any one of claims 1 to 17.
19. The data processing system includes: configured to apply a voltage to the two-dimensional systolic array, the applied voltage being below a predetermined critical voltage of the two-dimensional systolic array, the data processing system further comprising:
20. The data processing system of claim 18, configured to receive an indication of one or more errors from the computational unit and, in response, increase the applied voltage above the predetermined critical voltage of the two-dimensional systolic array.
20. the data processing system applies a voltage below the predetermined critical voltage of the two-dimensional systolic array to one or more of the first checksum circuit, the second checksum circuit, and the output checksum circuit; 20. The data processing system of claim 19, further configured to, in response to receiving the indication, continue to apply a voltage below the predetermined critical voltage to one or more of the first checksum circuit, the second checksum circuit, and the output checksum circuit.
21. 20. The data processing system of claim 18, wherein the data processing system is configured to send control signals to one or both of the first and second checksum circuits, the timing of the sending being based on a number of time steps for loading first or second input values across the two-dimensional systolic array.
22. One or more computer programs that, when executed by a computing unit, cause the computing unit to perform an operation, the computing unit including a two dimensional systolic array of processing elements, a first checksum circuit, a second checksum circuit, and an output checksum circuit; The operation includes: receiving, by the two-dimensional systolic array, a first input element from a first input matrix along a first direction of the two-dimensional systolic array; receiving, by the two-dimensional systolic array, second input elements from a second input matrix along a second direction of the two-dimensional systolic array; generating, by the first checksum circuitry, one or more groups of first checksums from the first input elements while the two-dimensional systolic array is receiving the first input elements; generating, by the second checksum circuitry, one or more groups of second checksums while the two-dimensional systolic array is receiving the second input elements; generating, by the two-dimensional systolic array, an output matrix from the first input matrix, the second input matrix, the one or more groups of first checksums, and the one or more groups of second checksums; receiving the output matrix by the output checksum circuit; and determining, by the output checksum circuitry, from the output matrix, the occurrence of one or more errors in the generation of the output matrix.
Citation Information
Patent Citations
Mapping method and device for virtualized wireless sensor network, and storage medium
CN110933728A
matrix multiplier
JP2021508125A