Error checking for systolic array computation

KR103015879B1Active Publication Date: 2026-09-04GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
KR1020237026835
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-08-24
Filing Date
2022-07-13
Publication Date
2026-09-04
Estimated Expiration
2042-07-13

Smart Images

  • Figure R1020237026835_ABST
    Figure R1020237026835_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure relate to a computation unit configured to implement a systoric array and to detect errors while processing data on the systoric array. A checksum circuit communicating with the systoric array is configured to compute checksums and perform error detection while the systoric array processes input data. Instead of pre-generating checksums of input matrices, input matrices may be fed directly to the systoric array through the checksum circuit. At the output side, the checksum circuit may generate checksums and compare them with the checksums of output matrices generated by the systoric array. The operations of checking for errors and generating output matrices may be performed without delaying the operations of the systoric array and without pre-processing the input matrices.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] This application is a continuation of U.S. Patent Application No. 17 / 410,558 filed on August 24, 2021, claiming priority to U.S. Provisional Application No. 63 / 222,549 filed on July 16, 2021, titled "Error Checking For Systolic Array Computation," the disclosures of which are incorporated herein by reference. Background Technology

[0002] Systolic arrays are arrays of processing elements, such as processors, microprocessors, or specialized circuits, configured to process data. Adjacent processing elements of a systolic array may be connected, for example, via one or more interconnects on a printed circuit board, such as wires or other physical connections.

[0003] Algorithm-based fault tolerance (ABFT) refers to methods or techniques for detecting and correcting errors during the execution of different types of arithmetic or logic algorithms, such as matrix multiplication and Fourier transforms. For matrix multiplication of input matrices A and B and output matrix C, e.g., A × B = C, ABFT for matrix multiplication involves generating a checksum row for A and a checksum column for matrix B. Each element of the checksum row for A is the result of a linear operation performed on the elements of individual columns of matrix A, for example, the result of adding each element of the column of matrix A to generate the checksum value of the checksum row of matrix A. Similarly, each checksum value of column B is the result of a linear operation performed on the elements of each row of matrix B.

[0004] After matrices A and B (with their corresponding checksum rows / columns) are multiplied, the output matrix includes not only a sub-matrix representing the product of multiplying matrices A and B, but also a checksum row and a checksum column. As part of ABFT for matrix multiplication, the checksum values of the checksum rows and columns of matrix C are compared with the result of performing the same linear operation on matrices A and B, but now on the sub-matrix of C. If the result of performing a linear operation on a row or column of the sub-matrix of C does not match the corresponding checksum value of C, the mismatch indicates that an error occurred during matrix multiplication. Prior Art Document M. Safarpour, R. Inanlou, O. Silven. “Algorithm Level Error Detection in Low Voltage Systolic Array,” IEEE. (2021.07.06.) <https: / / ieeexplore.ieee.org / document / <9475533>

[0005] Aspects of the present disclosure relate to a computation unit configured to implement a systoric array and to detect errors while processing data on the systoric array. A checksum circuit communicating with the systoric array is configured to compute checksums and perform error detection while the systoric array processes input data. Instead of generating checksums of input matrices in advance, input matrices may be fed directly to the systoric array through the checksum circuit. At the output side, according to one aspect, the checksum circuit may generate a checksum for an output to the systoric array when the systoric array completes a corresponding operation to generate an output, or while the output is being streamed from the systoric array. According to another aspect, at the output side, the checksum circuit may generate checksums and compare them with the checksums of an output matrix generated by the systoric array.

[0006] Aspects of the present disclosure provide numerous technical advantages. Error detection through checksums for the outputs and operands of a systoric array can be performed without pre-processing operations to generate checksums for inputs or post-processing operations to compare checksums for accuracy. To determine whether hardware defects exist in the array or in individual processing elements of the array, the systoric array on the chip can be tested rapidly according to various random or pseudo-random testing patterns. Since no known inputs or outputs are required and no signatures are needed for pseudo-random testing, many different test patterns can be generated and deployed as desired for testing. Furthermore, error checking of operations to generate output matrices can be performed without delaying the operations of the systoric array and without pre-processing the input matrices.

[0007] At runtime, a processor as implemented according to aspects of the present disclosure can detect soft errors or hard errors, such as errors not caused by hardware defects and errors caused by hardware defects. Rapid identification of errors may be particularly important when the processor is part of an accelerator performing critical and time-sensitive tasks. For example, soft error detection may be particularly important when the processor processes input to a machine learning model trained to perform tasks related to banking, autonomous vehicle navigation / control, airplane or spacecraft navigation, etc.

[0008] In some examples, a computation unit as described herein may be configured to detect timing violation errors that can be detected and resolved to provide better energy efficiency to the systoric array. The circuitry for performing error detection may be implemented to reduce the likelihood of timing errors, thereby enabling accurate error detection to continue even while the systoric array is operating at a supply voltage below a predetermined threshold voltage. Aspects of the present disclosure also provide for detecting errors in some types of operations, including matrix multiplication, which can be corrected to improve the accuracy and reliability of the processor. Different computation units may be tuned to different voltage and / or frequency levels to tune the units for better performance, for example, during inference of a machine learning model executed by the computation unit. Detecting errors can improve the accuracy of the computation overall.

[0009] Aspects of the present disclosure can be implemented with minimal overhead and without affecting the performance of the computational unit's syntonic array. Since no preprocessing of inputs is required, the efficiency of error checking for the computational unit is further improved compared to other approaches where inputs are preprocessed. In addition, inputs to checksum circuits for performing error detection as described herein do not delay inputs to the syntonic array for processing.

[0010] "Relaxed" fault tolerance can also be applied only where errors are actually detected to further reduce overhead, particularly software interaction with the computation unit. In some examples, relaxed fault tolerance can be advantageous when the primary use case is to detect the presence of errors without needing to pinpoint the exact source of the error.

[0011] As error detection is applied to the computation unit as described herein, error correction mechanisms, such as error correction processes based on ABFT for matrix multiplication, can be implemented to recover from hard errors detected during the runtime of the computation unit.

[0012] Aspects of the present disclosure relate to a computation unit, wherein the computation unit comprises: a two-dimensional sistoric array of processing elements—the sistoric array is configured to receive first input elements from a first input matrix along a first direction of the sistoric array and to receive second input elements from a second input matrix along a second direction of the sistoric array—; a first checksum circuit configured to generate one or more groups of first checksums from the first input elements while the sistoric array receives the first input elements; a second checksum circuit configured to generate one or more groups of second checksums while the sistoric array receives the second input elements—the sistoric array is further configured to generate an output matrix from the first input matrix, the second input matrix, one or more groups of first checksums, and one or more groups of second checksums—and an output checksum circuit configured to receive the output matrix and, from the output matrix, determine the occurrence of one or more errors in the generation of the output matrix.

[0013] The aforementioned aspect may include one of the following features alone or in any combination. In some examples, the aforementioned aspect includes all of the following features together.

[0014] The output matrix comprises a data sub-matrix containing values ​​generated by a systoritic array using first input elements and second input elements, an output checksum row, and an output checksum column; to determine the occurrence of one or more errors in the generation of the output matrix, the output checksum circuit is configured to generate a row checksum from at least one row of the data sub-matrix; compare the row checksum with the checksum of the output checksum column; and determine the occurrence of an error in the generation of the output matrix from the comparison of the row checksum and the checksum of the output checksum column.

[0015] To determine the occurrence of one or more errors in the generation of the output matrix, the output checksum circuit is further configured to generate a column checksum from at least one column of the data sub-matrix; compare the column checksum with the checksum of the output checksum row; and determine the occurrence of an error in the generation of the output matrix from the comparison of the column checksum and the checksum of the output checksum row.

[0016] To determine the occurrence of one or more errors in the generation of the output matrix, the output checksum circuit is further configured to generate a row checksum from at least one row of the data sub-matrix; compare the row checksum with the checksum of the output checksum column; and determine the occurrence of an error in the generation of the output matrix from the comparison of the row checksum and the checksum of the output checksum column.

[0017] To compare the row checksum with the checksum of the output checksum column, the output checksum circuit is further configured to determine whether the absolute difference between the row checksum and the checksum of the output checksum column is within a predetermined threshold.

[0018] The computation unit further includes one or more checksum processing elements configured to receive checksums from one or both of the first checksum circuit and the second checksum circuit.

[0019] One or both of the first checksum circuit and the second checksum circuit are configured to transmit the first checksums or the second checksums to a systolic array for processing based on a control signal.

[0020] The timing of the control signal for one or both of the first checksum circuit and the second checksum circuit is based on the number of time steps for loading the first input values ​​or the second input values ​​across the systolic array.

[0021] The computation unit is further configured to transmit an indication of the occurrence of one or more errors to one or more devices connected to the computation unit in response to a decision regarding the occurrence of one or more errors in the generation of the output matrix.

[0022] The systolic array is further configured to receive a regulated voltage from one or more devices after an indication of the occurrence of one or more errors is transmitted, and this regulated voltage is higher than the threshold voltage for the computation unit.

[0023] The systolic array is further configured to receive a first voltage lower than the threshold voltage for the computation unit until it receives a regulated voltage in response to transmitting a signal.

[0024] One or both of the first checksum circuit and the second checksum circuit are configured to receive a second voltage higher than the threshold voltage.

[0025] One or both of the first checksum circuit and the second checksum circuit include 2-input 2-stage pipeline adder circuits configured to delay the generation of one or both of the first checksums and the second checksums.

[0026] One or both of the first checksum circuit and the second checksum circuit comprises one or more 2-cycle adder circuits and a plurality of registers, wherein the plurality of registers comprises one or more first registers configured to receive and transmit data according to a first clock frequency, one or more second registers configured to receive and transmit data according to a second clock frequency, and one or more third registers configured to receive and transmit data according to a third clock frequency, wherein the first clock frequency, the second clock frequency, and the third clock frequency are all different frequencies.

[0027] In some examples, the sistoric array is weight stationary. In other examples, the sistoric array is an output stationary sistoric array, and both the first checksum circuit and the second checksum circuit are connected to a plurality of checksum processing elements, and the checksum processing elements are configured to receive checksums generated from one or both of the first checksum circuit and the second checksum circuit.

[0028] Multiple checksum processing elements are arranged along the periphery of the systoric array.

[0029] The computation unit is configured to generate checksums only from the first checksum circuit or to generate checksums only from the second checksum circuit.

[0030] Aspects of the present disclosure relate to a data processing system, wherein the data processing system comprises one or more processors, one or more memory devices, and a computation unit, wherein the computation unit comprises a two-dimensional sistoric array of processing elements—the sistoric array is configured to receive first input elements from a first input matrix along a first direction of the sistoric array and to receive second input elements from a second input matrix along a second direction of the sistoric array—; a first checksum circuit configured to generate one or more groups of first checksums from the first input elements while the sistoric array receives the first input elements; a second checksum circuit configured to generate one or more groups of second checksums while the sistoric array receives the second input elements—the sistoric array is further configured to generate an output matrix from the first input matrix, the second input matrix, one or more groups of first checksums, and one or more groups of second checksums—; It includes an output checksum circuit configured to receive an output matrix and, from the output matrix, determine the occurrence of one or more errors in the generation of the output matrix.

[0031] The aforementioned aspects may include one or more of the following features, either alone or in combination. In some examples, the aforementioned aspects may include all of the following features together.

[0032] The data processing system is configured to apply a voltage to a systolic array—the applied voltage is less than the threshold voltage of the systolic array—; receive indication of one or more errors from a computation unit, and in response, increase the applied voltage to a level higher than the threshold voltage of the systolic array.

[0033] The data processing system is further configured to apply a voltage below the threshold voltage of the systolic array to one or more of the first checksum circuit, the second checksum circuit, and the output checksum circuit; and, in response to receiving an indication, to continue applying a voltage below the threshold voltage to one or more of the first checksum circuit, the second checksum circuit, and the output checksum circuit.

[0034] The data processing system is configured to transmit a control signal to one or both of the first checksum circuit and the second checksum circuit, and the timing of transmission is based on the number of time steps for loading the first input values ​​or the second input values ​​across the systolic array.

[0035] An aspect of the present disclosure relates to one or more non-transient computer-readable storage media encoded with computer instructions, wherein the computer instructions, when executed by a computational unit comprising a two-dimensional sistoric array of processing elements, a first checksum circuit, a second checksum circuit, and an output checksum circuit, cause the computational unit to perform operations, the operations comprising: receiving first input elements from a first input matrix along a first direction of the sistoric array by the sistoric array; receiving second input elements from a second input matrix along a second direction of the sistoric array by the sistoric array by the sistoric array; generating one or more groups of first checksums from the first input elements by the first checksum circuit while the sistoric array receives the first input elements by the first checksum circuit; generating one or more groups of second checksums by the second checksum circuit while the sistoric array receives the second input elements by the sistoric array; The operation of generating an output matrix from a first input matrix, a second input matrix, one or more groups of first checksums, and one or more groups of second checksums by a systolic array; the operation of receiving the output matrix by an output checksum circuit; and the operation of determining the occurrence of one or more errors in the generation of the output matrix by the output checksum circuit and from the output matrix. Brief explanation of the drawing

[0036] FIG. 1 is a block diagram of a computation unit according to aspects of the present disclosure.

[0037] FIG. 2 illustrates exemplary input matrices processed by a computation unit to generate individual output matrices having corresponding checksums according to aspects of the present disclosure.

[0038] Figure 3 is a block diagram of a vertical checksum circuit and checksum processing elements.

[0039] FIG. 4 is a block diagram of a horizontal checksum circuit. The horizontal checksum circuit (115) includes a plurality of adder circuits, registers, and multiplexers.

[0040] Figure 5 is a block diagram of an output checksum circuit. The output checksum circuit is configured to generate a checksum from the rows and columns of the data sub-matrix of the output matrix C.

[0041] FIG. 6 is an exemplary computational unit having a vertical checksum circuit and an output checksum circuit, but not a horizontal checksum circuit.

[0042] FIG. 7 is an exemplary computational unit having an output static systolic array.

[0043] FIG. 8a is an exemplary vertical checksum circuit. The vertical checksum circuit is configured to generate checksums from the rows of an input matrix of length 8.

[0044] FIG. 8b is an exemplary vertical checksum circuit having 2-input 2-stage pipeline adder circuits.

[0045] FIG. 8c is a schematic diagram of another exemplary vertical checksum circuit having 2-cycle non-pipeline adder circuits.

[0046] FIG. 9a is a flowchart of an exemplary process for performing matrix multiplication using error detection on a computation unit according to aspects of the present disclosure.

[0047] FIG. 9b is a flowchart of an exemplary process for detecting the occurrence of errors during processing of a computation unit according to aspects of the present disclosure.

[0048] FIG. 10a is a block diagram of a data processing system implementing an exemplary computation unit.

[0049] FIG. 10b is a flowchart of an exemplary process for adjusting the voltage supplied to a computation unit according to aspects of the present disclosure.

[0050] FIG. 11 is a block diagram of an exemplary environment for implementing a data processing system including a computation unit (1000). Specific details for implementing the invention

[0051] FIG. 1 is a block diagram of a computation unit (100) according to aspects of the present disclosure. The computation unit (100) includes a syntonic array (110) of processing elements (110A-P), a horizontal checksum circuit (115), a vertical checksum circuit (120), an output checksum circuit (125), and checksum processing elements (130).

[0052] A sistoric array (110) can be implemented in various different ways and, generally, may include one or more data buses or interconnects connecting different processing elements to neighboring processing elements. For example, the sistoric array may include data buses along each row and / or column of the sistoric array (110), and the processing elements (110A-P) of the sistoric array (110) may be configured in a square or rectangular arrangement (however, in some examples, the processing elements (110A-P) may be configured in other geometric arrangements, such as a hexagonal arrangement). In some examples, the sistoric array (110) may include one or more data buses along processing elements positioned diagonally up or down relative to each other in the array. In some examples, the systoric array (110) may implement two-phase non-overlapping clocks and latches to leverage time borrowing.

[0053] Data buses enable data to flow from a processing element to a processing element, or from an external input to a processing element, or from processing to an external output. External inputs and outputs may include, for example, other devices coupled to the computation unit (100) for communication, or memory devices that are part of the computation unit (100) or located outside the unit (100). In some examples, the systoric array (110) may be configured to receive or generate an input or output from outside the systoric array (110) only through processing elements on the periphery of the systoric array (110). In other examples, the systoric array (110) may be configured to receive or generate an external input or output from any of the processing elements (110A-P).

[0054] A processing element of the systoric array (110) may be a processor, microprocessor, or any specialized circuit configured to receive one or more inputs, process the input(s), and generate one or more outputs. The processing elements may be, for example, data processing units (DPUs) configured to perform a specific type of operation. In some examples, the processing element may be configured to temporarily store memory, for example, in a register or using one or more circuits, such as latches or flip / flop circuits.

[0055] Processing elements (110A-P) can be arranged according to a number of topologies. For example, processing elements (110A-P) can be arranged as an array, a mesh, or a cube. Each processing element may have one or more channels connecting itself to neighboring processing elements. The one or more channels may be, for example, wires or other physical connections on a silicon wafer on which the systoric array (110) is implemented. Depending on the number of channels for each processing element, the systoric array (110) may have one or more dimensions. For the purposes of explanation, the systoric array (110) is assumed to be a two-dimensional mesh, but it is understood that in other implementations, the systoric array (110) may be configured according to other topologies without loss of universality.

[0056] It is understood that 16 processing elements are shown in the systoric array (110) and arranged as a 4 × 4 mesh, but in various examples the number of processing elements in the systoric array (110) may vary, for example, 65,536 processing elements are arranged as a 256 × 256 mesh or 16,384 processing elements are arranged as a 128 × 128 mesh.

[0057] Data may travel independently across each data bus or interconnect of the systoric array (110). In other words, different processing elements may receive different data at different points in time, measured, for example, as time steps, and transmit it to neighboring processing elements. A time step may refer to one or more clock cycles of the computation unit (100) or a larger system communicably coupled to the computation unit (100). For ease of description in this specification, a time step in this specification will refer to a single clock cycle unless otherwise indicated. For each time step, each processing element (110A-P) may be performing a different operation. For example, in a single time step, some processing elements may be receiving a new input, transmitting a generated output, processing a received input, or are idle due to a lack of data to receive or transmit.

[0058] The functions of the processing elements (110A-P) may correspond to the functions of the sistoric array (110), for example, operations configured to be performed by the sistoric array (110). For example, the sistoric array (110) may be configured to perform matrix multiplication on two input matrices, for example, a weight matrix for a neural network layer and an input matrix for that layer. In other examples, the sistoric array (110) may be configured to perform convolution operations as part of the received input to a convolutional neural network ("ConvNet"). In the example of matrix multiplication, each processing unit (110A0P) may be configured to perform a multiplication-and-accumulation ("MAC") operation. Processing elements configured to perform MAC operations may be referred to as MAC units. Aspects of the present disclosure relate to error checking and correction during matrix multiplication performed by a systorilic array of MAC units, and the following examples assume that processing elements (110A-P) are configured for MAC operations unless otherwise indicated.

[0059] For example, a processing element configured as a MAC unit can process data as follows: for each time step, the processing element can multiply the matrix element loaded into the unit with the incoming input matrix element and add the product of the multiplication to a running partial sum maintained by the MAC unit. The MAC unit can forward the input matrix element to a neighboring MAC unit along a data bus in a first direction (e.g., to the right) and transmit the updated partial sum to the MAC unit along the same or different data bus in a second direction (e.g., downward).

[0060] The sistoric array (110) may be configured to perform operations such as various levels of pipeline sub-operations, inputs of different sizes, debugging, etc. For example, one operation for the sistoric array (110) may be defined by operations performed individually by sub-arrangements of groups of processing elements, such as 2 × 2 or 4 × 4 processing elements. In other examples, the sistoric array (110) may be configured to perform more complex operations, such as discrete Fourier transforms, or to define different control flows that control how data is input to and output from the sistoric array (110). The sistoric array (110) may also perform operations such as matrix multiplication on blocks of inputs that are part of larger matrices to which the sistoric array (110) is tasked with multiplying. In other examples, different processing units of the systolic array (110) may be implemented as different types of circuits or combinations of circuits, such as adders, multipliers, accumulators, circuits configured to compare two inputs, etc.

[0061] Although the examples in this specification focus on algorithm-based fault tolerance for matrix multiplication on a systolic array (110), it is understood that aspects of this disclosure are applicable to error detection on systolic arrays configured to perform other types of operations under any of the various conditions, such as the operations and conditions described herein.

[0062] A sistoric array (110) is configured to perform matrix multiplication on input matrices having different data types for their individual elements. For example, the sistoric array (110) can perform matrix multiplication on binary matrices having 1s and 0s as elements; integer matrices having positive or negative integers as elements; and floating-point matrices having floating-point values ​​as elements, etc. The elements may have different levels of precision, for example, 8-bit integers, or the floating-point values ​​may have different ranges and levels of precision, etc.

[0063] As described in more detail herein, the vertical checksum circuit (120) and the horizontal checksum circuit (115) are configured to generate checksums for input matrices (140A, B). References to "vertical" and "horizontal" checksum circuits are not intended to be limited to the orientation of inputs received by the systoric array (110), but are provided merely for ease of explanation with the orientation of input matrices for exemplary circuits as illustrated in the drawings. In some examples, the vertical checksum circuit (120) as described herein may be oriented to provide inputs to the systoric array (110) along the columns of the array instead of columns, and the horizontal checksum circuit (115) may be oriented to provide inputs to the systoric array (110) along the rows of the array instead of columns. The checksum circuits (115, 120) may be arranged in any direction relative to each other to provide inputs to the systoric array (110).

[0064] The checksum processing elements (130) include additional processing elements for storing checksums generated for each column of the input matrix A (140A). As each column is processed by the vertical checksum circuit (120), the resulting checksums are pushed into the checksum processing elements (130) which can be communicateably coupled to the systoric array (110) and are included during the multiplication of the input matrices (140A-B) which are augmented by their individual checksum columns or rows by the checksum circuits (115, 120). As described in more detail herein, the orientation and position of the checksum processing elements (130) may also vary with respect to the orientation and position of the vertical checksum circuit (120).

[0065] The output checksum circuit (125) is configured to generate a checksum for the resulting output matrix, which is the product of the input matrix A (140A) (and its checksum) and the input matrix (140B) (and its checksum) multiplied by the systoric array (110). The output matrix includes a sub-matrix which is the product of the input matrix A and the input matrix B, along with the checksum column and checksum row generated by the output checksum circuit (125). The output checksum circuit (125) is configured to compare the sum of the individual columns or rows of the checksum column and row with the sum of the sub-matrix and to detect errors indicated by discrepancies between the checksum value and the computed sum compared.

[0066] FIG. 2 illustrates exemplary input matrices (140A, B) that are processed by a computation unit (100) to generate individual output matrices (140C) having corresponding checksums according to aspects of the present disclosure.

[0067] In this example, the input matrix A (140A) ("Matrix A") is multiplied by the input matrix B (140B) ("Matrix B"), which can be expressed as A × B. In FIG. 2, the input matrices are 4 × 4 and match the dimensions of the sistoric array (110). In other examples, the dimensions of the input matrices may vary. For instance, matrix A may be n × 4 and matrix B may be 4 × m, where n and m are positive integers. As described herein, the sistoric array (110) may have arbitrary dimensions, and thus the inner dimension (4 in this example) of matrices A and B may also vary based on the dimensions of the sistoric array (110).

[0068] Based on the exemplary orientation and position of the checksum circuits (115, 120), the input matrix A (140A) ("Input Matrix A") and the input matrix B (140B) ("Input Matrix B") are first manipulated to ensure that the order in which the elements of the matrices are received by the systoric array (110) allows for accurate multiplication according to the rules of matrix multiplication. A sequencer configured as part of a data processing system implementing the computation unit (100) may be implemented to perform some or all of the operations currently described in relation to manipulating matrices A and B and coordinating the timing for the elements of the input matrices to be provided to the systoric array (110). The sequencer and other components of an exemplary data processing system are described herein with reference to FIG. 10a.

[0069] The input matrix A is transposed (A T ), each of its columns is staggered by one time step, e.g., one clock cycle. The time steps may define one or more clock cycles or time periods that can be managed by a timing circuit, e.g., a timing circuit as described herein with reference to FIG. 10a. Transposed matrix (AT Each column of ) passes through individual processing elements of the vertical checksum circuit (120). As described herein with reference to FIG. 3, the vertical checksum circuit (120) is a transposed matrix (A T Generates a checksum for each column of ) (each of which corresponds to a row of the original input matrix A). The vertical checksum circuit (120) generates a checksum for the transposed matrix (A T After processing the received elements of ), the vertical checksum circuit (120) passes the received elements to the systoric array (110) to perform matrix multiplication. The generated checksums are also pushed to the checksum processing elements (130).

[0070] The computation unit (100) may be configured to receive input matrices of different block sizes. The computation unit (100) may receive multiple matrices that are sub-matrices of a larger input. For example, the computation unit (100) may be 128 × 128 elements, and may be configured to perform block matrix multiplication when the original input matrices are large, e.g., 4,096 × 4,096 elements. The "A" input matrices may be 128 × 128 elements, and the "B" input matrices for the computation unit (100) may be 1,024 × 128 elements.

[0071] The input matrix B is also pre-processed before its elements are received by the horizontal checksum circuit (115). The input matrix B is inverted such that the positions of each element in each row are swapped, for example, with the first element as the last element and the second element as the second-to-last element (B r ). Inverted matrix (B r ) is also staggered by one time step for each row. As described herein with reference to FIG. 4, the horizontal checksum circuit (115) is an inverted matrix (B rGenerates a checksum for each row of ). The horizontal checksum circuit (115) generates a checksum for the inverted matrix (B r After processing the received elements of ), the horizontal checksum circuit (115) passes the received elements to the systolic array (110) to perform matrix multiplication.

[0072] In some examples, elements of matrix A or elements of matrix B are loaded into a sistoric array (110), for example, one element is loaded into an individual processing element of the sistoric array (110), and then, elements of the other matrix are streamed into the sistoric array (110). For example, input matrix A may be a weighted matrix corresponding to the weights of a layer of a neural network, while input matrix B may contain input values ​​to be multiplied with the weights of a layer of a neural network. Elements of matrix A may be loaded into the sistoric array (110), for example, by latching or weighted loading, and elements of matrix B may be streamed into the sistoric array (110) for processing. The choice of streaming or loading elements of the input matrices may depend on the dimensions of the input matrices, for example, wider or larger matrices may be streamed into the sistoric array, while matrices that fit the sistoric array (110) may be loaded. The output matrix C (140C) ("matrix C") can be streamed from the sistoric array (110), or in some examples, can be fully pushed out of the sistoric array (110) only after the multiplication is completed.

[0073] For example, based on the number and configuration of processing elements of the systoric array for performing matrix multiplication, after a number of time steps have elapsed, the output checksum circuit (125) begins to receive elements of matrix C. Because input matrices A and B are staggered, the elements of matrix C will also be staggered. However, in some examples, the systoric array (110) may be configured as an output static systoric array, where the elements of output matrix C are retained in the systoric array (110) until the multiplication is completed, and then pop out through the output checksum circuit (125). As described in more detail with reference to FIG. 5, the output checksum circuit (125) generates checksums for the output matrix and compares such values ​​with checksums from the output matrix C. If the generated checksums do not match the checksums of the output matrix C, the output checksum circuit (125) may generate a flag or other indication indicating that an error has been detected.

[0074] The output matrix C has two parts: a data sub-matrix and a checksum part. The checksum part of the output matrix C is indicated by surrounding elements of the output matrix C, illustrated by dashed outlines for the checksum row (202) of the output matrix C and dotted outlines for the column checksum (204). The data sub-matrix (140D) of the checksum matrix is ​​indicated by cross-thatched elements. The data sub-matrix (140D) corresponds to the product of input matrices A and B without added checksums. The entire output matrix C (element (140C)) having the checksum row (202), column checksum (204), and data sub-matrix (140D) corresponds to the product of matrices A and B, each having its own individual checksum row or column as generated by the horizontal and vertical checksum circuits (115, 120).

[0075] Note that the checksums of the rows and columns (202, 204) share a common element indicated by the element in the upper right corner of the output matrix C. The checksums of the rows and columns (202, 204) for the output matrix C can be discarded after the output checksum circuit (125) performs checksum generation and comparison.

[0076] Overall, the introduction of horizontal, vertical, and output checksum circuits adds or does not add a low additional cost in terms of the additional time step(s) required to increase the time steps, that is, in the case of horizontal / vertical checksum circuits, to pass inputs to the systoric array (110) or to pass output elements from the systoric array (110) to their respective destinations from the output checksum circuit (125). The additional logic required is also scaled linearly for a two-dimensional systoric array, for example, only the length of the dimensions of the array increases instead of the number of processing elements of the array.

[0077] FIG. 3 is a block diagram of a vertical checksum circuit (120) and checksum processing elements (130). The checksum processing elements (130) include checksum processing elements (130A, B, C, and D) that can be configured similarly to the processing elements of the systoric array (110). In some examples, the error checking functionality provided by the vertical checksum circuit (120) may be disabled, for example, based on user input. In some cases, the checksum processing elements (130) may be used as part of performing matrix multiplication on the systoric array (110), even when the error checking is disabled.

[0078] The vertical checksum circuit (120) includes a plurality of adder circuits (305A, B, and C) (“adder circuits (305)”); as well as registers (310A, B, C, and D) (“registers (310)”). The registers (310) may be implemented according to any of the various techniques for storing data on the circuit, for example, as individual groups of flip / flop circuits. The adder circuits (305) may be implemented according to any of the various techniques for constructing circuits configured to receive two inputs and generate the resulting sum of the two inputs. The number of adder circuits (305) varies depending on the length of the row input of the input matrix. The example provided herein focuses on 4 × 4 input matrices, and thus the vertical checksum circuit (120) includes three adder circuits and four registers.

[0079] The vertical checksum circuit (120) receives an input matrix A, and the input matrix A is transposed as described in FIG. 2 and staggered by one time step (A T The inputs for the first row are the input matrix (A T It is received by the input matrix (140AT) (indicated as the input matrix in FIG. 3). The first row of the input matrix (140AT) corresponds to the first column of the input matrix A. The input matrix (140AT) includes elements a1, a2, a3, a4, a5, a6, a7, a8, a9, a10, a11, a12, a13, a14, a15, and a16.

[0080] Generally, the vertical checksum circuit (120) accumulates partial sums of checksums in registers (310A-C) and stores the calculated checksums in register (310D) before pushing checksums to checksum processing elements (130). Because the inputs are staggered, at the first time step, the vertical checksum circuit (120) receives input a1. Input a1 is stored in register (310A) and also passed to the systoric array (110) for processing, that is, matrix multiplication according to the matrix multiplication algorithm on the systoric array.

[0081] In the second time step, inputs a1 and a2 are received by an adder circuit (305A). The partial sum of a1 + a2 is stored in a register (310B). Input a2 is also passed to a systoritic array (110). Although not shown in FIG. 3, in the second time step, the next input a5 above a1 is also stored in a register (310A) and passed to a systoritic array (110).

[0082] In the third time step, the adder circuit (305B) adds the partial sum stored in the register (310B) and the input a3. The updated partial sum generated by the adder circuit (305B) is stored in the register (310C). The input a3 is also passed to the systolic array (110). Additionally, the next input (input a9) in the same column as a1 is passed to the register (310A), and the previous contents of the register (310A) are added to the next input a6 in the same column as a2 and stored in the register (310B).

[0083] In the fourth time step, the adder circuit (305C) adds the partial sum stored in the register (310C) and the input a4. Since the adder circuit (305C) is the last adder, this adder circuit (305C) adds the first column of A (i.e., A TA checksum is generated for the first row of the input matrix A. The checksum is stored in register (310D). Input a4 is also passed to the systolic array (110), and successive inputs (a7) of the input matrix A are passed to the systolic array (110) and added to update individual partial sums across registers (310A-C).

[0084] In the fifth time step, the checksum stored in the register (310D) is pushed down to the last available processing element among the checksum processing elements (130). For example, the checksum generated for the first column (inputs a1-a4) of the input matrix A is pushed to the processing element (130D). Also, in the fifth time step, input a8 is pushed, and the checksum for the second column of the input matrix A can be calculated and stored in the register (310D). In subsequent time steps, the checksum for the second column (inputs a5-a8) is pushed to the processing element (130C), then the third checksum for the third column (inputs a9-a12) of the processing element (130B) is pushed, and the checksum for the fourth column (inputs a13-a16) is pushed to the checksum processing element (130A).

[0085] After checksums are generated for an input matrix A, the systoric array (110) can generate an output matrix C having its own corresponding rows and columns of checksums by using the checksums stored in the checksum processing elements (130), for example, by multiplying each checksum with the corresponding row of the input matrix B.

[0086] The matrix in the previous example that is transposed (A TAlthough it has been described that checksums for columns of input matrix A are generated by summing the elements of each row of ), it is understood that the vertical checksum circuit (120) can be configured to perform any linear operation to generate checksums, for example, by adding and multiplying with constant values ​​or factors, or by multiplying checksum rows by a predetermined vector.

[0087] FIG. 4 is a block diagram of a horizontal checksum circuit (115). The horizontal checksum circuit (115) includes a plurality of adder circuits (405A-D) ("adder circuits (405)"); registers (410A-D) ("registers (410A-D)"); and multiplexers (415A-D) ("multiplexers (415)"). The adder circuits (405) and registers (410) may be implemented as described herein with reference to the adder circuits (305) and registers (310) shown in FIG. 3. The multiplexers (415) may be 2-to-1 multiplexers configured to transmit the value of a corresponding register or a corresponding input value to the systorilic array (110). Multiplexers can be replaced with any variation of a decision circuit section, such as one or more circuits or switches, controlled by one or more control signals to disable and enable specific parts of a horizontal checksum circuit (115) configured to select among a number of inputs and gate unselected inputs.

[0088] The horizontal checksum circuit (115) receives the input matrix B, and the input matrix B is inverted and staggered as described in FIG. 3 (B r ) . For ease of explanation, the inputs are presented in ascending order. Input matrix (B rThe inputs for the first column received by ) (i.e., the inputs for the last column of the input matrix B) include elements b1, b2, b3, b4, b5, b6, b7, b8, b9, b10, b11, b12, b13, b14, b15, and b16.

[0089] Generally, the horizontal checksum circuit (115) accumulates sums that are parts of checksums in the registers (410A-D) and, based on a control signal, determines when to push the contents of the registers (410A-D) as completed checksums to the systoric array (110). For example, a control signal, e.g., an is_checksum flag, can be propagated from the multiplexer (415A) to the multiplexer (415D) based on the size of the internal dimension of the input matrix A. As described herein, the input matrix B can be much longer than the input matrix A (e.g., 4 × m) (e.g., n × 4). Thus, checksums are generated across segments of the input matrix B and periodically transmitted to the systoric array (110). The control signal enables the horizontal checksum circuit (115) to determine when the data transmitted to the systoric array (110) is a regular input or generated checksum from matrix B. Based on a control signal, the horizontal checksum circuit (115) can pass the checksum and reset the accumulation of partial sums for a new set of inputs from matrix B. The control signal may be provided automatically by a sequencer communically coupled to the computation unit (100), and / or programmatically by software containing commands executed by the computation unit (100) to perform matrix multiplication on input matrices A and B.

[0090] For example, until a control signal is received, the multiplexer (415A) is configured to pass inputs from input matrix B to the systoritic array (110). When the control signal is received, the multiplexer (415A) instead passes the contents of register (410A) to the systoritic array (110) for processing. The timing of the control signal to the multiplexer (415A) must coincide with the completion of the checksum of the corresponding row of input matrix A. For example, the control signal may be passed to the multiplexer (415A) after a suitable number of time steps have elapsed to receive and accumulate each input of the row of input matrix B. An exemplary breakdown of operations performed on a time step basis for the horizontal checksum circuit (115) follows.

[0091] Because the inputs are staggered, at the first time step, the horizontal checksum circuit (115) receives input b13. Input b4 is added to the contents of the register (410A), i.e., becomes 0, because the register (410A) is initially empty. The multiplexer (415A) checks for a control signal, and in the absence of a control signal, transmits input b4 to the systolic array (110).

[0092] In the second time step, the horizontal checksum circuit (115) receives input b8. Input b8 is added to the contents of register (410B), i.e., becomes 0, because the registered 410B is initially empty. The multiplexer (415B) checks for a control signal, and in the absence of a control signal, passes input b8 to the systolic array (110). In the same time step, the horizontal checksum circuit (115) also receives the next input b3 following b4.

[0093] In the third time step, the horizontal checksum circuit (115) receives input b12. Input b12 is added to the contents of register (410C). The horizontal checksum circuit (115) also receives consecutive inputs from the upper rows of input matrix B and adds each received input to the individual partial sums stored in registers (410A-C). The horizontal checksum circuit (115) also receives inputs b7 and b2.

[0094] In the fourth time step, the horizontal checksum circuit (115) receives input b16. Input b16 is added to the contents of register (410D) and passed to the systoric array by the multiplexer (415D). The horizontal checksum circuit (115) also receives inputs b11, b6, and b1.

[0095] Additionally, during the fourth time step, the multiplexer (415A) receives a control signal indicating that the contents of the register (410A) are the checksum for the first row of matrix B. In this time step, the register (410A) contains the sum b4 + b3 + b2 + b1. The multiplexer (415A) pushes the checksum from the register (410A) to the systoric array (110) to be included in the matrix multiplication operation. Subsequently, the register (410A) is reset to 0 to begin calculating the checksum for the next set of inputs. The control signal is pushed down to the next multiplexer, namely the multiplexer (415B).

[0096] In the fifth time step, the multiplexer (415B) pushes the contents of the register (410B) into the systoric array (110). By the fifth time step, the register (410B) may contain the sum b8 + b7 + b6 + b5. The register (410B) is cleared, and the control signal is pushed down to the next multiplexer, the multiplexer (415C).

[0097] In the sixth time step, the multiplexer (415C) pushes the contents of the register (410C) into the systoric array (110). By the sixth time step, the register (410C) may contain the sum b12 + b11 + b10 + b9. The register (410C) is cleared, and the control signal is pushed down to the next multiplexer, the multiplexer (415D).

[0098] In the seventh time step, the multiplexer (415D) pushes the contents of the register (410D) into the systoric array (110). Up to the seventh time step, the register (410D) may contain the sum b16 + b15 + b14 + b13. The register (410D) is cleared. In subsequent time steps, the horizontal control circuit (115) may generate new individual checksums for new groups of inputs from matrix B in each row of matrix B.

[0099] In some examples, the vertical checksum circuit (120) consists of multiplexers that manage whether inputs from matrix A or generated checksums are passed to the systoric array (110). For example, the vertical checksum circuit (120) may receive an input matrix that is much longer than can be fitted into the systoric array (110) and whose contents are instead streamed to the systoric array (110). In such examples, the checksum processing elements (130) may not be used, or may instead be used by the horizontal checksum circuit (115). For example, the checksum processing elements (130) may be used by the horizontal checksum circuit (115) when receiving an input matrix that is not streamed to the systoric array (110) but is instead loaded directly into the processing elements (110A-P). In such examples, the checksum processing elements (130) may be arranged as additional rows for the systoric array (110). In other examples, both horizontal and vertical checksum circuits (115, 120) may be configured as multiplexers as described herein to handle streaming inputs. In some implementations, specific orientations and positions of the checksum circuits (115, 120) and checksum processing elements (130) may vary based on the direction(s) in which the systolic array (110) is configured to receive inputs.

[0100] Although the previous example described generating checksums for the rows of the input matrix B by summing the elements of each row, it is understood that the vertical checksum circuit (120) can be configured to perform any linear operation to generate checksums, for example, by adding and multiplying with constant values ​​or factors, or by multiplying the checksum rows by a predetermined vector. An additional consideration for the horizontal checksum circuit (115) is to generate checksums that are valid inputs to the sistoric array (110). While some data types, such as floating-point numbers, may have sufficient precision to represent even very large sums of individual elements of the input matrix, the sistoric array (110) may be configured, for example, for matrix multiplication low-precision integer values, for example, for matrices having 8-bit quantized values. The checksums calculated by the horizontal checksum circuit (115) may require higher precision than the inputs to the sistoric array (110).

[0101] To address these potential problems without changing the hardware or configuration of the systoric array, in some examples, the horizontal checksum circuit (115) may be configured to generate some value modulo the checksums, e.g., the maximum value supported by the systoric array (110). In other examples, Galois field arithmetic may be applied to make the precision required to represent the checksum equal to the precision required for the systoric array (110) to represent the elements of the matrices being multiplied. In such examples, the horizontal checksum circuit (115) may be composed of an additional circuit section configured to perform Galois field arithmetic.

[0102] It is understood that exemplary operations by horizontal and vertical checksum circuits are described as being performed over consecutive time steps, but some delay or idle time steps may be inserted as needed to synchronize the operations of the horizontal and vertical checksum circuits, and that the operations are performed by processing elements of the systoric array (110).

[0103] FIG. 5 is a block diagram of an output checksum circuit (140C). The output checksum circuit (140C) is configured to generate a checksum from the rows and columns of a data sub-matrix (140D) of an output matrix C. As described with reference to FIG. 2, the output matrix (C) includes a checksum row (202), a checksum column (204), and a data sub-matrix (140D), the latter of which corresponds to the product of matrices A and B without checksum rows and checksum columns generated by horizontal and vertical checksum circuits. The data sub-matrix (140D) has output elements d1, d2, d3, d4, d5, d6, d7, d8, d9, d10, d11, d12, d13, d14, d15, and d16. The checksum row (202) has checksums c1, c2, c3, c4, and c5. The checksum column (204) has checksums c5, c6, c7, c8, and c9.

[0104] In some examples, the output checksum circuit (140C) may be a combination of horizontal and vertical checksum circuits (115, 120). The registers (500A-D) and adder circuits (502A-C) may be configured similarly to the adder circuits and registers of the horizontal checksum circuit (115).

[0105] Because the input matrices A and B are staggered so that their individual elements are pushed into the systoric array (110), the input elements of the output matrix C are also staggered when they are pushed out of the systoric array (110). The output checksum circuit (140C) generates a checksum for each row of the output matrix C and compares it with the corresponding checksum of the checksum column (204). For example, outputs d13, d14, d15, and d16 are transmitted through the registers (500A-D) and adder circuits (502A-D) as described herein with reference to FIG. 3, and the row inputs of the input matrix (140AT). When 500D receives the checksum generated from the input, i.e., the outputs d13, d14, d15, and d16, the comparator circuit (505) compares the checksum of the register (500D) with the first checksum pushed from the checksum column (204), i.e., checksum c9.

[0106] If the checksum (c9) matches the contents of the register (500D), the error check is successful. If the checksum (c9) does not match the contents of the register (500D), the output checksum circuit (125) generates an error. In response to the generated error, the computation unit (100) may perform one or more of the various actions described herein with reference to FIG. 2. The output checksum circuit (125) may continue to calculate the checksums of each row of the data sub-matrix (140D) using the corresponding checksum of the checksum column (204). For example, the comparator circuit (505) compares the sum d9 + d10 + d11 + d12 with the checksum (c8); compares the sum d5 + d6 + d7 + d8 with the checksum (c7); compares the sum d1 + d2 + d3 + d4 with the checksum (c6); Then, compare the sum c1 + c2 + c3 + c4 with the checksum (c5).

[0107] Comparator circuits implemented as part of the output checksum circuit (125) can directly compare checksums or compute the absolute value of the difference between the compared checksums. Subsequently, the comparator circuits can determine whether the absolute value of the difference between the checksums is within a predetermined threshold. In some examples, the threshold may be defined programmatically, for example, based on usage cases with a higher or lower tolerance for error. In other examples, the threshold may be specified during the design of the computation unit (100). In some applications, for example, when running some machine learning models, absolute precision is not required, and some error in the calculated floating-point values ​​may be allowed, for example, within a predetermined threshold.

[0108] The output checksum circuit (125) generates checksums for each row of the data sub-matrix (140D), but the output checksum circuit (125) can also generate checksums for each column of the output matrix C and compare each generated checksum with the individual checksum of the checksum row (202). The output checksum circuit (125) may include adder circuits (504A-E); demultiplexers (503A-E); registers (507A-E); registers (510A-E); and comparator circuits (515A-E). It is understood that although the output checksum circuit (125) is illustrated as including demultiplexers, the demultiplexers (503A-E) may be replaced with any of the switches configured to transmit an input to one of a plurality of outputs based on control signals for disabling or enabling specific parts of the output checksum circuit (125), for example, control signals.

[0109] For example, in the first time step, the output d13 is transferred to a demultiplexer (503A) as well as a register (500A) (for computing the checksum for the first row of the output matrix C as described herein). The demultiplexers (503A-E) are configured to transfer the checksums of the checksum row (202) to individual registers (510A-E) while transferring the remainder of the columns of the output matrix C to individual adder circuits (504A-E) and registers (507A-E).

[0110] In the second time step, d13 is pushed to the adder circuit (504A), which adds the output d13 to the contents of the register (507A), which is initially 0. Also, during the second time step, the output (d14) is pushed to the adder circuit (502A) and also to the demultiplexer (503B). The output checksum circuit (125) can continue to receive and process the output elements of the output matrix C until it reaches the checksums of the checksum row (202). At the same time, the output checksum circuit (125) receives a control signal, e.g., a control signal used for the horizontal checksum circuit (115), and based on the presence of the control signal, can push the checksum values ​​of the checksum row (202) to the corresponding registers (510A-E).

[0111] When both contents of registers (510A) are loaded with corresponding checksums, e.g., checksum c1 for register (510A); checksum c2 for register (510B); checksum c3 for register (510C); checksum c4 for register (510D); and checksum c5 for register (510E), each comparator circuit compares the stored checksums with the checksums generated and stored in registers (507A-E). For example, comparator circuit (515A) compares the sum d1 + d5 + d9 + d13 with checksum c1; comparator circuit (515B) compares the sum d2 + d6 + d10 + d14 with checksum c2; and comparator circuit (515C) compares the sum d3 + d7 + d11 + d15 with checksum c3; The comparator circuit (515D) compares the sum d4 + d8 + d12 + d16 with the checksum c4; and the comparator circuit (515E) compares the sum c6 + c7 + c8 + c9 with the checksum c5.

[0112] In some examples, the output checksum circuit (125) may be configured to generate one or more measurements or indicators corresponding to the number and severity of different errors detected. For example, the output checksum circuit (125) may compute the absolute value of the difference between the compared checksums, which may be transmitted to a coupled external processor in addition to a occurrence flag. The absolute value of the difference between the compared checksums may be a measure of the severity of the errors detected by the output checksum circuit (125), for example, larger values ​​may indicate more severe errors than smaller values.

[0113] As in comparisons made by the comparator circuit (505), if any of the comparator circuits (515A-E) identifies any mismatches (or mismatches exceeding a predetermined threshold), the output checksum circuit (125) may generate a flag indicating that an error occurred during the matrix multiplication of input matrices A and B.

[0114] FIG. 6 is an exemplary computation unit (600) having a vertical checksum circuit (620) and an output checksum circuit (625), but not a horizontal checksum circuit. Input matrices (640A, 640B); an output matrix (640C) having a checksum column (602); and checksum processing elements (630) are also illustrated in FIG. 6. The output checksum circuit (625) may be configured to generate checksums for each row of the output matrix C and compare them with checksums generated by the vertical checksum circuit.

[0115] Omitting the horizontal checksum circuit still enables error detection and further reduces complexity and input to the computation unit (600) because, as described herein with reference to the horizontal checksum circuit (115), there is no requirement to manage the timing of control signals, for example, through a program executable by the computation unit (600). In some cases, where the correction actions taken do not require the exact location of the error, error detection without a specific indication of where the error occurred, as provided by comparing both the checksum row and the checksum column, may still be useful. For example, if the default correction action is to repeat the previous matrix multiplication or to cold restart the computation unit, the exact location of the error is not required.

[0116] In addition, omitting the horizontal checksum circuit allows the logic for generating checksums to be completely separated from the systoric array. Referring again to FIG. 4, the checksums generated by the horizontal checksum circuit (115) are pushed into the systoric array (110), and thus, for example, if the systoric array (110) is configured to receive input only up to a certain level of precision, it is required to follow the input specifications of the systoric array (110). As described in FIG. 4, since the computation unit (600) omits the horizontal checksum circuit, the remaining checksums are stored in checksum processing elements (630) that may be separate from the systoric array (610). In some examples, the checksums generated by the vertical checksum circuit (620) may have different precision measurements and / or may be of a completely different type from the elements of the input matrices.

[0117] The omission of the horizontal checksum circuit can allow for greater flexibility among different types of linear operations for checksum generation that can be implemented. For example, floating-point checksums can be used for 8-bit integer elements; 16-bit integers can be used for 8-bit integer elements; and Galois field arithmetic or modulo addition or multiplication operations can generally be used for checksum calculation.

[0118] In some implementations, the vertical checksum circuit (620) may be enhanced to generate multiple checksum columns based on different linear operations for generating checksums from the input elements of the rows of the input matrix A. Given that for each linear code defined for a finite field, there exists a corresponding linear real code having similar error detection and correction capabilities, additional error detection processes, such as N check-symbol Reed-Solomon codes, may be applied when using Galois field arithmetic to generate checksums. Otherwise, the corresponding real code may be used.

[0119] In some examples, when input matrix A is streamed and input matrix B is latched or preloaded into a systoric array, the computation unit for the systoric array may instead implement a suitably configured horizontal checksum circuit, that is, coupled to checksum processing elements and configured to generate checksum columns as described herein with reference to FIG. 3; and may implement a suitably configured output checksum circuit, that is, configured to compute column checksums and compare them with checksums generated by the horizontal checksum circuit. In other words, different implementations of the computation unit may have different orientations of the checksum circuit and positions of the checksum circuit with respect to how the systoric array receives the input, without affecting the overall error detection functionality as described herein.

[0120] A computation unit as described herein may be configured to operate using neither a horizontal checksum circuit nor a vertical checksum circuit, using either one of them, or using both. The computation unit may not use either the horizontal checksum circuit or the vertical checksum circuit, for example, because error detection is programmatically disabled. The same computation unit may also be configured to operate using only one of the horizontal and vertical checksum circuits to perform error detection, for example, as described herein with reference to FIG. 6. The same computation unit may also be configured to operate using both the horizontal and vertical checksum circuits to perform error detection, for example, as described herein with reference to FIG. 1 through 5.

[0121] FIG. 7 is an exemplary computation unit (700) having an output static sistoric array (710). In the output static sistoric array, each processing element holds individual output elements of an output matrix C (740C), and input matrices (A, B, 740A, B) can be streamed to the sistoric array (710). When both input elements of matrices A and B are processed, the output matrix C is pushed through an output checksum circuit (725). The computation unit (700) also includes a vertical checksum circuit (720) and a horizontal checksum (715). Both checksum circuits (715, 720) are coupled to checksum processing elements (730), which include rows (730A) of checksum processing elements and columns (730B) of checksum processing elements.

[0122] The vertical checksum circuit (720) is configured similarly to the vertical checksum circuit (120), (input matrix A and matrix (A TAs described herein with reference to ), checksums are generated from the rows of an input matrix A (740A), which may be a transposed version of the original input matrix. A column (730B) of checksum processing elements may correspond to checksum processing elements (130), which is configured to receive and store the generated checksums to be used for later comparison by an output checksum circuit (725).

[0123] The horizontal checksum circuit (715) can also be implemented similarly to the vertical checksum circuit (120), but with corresponding rows (730A) of checksum processing elements. Both the horizontal checksum circuit (715) and the rows (730A) are turned 90 degrees with respect to their counterpart circuits (720) and columns (730B) of the processing elements. However, the horizontal checksum circuit (715) can also be configured to generate checksums for corresponding columns (or corresponding rows if the input matrix B (740B) is transposed) of the input matrix B (740B).

[0124] The output checksum circuit (725) may be configured to generate and compare checksums as previously described with reference to the output checksum circuit (125) and FIG. 5. In this exemplary arrangement of horizontal and vertical checksum circuits, the checksum row (702) is pushed first before the values ​​of the data sub-matrix (740D) of the output matrix (740C). The output checksum circuit (725) may be configured to push each first output element of each individual column of the output matrix (740C) into the corresponding register of the comparator circuit and proceed to accumulate the remaining column output elements of the individual register coupled to the individual adder circuit.

[0125] In some examples, the output checksum circuit (725) may be configured to receive a checksum from a checksum row and subtract each output element of the output matrix C (740C) from the checksum, instead of adding separate partial sums to compare with the received checksum of the output matrix C (740C). After subtracting all the output elements, if the result is not zero or is close to zero within a predetermined threshold, the output checksum (725) may transmit an indication that an error has been detected. Depending on the results of error detection by comparing checksums within a row (704) of checksum processing elements (e.g., by subtracting the output elements of the corresponding rows, or by comparing each checksum with the corresponding accumulated checksum as described herein), the output checksum circuit (740) may also pinpoint exactly where the error occurred, for example, at the intersection of a column and a row having incorrect checksums.

[0126] In the configuration described with reference to FIG. 7, the computation unit (700) completely separates the error check logic implemented by checksum circuits (715, 720, and 725) and the data path of the systorilic array when, for example, receiving inputs and performing matrix multiplication. This allows for different linear operations of increased variety for ABFT for matrix multiplication without concerns about matching data types and precision, for example (as described with reference to the omission of the horizontal checksum circuit in the computation unit (600) of FIG. 6).

[0127] This is the inverse of the operation of the output checksum circuit (125) as described with reference to FIG. 5, where a control signal is used to indicate when the output checksum circuit (125) has received the checksum of the output matrix (140C). In some examples, the output checksum circuit (725) may also use a control signal, for example, when large input matrices (740A, B) are streamed to the systolic array (710). In such examples, the control signal may be transmitted to the output checksum circuit (125) at the beginning of each group of output elements for checksum generation and comparison, rather than at the end.

[0128] Systoric arrays can operate according to different voltage levels. The threshold supply voltage is a value indicating a voltage sufficient for the accurate operation of the systoric array in the face of different environmental or process-related variability that may affect circuit performance, for example. Voltage supplied to a systoric array at a level lower than the threshold supply voltage may be more energy-efficient because, for example, less energy is required to operate the systoric array, but it runs the risk of errors, such as timing errors, in the face of the aforementioned variability. If timing errors are not corrected or resolved, they can rapidly cascade into more serious errors, and in particular, in computation units using systoric arrays where proper timing of inputs and outputs across processing elements is critical. A data processing system implementing a computation unit as described herein may supply a reduced voltage to the computation unit and, if necessary, raise the voltage in response to receiving error detection flags from the computation unit. Compared to the normal operation of a computation unit having a systolic array, and especially when multiple computation units are executed in parallel, errors are relatively rare; therefore, generally, accidental errors and correction actions can be outweighed by running the systolic array at a supply voltage below a threshold supply voltage level.

[0129] The computation unit can detect errors related to low voltage according to the same mechanisms described herein with reference to FIGS. 1 through 6. While other approaches require additional logic in a number of different latches, shadow latches, or flip / flop circuits, the horizontal, vertical, and output checksum circuits are independent of the systoric array and do not require delaying the execution of matrix multiplication or other operations on the array. Additionally, the predetermined thresholds used by the comparator circuits of the output checksum circuits to compare the generated checksum with the received checksum can be tuned to allow smaller errors caused by the reduced voltage supplied to the systoric array.

[0130] For example, the control logic implemented by the data processing system implementing the computation unit can also further improve the energy efficiency of the computation unit by adjusting the supplied voltage or frequency, such as the clock frequency at which the data processing system operates, in relation to the observed rates at which errors occur.

[0131] Timing errors can occur in the processing of the systolic array as well as in the execution of error detection logic provided by the horizontal checksum circuit, the vertical checksum circuit, and / or the output checksum circuit. Timing errors in these circuits can be resolved under a number of different approaches. In some examples, a higher voltage may be supplied to one or more of the horizontal checksum circuit, the vertical checksum circuit, and the output checksum circuit. In other examples, specialized transistors or other components may be implemented in the vertical, horizontal, and / or output checksum circuits to improve the rate at which the error check operations described herein are performed, e.g., the rate at which checksums are generated and checksums are compared.

[0132] In other examples, as described herein with reference to FIGS. 8a through 8c, the horizontal checksum circuit may be modified to allow a wider timing margin when different operations are performed at time steps. The increased margin can mitigate the possibility of timing errors occurring in the error detection logic itself, and accordingly, reduce the possibility of inaccurate error detection or missing undetected errors caused by operating the computation unit at a lower supply voltage.

[0133] FIG. 8a is an exemplary vertical checksum circuit (800A). The vertical checksum circuit (800A) is configured to generate checksums from rows of an input matrix of length 8. The vertical checksum circuit (800A) can receive inputs a0-a7 and may include registers (805A-H) and adder circuits (804A-H). The generated checksum stored in register (805H) can be pushed to a checksum processing element (830). The inputs a0-a7 are also pushed to processing elements (803A-H) of a systorilic array.

[0134] FIG. 8b is an exemplary vertical checksum circuit (800B) having 2-input 2-stage pipeline adder circuits. Circuit (800B) is a modified version of circuit (800A) in which adder circuits (804A-H) are replaced with 2-input 2-stage pipeline adder circuits (820A, B, and C) ("adder circuits (820)"). Circuit (800B) may be configured to delay the generation of checksums. The adder circuits (820A, B, C) have a latency of two time steps, e.g., two clock cycles, and a throughput of one sum or add operation performed per clock cycle. Each adder circuit (820) may include a 2-to-1 adder circuit comprising intermediate adder circuits (A and B), and a register between the two intermediate adder circuits for storing intermediate sums.

[0135] The vertical checksum circuit (800B) divides the processing of input elements a0-a7 across two pipeline stages (806A and 806B). In stage (806A), odd-numbered input elements (e.g., a1, a3, a5, and a7) are accumulated across adder circuits (820A), while in stage (806B), even-numbered input elements (e.g., a0, a2, a4, and a6) are accumulated across adder circuits (820B). The partial sums generated through stages (806A, B) are added together in adder circuit (806C), and the results are passed to the checksum processing element (830).

[0136] The additional latency provided by the adder circuits (820) increases the timing margin for the accurate operation of the vertical checksum circuit (800B), thereby reducing the possibility or miscalculations resulting from timing errors caused by operating at a lower level supply voltage even when the vertical checksum circuit (800B) operates below a threshold voltage level. Compared to the vertical checksum circuit (800A), the vertical checksum circuit (800B) requires three additional clock cycles to generate the checksum, but the additional latency of the vertical checksum circuit (800B) does not affect the throughput of the input elements a0-a7 to the systoric array connected to the vertical checksum circuit (800B), nor does it affect the processing latency, for example, by not increasing the number of clock cycles in which the systoric array processes the input matrices.

[0137] In other examples of the vertical checksum circuit (800B), the vertical checksum circuit (800B) may include different numbers of stages and / or different numbers of adder circuits or other processing circuits configured to generate checksums according to a specific linear operation over different numbers of clock cycles.

[0138] FIG. 8c is a schematic diagram for another exemplary vertical checksum circuit (800C) having 2-cycle non-pipeline adder circuits (850). The vertical checksum circuit (800C) includes a number of 2-cycle adder circuits (850) as well as flip / flop circuits (860, 865, and 870). The 2-cycle adder circuits are configured to receive two inputs and generate the sum of such two inputs over two clock cycles. As indicated by the legend (880), in the schematic diagram for the vertical checksum circuit (800C), the flip / flop circuits (860) are colored with solid lines, the flip / flop circuits (865) have vertical hatch marks, and the flip / flop circuits (870) have diagonal hatch marks. In some examples, it is understood that other types of circuits configured to store data may be used instead of flip / flop circuits (860, 865, 870). Flip / flop circuits (860, 865, and 870) operate at different clock frequencies.

[0139] Flip / flop circuits (860) have a clock frequency ( It operates at ). The clock frequency can be set by a timing circuit connected to a computation unit implementing a vertical checksum circuit (800C), as described herein with reference to FIGS. 10 and 11. The flip / flop circuits (865 and 870) operate at different individual clock frequencies ( and It can operate at ). Clock frequencies ( and ) may have different phases as indicated by chart (890). The vertical checksum circuit (800C) also has an increased timing margin compared to the circuit (800A) to mitigate the possibility of timing errors affecting the error check functionality of the computation unit implementing the vertical checksum circuit (800C), as described herein with reference to the vertical checksum circuit (800B).

[0140] In other examples of the vertical checksum circuit (800C), the vertical checksum circuit (800B) may include different flip / flop circuits operating at different frequencies over different numbers of clock cycles, and / or different numbers of adder circuits or other processing circuits configured to generate checksums according to a specific linear operation.

[0141] FIG. 9a is a flowchart of an exemplary process for performing matrix multiplication using error detection on a computation unit according to aspects of the present disclosure.

[0142] According to block (905A), the sistoric array of the computation unit receives first input elements from a first input matrix along a first direction of the sistoric array. For example, the first input elements may be input elements of input matrix A received from around the top of the sistoric array, for example, as described herein with reference to FIG. 3. It is understood that the sistoric array may be configured to receive input from any direction, and that the corresponding checksum circuit is implemented in accordance with aspects of the present disclosure.

[0143] According to block (910A), the sistoric array receives second input elements from a second input matrix along a second direction of the sistoric array. The second direction may be a horizontal direction with respect to the fixed orientation of the sistoric array, e.g., from left to right, while the first direction may be a vertical direction with respect to the fixed orientation of the sistoric array, e.g., from top to bottom. The second input elements may be input elements from an input matrix B, for example, as described herein with reference to FIG. 4.

[0144] According to block (915A), the first checksum circuit of the computation unit generates one or more groups of first checksums from the first input elements while the systoric array receives the first input elements. The first checksum circuit may be a vertical checksum circuit, e.g., the vertical checksum circuit (120) of FIG. 2. The one or more groups of first checksums may be columns of checksums generated in checksum processing elements, e.g., the checksum processing elements (130) of FIG. 1, which are later stored. In some examples, the one or more groups of checksums may refer to multiple columns or rows of checksums generated according to other linear operations. In some examples, the computation unit includes one or more checksum processing elements configured to receive checksums from one or both of the first checksum circuit and the second checksum circuit.

[0145] According to block (920A), the second checksum circuit of the computation unit generates one or more groups of second checksums from the second input elements while the systoric array receives the second input elements. The second checksum circuit may be a horizontal checksum circuit, e.g., the horizontal checksum circuit (115) of FIG. 1. The second checksum circuit may generate checksums as described herein with reference to, e.g., the horizontal checksum circuit (115) and FIG. 4. In some examples, when an input matrix for the second checksum circuit is streamed to the systoric array, the second checksum circuit may push a control signal across the circuit to indicate when checksums can be pushed to the systoric array, and new checksums may be generated. The timing of the control signal for one or both of the first checksum circuit and the second checksum circuit may be based on the number of time steps for loading the first input values ​​or the second input values ​​across the systoric array. A time step can be one or more clock cycles.

[0146] In some examples, the systoric array is an output static systoric array, and both the first checksum circuit and the second checksum circuit are connected to a plurality of checksum processing elements, and the checksum processing elements are configured to receive checksums generated from one or both of the first checksum circuit and the second first checksum circuit. The plurality of checksum processing elements may be located on the periphery of the systoric array, for example, as described herein with reference to FIG. 7.

[0147] According to block (925A), the systoric array generates an output matrix from a first input matrix, a second input matrix, one or more groups of first checksums, and one or more groups of second checksums. The output matrix may be, for example, an output matrix C as described herein with reference to FIG. 2.

[0148] According to block (930A), the output checksum circuit receives the output matrix. As described herein with reference to FIG. 5, the output checksum circuit may receive the output matrix when the output matrix is ​​pushed out of the systoric array. In some examples, the computation unit generates checksums only from the first checksum unit or generates checksums only from the second checksum unit.

[0149] According to the diamond (935A), the output checksum circuit determines whether an error is detected in the computation of the output matrix. If an error is detected ("Yes"), according to the block (940A), the output checksum circuit transmits an indication of the occurrence of an error (or, if applicable, an error exceeding). Otherwise ("No"), the process (900A) can continue with new input elements to process and perform error detection.

[0150] Until it receives a response to an indication of error detection, the systoric array may receive a first voltage below a threshold voltage that can be predetermined as described herein with reference to the preceding descriptions and FIGS. 8a through 8c. In response to transmitting an indication or identifying one or more errors, the systoric array may begin to receive a second voltage higher than the threshold voltage. The systoric array may receive the second voltage automatically, or the second voltage may be applied to the systoric array by another device. In some examples, the systoric array automatically returns to receiving the first voltage after a certain period of time has elapsed without detecting any errors.

[0151] One or more of the first, second, and output checksum circuits may continue to receive a voltage lower than the threshold voltage of the systolic array. The first, second, and output checksum circuits may be implemented according to aspects of the present disclosure to increase the timing margin of the operation of the circuits in order to mitigate timing errors resulting from the reduced voltage. For example, one or both of the first checksum circuit and the second checksum circuit may include 2-input 2-stage pipeline adder circuits as described herein with reference to Drawing 800B. As another example, one or both of the first checksum circuit and the second checksum circuit may include a plurality of registers, each of which may include one or more 2-cycle adder circuits and one or more flip / flop circuits operating at different clock frequencies. For example, some registers may operate at a first frequency, some registers may operate at a second frequency, and others may operate at a third frequency. The second and third frequencies may be half of the first frequency, having different phases.

[0152] FIG. 9b is a flowchart of an exemplary process (900B) for detecting the occurrence of errors during processing of a computation unit according to aspects of the present disclosure. For example, the process (900B) may be performed as part of a decision regarding whether errors have been detected according to the diamond (935A) of FIG. 9a.

[0153] The output checksum circuit, according to block (910B), generates a column checksum from at least one row of the data sub-matrix. The data sub-matrix may be the output from processing both input matrices without their corresponding checksums, for example, processing the data sub-matrix (140D) as described with reference to FIG. 2. For example, the output checksum circuit may accumulate output elements in the received row of the output matrix and store the checksum in a register, for example, a register (500D) as described herein and illustrated in FIG. 5. Additionally, according to block (910B), the output checksum circuit generates a column checksum from at least one column of the data sub-matrix. The column checksums may be generated independently of the row checksums and may be generated in parallel for each column of the output matrix. For example, the output checksum circuit may generate column checksums and store them in individual registers of the circuit, such as registers (510A-E) as described and illustrated with reference to FIG. 5. In some examples, both column and row checksums are generated by the output checksum circuit. In other examples, only column checksums and only row checksums are generated.

[0154] According to block (920B), the output checksum circuit compares the row checksum with the checksum of the output checksum column of the output matrix. For example, the output checksum column may correspond to the checksum column (204). The comparison by the output checksum circuit may be performed by one or more comparator circuits, for example, the comparator circuit (505) described and illustrated with reference to FIG. 5. Additionally, according to block (920B), the output checksum circuit compares the column checksum with the checksum of the output checksum row of the output matrix. The output checksum row may be, for example, the checksum row (202) of the output matrix C as illustrated in FIG. 2. One or more output checksum rows may be checked in parallel by the output checksum circuit using comparator circuits (515A-E) described and illustrated in FIG. 5.

[0155] Comparisons between checksums within output checksum columns / rows and row / column checksums can be performed in parallel or sequentially, for example. In some examples, comparisons using row checksums are performed only on the checksums of the output checksum columns, and vice versa for column checksums and output checksum rows.

[0156] According to block (930B), the output checksum circuit determines the occurrence of an error in the generation of the output matrix by comparing the checksum of the output checksum column with the row checksum. The output checksum circuit may determine, for example, whether the compared checksums match or whether they match within a predetermined threshold. If not, the output checksum circuit may transmit an indication that an error has occurred, for example, as described herein with reference to block (940A) and FIG. 9a. Additionally, according to block (930B), the output checksum circuit determines the occurrence of an error in the generation of the output matrix by comparing the column checksum with the checksum of the output checksum row.

[0157] Decisions from comparisons between row / column checksums and the corresponding checksums of the output checksum column / row may occur, for example, simultaneously, in parallel, or sequentially. In some examples, the output checksum circuit may transmit a indication after one error is detected or after a threshold number of errors is indicated. In some examples, the indication may include a measure of the severity of the error, e.g., the absolute difference between the compared checksums. In other examples, the indication may include information regarding the source of the error, e.g., the intersection of the errors detected for the checksum of the checksum row and the checksum for the checksum column.

[0158] FIG. 10a is a block diagram of a data processing system (1001) implementing an exemplary computation unit (1000). The computation unit (1000) may be any of various different computation units, e.g., the computation unit (100) described herein with reference to FIG. 1 through 5. The computation unit (1000) may implement any of various combinations of horizontal, vertical, and output checksum circuits as described throughout this specification.

[0159] The data processing system may include a host interface (1005), a sequencer circuit (1010), one or more processor(s) (1015), a memory (1020), and a timing circuit (1025). The data processing system (1001) may be implemented on one or more devices across one or more physical locations, as described herein with reference to FIG. 11. In some examples, the components of the described data processing system (1001) may be implemented on one or more chips capable of interfacing with a host device according to any of various data buses or other physical interconnection interfaces. In some examples, the data processing system (1001) may be implemented on one or more devices on a network, such as one or more servers of a cloud platform.

[0160] The processor(s) (1015) and memory (1020) may be any of the various different types of processors and memories as described herein with reference to FIG. 11. In some examples, the processor(s) (1015) receive instructions executable by the computation unit (1000) to process data. For example, the instructions may be part of a computer program written to perform operations using the computation unit (1000).

[0161] The sequencer circuit (1010) can convert received commands into one or more signals understood by the computation unit (100), which causes the computation unit (1000) to perform any of various pre-configured operations. These operations may include, for example, loading data from memory (1020) into a systoritic array (not shown) of the computation unit (1000), moving data to one or more of the processing elements of the systoritic array, processing data by one or more processing elements, and pushing data from the systoritic array. The sequencer circuit (1010) may also be configured to generate one or more control signals to control when checksums are pushed to the computation unit (1000), for example, as described herein with reference to the horizontal checksum circuit (115) and FIG. 4.

[0162] The host interface (1005) may be configured to receive data from outside the data processing system, for example, from a processor or other device, and to transmit data generated by the computation unit (1000), for example, the product of matrix multiplication, to one or more devices or processors.

[0163] The timing circuit (1025) may be configured to control the timing of the computation unit, such as the clock frequency or clock rate of the computation unit. The time steps described herein with reference to the operations of the checksum circuits may be measured in terms of clock cycles managed by the timing circuit (1025).

[0164] The data processing system (1001) may also be connected to a power source (1030). The power source (1030) may be a battery or other form of power available on a host device implementing the data processing system, or a source outside the host device, and may be connected to the data processing system (1001) and the host device via some wireless or physical connection, for example, via wires. The power source (1030) may supply voltage to the computation unit (1000), and the voltage may be managed by the processor(s) (1015), for example, adjusted to be higher or lower.

[0165] FIG. 10b is a flowchart of an exemplary process for adjusting the voltage supplied to a computation unit according to aspects of the present disclosure.

[0166] According to block (1050), a data processing system implementing a computation unit may apply voltage to a systoric array of the computation unit. The supplied voltage may be smaller than a threshold voltage of the systoric array, which can be predetermined based on the operation of the computation unit when used, environmental factors and / or architectural features of the systoric array, etc.

[0167] According to block (1060), the data processing system receives indications of one or more errors from the computation unit and, in response, increases the applied voltage to a level higher than the threshold voltage of the systoric array. The computation unit may transmit indications of detected errors as part of performing the processes (900A-B) described, for example, in FIGS. 9a and 9b. The data processing system may increase the voltage supplied to the systoric array to reduce the risk of additional errors occurring as a result of timing violations.

[0168] According to block (1070), the data processing system may continue to apply a voltage below the threshold voltage of the systolic array to one or more of the first checksum circuit, the second checksum circuit, and the output checksum circuit of the computation unit in response to receiving a sign. As described above, the horizontal, vertical, and output checksum circuits may be configured according to one or more of various different approaches to increase the timing margins of operations performed by the circuits, such as the generation of checksums, the storage of checksums, and the comparison of checksums to reduce the risk of timing errors.

[0169] In some examples, the data processing system may reduce the voltage to the systoric array after a certain period of time or after a condition is satisfied, for example, after a certain period of time has elapsed without receiving additional indications of errors from the computation unit. In some examples, the data processing system may cause the computation to re-execute one or more operations, etc., by performing any of various actions in response to receiving an indication of an error, such as rolling back the operations performed by the computation unit to a previous checkpoint.

[0170] FIG. 11 is a block diagram of an exemplary environment (1100) for implementing a data processing system (1001) including a computation unit (1000). The system (1001) may be implemented on one or more devices having one or more processors at one or more locations, such as a server computing device (1105). The user computing device (1112) and the server computing device (1105) may be communicably coupled to one or more storage devices (1130) via a network (1160). The storage device(s) (1130) may be a combination of volatile and non-volatile memory and may be located at the same or different physical locations as the computing devices (1112, 1105). For example, the storage device(s) (1130) may include any type of non-transient computer-readable medium capable of storing information, such as a hard drive, a solid-state drive, a tape drive, an optical storage device, a memory card, a ROM, a RAM, a DVD, a CD-ROM, a writeable and read-only memory.

[0171] A server computing device (1105) may include one or more processors (1113) and memory (1114). Memory (1114) may store information accessible by the processor(s) (1113), including instructions (1121) that can be executed by the processor(s) (1113). Memory (1114) may also include data (1123) that can be retrieved, manipulated, or stored by the processor(s) (1113). Memory (1114) may be a type of non-transient computer-readable medium capable of storing information accessible by the processor(s) (1113), such as volatile and non-volatile memory. The processor(s) (1113) may include one or more CPUs (central processing units), GPUs (graphic processing units), FPGAs (field-programmable gate arrays), and / or ASICs (application-specific integrated circuits) such as TPUs (tensor processing units).

[0172] The instructions (1121) may include one or more instructions that, when executed by the processor(s) (1113), cause one or more processors to perform actions defined by the instructions. The instructions (1121) may be stored in an object code format for direct processing by the processor(s) (1113), or in other formats including a collection of independent source code modules or interpretable scripts that are interpreted or pre-compiled as required. The instructions (1121) may include instructions for implementing a system (400) consistent with aspects of the present disclosure. The system (400) may be executed using the processor(s) (1113) and / or other processors located remotely from the server computing device (1105).

[0173] Data (1123) may be retrieved, stored, or modified by processor(s) (1113) according to instructions (1121). Data (1123) may be stored in computer registers, in relational or non-relational databases, as tables having multiple different fields and records, or as JSON, YAML, proto, or HTML documents. Data (1123) may also be formatted into computer-readable formats such as (but not limited to) binary values, ASCII, or Unicode. Furthermore, data (1123) may contain information sufficient to identify relevant information, such as numbers, descriptive text, proprietary codes, pointers, references to data stored in other memories including other network locations, or information used by a function to compute relevant data.

[0174] The user computing device (1112) may also be configured similarly to the server computing device (1105) to have one or more processors (1116), memory (1117), instructions (1118), and data (1119). The user computing device (1112) may also include user output (1126) and user input (1124). The user input unit (1124) may include any suitable mechanism or technique for receiving input from a user, such as a keyboard, a mouse, mechanical actuators, soft actuators, touchscreens, microphones, and sensors.

[0175] The server computing device (1105) may be configured to transmit data to the user computing device (1112), and the user computing device (1112) may be configured to display at least a portion of the received data on a display implemented as part of the user output (1126). The user output (1126) may also be used to display an interface between the user computing device (1112) and the server computing device (1105). Alternatively or additionally, the user output (1126) may include one or more speakers, transducers or other audio outputs, a haptic interface or other tactile feedback that provide non-visual and non-audible information to the platform user of the user computing device (1112).

[0176] FIG. 11 illustrates processors (1113, 1116) and memories (1114, 1117) located within computing devices (1105, 1112), but the components described herein, including processors (1113, 1116) and memories (1114, 1117), may include a plurality of processors and memories that can operate at different physical locations rather than within the same computing device. For example, some of the instructions (1121, 1118) and data (1123, 1119) may be stored on a removable SD card, while others may be stored within a read-only computer chip. Some or all of the instructions and data may be stored at a location that is physically remote from the processors (1113, 1116) but is still accessible by the processors (1113, 1116). Similarly, processors (1113, 1116) may include a collection of processors capable of performing concurrent and / or sequential operations. Each of the computing devices (1105, 1112) may include one or more internal clocks that provide timing information that can be used for timing operations and programs executed by the computing devices (1105, 1112).

[0177] The server computing device (1105) may be configured to receive requests for processing data from the user computing device (1112). For example, the environment (1100) may be part of a computing platform configured to provide various services to users through various user interfaces and / or APIs that expose platform services. One or more services may be a set of machine learning frameworks or tools for creating neural networks or other machine learning models according to specific tasks and training data. The user computing device (1112) may receive and transmit data specifying actions to be performed by the computation unit (1000).

[0178] Devices (1112, 1105) may be capable of direct and indirect communication through the network (1160). Devices (1105, 1112) may set up listening sockets capable of accepting initiation connections for transmitting and receiving information. The network (1160) itself may include various configurations and protocols, including the Internet, the World Wide Web, intranets, virtual private networks, wide area networks, local networks, and private networks using communication protocols proprietary to one or more companies. The network (1160) may support various short-range and long-range connections. Short-range and long-range connections may be made through different bandwidths, such as 2.402 GHz to 2.480 GHz (generally associated with Bluetooth® standards), 2.4 GHz and 11 GHz (generally associated with Wi-Fi® communication protocols); or various communication standards such as LTE® standards for wireless broadband communication. In addition or alternatively, the network (1160) may also support wired connections between devices (1112, 1105), including through various types of Ethernet connections.

[0179] Although a single server computing device (1105), a user computing device (1112), and a data processing system (1001) are illustrated in FIG. 11, it is understood that aspects of the present disclosure may be implemented according to various different configurations and quantities of computing devices, including paradigms for sequential or parallel processing or through a distributed network of multiple devices. In some implementations, aspects of the present disclosure may be performed on a single device and any combination thereof. In some examples, one or more devices implement one or more data processing systems, and each data processing system includes one or more computation units according to aspects of the present disclosure. In some examples, a single device may implement multiple computation units, and each of the multiple computation units is configured to communicate with at least one other computation unit to perform distributed data processing tasks, for example, sequential or parallel processing.

[0180] Aspects of the present disclosure may be implemented as digital circuits, one or more computer programs on computer-readable storage media, or a combination of one or more of the foregoing. Computer-readable storage media may be non-transient, for example, as one or more instructions that are executable by a cloud computing platform and stored on a tangible storage device.

[0181] In this specification, the phrase “configured to” is used in different contexts relating to computer systems, hardware, or parts of computer programs, engines, or modules. When it is said that a system is configured to perform one or more operations, it means that the system has appropriate software, firmware, and / or hardware installed on the system that causes the system to perform one or more operations during operation. When it is said that some hardware is configured to perform one or more operations, it means that the hardware includes one or more circuits that receive input during operation and generate output corresponding to one or more operations according to the input. When it is said that a computer program, engine, or module is configured to perform one or more operations, it means that the computer program includes one or more program instructions that cause one or more computers to perform one or more operations when executed by one or more computers.

[0182] It is understood that although the operations illustrated in the drawings and mentioned in the claims are illustrated in a specific order, the operations may be performed in a different order than illustrated, some operations may be omitted, performed more than once, and / or performed in parallel with other operations. Additionally, the separation of different system components configured to perform different operations should not be understood as requiring the components to be separated. The described components, modules, programs, and engines may be integrated together as a single system or be part of multiple systems.

[0183] Unless otherwise noted, the foregoing alternative examples are not mutually exclusive but may be implemented in various combinations to achieve distinct advantages. Since these and other variations and combinations of the features discussed above may be utilized without exceeding the scope of the claims defined by the claims, the foregoing description of the examples should be taken as an example rather than as a limitation of the scope of the claims defined by the claims. Furthermore, provisions expressed as "for instance," "comprising," etc., as well as the provision of the examples described herein, should not be interpreted as limiting the scope of the claims to specific examples; rather, the examples are intended to illustrate only one of many possible implementations. Additionally, the same reference numbers in different drawings may identify the same or similar elements.

Claims

Claim 1 As a computation unit, a two-dimensional systolic array of processing elements ― said systolic array is configured to receive first input elements from a first input matrix along a first direction of said systolic array and to receive second input elements from a second input matrix along a second direction of said systolic array ―; a first checksum circuit configured to generate one or more groups of first checksums from the first input elements while said systolic array receives the first input elements; a second checksum circuit configured to generate one or more groups of second checksums while said systolic array receives the second input elements ― said systolic array is further configured to generate an output matrix from the first input matrix, the second input matrix, one or more groups of the first checksums, and one or more groups of the second checksums ―; A computation unit comprising an output checksum circuit configured to receive the output matrix and, from the output matrix, determine the occurrence of one or more errors in the generation of the output matrix. Claim 2 In claim 1, the output matrix comprises a data sub-matrix including values ​​generated by the systoric array using the first input elements and the second input elements, an output checksum row, and an output checksum column; and to determine the occurrence of one or more errors in the generation of the output matrix, the output checksum circuit is configured to generate a row checksum from at least one row of the data sub-matrix; compare the row checksum with the checksum of the output checksum column; and determine the occurrence of an error in the generation of the output matrix from the comparison of the row checksum and the checksum of the output checksum column. Claim 3 In claim 2, to determine the occurrence of one or more errors in the generation of the output matrix, the output checksum circuit is further configured to generate a column checksum from at least one column of the data sub-matrix; compare the column checksum with the checksum of the output checksum row; and determine the occurrence of an error in the generation of the output matrix from the comparison of the column checksum and the checksum of the output checksum row. Claim 4 In claim 3, to determine the occurrence of one or more errors in the generation of the output matrix, the output checksum circuit is further configured to generate a row checksum from at least one row of the data sub-matrix; compare the row checksum with the checksum of the output checksum column; and determine the occurrence of an error in the generation of the output matrix from the comparison of the row checksum and the checksum of the output checksum column. Claim 5 In claim 2, for comparing the row checksum with the checksum of the output checksum column, the output checksum circuit is further configured to determine whether the absolute difference between the row checksum and the checksum of the output checksum column is within a predetermined threshold, a computational unit. Claim 6 A computational unit according to any one of claims 1 to 5, wherein the computational unit further comprises one or more checksum processing elements configured to receive checksums from one or both of the first checksum circuit and the second checksum circuit. Claim 7 A computational unit according to any one of claims 1 to 5, wherein one or both of the first checksum circuit and the second checksum circuit are configured to transmit the first checksums or the second checksums to the systoric array for processing based on a control signal. Claim 8 In claim 7, the timing of the control signal for one or both of the first checksum circuit and the second checksum circuit is based on the number of time steps for loading the first input values ​​or the second input values ​​across the systoric array, a computational unit. Claim 9 A computation unit according to any one of claims 1 to 5, wherein the computation unit is further configured to transmit an indication of the occurrence of one or more errors to one or more devices connected to the computation unit in response to a determination of the occurrence of one or more errors in the generation of the output matrix. Claim 10 In claim 9, the computation unit is further configured such that the systoric array receives, after an indication of the occurrence of the one or more errors is transmitted, a regulated voltage higher than the threshold voltage for the computation unit from the one or more devices. Claim 11 In claim 10, the computational unit is further configured such that the systoric array receives a first voltage lower than the threshold voltage for the computational unit until it receives the adjusted voltage in response to transmitting the indication. Claim 12 In claim 11, a computation unit wherein one or both of the first checksum circuit and the second checksum circuit are configured to receive a second voltage higher than the threshold voltage. Claim 13 A computation unit according to claim 10, wherein one or both of the first checksum circuit and the second checksum circuit comprises 2-input 2-stage pipeline adder circuits configured to delay the generation of one or both of the first checksums and the second checksums. Claim 14 A computational unit according to claim 10, wherein one or both of the first checksum circuit and the second checksum circuit comprises one or more 2-cycle adder circuits and a plurality of registers, wherein the plurality of registers comprises one or more first registers configured to receive and transmit data according to a first clock frequency, one or more second registers configured to receive and transmit data according to a second clock frequency, and one or more third registers configured to receive and transmit data according to a third clock frequency, wherein the first clock frequency, the second clock frequency, and the third clock frequency are all different frequencies. Claim 15 A computational unit according to any one of claims 1 to 5, wherein the sistoric array is an output static sistoric array, and both the first checksum circuit and the second checksum circuit are connected to a plurality of checksum processing elements, and the checksum processing elements are configured to receive checksums generated from one or both of the first checksum circuit and the second checksum circuit. Claim 16 In claim 15, the plurality of checksum processing elements are computation units arranged along the periphery of the systoric array. Claim 17 A computation unit according to any one of claims 1 to 5, wherein the computation unit is configured to generate checksums only from the first checksum circuit or to generate checksums only from the second checksum circuit. Claim 18 As a data processing system, the data processing system comprises one or more processors, one or more memory devices and a computation unit, wherein the computation unit comprises a two-dimensional sistoric array of processing elements—the sistoric array is configured to receive first input elements from a first input matrix along a first direction of the sistoric array and to receive second input elements from a second input matrix along a second direction of the sistoric array—; a first checksum circuit configured to generate one or more groups of first checksums from the first input elements while the sistoric array receives the first input elements; a second checksum circuit configured to generate one or more groups of second checksums while the sistoric array receives the second input elements—the sistoric array is further configured to generate an output matrix from the first input matrix, the second input matrix, one or more groups of first checksums, and one or more groups of second checksums—; A data processing system comprising an output checksum circuit configured to receive the output matrix and, from the output matrix, determine the occurrence of one or more errors in the generation of the output matrix. Claim 19 In claim 18, the data processing system is configured to apply a voltage to the systoric array—the applied voltage is less than the threshold voltage of the systoric array—and to receive indication of one or more errors from the computation unit and, in response, increase the applied voltage to a level higher than the threshold voltage of the systoric array. Claim 20 In claim 19, the data processing system is further configured to apply a voltage below a threshold voltage of the systolic array to one or more of the first checksum circuit, the second checksum circuit, and the output checksum circuit; and, in response to receiving the indication, to continue applying a voltage below the threshold voltage to one or more of the first checksum circuit, the second checksum circuit, and the output checksum circuit. Claim 21 A data processing system according to any one of claims 18 to 20, wherein the data processing system is configured to transmit a control signal to one or both of the first checksum circuit and the second checksum circuit, and the timing of the transmission is based on the number of time steps for loading the first input values ​​or the second input values ​​across the systoric array. Claim 22 One or more non-transient computer-readable storage media are encoded with computer instructions that cause the computation unit to perform operations when executed by the computation unit, which includes a two-dimensional sistoric array of processing elements, a first checksum circuit, a second checksum circuit, and an output checksum circuit, said operations, said operations include: receiving first input elements from a first input matrix along a first direction of the sistoric array by the sistoric array; receiving second input elements from a second input matrix along a second direction of the sistoric array by the sistoric array; generating one or more groups of first checksums from the first input elements by the first checksum circuit while the sistoric array receives the first input elements; and generating one or more groups of second checksums by the second checksum circuit while the sistoric array receives the second input elements. One or more non-transient computer-readable storage media comprising: generating an output matrix from the first input matrix, the second input matrix, one or more groups of the first checksums, and one or more groups of the second checksums by the above systorilic array; receiving the output matrix by the output checksum circuit; and determining the occurrence of one or more errors in the generation of the output matrix by the output checksum circuit and from the output matrix.

Citation Information

Patent Citations

  • Method and equipment for analysing information included in data structure

    JP1995095197A

  • Systolic array and processing system

    JP2020144843A

  • System and method for reconfigurable systolic array with partial read / write

    KR1020210084220A

  • Systolic array and processing system

    KR1020200107295A