Matrix multiplier and operation method of matrix multiplication device including the same
By using BCQ technology to quantize the weight matrix in the matrix multiplication device, quantized code values and quantized scale factors are generated, and the problems of high calculation speed and calculation amount in the existing technology are solved, and efficient calculation of matrix multiplication is realized.
Patent Information
- Application Number
- JP2024155446
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-25
- Filing Date
- 2024-09-10
- Publication Date
- 2025-05-12
AI Technical Summary
The prior art has a large calculation speed and calculation amount when performing matrix multiplication, which is difficult to meet the fast computing needs of artificial intelligence models.
The matrix multiplication device including an input vector scaler, a fixed point data type converter and processing elements is used to quantize the weight matrix through binary coded quantization (BCQ), generate quantized code values and quantized scale factors, and then perform matrix multiplication operations.
Quantitative processing reduces the calculation amount and time of matrix multiplication, improves the efficiency of matrix multiplication, and reduces the computational cost of artificial intelligence models.
Smart Images

Figure 2025073072000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a semiconductor device. More specifically, the present disclosure relates to a matrix multiplier that performs matrix multiplication and / or a matrix multiplication device including the same. [Background technology]
[0002] Recently, as artificial intelligence technology has developed, the amount of calculation required for an artificial intelligence model has increased dramatically. As a result, various techniques for shortening the operation time of an artificial intelligence model have been researched.
[0003] In general, most of the operation time of an AI model is used for matrix multiplication. For example, an AI model uses most of its operation time for multiplying an input matrix and a weight matrix to calculate an output matrix. For this reason, various algorithms such as BCQ (Binary Coding Quantization) are being researched to multiply an input matrix and a weight matrix with less computational effort. Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure is intended to solve the above-mentioned technical problems. More specifically, an object of the present disclosure is to provide a matrix multiplier and a matrix multiplication device including the same, configured to perform matrix multiplication at a higher speed and with a smaller amount of calculations. [Means for solving the problem]
[0005] A matrix multiplier according to an embodiment of the present disclosure may include an input vector scaler that generates a first scaled input vector based on a first input vector and a plurality of quantization scale factors; a first data type converter that generates a first fixed-point scaled input vector based on the first scaled input vector; a processing element array including a first processing element that generates first fixed-point output elements based on the first fixed-point scaled input vector and a first plurality of quantization code values and a second processing element that generates second fixed-point output elements based on the first fixed-point scaled input vector and a second plurality of quantization code values; and a second data type converter that converts data types of the first and second fixed-point output elements to generate first and second output elements, respectively, and outputs a first output vector including the first and second output elements.
[0006] A matrix multiplier according to an embodiment of the present disclosure may include an input vector scaler that generates a first plurality of scaled input elements based on a first input element and a first plurality of quantization scale factors and generates a second plurality of scaled input elements based on a second input element and a second plurality of quantization scale factors, a first data type converter that generates a first plurality of fixed-point scaled input elements based on the first plurality of scaled input elements and generates a second plurality of fixed-point scaled input elements based on the second plurality of scaled input elements, a first processing element that accumulates the first plurality of fixed-point scaled input elements and the second plurality of fixed-point scaled input elements based on a plurality of quantization code values to generate a first fixed-point output element, and a second data type converter that converts a data type of the first fixed-point output elements to generate a first output element.
[0007] A method of operating the matrix multiplication device according to an embodiment of the present disclosure may include the steps of receiving first through N-th weights from an external device, performing binary coding quantization (BCQ) on the first through N-th weights to generate first through (N×R) quantization code values and first through (N×R) quantization scale coefficients, receiving first through N-th input elements from the external device, scaling the first through N-th input elements based on the first through (N×R) quantization scale coefficients to generate first through (N×R) scaled input elements, and outputting a first output element generated by accumulating the first through (N×R) scaled input elements based on the first through (N×R) quantization code values.
[0008] According to an embodiment of the present disclosure, a matrix multiplication device that receives a weight matrix and a first input vector from an external device includes: a BCQ circuit that performs binary coding quantization on the weight matrix to generate a plurality of quantization code values and a plurality of quantization scale coefficients; and a matrix multiplier that calculates a first output vector corresponding to a product of the first input vector and the weight matrix based on the plurality of quantization code values and the plurality of quantization scale coefficients, the matrix multiplier including an input vector scaler that scales the first input vector based on the plurality of quantization scaling values to generate a first scaled input vector; a first data type converter that generates a first fixed-point scaled input vector based on the first scaled input vector; a processing element array that calculates a first fixed-point output vector based on the plurality of quantization code vectors and the first fixed-point scaled input vector; and a second data type converter that converts a data type of the first fixed-point output vector to generate the first output vector.
[0009] According to an embodiment of the present disclosure, a matrix multiplication device that receives an n-dimensional input vector and a weight matrix having 'n by m' dimensions and outputs an m-dimensional output vector includes a BCQ (binary coding quantization) circuit that generates 1st to Rth quantization code matrices having 'n by m' dimensions and 1st to (n×R)th quantization scale coefficients corresponding to different rows of the 1st to Rth quantization code matrices, respectively, based on the weight matrix, an input vector scaler that scales elements of the n-dimensional input vector based on the 1st to (n×R)th quantization scale coefficients, and a processing element array including a plurality of processing elements, each of which is configured to accumulate elements of the input vector scaled based on the 1st to Rth quantization code matrices to output different output elements included in the n-dimensional output vector. Effect of the Invention
[0010] Therefore, according to the embodiment of the present disclosure, the amount of calculation of the matrix multiplier can be reduced. [Brief description of the drawings]
[0011] Various exemplary embodiments are described below with reference to one or more drawings. [Figure 1] 1 is a block diagram illustrating a matrix multiplication device according to an embodiment of the present disclosure. [Diagram 2] 2 illustrates the operation of a matrix multiplication device implemented to directly multiply the input matrix and weight matrix of FIG. 1; [Diagram 3] FIG. 2 is a diagram illustrating the operation of the BCQ circuit of FIG. [Figure 4] 2 is a diagram illustrating the operation of the BCQ circuit of FIG. 1, which performs binary coding quantization operation for each column of a weight matrix. [Diagram 5]5 illustrates the operation of a matrix multiplier according to the embodiment of FIG. 4. [Figure 6] 5 is a block diagram showing a configuration of the matrix multiplier of FIG. 1 according to the embodiment of FIG. 4; [Figure 7] FIG. 7 is a diagram illustrating a configuration of the first data type converter in FIG. 6. [Figure 8] FIG. 8 is a diagram illustrating the operation of the exponent extraction circuit of FIG. [Figure 9] 8 is a diagram illustrating the operation of the data type conversion circuit of FIG. 7. [Figure 10] FIG. 7 is a block diagram showing in more detail the configuration of the processing element array of FIG. 6. [Figure 11] FIG. 7 is a block diagram showing in more detail a portion of the operation of the matrix multiplier of FIG. 6. [Figure 12] FIG. 12 is a diagram illustrating the operation of the processing element of FIG. [Figure 13] FIG. 13 illustrates a configuration of one of the processing elements of FIG. 12 implemented in accordance with one embodiment. [Figure 14] FIG. 7 illustrates an operation of the second data type converter in FIG. 6. [Figure 15] 2 illustrates the operation of the BCQ circuit of FIG. 1 according to an embodiment of the present disclosure. [Figure 16] FIG. 16 illustrates a weight matrix approximated by the embodiment of FIG. 15. [Figure 17] FIG. 17 is a diagram showing the quantization code matrix of FIG. 16. [Figure 18] FIG. 2 is a block diagram showing a configuration of the matrix multiplier of FIG. 1 implemented according to an embodiment of the present disclosure. [Figure 19] FIG. 19 is a block diagram showing the configuration of the input vector scaler of FIG. 18. [Figure 20] FIG. 20 illustrates the operation of the input vector scaling circuit of FIG. 19 in more detail. [Figure 21] FIG. 19 is a block diagram showing a configuration of the first data type converter in FIG. 18. [Figure 22]FIG. 19 is a block diagram showing in more detail the configuration of the processing element array of FIG. 18. [Diagram 23] FIG. 23 is a block diagram showing the output of the processing element of FIG. 22 in more detail. [Figure 24] FIG. 23 is a block diagram showing in more detail the operation of the first processing element row of FIG. 22. [Diagram 25] FIG. 25 illustrates a configuration of the processing element of FIG. 24 according to one embodiment. [Figure 26] FIG. 19 illustrates the operation of the second data type converter in FIG. 18. [Figure 27] 2 is a flowchart showing the operation of the matrix multiplication device of FIG. 1; [Figure 28] 28 is a flowchart showing step S150 of FIG. 27 in more detail. [Figure 29] 2 is a flowchart showing the operation of the matrix multiplication device of FIG. 1; [Diagram 30] 30 is a flowchart showing step S250 of FIG. 29 in more detail. [Diagram 31] FIG. 23 is a block diagram showing the processing element array of FIG. 22 implemented in a systolic array fashion. [Diagram 32] FIG. 32 is a diagram showing the configuration of the processing element in FIG. 31 in more detail. [Diagram 33] 20 illustrates the operation of the input vector scaling circuit of FIG. 19 in accordance with one embodiment. [Diagram 34] 2 illustrates the operation of the matrix multiplication device of FIG. 1 according to one embodiment. [Diagram 35] FIG. 35 shows the full input matrix of FIG. 34. [Diagram 36] FIG. 35 shows the full-weight matrix of FIG. 34. [Figure 37] FIG. 35 shows the full output matrix of FIG. 34. [Figure 38] 1 is a block diagram illustrating a matrix multiplication device according to one embodiment. [Figure 39]FIG. 1 is a block diagram illustrating a neural processing system implemented in accordance with one embodiment. [Diagram 40] FIG. 40 is a block diagram showing an artificial intelligence model driven by the neural processing system of FIG. 39. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] Hereinafter, the embodiments of the present disclosure will be described clearly and in detail to such an extent that a person having ordinary skill in the art of the present disclosure can easily carry out the present disclosure. Details such as detailed configurations and structures are provided simply to aid in the overall understanding of the embodiments of the present disclosure. Therefore, those skilled in the art can make modifications to the embodiments described herein without departing from the technical spirit and scope of the present disclosure. Furthermore, descriptions of well-known functions and structures are omitted for clarity and conciseness. The configurations in the following drawings or detailed description may be associated with other components in addition to those shown in the drawings or described in the detailed description. The terms used in the present disclosure are defined in consideration of the functions of the present disclosure and are not limited to specific functions. The definitions of the terms are determined based on the matters described in the detailed description.
[0013] The components described with reference to terms such as driver or block in the detailed description can be realized in the form of software, hardware, or a combination thereof. Exemplarily, software may be machine code, firmware, embedded code, and application software. For example, hardware may include electric circuits, electronic circuits, processors, computers, integrated circuit cores, pressure sensors, inertial sensors, MEMS (Micro Electro Mechanical System), manual elements, or a combination thereof.
[0014] For the sake of simplicity, matrices are referred to below through square brackets “[”, “]” and sets are referred to through curly brackets “{”, “}”. However, the scope of the present disclosure is not limited to such notations.
[0015] 1 is a block diagram showing a matrix multiplication device according to an embodiment of the present disclosure. Referring to FIG. 1, the matrix multiplication device MMD can include a matrix multiplier 100 and a binary coding quantization circuit 200.
[0016] The matrix multiplication device MMD can receive an input matrix XM. The input matrix XM can include a plurality of input vectors. Each of the plurality of input vectors can include a plurality of input elements. For example, the input matrix XM can be expressed by the following Equation 1:
[0017]
number
[0018]
number
[0019]
number
[0020]
number
[0021] For a simpler description, an embodiment in which the dimension of each input vector included in the input matrix XM is 'n' will be representatively described below. That is, an embodiment in which each input vector includes 'n' input elements will be representatively described below. In other words, an embodiment in which the input matrix XM includes 'n' columns will be representatively described below.
[0022] In one embodiment, each of the input elements included in the input matrix XM may have a 16-bits floating point (FP16) and / or 32-bits floating point (FP32) data type, although the scope of the present disclosure is not limited in this respect.
[0023] The matrix multiplication device MMD may receive a weight matrix WM. The weight matrix WM may include a plurality of weights. For example, the weight matrix WM may be expressed by the following Equation 2.
[0024]
number
[0025] In one embodiment, each of the weights included in the weight matrix WM may have a 16-bit floating point (FP16) or 32-bit floating point (FP32) data type, although the scope of the present disclosure is not limited in this respect.
[0026] The BCQ circuit 200 may perform binary coding quantization (BCQ) on the weight matrix WM. For example, the BCQ circuit 200 may determine a plurality of quantization sign values QSV and a plurality of quantization scale coefficients QSC based on the weight matrix WM.
[0027] More specifically, the BCQ circuit 200 can convert each of the multiple weights into multiple 'quantization scale factor QSC-quantization code value QSV pairs.' That is, the BCQ circuit 200 can approximate each of the weights in the weight matrix WM with multiple 'quantization scale factor QSC and quantization code value QSV pairs.'
[0028] In one embodiment, each of the multiple quantization code values QSV can indicate '-1' or '+1'.
[0029] In one embodiment, each of the quantization scale factors QSC can have the same data type as the weight of the weight matrix WM. For example, each of the quantization scale factors QSC can have an FP16 or FP32 data type. However, the scope of the present disclosure is not limited thereto. The specific operation of the BCQ circuit 200 will be described in more detail with reference to FIG. 3 below.
[0030] The matrix multiplier 100 may receive a plurality of quantization code values QSV and a plurality of quantization scale factors QSC. The matrix multiplier 100 may perform matrix multiplication on the input matrix XM and the weight matrix WM based on the plurality of quantization code values QSV and the plurality of quantization scale factors QSC. For example, the matrix multiplier 100 may multiply the input matrix XM by the weight matrix WM that is approximated based on the plurality of quantization code values QSV and the plurality of quantization scale factors QSC to generate an output matrix YM. The input matrix XM is provided to the matrix multiplier 100 in the form of a quantization sign bit having a 1-bit code length. However, the scope of the present disclosure is not limited thereto.
[0031] The output matrix YM may include a plurality of output vectors, each of which may include a plurality of output elements. For example, the output matrix YM may be expressed by the following Equation 3:
[0032]
number
[0033]
number
[0034]
number
[0035] In one embodiment, 'n' and 'm' may be integers that are the same as each other. For example, the weight matrix WM may be implemented as a square matrix. In this case, the dimensions of the output vectors included in the output matrix YM may be the same as the dimensions of the input vectors. However, the scope of the present disclosure is not limited thereto.
[0036] In one embodiment, when the matrix multiplication device MMD directly multiplies the input matrix XM and the weight matrix WM to calculate the output matrix YM, the matrix multiplication device MMD may have to process a very large amount of calculations. In this case, the operation speed of the matrix multiplication device MMD may be reduced. The operation of the matrix multiplication device MMD that directly multiplies the input matrix XM and the weight matrix WM to calculate the output matrix YM will be described in more detail with reference to FIG. 2 below.
[0037] On the other hand, when the matrix multiplication device MMD multiplies the input vector matrix XM by a weight matrix WM approximated based on a plurality of quantization code values QSV and a plurality of quantization scale coefficients QSC to calculate the output matrix YM, the amount of calculation of the matrix multiplication device MMD is reduced (e.g., significantly reduced). The operation of the matrix multiplication device MMD that calculates the output matrix YM based on a plurality of quantization code values QSV and a plurality of quantization scale coefficients QSC will be described in more detail with reference to the following drawings.
[0038] Fig. 2 is a diagram showing the operation of a matrix multiplication device realized to directly multiply the input matrix and the weight matrix in Fig. 1. Referring to Fig. 1 and Fig. 2, the matrix multiplication device MMD can directly multiply the input matrix XM and the weight matrix WM to calculate the output matrix YM.
[0039] The matrix multiplication device MMD may need to perform 'n' floating point multiplications and then 'n-1' floating point summations to calculate one output element included in the output matrix YM. For example, the matrix multiplication device MMD may need to perform 'n' floating point multiplications and then 'n-1' floating point summations to calculate one output element included in the output matrix YM in a manner similar to the following Equation 4. 11 ) can be calculated.
[0040]
number
[0041]
number
[0042]
number
[0043] Fig. 3 is a diagram showing the operation of the BCQ circuit of Fig. 1. Referring to Fig. 1 and Fig. 3, the horizontal axis indicates the magnitude of the weights included in the weight matrix WM.
[0044] The BCQ circuit 200 can approximate one or more of the weights included in the weight matrix WM to a number of quantum levels QL. The BCQ circuit 200 can determine the number of the quantum levels QL based on a BCQ resolution determined in advance. For example, the BCQ circuit 200 can approximate each of the weights to 2 R (where R is the BCQ resolution) quantum levels QL can be approximated. In this case, 2 R Each of the quantization levels QL can be determined based on a combination of the R quantization scale factors QSC and the R quantization code values QSV.
[0045] However, for the sake of simplicity, an embodiment in which R is '3' will be representatively described below. For example, the BCQ circuit 200 can approximate each of the weights included in the weight matrix WM to the first quantum level QL1 to the eighth quantum level QL8. In this case, each of the first quantum level QL1 to the eighth quantum level QL8 can be determined based on the following Equation 5.
[0046]
number
[0047] For example, the first quantum level QL1 is b 1 ~b 3 In this case, the first quantum level QL1 corresponds to the case where all of the 1 -a 2 -a 3 ". Similarly, the eighth quantum level QL8 corresponds to b 1 ~b 3 In this case, the eighth quantum level QL8 corresponds to the case where all of the 1 +a 2 +a 3 In this way, the seventh quantum level QL7 can correspond to b 1 ~b 3 In this case, the seventh quantum level QL7 corresponds to the case where “+a 1 +a 2 -a 3 ". In some embodiments, the sign of such a sum is determined based on a binary encoding of the quantum level QL, although the scope of the disclosure is not limited in this respect.
[0048] The BCQ circuit 200 can approximate each weight included in the weight matrix WM to a quantum level having a closest value among the first quantum level QL1 to the eighth quantum level QL8. For example, the weight “w 11 When the magnitude of the weight “w” is closest to the seventh quantum level QL7 among the first quantum level QL1 to the eighth quantum level QL8, the BCQ circuit 200 11 " can be approximated to the seventh quantum level QL7. In this case, the BCQ circuit 200 uses the weight "w 11" to a plurality of quantization scale coefficients QSC (e.g., a 1 ~a 3 ) and a combination of multiple quantization code values QSV (e.g., '+1', '+1', '-1'). Similarly, the BCQ circuit 200 can approximate each of the weights included in the weight matrix WM based on multiple quantization code values QSV and multiple quantization scale factors QSC.
[0049] In one embodiment, the quantization scale factors QSC may have different magnitudes. For example, 1 The size of a 2 It may be larger than the size of a 2 The size of a 3 may be larger than the size of
[0050] For the sake of simpler explanation, an embodiment in which the BCQ resolution is '3' is representatively described in Fig. 3, but the scope of the present disclosure is not limited thereto. For example, the BCQ resolution can be determined to be an integer equal to or greater than '2' depending on the type of artificial intelligence model driven based on the matrix multiplication device MMD.
[0051] In one embodiment, when the artificial intelligence model driven based on the matrix multiplication device MMD is a large language model (LLM), the BCQ resolution may be '3'. However, the scope of the present disclosure is not limited thereto.
[0052] In one embodiment, when an artificial intelligence model driven based on the matrix multiplication device MMD is an image object identification model, the BCQ resolution may be '2', but the scope of the present disclosure is not limited thereto.
[0053] In one embodiment, the BCQ circuit 200 may perform a binary coding quantization operation for each column of the weight matrix WM. For example, the BCQ circuit 200 may determine a plurality of different quantization scale coefficients QSC for each column of the weight matrix WM. In this case, the quantization scale coefficients for the weights included in the first column of the weight matrix WM may be different from the quantization scale coefficients for the weights included in the second column of the weight matrix WM. An embodiment in which the BCQ circuit 200 performs a binary coding quantization operation for each column of the weight matrix WM will be described in more detail with reference to FIGS. 4 to 14 below. However, the scope of the present disclosure is not limited thereto.
[0054] In one embodiment, the BCQ circuit 200 may perform a binary coding quantization operation for each row of the weight matrix WM. For example, the BCQ circuit 200 may determine a plurality of quantization scale coefficients QSC that are different from each other for each row of the weight matrix WM. In this case, the quantization scale coefficients for the weights included in the first row of the weight matrix WM may be different from the quantization scale coefficients for the weights included in the second row of the weight matrix WM. An embodiment in which the BCQ circuit 200 performs a binary coding quantization operation for each row of the weight matrix WM will be described in more detail with reference to FIGS. 15 to 40 below. However, the scope of the present disclosure is not limited thereto.
[0055] In one embodiment, the BCQ circuit 200 can approximate each of the weights to a first quantum level QL1 to an eighth quantum level QL8 based on a uniform BCQ algorithm. That is, the BCQ circuit 200 can approximate the weights based on the first quantum level QL1 to the eighth quantum level QL8 having uniform intervals from each other. In this case, the magnitudes of the quantization scale coefficients QSC can be realized as a geometric sequence with a common ratio of '2'. For example, 1 The size of a2 may be twice as large as a 2 The size of a 3 However, the scope of the present disclosure is not limited in this respect.
[0056] 4 is a diagram illustrating the operation of the BCQ circuit of FIG. 1 performing a binary coding quantization operation for each column of a weight matrix WM. Referring to FIG. 1 and FIG. 3 to FIG. 4, the BCQ circuit 200 can perform a binary coding quantization operation for each column of a weight matrix WM.
[0057] First, the weight matrix WM can be expressed as Equation 6 below.
[0058]
number
[0059]
number
[0060]
number
[0061] The BCQ circuit 200 may perform a binary coding quantization operation for each column of the weight matrix WM based on the following Equation 7.
[0062]
number
[0063]
number
[0064]
number
[0065]
number
[0066]
number
[0067]
number
[0068] That is, the BCQ circuit 200 can approximate each weight of the weight matrix WM based on a plurality of quantization code values QSV (ie, "b" values) and a plurality of quantization scale factors QSC (ie, "a" values).
[0069] The BCQ circuit 200 can provide a plurality of quantization code values QSV (i.e., “b” values) and a plurality of quantization scale factors QSC (i.e., “a” values) to the matrix multiplier 100. In the following, the operation of the matrix multiplier 100, which performs a matrix multiplication operation based on the plurality of quantization code values QSV and the plurality of quantization scale factors QSC, will be described.
[0070] Fig. 5 is a diagram illustrating the operation of the matrix multiplier according to the embodiment of Fig. 4. Referring to Figs. 1, 3 to 5, the matrix multiplier 100 can calculate an output matrix YM based on a plurality of quantization code values QSV and a plurality of quantization scale coefficients QSC.
[0071] In the following, for the sake of simplicity, the first input vector (i.e.
[0072]
number
[0073]
number
[0074] The matrix multiplier 100 calculates y 11 can be calculated.
[0075]
number
[0076]
number
[0077]
number
[0078]
number
[0079]
number
[0080]
number
[0081]
number
[0082]
number
[0083] In one embodiment, when the data types of the input elements are converted to fixed point, the matrix multiplier 100 can calculate the partial sums of Equation 10 with a smaller amount of calculations. For example, when the data types of the input elements are fixed point corresponding to the same exponent value, the matrix multiplier 100 must perform R×(n−1) fixed point summations to calculate all the partial sums of Equation 10. The detailed operation of the matrix multiplier 100 for converting the data types of the input elements to fixed point will be described in more detail with reference to FIGS. 6 to 15 below.
[0084] The matrix multiplier 100 multiplies each of the partial sums in Equation 10 by the quantization scale coefficient QSC, and accumulates the multiplied values to calculate one output element. For example, the matrix multiplier 100 performs 'R' floating point multiplications and then performs 'R-1' floating point sums to calculate one output element (e.g., y 11) can be calculated. In this case, unlike the one previously described with reference to FIG. 2, the number of floating-point multiplications performed by the matrix multiplier 100 can be minimized (i.e., instead of performing 'm×n' floating-point multiplications and 'm×(n-1)' floating-point sums, the output element can be calculated by performing 'R' floating-point multiplications and 'R-1' floating-point sums). Therefore, according to the embodiment of FIGS. 4 and 5, the operating speed of the matrix multiplication device MMD can be improved.
[0085] Fig. 6 is a block diagram showing a configuration of the matrix multiplier of Fig. 1 according to the embodiment of Fig. 4. With reference to Figs. 1 and 3 to 6, the matrix multiplier 100 of Fig. 1 can be realized by a matrix multiplier 10.
[0086] The matrix multiplier 10 may include a first data type converter 11 (data type converter #1), a quantization sign value buffer 12, a processing element array 13, a second data type converter 14 (data type converter #2), a quantization scale coefficient buffer 15, a partial sum scaler 16, and an accumulator 17.
[0087] The first data type converter 11 may receive an input matrix XM. For example, the first data type converter 11 may receive a plurality of input vectors (e.g.,
[0088]
number
[0089] The first data type converter 11 can extract an exponent EXP from each of a plurality of input vectors. For example, the first data type converter 11 extracts an exponent EXP from a first input vector (
[0090]
number
[0091]
number
[0092] The first data type converter 11 can convert the data type of each of the plurality of input vectors to a fixed point based on the extracted exponent. For example, the first data type converter 11 can convert the data type of each of the plurality of input elements to a fixed point. That is, the first data type converter 11 can receive the input matrix XM and output a fixed point input matrix XM_fxp.
[0093] The fixed-point input matrix XM_fxp may include a plurality of fixed-point input vectors. Each of the plurality of fixed-point input vectors may include a plurality of fixed-point input elements. For example, the fixed-point input matrix XM_fxp may be expressed as follows:
[0094]
number
[0095]
number
[0096] The quantized code value buffer 12 can store a plurality of quantized code values QSV provided from the BCQ circuit 200. The quantized code value buffer 12 can provide a plurality of quantized code values QSV to the processing element array 13.
[0097] The processing element array 13 may receive a plurality of quantization code values QSV and a fixed-point input matrix XM_fxp. The processing element array 13 may calculate a plurality of fixed-point partial sums PSM_fxp based on the plurality of fixed-point input elements and the plurality of quantization code values QSV included in the fixed-point input matrix XM_fxp. For example, the processing element array 13 may receive a first input vector (i.e.,
[0098]
number
[0099]
number
[0100] The processing element array 13 may include a plurality of processing elements arranged in row and column directions. Each of the processing elements may calculate a different fixed-point partial sum PSM_fxp, as previously described with reference to Equation 12. The configuration and operation of each of the processing elements will be described in more detail with reference to FIGS. 10 to 13 below.
[0101] The second data type converter 14 can receive a plurality of exponents EXP from the first data type converter 11. The second data type converter 14 can receive a plurality of fixed-point partial sums PSM_fxp from the processing element array 13. The second data type converter 14 can convert the data type of the plurality of fixed-point partial sums PSM_fxp to floating point based on the plurality of exponents EXP. That is, the second data type converter 14 can output a plurality of partial sums PSM having a floating point format. For example, the second data type converter 14 can convert the data type of the plurality of fixed-point partial sums PSM_fxp to floating point based on the plurality of exponents EXP. 1_c1_X1 ~PSM' R_c1_X1 PSM 1_c1_X1 ~PSM 2_c1_X1 A more detailed configuration and operation of the second data type converter 14 will be described in more detail with reference to FIG.
[0102] The quantization scale factor buffer 15 can store a plurality of quantization scale factors QSC provided by the BCQ circuit 200. The quantization scale factor buffer 15 can provide a plurality of quantization scale factors QSC to the partial sum scaler 16.
[0103] The partial sum scaler 16 may receive a plurality of quantization scale factors QSC and a plurality of partial sum PSMs. The partial sum scaler 16 may scale the plurality of partial sum PSMs based on the plurality of quantization scale factors QSC. For example, the partial sum scaler 16 may multiply each of the plurality of partial sum PSMs by a corresponding quantization scale factor QSC to generate a plurality of scaled partial sum SCPSMs. A more detailed configuration and operation of the partial sum scaler 16 will be described in more detail with reference to FIG. 11 below.
[0104] In one embodiment, the partial sum scaler 16 may temporarily store the multiple scaled partial sums SCPSM in a volatile memory device (eg, a static random access memory (SRAM) device) external to the matrix multiplication device MMD.
[0105] The accumulator 17 can receive multiple scaled partial sums SCPSM. The accumulator 17 can calculate multiple output elements based on the multiple scaled partial sums SCPSM. For example, the accumulator 17 can add the scaled partial sums SCPSM to calculate one output element. That is, the accumulator 17 can add R scaled partial sums SCPSM to calculate one output element. In this manner, the accumulator 17 can calculate multiple output elements and output the output matrix YM. A more detailed operation of the accumulator 17 will be described in more detail with reference to FIG. 11 below.
[0106] In one embodiment, the accumulator 17 can read a number of scaled partial sums SCPSM from a volatile memory device external to the matrix multiplication device MMD.
[0107] In one embodiment, the operating speed of the processing element array 13 may be faster than the speed at which the matrix multiplication device MMD accesses an external volatile memory device. In this case, a bottleneck phenomenon may occur in the operating speed of the matrix multiplication device MMD depending on the speed at which the scaled partial sums SCPSM are stored in and read from the volatile memory device external to the matrix multiplication device MMD. The configuration and operation of the matrix multiplication device MMD in which access to the external volatile memory device is minimized will be described with reference to the following Figures 15 to 40.
[0108] Fig. 7 is a diagram showing the configuration of the first data type converter in Fig. 6. Referring to Figs. 1 and 3 to 7, the first data type converter 11 can include a first exponent extraction circuit 11a_1 through an (h)th exponent extraction circuit 11a_h (exponent extract circuit) and a first data type conversion circuit 11b_1 through an (h)th data type conversion circuit 11b_h.
[0109] The first exponent extraction circuit 11a_1 to the (h)th exponent extraction circuit 11a_h can receive input vectors different from each other. For example, the first exponent extraction circuit 11a_1 to the (h)th exponent extraction circuit 11a_h can receive the first input vector to the (h)th input vector (i.e.,
[0110]
number
[0111] Each of the first exponent extraction circuit 11a_1 to the (h)th exponent extraction circuit 11a_h can extract exponents from a plurality of input elements included in a received input vector. For example, the first exponent extraction circuit 11a_1 extracts an exponent from an input element (i.e., x 11 ~x 1n ), and the second exponent extraction circuit 11a_2 can extract the first exponent EXP1 from the input element (i.e., x 21 ~x 2n) to extract the second exponent EXP2. In this manner, the first exponent extraction circuit 11a_1 through the (h)th exponent extraction circuit 11a_h can extract the first exponent EXP1 through the (h)th exponent through EXPh, respectively.
[0112] The first exponent extraction circuit 11a_1 through the (h)th exponent extraction circuit 11a_h can provide the extracted exponents to the first data type conversion circuit 11b_1 through the (h)th data type conversion circuit 11b_h, respectively. Also, the first exponent extraction circuit 11a_1 through the (h)th exponent extraction circuit 11a_h can provide the extracted exponents to the second data type converter 14. More detailed operations of the first exponent extraction circuit 11a_1 through the (h)th exponent extraction circuit 11a_h will be described in more detail with reference to FIG. 8 below.
[0113] The first data type conversion circuit 11b_1 to the (h)th data type conversion circuit 11b_h can receive the first exponent EXP1 to the (h)th exponent EXPh, respectively.
[0114]
number
[0115] The first data type conversion circuit 11b_1 to the (h)th data type conversion circuit 11b_h can convert the data type of the received input vector into a fixed point based on the received exponent. The first data type conversion circuit 11b_1 to the (h)th data type conversion circuit 11b_h can convert the data type of the received input vector into a fixed point based on the received exponent.
[0116]
number
[0117] Fig. 8 is a diagram showing the operation of the exponent extraction circuit of Fig. 7. In the following, for the sake of simpler explanation, the operation of the first exponent extraction circuit 11a_1 will be representatively described. However, the scope of the present disclosure is not limited thereto.
[0118] 1, 3 to 8, the first exponent extraction circuit 11a_1 can receive a plurality of input elements. For example, the first exponent extraction circuit 11a_1 receives x 11 ~x 1n can be received.
[0119] The data type of each of the input elements may be floating point. For example, x 11 ~x 1n Each may include a sign part SP, an exponent part EXPP, and a mantissa part MTSP.
[0120] The first exponent extraction circuit 11a_1 can determine the largest value among the values of the exponent parts EXPP of the received input elements. In this case, the first exponent extraction circuit 11a_1 can determine the exponent of the determined input element as the first exponent EXP1. That is, the first exponent extraction circuit 11a_1 can extract x 11 ~x 1n The largest exponent among the exponents can be extracted.
[0121] For the sake of simpler explanation, an embodiment in which the first exponent extraction circuit 11a_1 extracts the largest value among the values of the exponent parts of a plurality of input elements is representatively described in Fig. 8, but the scope of the present disclosure is not limited thereto. For example, the first exponent extraction circuit 11a_1 can be realized to extract the smallest value among the values of the exponent parts of a plurality of input elements.
[0122] Fig. 9 is a diagram showing the operation of the data type conversion circuit of Fig. 7. In the following, for the sake of simpler explanation, the operation of the first data type conversion circuit 11b_1 is representatively described. However, the scope of the present disclosure is not limited thereto.
[0123] 1, 3 to 9, the first data type converter circuit 11b_1 can receive a plurality of input elements. For example, the first data type converter circuit 11b_1 can receive x 11 ~x 1n can be received.
[0124] The first data type conversion circuit 11b_1 can receive the first exponent EXP1. The first data type conversion circuit 11b_1 can convert each of the received input elements into a fixed-point data type based on the first exponent EXP1. That is, the first data type conversion circuit 11b_1 can convert x 11 ~x 1n x' 11 ~x' 1n However, in the following, for the sake of simplicity, we use the input element “x 11 " to the fixed-point input element "x' 11 The operation of the first data type converter circuit 11b_1 that converts the data into
[0125] The first data type conversion circuit 11b_1 converts the first exponent EXP1 and x 11 The difference between the exponents of the input element “x 11The mantissa of x (hereinafter referred to as the first mantissa MTSPa) can be shifted in the direction of the least significant bit (LSB). For example, the first exponent EXP1 and the mantissa of x 11 If the difference between the values of the exponent parts EXPP of the first and second mantissa parts is '4', the first data type conversion circuit 11b_1 can insert '4' pieces of '0' bits into the most significant bit place (MSB) of the first mantissa part MTSPa.
[0126] The first data type conversion circuit 11b_1 converts the fixed-point input element "x' 11 For example, the first data type conversion circuit 11b_1 may cut off the lower bits of the shifted first mantissa part MTSPa according to the code length of the second mantissa part MTSPb. Alternatively, the first data type conversion circuit 11b_1 may determine the lower bits of the shifted first mantissa part MTSPa based on various types of rounding algorithms such as “nearest even rounding” according to the code length of the second mantissa part MTSPb. However, the scope of the present disclosure is not limited thereto.
[0127] In one embodiment, the code length of the first mantissa part MTSPa may be '10-bit' or '23-bit', but the scope of the present disclosure is not limited thereto.
[0128] In one embodiment, the code length of the second mantissa part MTSPb may be '7-bit', but the scope of the present disclosure is not limited thereto.
[0129] In one embodiment, the data type of each of the multiple fixed-point input elements may be INT8 (8-bit integer), although the scope of the present disclosure is not limited in this respect.
[0130] FIG. 10 is a block diagram showing the configuration of the processing element array of FIG. 6 in more detail. Referring to FIGS. 1 and 3 to 10, the processing element array 13 can include a plurality of processing elements PE arranged in row and column directions. For a simpler explanation, it is assumed below that the plurality of processing elements PE are arranged in (p) rows and (q) columns. Also, a processing element arranged in the (i)th row and (j)th column of the processing element array 13 is referred to as "PEij". For example, a processing element arranged in the first row and second column of the processing element array 13 is referred to as "PE12".
[0131] The processing element array 13 can include a first processing element row PER1 to a (p)th processing element row PERp. Each of the first processing element row PER1 to the (p)th processing element row PERp can include a plurality of processing elements PE that are different from each other. For example, the first processing element row PER1 can include processing elements PE11 to PE1q.
[0132] In one embodiment, 'p' may be an integer of the same or smaller magnitude than 'h', although the scope of the present disclosure is not limited in this respect.
[0133] The first processing element row PER1 to the (p)th processing element row PERp may receive different fixed-point input vectors from each other. For example, the first processing element row PER1 to the (p)th processing element row PERp may receive the first to (p)th fixed-point input vectors (i.e.,
[0134]
number
[0135] Each of the multiple processing elements PE can receive multiple quantized code values QSV from the quantized code value buffer 12.
[0136] The processing elements arranged in the same column of the processing element array 13 may receive the same quantization code value. In a more detailed example, the quantization code value received by the processing element PE11 may be the same as the quantization code value received by the processing element PE21.
[0137] Processing elements arranged in different columns of the processing element array 13 can receive different quantization code values. For example, the quantization code value received by the processing element PE11 may be different from the quantization code value received by the processing element PE12.
[0138] The quantized code values provided to different columns of the processing element array 13 are described in more detail with reference to FIG. 12 below.
[0139] Each of the multiple processing elements PE can calculate a different fixed-point partial sum PSUM_fxp based on the received fixed-point input vector and the quantization code value QSV. The multiple processing elements PE can provide the calculated fixed-point partial sum PSUM_fxp to the second data type converter 14. A specific manner in which the multiple processing elements PE calculate different fixed-point partial sums PSUM_fxp will be described in more detail with reference to Figures 11 and 13 below.
[0140] FIG. 11 is a block diagram showing in more detail a portion of the operation of the matrix multiplier of FIG. 6. In the following, the first fixed-point input vector (
[0141]
number
[0142] 1, 3 to 11, the first processing element row PER1 receives a first fixed-point input vector (i.e.,
[0143]
number
[0144] The processing elements PE11 to PE1q can receive quantization code vectors different from each other. For example, the processing elements PE11 to PE1R receive quantization code vectors (i.e.,
[0145]
number
[0146]
number
[0147] Each of the processing elements included in the first processing element row PER1 can calculate a different fixed-point weighted sum PSM_fxp. For example, each of the processing elements PE11 to PE1R calculates PSM' 1_c1_X1 ~PSM' R_c1_X1 Similarly, the processing elements PE1(R+1) to PE1(2R) can calculate PSM' 1_c2_X1 ~PSM' R_c2_X1 can be calculated.
[0148] The second data type converter 14 may receive a first exponent EXP1. The second data type converter 14 may receive a plurality of fixed-point weighted sums PSM_fxp from a first processing element row PER1. The second data type converter 14 may convert the received fixed-point weighted sums PSM_fxp to a floating-point data type based on the first exponent EXP1. For example, the second data type converter 14 may convert PSM' 1_c1_X1 ~PSM' R_c1_X1 PSM 1_c1_X1 ~PSM R_c1_X1 Similarly, the second data type converter 14 can convert PSM' 1_c2_X1 ~PSM' R_c2_X1 PSM 1_c2_X1 ~PSM R_c2_X1 can be converted to
[0149] The partial sum scaler 16 may include a plurality of multiplier circuits MUL. The plurality of multiplier circuits MUL may receive different weighted sums. For easier explanation, the multiplier circuit MUL that receives the weighted sum PSM generated based on the processing elements arranged in the (j)th column of the processing element array 13 will be referred to as "MUL_j". For example, the multiplier circuits MUL_1 to MUL_R may receive the PSM 1_c1_X1 ~PSM R_c1_X1 The multiplication circuits MUL_R+1 to MUL_2R can receive the PSM 1_c2_X1 ~PSM R_c2_X1 can be received respectively.
[0150] Each of the multiple multiplication circuits MUL can receive one quantization scale coefficient QSC corresponding to the received weighted sum PSM. For example, the multiplication circuits MUL_1 to MUL_R are 1_c1 ~a R_c1 Each of the multiplication circuits MUL_R+1 to MUL_2R receives a 1_c2 ~a R_c2 can be received respectively.
[0151] Each of the multiple multiplication circuits MUL can multiply the received weighted sum PSM and the quantization scale coefficient QSC to output a scaled weighted sum SCPSM. For example, the multiplication circuits MUL_1 to MUL_R each have a SCPSM 1_c1_X1 ~SCPSM R_c1_X1 The multiplication circuits MUL_R+1 to MUL_2R can output SCPSM 1_c2_X1 ~SCPSM R_c2_X1 can be output.
[0152] The accumulator 17 may receive the plurality of scaled weighted sums SCPSM from the partial sum scaler 16. The accumulator 17 may accumulate scaled weighted sums corresponding to the same input vectors and the same column vectors among the plurality of scaled weighted sums SCPSM to calculate an output element. For example, the accumulator 17 may accumulate the scaled weighted sums corresponding to the first input vector (i.e.,
[0153]
number
[0154]
number
[0155]
number
[0156]
number
[0157] That is, the matrix multiplier 10 uses the processing elements included in the first processing element row PER1 to multiply the first input vector (i.e.,
[0158]
number
[0159]
number
[0160] In one embodiment, when the product of the number of columns (i.e., 'm') of the weight matrix WM and the BCQ resolution (i.e., 'R') is greater than the number of columns (i.e., 'q') of the processing element array 13, the matrix multiplier 10 can compute the output elements based on various tiling techniques. Tiling techniques are described in more detail with reference to Figures 34-37 below.
[0161] In one embodiment, the product of the number of columns (i.e., 'm') of the weight matrix WM and the BCQ resolution (i.e., 'R') may not be an integer multiple of the number of columns (i.e., 'q') of the processing element array 13. In this case, some columns of the processing element array 13 may not need to perform the partial sum calculation operation described above. That is, according to the embodiment of Figures 4 to 11, the operating efficiency of the processing element array 13 may decrease, and the operating speed of the matrix multiplication device MMD may decrease.
[0162] Fig. 12 is a diagram showing the operation of the processing element in Fig. 11. Referring to Fig. 1 and Fig. 3 to Fig. 12, the first processing element row PER1 can include processing elements PE11, PE12, PE1R, and PE1(R+1). For the sake of simpler explanation, the operations of the processing elements PE11, PE12, PE1R, and PE1(R+1) will be representatively explained below, but the scope of the present disclosure is not limited thereto.
[0163] Each of the processing elements included in the first processing element row PER1 receives a first fixed-point input vector (i.e.,
[0164]
number
[0165] Each of the processing elements PE11, PE12, PE1R, and PE1(R+1) can receive a different quantized code vector. For example, the processing element PE11 receives
[0166]
number
[0167]
number
[0168]
number
[0169]
number
[0170] Each of the processing elements PE11, PE12, PE1R, and PE1(R+1) can calculate a different fixed-point partial sum PSM_fxp based on the order in which the fixed-point input elements and the quantization code values are received. For example, the processing element PE11 calculates PSM' 1_c1_X1 can be calculated.
[0171]
number
[0172] 12 shows an embodiment in which identical fixed-point input elements are individually provided to each of the processing elements included in the first processing element row PER1 for the sake of simpler description, but the scope of the present disclosure is not limited thereto. For example, each of the processing elements included in the first processing element row PER1 may be implemented to sequentially transmit received fixed-point input elements to adjacent processing elements in the row direction. However, the scope of the present disclosure is not limited thereto.
[0173] In one embodiment, the processing elements arranged in the same column of the processing element array 13 may receive the same quantized code vector. In this case, the processing elements arranged in the same column may sequentially transmit the quantized code vector to the processing elements adjacent to each other in the column direction. However, the scope of the present disclosure is not limited thereto.
[0174] Fig. 13 is a diagram showing a configuration of one of the processing elements of Fig. 12, which is realized according to an embodiment. Referring to Figs. 1 and 3 to 13, the processing element PE may include an arithmetic logic unit (ALU).
[0175] The arithmetic logic unit ALU may include a first input terminal TI1 to a third input terminal TI3 (input terminal) and an output terminal TO (output terminal). The first input terminal TI1 is a fixed-point input element IE_fxp (for example, x' 11 The second input terminal TI2 can receive the quantized code value QSV. The third input terminal TI3 can be connected to the output terminal TO.
[0176] The arithmetic logic unit ALU can output at the output terminal TO a value obtained by adding the product of the values received at the first input terminal TI1 and the second input terminal TI2 to the value received at the third input terminal TI3, although the scope of the present disclosure is not limited thereto.
[0177] 14 is a diagram illustrating the operation of the second data type converter in FIG. 6. In the following, for the sake of simpler explanation, the first input vector (i.e.,
[0178]
number
[0179] 1 and 3 to 14, the second data type converter 14 can receive a first exponent EXP1 from the first data type converter 11. The second data type converter 14 can receive a fixed-point partial sum PSM_fxp from one processing element PE.
[0180] The second data type converter 14 can convert the data type of the fixed-point partial sum PSM_fxp into floating-point to generate the partial sum PSM.
[0181] The second data type converter 14 can add the exponent part EXPP to the fixed-point partial sum PSM_fxp. The second data type converter 14 can determine the exponent part EXPP of the partial sum PSM as the first exponent EXP1.
[0182] The second data type converter 14 can add multiple '0' bits to the LSB position of the mantissa of the partial sum PSM according to the difference in code length between the mantissa of the partial sum PSM and the mantissa of the fixed-point partial sum PSM_fxp, but the scope of the present disclosure is not limited thereto.
[0183] 15 is a diagram illustrating the operation of the BCQ circuit of FIG 1 according to an embodiment of the present disclosure. Referring to FIG 1, FIG 3, and FIG 15, the BCQ circuit 200 can perform a binary coding quantization operation for each row of the weight matrix WM.
[0184] First, the weight matrix WM can be expressed as in Equation 14 below.
[0185]
number
[0186]
number
[0187]
number
[0188] The BCQ circuit 200 may perform a binary coding quantization operation for each row of the weight matrix WM according to the following Equation 15.
[0189]
number
[0190]
number
[0191]
number
[0192] Fig. 16 is a diagram showing a weight matrix approximated by the embodiment of Fig. 15. Referring to Figs. 1, 3, and 15 to 16, the quantized code vector
[0193]
number
[0194]
number
[0195]
number
[0196] In this manner, the BCQ circuit 200 can approximate each row vector included in the weight matrix WM with a combination of a plurality of quantized code values QSV and a plurality of quantized code coefficients QSC.
[0197] For the sake of simplicity, in the following, the output element in the first row and first column of the output matrix YM (i.e., y 11 ) will be representatively described, however, the scope of the present disclosure is not limited thereto.
[0198] The BCQ circuit 200 may provide a plurality of quantization code values QSV (ie, “b” values) and a plurality of quantization scale factors QSC (ie, “a” values) to the matrix multiplier 100 .
[0199] The matrix multiplier 100 can calculate an output matrix YM based on a plurality of quantization code values QSV and a plurality of quantization scale coefficients QSC. For example, the matrix multiplier 100 can calculate an output matrix YM by multiplying an input matrix XM by a weight matrix WM that is approximated by the method described above with reference to Equations 14 to 16 and FIG. 16.
[0200] In a more detailed example, the matrix multiplier 100 calculates y 11 can be calculated.
[0201]
number
[0202]
number
[0203] Meanwhile, referring to FIG. 16 and Equations 14 to 16, the first quantization scale coefficient (i.e., a 1_r1 ~a 1_rn ) is multiplied by the quantized code vector (i.e.,
[0204]
number
[0205]
number
[0206]
number
[0207] In one embodiment, the first quantization code matrix QSM_1 to the (R)th quantization code matrix QSM_R can be realized with the same number of rows and columns as the weight matrix WM. For example, the number of rows of the first quantization code matrix QSM_1 may be 'n' and the number of columns may be 'm'. The first quantization code matrix QSM_1 to the (R)th quantization code matrix QSM_R will be described in more detail with reference to FIG. 17 below.
[0208] Figure 17 is a diagram showing the quantization code matrix of Figure 16. Referring to Figures 1, 3, and 15 to 17, the BCQ circuit 200 can perform binary coding quantization (BCQ) of the weight matrix WM with a plurality of quantization code matrices QSM and a plurality of quantization scale coefficients QSC.
[0209] The number of quantization code matrices QSM may be determined by the BCQ resolution (i.e., 'R'). For example, the BCQ circuit 200 may perform binary coding quantization on the weight matrix WM to generate the first quantization code matrix QSM_1 through the (R)th quantization code matrix QSM_R.
[0210] Each of the first quantization code matrix QSM_1 to the (R)th quantization code matrix QSM_R can be realized with the same number of rows and columns as the weight matrix WM. For example, the number of rows of each of the first quantization code matrix QSM_1 to the (R)th quantization code matrix QSM_R may be 'n' and the number of columns may be 'm'. In this case, the quantization code value QSV arranged in the (i)th row and the (j)th column of the (k)th quantization code matrix QSM_k is "b j_k_ri "It could be.
[0211] Each weight of the weight matrix WM is approximated based on the quantization code value QSV arranged at the corresponding position of the first quantization code matrix QSM_1 to the (R)th quantization code matrix QSM_R. For example, the weight arranged in the (i)th row and (j)th column of the weight matrix WM (i.e., w ij ) is a quantization code value (i.e., b j_1_ri ~b j_R_ri ) is approximated based on
[0212] Each of the first quantization code matrix QSM_1 through the (R)th quantization code matrix QSM_R may correspond to a plurality of quantization scale coefficients. That is, as described above with reference to Figs. 15 and 16, when the BCQ circuit 200 performs a binary coding quantization operation for each row of the weight matrix WM, different rows of the weight matrix WM are approximated based on different quantization scale coefficients, and weights included in the same row of the weight matrix WM are approximated based on the same quantization scale coefficient.
[0213] For example, the quantization code value included in the (i)th row of the (k)th quantization code matrix QSM_k (i.e., b j_k_ri ~b m_k_ri ) are all quantization scale factors “a k_riThat is, the first row to the (n)th row of the (k)th quantization code matrix QSM_k may correspond to the quantization scale coefficient “a k_r1 "~"a k_rn " can correspond to each of the above.
[0214] In more detail, the first row to the (n)th row of the first quantization code matrix QSM_1 are each represented as “a 1_r1 "~"a 1_rn In this case, all the quantization code values arranged in the first row of the first quantization code matrix QSM_1 correspond to a 1_r1 It can correspond to.
[0215] In this manner, the weights (i.e., w ij ) is a quantization scale coefficient QSC (i.e., a) corresponding to the quantization code value arranged in the (i)th row of the first quantization code matrix QSM_1 to the (R)th quantization code matrix QSM_R. 1_ri ~a R_ri ) (That is, according to the embodiment of the present disclosure, the quantization scale factor is defined regardless of the column number of the weight matrix.) Therefore, each of the multiple weights is approximated by multiple quantization scale factors QSC and multiple quantization code values QSV according to Equation 19 below.
[0216]
number
[0217] Fig. 18 is a block diagram showing a configuration of the matrix multiplier of Fig. 1 realized by an embodiment of the present disclosure. Referring to Figs. 1, 3, and 15 to 18, the matrix multiplier 100 may include a quantization scale coefficient buffer 110, an input vector scaler 120, a first data type converter 130 (data type converter #1), a quantization sign value buffer 140, a processing element array 150, and a second data type converter 160 (data type converter #2).
[0218] The quantization scale factor buffer 110 may store a number of quantization scale factors QSC provided by the BCQ circuit 200. The quantization scale factor buffer 15 may provide a number of quantization scale factors QSC to the input vector scaler 120.
[0219] The input vector scaler 120 may receive an input matrix XM. For example, the input vector scaler 120 may receive multiple input vectors that include multiple input elements (e.g.,
[0220]
number
[0221] The input vector scaler 120 may scale the input matrix XM based on multiple quantization scale factors QSC. For example, the input vector scaler 120 may generate multiple scaled input vectors SCX based on multiple input vectors. In this case, the multiple scaled input vectors SCX may be generated based on multiple input vectors (e.g.,
[0222]
number
[0223]
number
[0224]
number
[0225] In one embodiment, the multiple scaled input vectors SCX are included in a scaled input matrix, where the row size of the scaled input matrix may be an integer multiple of the row size of the input matrix XM, and the column size of the scaled input matrix may be the same as the column size of the input matrix XM, although the scope of the present disclosure is not limited in this respect.
[0226] Each of the scaled input vectors SCX can be realized as a row vector having a dimension R times that of the corresponding input vector. For example, the first input vector (i.e.,
[0227]
number
[0228]
number
[0229]
number
[0230] In this manner, the second through (h) scaled input vectors (i.e.,
[0231]
number
[0232]
number
[0233] In one embodiment, the data types of the input elements and the quantization scale factors QSC may be floating point, in which case the data type of each of the scaled input elements may be floating point, although the scope of the present disclosure is not limited in this respect.
[0234] In one embodiment, the code length of each of the multiple input elements may be 16-bits or 32-bits, although the scope of the present disclosure is not limited in this respect.
[0235] In one embodiment, the code length of each of the multiple quantization scale factors QSC may be 16-bit or 32-bit, although the scope of the present disclosure is not limited in this respect.
[0236] The first data type converter 130 may receive a plurality of scaled input vectors SCX. For example, the first data type converter 13 may receive first to (h)-th scaled input vectors SCX (i.e.,
[0237]
number
[0238] The first data type converter 130 can extract an exponent EXP from each of the multiple scaled input vectors SCX. For example, the first data type converter 130 can extract an exponent EXP from each of the multiple scaled input vectors SCX.
[0239]
number
[0240]
number
[0241] The first data type converter 130 can convert the data type of the plurality of scaled input vectors SCX to fixed point. For example, the first data type converter 130 can receive the plurality of scaled input vectors SCX and output the plurality of fixed point scaled input vectors SCX_fxp. That is, the first data type converter 130 can convert the first through (h)th fixed point scaled input vectors (i.e.,
[0242]
number
[0243] More specifically, the first data type converter 130 may convert the data type of each of the scaled input elements included in the plurality of scaled input vectors SCX to fixed point based on the extracted exponents. In this case, the first fixed point scaled input vector (i.e.,
[0244]
number
[0245] The quantization code value buffer 140 can store a plurality of quantization code values QSV provided from the BCQ circuit 200. The quantization code value buffer 140 can provide a plurality of quantization code values QSV to the processing element array 150.
[0246] The processing element array 150 may receive a plurality of quantization code values QSV and a plurality of fixed-point scaled input vectors SCX_fxp. The processing element array 150 may generate a fixed-point output matrix YM_fxp based on the plurality of fixed-point scaled input elements (i.e., the fixed-point input vectors SCX_fxp) and the plurality of quantization code values QSV. The fixed-point output matrix YM_fxp may be expressed as Equation 21 below.
[0247]
number
[0248]
number
[0249] The processing element array 150 may include a plurality of processing elements arranged in row and column directions. Each of the plurality of processing elements may calculate and output a different fixed-point output element of the above-mentioned Equation 21. The configuration and operation of each of the processing elements will be described in more detail with reference to the following FIGS. 22 to 25.
[0250] The second data type converter 160 may receive a plurality of exponents EXP from the first data type converter 130. The second data type converter 160 may receive a fixed-point output matrix YM_fxp from the processing element array 150. For example, the second data type converter 160 may receive a plurality of fixed-point output elements from the processing element array 150.
[0251] The second data type converter 160 can convert the data type of the fixed-point output matrix YM_fxp to floating point based on the multiple exponents EXP. That is, the second data type converter 160 can output the output matrix YM having a floating point data type. For example, the second data type converter 160 can convert the data type of each of the multiple received fixed-point output elements to floating point to generate multiple output elements. A more detailed configuration and operation of the second data type converter 160 will be described in more detail with reference to FIG. 26 below.
[0252] Figure 19 is a block diagram showing the configuration of the input vector scaler of Figure 18. Referring to Figures 1, 3, and 15 to 19, the input vector scaler 120 can include a first input vector scaling circuit 121 through a (h)th input vector scaling circuit 12h.
[0253] The first input vector scaling circuit 121 to the (h)th input vector scaling circuit 12h can receive different input vectors from each other. For example, the first input vector scaling circuit 121 to the (h)th input vector scaling circuit 12h can receive the first to (h)th input vectors (i.e.,
[0254]
number
[0255] Each of the first input vector scaling circuit 121 through the (h)th input vector scaling circuit 12h can sequentially receive multiple input elements. For example, the first input vector scaling circuit 121 receives x 11 ~x 1n , and the (h) input vector scaling circuit 12h receives x h1 ~ hn can be received in sequence.
[0256] Each of the first input vector scaling circuit 121 through the (h)th input vector scaling circuit 12h can sequentially receive a plurality of quantization scale coefficients QSC from the quantization scale coefficient buffer 110.
[0257] The quantization scale coefficients QSC that each of the first input vector scaling circuit 121 through the (h)th input vector scaling circuit 12h receives from the quantization scale coefficient buffer 110 may be identical to one another. For example, the quantization scale coefficients QSC that the first input vector scaling circuit 121 sequentially receives may be identical to the quantization scale coefficients QSC that the second input vector scaling circuit 122 sequentially receives.
[0258] The order in which the first input vector scaling circuit 121 through the (h)th input vector scaling circuit 12h receive the multiple quantization scale coefficients QSC may be the same as each other. For example, the quantization scale coefficient QSC that the first input vector scaling circuit 121 receives first may be the same as the quantization scale coefficient QSC that the second input vector scaling circuit 122 receives first. Similarly, the quantization scale coefficient QSC that the first input vector scaling circuit 121 receives second may be the same as the quantization scale coefficient QSC that the second input vector scaling circuit 122 receives second.
[0259] The first input vector scaling circuit 121 through the (h)th input vector scaling circuit 12h respectively generate first through (h)th scaled input vectors (i.e.
[0260]
number
[0261]
number
[0262] Fig. 20 is a diagram showing in more detail the operation of the input vector scaling circuit of Fig. 19. For the sake of simpler explanation, the operation of the first input vector scaling circuit 121 will be representatively described below with reference to Figs. 1, 3, and 15 to 20. However, the scope of the present disclosure is not limited thereto, and the second input vector scaling circuit 122 to the (h)th input vector scaling circuit 12h may operate in a similar manner.
[0263] The first input vector scaling circuit 121 scales the first input vector (i.e.,
[0264]
number
[0265] The first input vector scaling circuit 121 may receive a plurality of quantization scale factors QSC in sequence. For example, the first input vector scaling circuit 121 may receive a first row vector of the weight matrix WM (i.e.,
[0266]
number
[0267]
number
[0268] In one embodiment, the second input vector scaling circuit 122 to the (h)th input vector scaling circuit 12h are 1_r1 ~a R_rn can be received in sequence.
[0269] In one embodiment, the first input vector scaling circuit (121) to the (h)th input vector scaling circuit (12h) each include the above-mentioned 1_r1 ~a R_rn may be the same as each other, but the scope of the present disclosure is not limited in this respect.
[0270] The first input vector scaling circuit 121 generates a first scaled input vector based on the order in which the input elements and the plurality of quantization scale factors QSC are received.
[0271]
number
[0272] More specifically, the first input vector scaling circuit 121 scales x 11 Let us use the first row vector of the weight matrix WM (i.e.,
[0273]
number
[0274] Then, the first input vector scaling circuit 121 scales x 12 Let us use the second row vector of the weight matrix WM (i.e.,
[0275]
number
[0276] In this manner, the first input vector scaling circuit 121 scales x 13 ~x 1n A number of scaled input elements SCIE corresponding to the respective scaled input elements S C can be sequentially computed.
[0277] The first input vector scaling circuit 121 can sequentially output the multiple scaled input elements SCIE that have been calculated.
[0278] That is, according to an embodiment of the present disclosure, the first input vector scaling circuit 121 scales one input element (e.g., x 11 ) based on multiple scaled input elements (e.g., scaled input elements illustrated with diagonal stripes (SCIE for x 11) can be generated. In other words, the first input vector scaling circuit 121 can generate multiple scaled input elements by repeatedly using one input element. Therefore, according to the embodiment of the present disclosure, the input reuse of the matrix multiplier 100 is maximized, and thus the number of times that the matrix multiplier 100 receives input elements from the outside is minimized. In this case, the number of times that the matrix multiplier 100 accesses an external memory device that stores the input elements is minimized, and thus the operating efficiency and operating speed of the matrix multiplication device MMD can be improved.
[0279] In one embodiment, the BCQ circuit 200 can generate a plurality of quantization scale factors QSC and a plurality of quantization code values QSV from a plurality of weights based on a uniform BCQ algorithm. In this case, the multiplication of the quantization scale factors QSC and the input elements shown in FIG. 20 can be performed with a smaller amount of calculations. For example, when the magnitude ratio between the plurality of quantization scale factors is a 'power of 2', the mantissas of the plurality of quantization scale factors QSC may be the same, and the exponents of the plurality of quantization scale factors QSC may differ from each other by '1'. In this case, a single input element (e.g., x 11 ) for multiple different quantization scale factors (e.g., a 1_r1 ~a R_r1 ) each product is more easily performed. However, the scope of the present disclosure is not limited in this respect.
[0280] Fig. 21 is a block diagram showing a configuration of the first data type converter of Fig. 18. With reference to Figs. 1, 3, 15 to 18, and 21, the first data type converter 130 can include a first exponent extraction circuit 131_1 through an (h)th exponent extraction circuit 131_h and a first data type conversion circuit 132_1 through an (h)th data type conversion circuit 132_h.
[0281] The first exponent extraction circuit 131_1 to the (h)th exponent extraction circuit 131_h can receive the scaled input vector SCX different from each other. For example, the first exponent extraction circuit 131_1 to the (h)th exponent extraction circuit 131_h can receive the first to the (h)th scaled input vectors (i.e.,
[0282]
number
[0283] Each of the first exponent extraction circuit 131_1 to the (h)th exponent extraction circuit 131_h can extract exponents from a plurality of scaled input elements SCIE included in the received scaled input vector SCX. For example, the first exponent extraction circuit 131_1 extracts an exponent from the first scaled input vector (
[0284]
number
[0285]
number
[0286] The first exponent extraction circuit 131_1 to the (h)th exponent extraction circuit 131_h can provide the extracted exponents to the first data type conversion circuit 132_1 to the (h)th data type conversion circuit 132_h, respectively. Also, the first exponent extraction circuit 131_1 to the (h)th exponent extraction circuit 131_h can provide the extracted exponents to the second data type converter 160, respectively.
[0287] The specific manner in which each of the first exponent extraction circuit 131_1 to the (h)th exponent extraction circuit 131_h extracts exponents from the received elements is similar to the operation of the exponent extraction circuit previously described with reference to Figures 7 to 8, so detailed description thereof will be omitted.
[0288] The first data type conversion circuit 132_1 to the (h)th data type conversion circuit 132_h can receive the first exponent EXP1 to the (h)th exponent EXPh, respectively. The first data type conversion circuit 132_1 to the (h)th data type conversion circuit 132_h can receive the first to the (h)th scaled input vectors (i.e.,
[0289]
number
[0290] Each of the first data type conversion circuit 132_1 to the (h)th data type conversion circuit 132_h can convert the data type of the received scaled input vector into a fixed point based on the received exponent. That is, the first data type conversion circuit 132_1 to the (h)th data type conversion circuit 132_h can convert the data type of the received scaled input vector into a fixed point based on the received exponent.
[0291]
number
[0292]
number
[0293] The specific manner in which each of the first data type conversion circuit 132_1 to the (h)th data type conversion circuit 132_h converts the data type of the received element into a fixed point based on the received exponent is similar to the operation of the data type conversion circuit previously described with reference to Figures 7 and 9, so detailed description thereof will be omitted.
[0294] FIG. 22 is a block diagram showing the configuration of the processing element array of FIG. 18 in more detail. Referring to FIG. 1, FIG. 3, and FIG. 15 to FIG. 22, the processing element array 150 can include a plurality of processing elements PE arranged in row and column directions. For the sake of simpler explanation, it is assumed below that the plurality of processing elements PE are arranged along (h) rows and (m) columns. Also, a processing element arranged in the (i)th row and (j)th column of the processing element array 150 is referred to as "PEij". For example, a processing element arranged in the first row and second column of the processing element array 150 is referred to as "PE12".
[0295] The processing element array 150 can include a first processing element row PER1 to a (h)th processing element row PERh. Each of the first processing element row PER1 to the (h)th processing element row PERh can include a plurality of processing elements PE. For example, the first processing element row PER1 can include processing elements PE11 to PE1q.
[0296] The processing element array 150 can include a first processing element column PEC1 to an (m)th processing element column PECm. Each of the first processing element column PEC1 to the (m)th processing element column PECm can include a plurality of processing elements PE. For example, the first processing element column PEC1 can include processing elements PE11 to PEh1.
[0297] Different processing element rows may receive different fixed-point input vectors SCX_fxp. For example, the first processing element row PER1 through the (h)th processing element row PERh may receive the first through the (h)th fixed-point scaled input vectors SCX_fxp, respectively (i.e.,
[0298]
number
[0299] The processing elements included in the same processing element row may receive the same fixed-point input vector SCX_fxp. For example, each of the processing elements PE11 to PE1m receives a first fixed-point scaled input vector (i.e.,
[0300]
number
[0301] Different processing element columns can receive different quantized code values QSVs. For example, the first processing element column PEC1 to the (m)th processing element column PECm can receive the first plurality of quantized code values QSVs_1 to the (m)th plurality of quantized code values QSVs_m, respectively.
[0302] The processing elements arranged in the same processing element column can receive the same multiple quantization code values QSVs. For example, each of the processing elements PE11 to PEh1 can receive the first multiple quantization code values QSVs_1, and each of the processing elements PE12 to PEh2 can receive the second multiple quantization code values QSVs_2.
[0303] Each of the processing elements PE can calculate a different fixed-point output element based on the received fixed-point scaled input element SCIE and the multiple quantization code values QSVs. That is, according to the embodiment of the present disclosure, one processing element PE can calculate one fixed-point output element. For example, the processing element PEij calculates y' ij In the following, the fixed-point output elements calculated in each processing element PE will be described in more detail.
[0304] The first processing element row PER1 to the (h)th processing element row PERh can respectively calculate different fixed-point output vectors. For example, the first processing element row PER1 to the (h)th processing element row PERh can respectively calculate the first to (h)th fixed-point output vectors (i.e.,
[0305]
number
[0306] The processing elements arranged in the same processing element row and different processing element columns can operate on different fixed-point output elements. For example, the processing elements PE11 to PE1m operate on y' 11 ~y'1m In a similar manner, the processing elements PE21 to PE2m can calculate y' 21 ~y' 2m The processing elements PEh1 to PEhm can respectively calculate y' h1 ~y' hm can be calculated respectively.
[0307] Fig. 23 is a block diagram showing in more detail the output of the processing element of Fig. 22. Referring to Figs. 1, 3, and 15 to 23, each of the multiple processing elements PE can provide a calculated fixed-point output element to the second data type converter 160.
[0308] For a simpler description, an embodiment in which each of the plurality of processing elements PE directly provides a calculated fixed-point output element to the second data type converter 160 as shown in FIG. 23 will be representatively described. However, the scope of the present disclosure is not limited thereto, and each of the plurality of processing elements PE may transmit the calculated fixed-point output element to the second data type converter 160 in a systolic array manner. An embodiment in which the plurality of processing elements PE are realized in a systolic array manner will be described in more detail with reference to the following FIGS. 31 and 32.
[0309] Fig. 24 is a block diagram showing in more detail the operation of the first processing element row in Fig. 22. That is, hereinafter, the operation of the first processing element row PER1 will be representatively described with reference to Figs. 1, 3, and 15 to 24. However, the scope of the present disclosure is not limited thereto, and the second processing element row PER2 to the (h)th processing element row PERh may also operate in a similar manner.
[0310] The first processing element row PER1 receives the first fixed-point scaled input vector (i.e.,
[0311]
number
[0312] The processing elements PE11 to PE1m can receive the first plurality of quantization code values QSV_1 to the (m)th plurality of quantization code values QSV_m, respectively. For example, the processing element PE11 can receive the first plurality of quantization code values QSVs_1, and the processing element PE12 can receive the second plurality of quantization code values QSVs_2.
[0313] The first plurality of quantization code values QSVs_1 may include quantization code values QSV arranged in the first column of the first quantization code matrix QSM_1 to the (R)th quantization code matrix QSM_R previously described with reference to FIG. 17. For example, the first plurality of quantization code values QSVs_1 may include quantization code values (i.e., b 1_1_r1 ~b 1_R_rn ).
[0314] More specifically, the processing element PE11 processes the quantization code values (i.e., b 1_1_r1 ~b 1_R_r1 ) are sequentially received, and then the quantization code values (i.e., b 1_1_r2~b 1_R_r2 In this manner, the processing element PE11 sequentially receives the quantization code values (i.e., b 1_1_rn ~b 1_R_rn ) can be received in sequence.
[0315] The processing element PE11 selects a fixed-point output element (e.g., y') based on the order in which the quantized code value QSV and the plurality of fixed-point scaled input elements SCIE_fxp are received. 11 For example, the processing element PE11 can calculate the input element “x 11 The multiple fixed-point scaled input elements (i.e., “SCIE_fxp for x” in Figure 24) are generated based on the 11 The fixed-point scaled input elements indicated by " " are used as input elements for the first quantization code matrix QSM_1 to the (R)th quantization code matrix QSM_R. 1_1_r1 ~b 1_R_r1 In this way, the processing element PE11 can accumulate the values multiplied by each of the fixed-point formats. 1_r1 x 11 " to "b 1_1_r1 " multiplied by "a" converted to fixed-point format R_rn x 1n " to "b 1_R_rn The values multiplied by y' are sequentially accumulated to produce the fixed-point output element "y' 11 " can be calculated and output.
[0316] Similarly, the processing element PE1j processes the quantization code values (i.e., b j_1_r1 ~b j_R_r1In this case, the processing element PE1j may receive the first fixed-point scaled input vector (i.e.,
[0317]
number
[0318] Fig. 25 is a diagram showing a configuration of the processing element of Fig. 24 according to an embodiment. Referring to Figs. 1, 3, and 15 to 25, the processing element PE may include an arithmetic logic unit ALU and an accumulation register REG_ACC.
[0319] The arithmetic logic unit ALU may include a first input terminal TI1 to a third input terminal TI3 and an output terminal TO.
[0320] The arithmetic logic unit ALU may receive the fixed-point scaled input element SCIE_fxp via a first input terminal TI1. For example, the arithmetic logic unit ALU may sequentially receive a plurality of fixed-point scaled input elements SCIE_fxp via a first input terminal TI1.
[0321] In one embodiment, each fixed-point scaled input element SCIE_fxp has a code length of 8-bits, and the arithmetic logic unit ALU is configured to receive data in 8-bit units via a first input terminal TI1.
[0322] The arithmetic logic unit ALU may receive the quantized code value QSV via a second input terminal TI2, for example, the arithmetic logic unit ALU may sequentially receive a plurality of quantized code values QSV via the second input terminal TI2.
[0323] In one embodiment, the arithmetic logic unit ALU may receive each quantization code value QSV through the second input terminal TI2 in the form of a control signal indicating one of logic low or logic high. For example, if the control signal provided to the second input terminal TI2 is logic high, the arithmetic logic unit ALU may determine that a quantization code value QSV indicating '+1' has been received. Conversely, if the control signal received through the second input terminal TI2 is logic low, the arithmetic logic unit ALU may determine that a quantization code value QSV indicating '-1' has been received. However, the scope of the present disclosure is not limited thereto.
[0324] The third input terminal TI3 may be connected to the accumulation register REG_ACC, and the arithmetic logic unit ALU may receive data stored in the accumulation register REG_ACC through the third input terminal TI3.
[0325] The arithmetic logic unit ALU may calculate a value obtained by adding the product of the quantization code value QSV received through the second input terminal TI2 and the fixed-point scaled input element SCIE_fxp received through the first input terminal TI1 to the data received through the third input terminal TI3. The arithmetic logic unit ALU may provide the calculated value to the accumulation register REG_ACC through the output terminal TO to update the value stored in the accumulation register REG_ACC. That is, the arithmetic logic unit ALU may update the accumulation register REG_ACC with a value calculated according to the following Equation 22.
[0326]
number
[0327] That is, the arithmetic logic unit ALU can accumulate a plurality of fixed-point scaled input elements SCIE_fxp sequentially received through the first input terminal TI1 based on a plurality of quantization code values QSV sequentially received through the second input terminal TI2. In this manner, the arithmetic logic unit ALU accumulates a plurality of fixed-point scaled input elements SCIE_fxp in the accumulation register REG_ACC based on the plurality of quantization code values QSV, thereby generating a fixed-point output element OE_fxp (e.g., y' 11 ) can be calculated.
[0328] That is, the arithmetic logic unit ALU may store the fixed-point output element OE_fxp in the accumulation register REG_ACC, which may provide the fixed-point output element OE_fxp to the second data type converter 160.
[0329] In one embodiment, the fixed-point output element OE_fxp may have a code length of 8-bits or more, for example, the fixed-point output element OE_fxp may have a code length long enough to represent the accumulated magnitude of multiple fixed-point scaled input elements SCIE_fxp.
[0330] In one embodiment, the accumulation register REG_ACC may have a size of 8-bits or more, for example, the accumulation register REG_ACC may be large enough to store the fixed-point output element OE_fxp.
[0331] In one embodiment, the fixed-point output element OE_fxp may have a code length of 10-bit to 12-bit, although the scope of the present disclosure is not limited in this respect.
[0332] In one embodiment, the accumulation register REG_ACC may have a size of 10-bits to 12-bits, although the scope of the present disclosure is not limited in this respect.
[0333] In one embodiment, the multiple fixed-point scaled input elements SCIE_fxp received by the arithmetic logic unit ALU may correspond to the same exponent value. In this case, the arithmetic logic unit ALU can perform the operation of Equation 22 without considering the place values of each of the multiple fixed-point scaled input elements SCIE_fxp. Therefore, the arithmetic logic unit ALU can operate the fixed-point output element OE_fxp with a minimum amount of operations.
[0334] FIG. 26 is a diagram illustrating the operation of the second data type converter in FIG. 18. For the sake of simpler explanation, the operation of the second data type converter 160 for one fixed-point output element OE_fxp will be representatively described below. However, the scope of the present disclosure is not limited thereto. For example, the second data type converter 160 may operate in a similar manner for any fixed-point output element OE_fxp.
[0335] 1, 3, and 15 to 26, the second data type converter 160 may receive a fixed point output element OE_fxp. For example, the second data type converter 160 may receive y' 11 ~y' hn One of the following can be received:
[0336] The second data type converter 160 may receive the first exponent EXP1 through the (h)th exponent EXPh from the first data type converter 130. In this case, the first exponent EXP1 through the (h)th exponent EXPh are respectively a first fixed-point output vector (i.e.,
[0337]
number
[0338]
number
[0339] The fixed-point output element OE_fxp may include an exponent part SP and a mantissa part MTSP. The second data type converter 14 may convert the data type of the fixed-point output element OE_fxp into a floating point to generate the output element OE. A specific manner in which the second data type converter 160 converts the data type of the fixed-point output element OE_fxp will be described below.
[0340] The second data type converter 160 may add an exponent portion EXPP to the fixed-point output element OE_fxp. For example, the second data type converter 160 may determine an exponent corresponding to the fixed-point output element OE_fxp among the exponents received from the first data type converter 130 by using the exponent portion EXPP of the output element OE. In a more detailed example, when the fixed-point output element OE_fxp is added to the first fixed-point output vector (
[0341]
number
[0342]
number
[0343] The second data type converter 160 can add multiple '0' bits to the LSB position of the mantissa of the output element OE according to the difference in code length between the mantissa of the fixed-point output element OE_fxp and the mantissa of the output element OE. However, the scope of the present disclosure is not limited thereto.
[0344] Fig. 27 is a flowchart showing the operation of the matrix multiplication device of Fig. 1. In the following, with reference to Figs. 1, 3, and 15 to 27, one input vector (for example,
[0345]
number
[0346]
number
[0347] In step S110, the matrix multiplication device MMD may receive a weight matrix WM. For example, the BCQ circuit 200 may receive a plurality of weights (e.g., w 11 ~w nm ) can be received.
[0348] In step S120, the matrix multiplication device MMD may perform binary coding quantization on the weight matrix WM to generate a plurality of quantization scale factors QSC and a plurality of quantization code values QSV. For example, the BCQ circuit 200 may convert each of the plurality of weights into two or more quantization scale factor QSC-quantization code value QSV pairs. The BCQ circuit 200 may provide the generated plurality of quantization scale factors QSC and a plurality of quantization code values QSV to the matrix multiplier 100. In this case, the plurality of quantization scale factors QSC are stored in the quantization scale factor buffer 110, and the plurality of quantization code values QSV are stored in the quantization code value buffer 140. However, the scope of the present disclosure is not limited thereto.
[0349] In step S130, the matrix multiplication device MMD multiplies an input vector (e.g.,
[0350]
number
[0351] In one embodiment, the matrix multiplication device MMD may perform step S130 regardless of the order of steps S110 to S120. For example, the matrix multiplication device MMD may perform step S130 before steps S110 to S120, or between steps S110 and S120.
[0352] In step S140, the matrix multiplication device MMD may scale the input vector based on the plurality of quantization scale factors QSC to generate a scaled input vector SCX. For example, the input vector scaler 120 may scale each of the plurality of input elements based on the plurality of quantization scale factors QSC provided from the quantization scale factor buffer 110.
[0353] In step S150, the matrix multiplication device MMD accumulates elements of the input vector SCX scaled based on the multiple quantization code values QSV to generate an output vector (e.g.,
[0354]
number
[0355] Figure 28 is a flowchart showing in more detail step S150 of Figure 27. Referring to Figures 1, 3, and 15 to 28, step S150 may include steps S151 to S153.
[0356] In operation S151, the matrix multiplier 100 may convert the data type of the scaled input vector SCX to a fixed point. That is, the matrix multiplier 100 may generate a fixed-point scaled input vector SCX_fxp based on the scaled input vector SCX. For example, the first data type converter 130 may convert the data type of each of the plurality of scaled input elements SCIE to a fixed point to generate a plurality of fixed-point scaled input elements SCIE_fxp.
[0357] In step S152, the matrix multiplier 100 accumulates elements of the scaled input vector SCX (i.e., the fixed-point scaled input vector SCX_fxp) that has been converted to a fixed-point data type based on a plurality of quantization code values QSV to generate a fixed-point output vector (e.g.,
[0358]
number
[0359]
number
[0360] In step S153, the matrix multiplier 100 converts a fixed-point output vector (e.g.,
[0361]
number
[0362] Fig. 29 is a flowchart showing the operation of the matrix multiplication device of Fig. 1. In the following, with reference to Figs. 1, 3, 15 to 26, and 29, one input vector (for example,
[0363]
number
[0364] In step S210, the matrix multiplication device MMD may receive the first to n-th weights. For example, the BCQ circuit 200 may receive a weight (e.g., w 11 ~w n1 ) can be received.
[0365] In step S220, the matrix multiplication device MMD may perform binary coding quantization on the first through (n)-th weights to generate the first through (n×R)-th quantization scale coefficients QSC and the first through (n×R)-th quantization code values QSV. For example, the BCQ circuit 200 may generate 'R' quantization scale coefficients QSC and 'R' quantization code values QSV per weight.
[0366] In step S230, the matrix multiplication device MMD may receive the first to n-th input elements. For example, the matrix multiplier 100 may multiply a plurality of input elements (e.g., x 11 ~x 1n ) can be received.
[0367] In one embodiment, the matrix multiplication device MMD may perform step S230 regardless of the order of steps S210 to S220. For example, the matrix multiplication device MMD may perform step S230 before steps S210 to S220, or between steps S210 and S220.
[0368] In step S240, the matrix multiplication device MMD may scale the first through (n)th input elements based on the first through (n×R)th quantization scale coefficients QSC to generate the first through (n×R)th scaled input elements SCIE. For example, the input vector scaler 120 may generate 'R' scaled input elements SCIE per input element based on the 'R' quantization scale coefficients QSC.
[0369] In step S250, the matrix multiplication device MMD accumulates the first through (n×R) scaled input elements SCIE based on the first through (n×R) quantized code values QSV to generate one output element (e.g., y 11 For example, the matrix multiplier 100 may generate the first through (n×R) scaled input elements SCIE and the first through (n×R) quantized code values QSV (e.g., b 1_1_r1 ~b 1_R_rn ) Each product is accumulated to produce a single output element, “y 11 " can be generated.
[0370] Figure 30 is a flowchart showing in more detail step S250 of Figure 29. Referring to Figures 1, 3, 15 to 26, and 29 to 30, step S250 may include steps S251 to S253.
[0371] In operation S251, the matrix multiplier 100 may convert the data types of the first through (n×R)-th scaled input elements SCIE into fixed-point data. For example, the first data type converter 130 may convert the first through (n×R)-th scaled input elements SCIE into the first through (n×R)-th fixed-point scaled input elements SCIE_fxp, respectively.
[0372] In operation S252, the matrix multiplier 100 may accumulate the first through (n×R) quantized code values QSV based on the first through (n×R) quantized code values QSV to generate a fixed-point output element OE_fxp. For example, one processing element PE may sequentially receive the first through (n×R) quantized code values QSV and sequentially receive the first through (n×R) fixed-point scaled input elements SCIE_fxp. The processing element PE may accumulate the products of the first through (n×R) quantized code values QSV and the first through (n×R) fixed-point scaled input elements SCIE_fxp to generate one fixed-point output element OE_fxp (e.g., y' 11 ) can be generated.
[0373] In step S253, the matrix multiplier 100 may convert the data type of the fixed-point output element OE_fxp into a floating-point data type. For example, the second data type converter 160 converts the data type of the fixed-point output element OE_fxp (e.g., y' 11 ) and receives an output element OE (e.g., y 11 ) can be output.
[0374] Fig. 31 is a block diagram showing the processing element array of Fig. 22 realized in a systolic array manner. With reference to Figs. 1, 3, 15 to 22, and 31, the processing element array 150 can be realized by the processing element array 250 of Fig. 31.
[0375] The processing element array 250 may include a plurality of processing elements PE arranged in a row direction and a column direction. The plurality of processing elements PE may operate in a systolic array manner.
[0376] The processing element array 250 may be implemented to sequentially propagate a plurality of fixed-point scaled input elements SCIE_fxp in a row direction. For example, a first processing element row PER1 may receive a first fixed-point scaled input vector (i.e.,
[0377]
number
[0378] For example, the processing element PE11 may receive one fixed-point scaled input element SCIE_fxp at a first time point. The processing element PE11 may transfer the fixed-point scaled input element SCIE_fxp to the processing element PE12 arranged adjacent to the processing element PE11 in the row direction at a second time point after the first time point. In this manner, the processing element PE included in the first processing element row PER1 may sequentially transfer the multiple fixed-point scaled input elements SCIE_fxp provided from the first data type converter 130 to the adjacent processing elements.
[0379] The processing element array 250 can be realized to sequentially propagate the plurality of quantized code values QSV in the column direction. For example, the first processing element column PEC1 can sequentially propagate the first plurality of quantized code values QSVs_1 in the column direction.
[0380] For example, the processing element PE11 may receive one quantization code value QSV at a first time point. The processing element PE11 may transmit the quantization code value QSV to the processing element PE21 disposed adjacent to the processing element PE11 in the column direction at a second time point after the first time point. In this manner, the processing element PE included in the first processing element column PEC1 may sequentially transmit the first plurality of quantization code values QSVs_1 provided from the quantization code value buffer 140 to the adjacent processing elements.
[0381] Each of the multiple processing elements PE can generate a different fixed-point output element OE_fxp in the same manner as previously described with reference to Figures 18 to 24. For a simpler explanation, a detailed explanation of the manner in which each of the multiple processing elements PE generates the fixed-point output element OE_fxp will be omitted.
[0382] The processing element array 250 can be realized to sequentially propagate the fixed-point output elements OE_fxp in the column direction. For example, the fixed-point output element OE_fxp (i.e., y') calculated from the processing element PE11 is 11 ) can be propagated sequentially in the column direction. In this manner, the fixed-point output element OE_fxp is transferred to the second data type converter 160. The manner in which the fixed-point output element OE_fxp is propagated is similar to the manner in which the quantization code value QSV is propagated, and therefore a detailed description thereof will be omitted.
[0383] That is, each of the processing elements included in the processing element array 250 is realized to receive one or more of the fixed-point output element OE_fxp, the quantization code value QSV, and the fixed-point scaled input element SCIE_fxp from the adjacently arranged processing element. Conversely, each of the processing elements included in the processing element array 250 is realized to transmit one or more of the fixed-point output element OE_fxp, the quantization code value QSV, and the fixed-point scaled input element SCIE_fxp to the adjacently arranged processing element. A more detailed configuration of the processing element PE operating in a systolic array manner will be described in more detail with reference to FIG. 32 below.
[0384] For the sake of simpler explanation, an embodiment in which the fixed-point output element OE_fxp, the quantization code value QSV, and the fixed-point scaled input element SCIE_fxp are each propagated in a systolic array manner is representatively shown in Fig. 31, but the scope of the present disclosure is not limited thereto. For example, the processing element array 250 may be realized so that only one or two of the fixed-point output element OE_fxp, the quantization code value QSV, and the fixed-point scaled input element SCIE_fxp are propagated in a systolic array manner.
[0385] In one embodiment, each of the processing elements included in the processing element array 250 may operate in response to the same control clock signal. In this case, each of the multiple processing elements may transmit the fixed-point output element OE_fxp, the quantization code value QSV, and / or the fixed-point scaled input element SCIE_fxp to the other processing elements at the same time. However, the scope of the present disclosure is not limited in this respect.
[0386] Figure 32 is a diagram showing in more detail the configuration of the processing element of Figure 31. With reference to Figures 1, 3, 15 to 22, and 31 to 32, the processing element PE may include an arithmetic logic unit ALU, an accumulation register REG_ACC, a scaled input element register REG_SCIE, and a quantization code value register REG_QSV. For a simpler explanation, detailed explanations regarding the configuration and operation of the arithmetic logic unit ALU and the accumulation register REG_ACC previously described with reference to Figure 25 will be omitted.
[0387] In the following, an embodiment in which the arithmetic logic unit ALU, the accumulation register REG_ACC, the scaled input element register REG_SCIE, and the quantization code value register REG_QSV operate in response to the same control clock signal will be described as a representative example, but the scope of the present disclosure is not limited thereto.
[0388] The scaled input element register REG_SCIE can sequentially receive a plurality of fixed-point scaled input elements SCIE_fxp. The scaled input element register REG_SCIE can receive one fixed-point scaled input element SCIE_fxp and transmit it to the adjacent processing element PE and the first input terminal TI1 after one period of the control clock signal has elapsed.
[0389] The quantization code value register REG_QSV can sequentially receive a plurality of quantization code values QSV. The quantization code value register REG_QSV can receive one quantization code value QSV and transmit it to the adjacent processing element PE and the second input terminal TI2 after one period of the control clock signal has elapsed.
[0390] The accumulation register REG_ACC can store the fixed-point output element OE_fxp provided by the arithmetic logic unit ALU. For example, the accumulation register REG_ACC can store the fixed-point output element OE_fxp calculated by the arithmetic logic unit ALU in a manner similar to that previously described with reference to FIG.
[0391] The accumulation register REG_ACC can transfer the fixed-point output element OE_fxp to an adjacent processing element PE or the second data type converter 160. For example, the accumulation register REG_ACC can transfer the fixed-point output element OE_fxp to an accumulation register of an adjacent processing element PE. More specifically, the fixed-point output element OE_fxp (i.e., y') calculated by the processing element PE11 can be transferred to the accumulation register of the adjacent processing element PE. h1 ) can be transmitted to the second data type converter 160 via the processing elements PE21 to PEh1 in sequence. However, the scope of the present disclosure is not limited to this.
[0392] For the sake of simplicity, FIG. 32 illustrates an embodiment in which each of the registers included in the processing element PE receives and outputs data every cycle of the control clock signal, but the scope of the present disclosure is not limited to the specific operation method of the registers in response to the control clock signal.
[0393] FIG 33 illustrates the operation of the input vector scaling circuit of FIG 19 according to one embodiment. Referring to FIGS. 1, 3, 15-19, and 33, the first input vector scaling circuit 121 may be implemented as a first input vector scaling circuit 221. However, the scope of the present disclosure is not limited in this respect, and the second input vector scaling circuit 122 through the h-th input vector scaling circuit 12h may operate in a similar manner. The following mainly describes the differences between the first input vector scaling circuit 121 and the first input vector scaling circuit 221.
[0394] The first input vector scaling circuit 221 scales the first input vector (i.e.,
[0395]
number
[0396] The first input vector scaling circuit 221 may sequentially receive a plurality of quantization scale factors QSC. For example, the first input vector scaling circuit 221 may sequentially receive a quantization scale factor corresponding to a first quantization code matrix QSM_1 (i.e., a 1_r1 ~a 1_rn ), the quantization scale coefficient corresponding to the second quantization code matrix QSM_2 (i.e., a 2_r1 ~a 2_rn In this manner, the first input vector scaling circuit 221 can sequentially receive all of the quantization scale coefficients QSC for computing the aforementioned 'scaled input element SCIE'.
[0397] The first input vector scaling circuit 221 generates a first scaled input vector (
[0398]
number
[0399] In one embodiment, when the input vector scaler 120 is implemented based on the first input vector scaling circuit 221 described above, the processing element array 150 may receive the quantization code values QSV in an order different from that previously described with reference to FIG. 24. For example, the first processing element column PEC1 may receive the first plurality of quantization code values QSVs_1 in an order different from that previously described with reference to FIG. 24. In a more detailed example, the processing element PE11 may receive the quantization code values (i.e., b 1_1_r1 ~b 1_1_rn ) is sequentially received, and then the quantization code value (i.e., b 1_2_r1 ~b 1_2_rn ) in sequence, that is, the order in which the processing element array 150 receives the quantization code values QSV is determined by the order in which the input vector scaler 120 receives the quantization scale factors QSC.
[0400] Figure 34 is a diagram illustrating the operation of the matrix multiplication device of Figure 1 according to one embodiment. With reference to Figures 1 to 3 and 15 to 34, the matrix multiplication device MMD can receive a full input matrix FXM. The full input matrix FXM may include the input matrix XM previously described with reference to Figures 1 to 34. For example, the full input matrix FXM may include a plurality of input matrices XM.
[0401] The matrix multiplication device MMD may receive a full weight matrix FWM. The full weight matrix FWM may include the weight matrix WM previously described with reference to Figures 1 to 34. For example, the full weight matrix FWM may include a plurality of weight matrices WM.
[0402] The BCQ circuit 200 can perform a binary coding quantization operation on each of the weight matrices WM included in the full-weight matrix FWM. For example, the BCQ circuit 200 can generate a plurality of quantization code values QSV and a plurality of quantization scale factors QSC from each of the weight matrices WM.
[0403] The matrix multiplier 100 may receive a plurality of quantization code values QSV and a plurality of quantization scale factors QSC. The matrix multiplier 100 may perform matrix multiplication on the full-input matrix FXM and the full-weight matrix FWM based on the plurality of quantization code values QSV and the plurality of quantization scale factors QSC.
[0404] The matrix multiplier 100 may perform matrix multiplication on the full-input matrix FXM and the full-weight matrix FWM through one of various tiling techniques. For example, the matrix multiplier 100 may calculate the full-output matrix FYM by sequentially calculating the products of a plurality of input matrices XM and a plurality of weight matrices WM and then combining the calculated results.
[0405] FIG. 35 is a diagram showing the full-input matrix of FIG. 34. Referring to FIG. 1 to FIG. 3 and FIG. 15 to FIG. 35, the full-input matrix FXM may include a plurality of input matrices XM. In other words, the full-input matrix FXM is tiled with a plurality of input matrices XM arranged in row and column directions. For easier explanation, hereinafter, an input matrix arranged in the (i)th row and (j)th column of the full-input matrix FXM is referred to as "XM_ij".
[0406] In one embodiment, the input matrix XM described above with reference to Figures 1 to 3 and Figures 15 to 33 may be one of a plurality of input matrices XM included in the full-input matrix FXM.
[0407] In one embodiment, each of the input matrices XM included in the full-input matrix FXM may have the same row size and column size. For example, each of the input matrices XM may include 'n' input elements per row. Each of the input matrices XM may include 'h' input elements per column.
[0408] The row size of the full-input matrix FXM may be an integer multiple of the row size of each of the multiple input matrices XM. For example, one row of the full-input matrix FXM may include 'N' input elements, where 'N' may be an integer multiple of 'n'.
[0409] The column size of the full-input matrix FXM may be an integer multiple of the column size of each of the multiple input matrices XM. For example, one column of the full-input matrix FXM may include 'H' input elements, where 'H' may be an integer multiple of 'h'.
[0410] FIG. 36 is a diagram showing the full-weight matrix of FIG. 34. Referring to FIG. 1 to FIG. 3 and FIG. 15 to FIG. 36, the full-weight matrix FWM may include a plurality of weight matrices WM. In other words, the full-weight matrix FWM is tiled with a plurality of weight matrices WM arranged in the row direction and the column direction. For a simpler explanation, the weight matrix arranged in the (i)th row and (j)th column of the full-weight matrix FWM will be referred to as "WM_ij" below.
[0411] In one embodiment, the weight matrix WM previously described with reference to Figures 1 to 3 and Figures 15 to 33 may be one of a plurality of weight matrices WM included in the full-input matrix FXM.
[0412] In one embodiment, each of the weight matrices WM included in the full-weight matrix FWM may have the same row size and column size. For example, each of the weight matrices WM may include 'm' weights for each row. Each of the weight matrices WM may include 'n' weights for each column.
[0413] The row size of the full-weight matrix FWM may be an integer multiple of the row size of each of the weight matrices WM. For example, one row of the full-weight matrix FWM may include 'M' weights. In this case, 'M' may be an integer multiple of 'm'.
[0414] The column size of the full-weight matrix FWM may be an integer multiple of the row size of each of the weight matrices WM. For example, one column of the full-weight matrix FWM may include 'N' weights. In this case, 'N' may be an integer multiple of 'n'.
[0415] The BCQ circuit 200 may perform a binary coding quantization operation on each of a plurality of weight matrices WM included in the full-weight matrix FWM. In this case, the quantization scale coefficient generated based on the weight matrix WM_11 may be different from the quantization scale coefficient generated based on the weight matrix WM_12. Similarly, the quantization code value generated based on the weight matrix WM_11 may be different from the quantization code value generated based on the weight matrix WM_12. A specific method in which the BCQ circuit 200 performs a binary coding quantization operation on each weight matrix WM is similar to that described above with reference to Figs. 15 to 17, and therefore a detailed description thereof will be omitted.
[0416] Figure 37 is a diagram showing the full-output matrix of Figure 34. Referring to Figures 1-3 and 15-37, the full-output matrix FYM may correspond to the product of the full-input matrix FXM and the full-weight matrix FWM.
[0417] The full-output matrix FYM may include multiple sub-matrices FYM_sub arranged in row and column directions. For easier explanation, hereinafter, the sub-matrix arranged in the (i)th row and (j)th column of the full-output matrix FYM is referred to as “FYM_sub_ij”.
[0418] Each of the sub-matrices FYM_sub may have the same row size and column size as each other. The row size of each of the sub-matrices FYM_sub may be the same as the row size of the input matrix XM. The column size of each of the sub-matrices FYM_sub may be the same as the column size of the weight matrix WM. For example, each of the sub-matrices FYM_sub may include 'm' output elements per row. Each of the sub-matrices FYM_sub may include 'h' output elements per column.
[0419] The row size of the full-output matrix FYM may be the same as the row size of the full-weight matrix FWM, for example, the row size of the full-output matrix FYM may be 'M'.
[0420] The column size of the full-output matrix FYM may be the same as the column size of the full-input matrix FXM, for example, the column size of the full-output matrix FYM may be 'H'.
[0421] The matrix multiplication device MMD can calculate the full-output matrix FYM in units of sub-matrix FYM_sub. For example, the matrix multiplication device MMD can calculate one sub-matrix FYM_sub by adding the products of a plurality of tiled input matrices XM and a plurality of tiled weight matrices WM.
[0422] To give a more detailed example, when 'N' is three times 'n', the matrix multiplier 100 can calculate the sub-matrix FYM_sub_11 by sequentially calculating the product of the input matrix XM_11 and the weight matrix WM_11, the product of the input matrix XM_12 and the weight matrix WM_21, and the product of the input matrix XM_13 and the weight matrix WM_31, and then adding them. In this case, each product of the tiled input matrix and the tiled weight matrix may correspond to the output matrix YM previously described with reference to Figures 1 to 3 and Figures 15 to 33. That is, the matrix multiplication device MMD can calculate a first output matrix based on the product of the input matrix XM_11 and the weight matrix WM_11, calculate a second output matrix based on the product of the input matrix XM_12 and the weight matrix WM_21, and calculate a third output matrix based on the product of the input matrix XM_13 and the weight matrix WM_31. Thereafter, the matrix multiplication device MMD can add the first to third output matrices described above to calculate the sub-matrix FYM_sub_11, although the scope of the present disclosure is not limited thereto.
[0423] That is, the matrix multiplication device MMD can be implemented to accumulate a plurality of output matrices to calculate one sub-matrix (i.e., a portion of the full-output matrix FYM). For example, the matrix multiplication device MMD can be implemented to temporarily store a plurality of output matrices in an external volatile memory device (e.g., an SRAM device) and then accumulate the output matrices to calculate one sub-matrix. In this manner, the matrix multiplication device MMD can calculate the full-output matrix FYM by sequentially calculating a plurality of sub-matrices FYM_sub.
[0424] 38 is a block diagram showing a matrix multiplication device according to an embodiment. Referring to FIG. 38, the matrix multiplication device MMD may include a matrix multiplier array 1000, a BCQ circuit 1200, and an output matrix accumulator 3000.
[0425] The BCQ circuit 1200 can receive a weight matrix WM. The BCQ circuit 1200 can perform binary coding quantization on the weight matrix WM at a BCQ resolution of 'R' to generate a plurality of quantization code matrices QSM and a plurality of quantization scale coefficients QSC. The operation of the BCQ circuit 1200 is similar to that of the BCQ circuit 200 previously described with reference to Figures 1 to 3 and Figures 15 to 33, and therefore a detailed description thereof will be omitted.
[0426] For the sake of simpler explanation, the first quantization code matrix QSM_1 through the (R)th quantization code matrix QSM_R and the corresponding quantization scale coefficients are referred to as the first plurality of quantization scale coefficients QSCs_1 through the (R)th plurality of quantization scale coefficients QSCs_R, respectively. For example, the first plurality of quantization scale coefficients QSCs_1 may refer to the quantization scale coefficients previously shown in FIG. 24 with diagonal stripes. Similarly, the second plurality of quantization scale coefficients QSCs_2 may refer to the quantization scale coefficients previously shown in FIG. 24 with dot-patterns.
[0427] The matrix multiplier array 1000 can include a first matrix multiplier 1110 through an (R)th matrix multiplier 11R0.
[0428] Each of the first matrix multiplier 1110 to the (R)th matrix multiplier 11R0 can receive an input matrix XM. That is, each of the first matrix multiplier 1110 to the (R)th matrix multiplier 11R0 can receive the same input matrix XM.
[0429] The first matrix multiplier 1110 through the (R)th matrix multiplier 11R0 can receive the first plurality of quantization scale coefficients QSCs_1 through the (R)th plurality of quantization scale coefficients QSCs_R, respectively. The first matrix multiplier 1110 through the (R)th matrix multiplier 11R0 can receive the first quantization code matrix QSM_1 through the (R)th quantization code matrix QSM_R, respectively.
[0430] Each of the first matrix multiplier 1110 to the (R)th matrix multiplier 11R0 can be realized in a manner similar to that previously described with reference to Figures 1 to 3 and 15 to 33. For example, the first matrix multiplier 1110 can generate the first sub-output matrix YM_sub_1 by scaling the input matrix XM based on a first plurality of quantization scale coefficients QSCs_1 and then accumulating the scaled matrix based on a first quantization code matrix QSM_1. In this manner, the first matrix multiplier 1110 to the (R)th matrix multiplier 11R0 can generate the first sub-output matrix YM_sub_1 to the (R)th sub-output matrix YM_sub_R, respectively.
[0431] The output matrix accumulator 1300 may receive the first sub-output matrix YM_sub_1 through the (R)th sub-output matrix YM_sub_R. The output matrix accumulator 1300 may add the first sub-output matrix YM_sub_1 through the (R)th sub-output matrix YM_sub_R to generate the output matrix YM.
[0432] That is, according to the embodiment of Fig. 35, a plurality of matrix multipliers generate sub-output matrices based on different quantization code matrices, and then the output matrix is calculated by accumulating the sub-output matrices. In this case, even if the weight matrix WM is binary-coded quantized at a high BCQ resolution, a plurality of matrix multipliers calculate sub-output matrices corresponding to different quantization code matrices in parallel, so that the operating speed of the matrix multiplication device MMD can be improved.
[0433] For the sake of simpler explanation, an embodiment in which the matrix multiplier array 1000 includes matrix multipliers whose number corresponds to the BCQ resolution (i.e., 'R') is representatively described in Fig. 35, but the scope of the present disclosure is not limited thereto. For example, each matrix multiplier can be realized to operate based on multiple quantization code matrices, similar to those previously described with reference to Figs. 1 to 3 and Figs. 15 to 33.
[0434] Figure 39 is a block diagram showing a neural processing system implemented according to an embodiment. Referring to Figure 39, the neural processing system 2000 may include a central processing unit 2100, a neural processing unit 2200, a volatile memory device 2300, a nonvolatile memory device 2400, and a user interface 2500. The central processing unit 2100, the neural processing unit 2200, the volatile memory device 2300, the nonvolatile memory device 2400, and the user interface 2500 may be connected via a bus.
[0435] The central processing unit 2100 can control the overall operation of the neural processing system 2000. For example, the central processing unit 2100 can control each component of the neural processing system 2000 to drive an artificial intelligence model.
[0436] In one embodiment, the artificial intelligence model implemented by the neural processing system 2000 may be any type of artificial intelligence model, such as a language model, an image identification model, an image generation model, a weather analysis model, etc. For example, the artificial intelligence model implemented by the neural processing system 2000 may be any type of artificial intelligence model, such as GPT-3, GPT-4, Pangu, GShard, Megatron-LM, etc. However, the scope of the present disclosure is not limited thereto.
[0437] In one embodiment, the artificial intelligence models executed by the neural processing system 2000 are capable of performing inference and / or training operations, although the scope of the present disclosure is not limited in this respect.
[0438] Each of the artificial intelligence models may include a number of processing layers. Each of the multiple processing layers may be implemented to receive layer input data and generate layer output data. In this case, the generated layer output data may be used as layer input data for another processing layer. For example, the generated layer output data from a first processing layer may be used as layer input data for a second processing layer. A more detailed description of the artificial intelligence models and processing layers is provided with reference to FIG. 40 below.
[0439] Each of the multiple processing layers may transform layer input data into layer output data based on a matrix multiplication operation. For example, each of the multiple processing layers may multiply an input matrix corresponding to the layer input data by a weight matrix to generate an output matrix corresponding to the layer output data. However, the scope of the present disclosure is not limited thereto, and each of the multiple processing layers may transform an input matrix corresponding to the layer input data in an arbitrary manner to generate output data. For example, each of the multiple processing layers may be realized to sequentially multiply an input matrix corresponding to the layer input data by a plurality of weight matrices to generate layer output data, or to transform an input matrix into layer output data based on an arbitrary transformation parameter. That is, the scope of the present disclosure is not limited to a specific manner in which each of the multiple processing layers transforms layer input data.
[0440] The neural processing unit 2200 may include a matrix multiplication device 2210. The matrix multiplication device 2210 may perform at least some of the operations included in the multiple processing layers. For example, the matrix multiplication device 2210 may perform matrix multiplication operations included in the multiple processing layers.
[0441] In one embodiment, matrix multiplication operations may account for a large portion of the processing load used by the neural processing system 2000 to execute each of the multiple processing layers.
[0442] In one embodiment, the matrix multiplication device 2210 can be realized by the matrix multiplication device MMD described above with reference to Figures 1 to 38. In this case, the neural processing unit 2200 can execute operations included in multiple processing layers with a smaller amount of calculations. Therefore, the neural processing system 2000 including the matrix multiplication device MMD according to the embodiment of the present disclosure can run an artificial intelligence model at a faster speed.
[0443] The volatile memory device 2300 can be used as an operating memory for the neural processing unit 2200. For example, the volatile memory device 2300 can temporarily store data generated during the operation of the neural processing unit 2200.
[0444] In one embodiment, the neural processing unit 2200 can access the volatile memory device 2300 to perform operations included in the multiple processing layers. For example, the neural processing unit 2200 can be implemented to read parameters stored in the volatile memory device 2300 and perform operations on layer input data, or to temporarily store intermediate data generated during operations in the volatile memory device 2300.
[0445] In one embodiment, the calculation speed of the neural processing unit 2200 may be faster than the access speed of the neural processing unit 2200 to the volatile memory device 2300. As a result, a bottleneck phenomenon may occur in the operation speed of the artificial intelligence model due to the communication speed between the neural processing unit 2200 and the volatile memory device 2300.
[0446] In one embodiment, when the matrix multiplication device 2210 is realized by the matrix multiplication device MMD previously described with reference to FIGS. 1 to 38, each of the processing elements PE included in the matrix multiplier MMD can calculate one output element. In this case, the neural processing unit 2200 temporarily stores the partial sum (PSUM) in the volatile memory device 2300, and then the output element is calculated without re-reading the stored partial sum (PSUM) to calculate the output element, thereby minimizing the number of accesses to the volatile memory device 2300 of the neural processing unit 2200. Therefore, according to the embodiment of the present disclosure, a bottleneck phenomenon in the operation speed of the artificial intelligence model caused by the access to the volatile memory device 2300 of the neural processing unit 2200 is minimized.
[0447] In one embodiment, when the matrix multiplication device 2210 is implemented as the matrix multiplication device MMD previously described with reference to Figures 1 to 39, each processing element PE can operate one output element. In this case, the artificial intelligence model can be driven by a smaller size processing element array 130 (i.e., a smaller number of processing elements PE). Therefore, according to the embodiment of the present disclosure, the production cost of the neural processing system 2000 for driving the artificial intelligence model can be minimized.
[0448] In one embodiment, the volatile memory device 2300 may be implemented with any type of volatile memory, such as dynamic random access memory (DRAM) or static random access memory (SRAM).
[0449] In one embodiment, the volatile memory device 2300 may be used as a buffer memory, working memory, or cache memory for the central processing unit 2100. However, the scope of the present disclosure is not limited in this respect.
[0450] The non-volatile memory device 2400 may store data for operation of the neural processing system 2000. For example, the non-volatile memory device 2400 may store various types of data, such as parameters for driving an operating system (OS) or an artificial intelligence model of the neural processing system 2000. However, the scope of the present disclosure is not limited thereto.
[0451] The central processing unit 2100 may communicate with a user through a user interface 2500. The central processing unit 2100 may provide model input data provided by a user through the user interface 2500 to the volatile memory device 2300 or the neural processing unit 2200. The central processing unit 2100 may return model output data generated by the artificial intelligence model based on the model input data to the user through the user interface 2500.
[0452] Fig. 40 is a block diagram showing an artificial intelligence model driven by the neural processing system of Fig. 39. Referring to Figs. 39-40, the neural processing system 2000 can drive an artificial intelligence model AIM.
[0453] The artificial intelligence model AIM can receive the model input data MID. The artificial intelligence model AIM can include a first processing layer PL_1 to an L-th processing layer PL_L.
[0454] The artificial intelligence model AIM may sequentially convert the model input data MID through the first processing layer PL_1 to the L-th processing layer PL_L to generate the model output data MOD. For example, the first processing layer PL_1 may receive the model input data MID and generate the second layer input data LID_2. The second processing layer PL_2 may receive the second layer input data LID_2 and generate the third layer input data LID_3. In this manner, the L-th processing layer PL_L may receive the L-th layer input data LID_L and generate the model output data MOD.
[0455] Each of the first processing layer PL_1 to the Lth processing layer PL_L may convert received data into output data through various types of operations. For example, the operations performed by the first processing layer PL_1 to convert the model input data MID into the second layer input data LID_2 include a matrix multiplication operation. Similarly, each of the first processing layer PL_1 to the Lth processing layer PL_L may need to perform a matrix multiplication operation to convert the received layer input data. However, the scope of the present disclosure is not limited thereto, and some of the first processing layer PL_1 to the Lth processing layer PL_L may not need to perform a matrix multiplication operation.
[0456] In one embodiment, the matrix multiplication operations performed by each of the first processing layer PL_1 to the L-th processing layer PL_L may be performed through a matrix multiplication device 2210.
[0457] In one embodiment, when the matrix multiplication device 2210 is implemented by the matrix multiplication device MMD previously described with reference to Figures 1 to 39, the matrix multiplication device 2210 can output the result of the matrix multiplication operation at a faster speed. Therefore, according to the embodiment of the present disclosure, the operating speed of the artificial intelligence model AIM can be improved.
[0458] For the sake of simpler explanation, FIG. 40 representatively describes an embodiment in which the artificial intelligence model AIM is composed of a plurality of processing layers operating in series, but the scope of the present disclosure is not limited thereto. For example, the artificial intelligence model AIM may further include a processing layer that operates in parallel with at least a part of the above-mentioned first processing layer PL_1 to Lth processing layer PL_L. That is, the scope of the present disclosure is not limited to a specific implementation method of the artificial intelligence model AIM.
[0459] The above is a specific embodiment for carrying out the present disclosure. The present disclosure includes not only the above-mentioned embodiment, but also embodiments that can be simply modified or easily modified. The present disclosure also includes techniques that can be easily modified and carried out using the embodiment. Therefore, the scope of the present disclosure should not be limited to the above-mentioned embodiment, but should be determined not only by the claims described below but also by equivalents to the claims of the present disclosure. [Explanation of symbols]
[0460] MMD: Matrix Multiplication Device XM: Input Matrix WM: Weight matrix YM: Output matrix 100: Matrix multiplier 200:BCQ circuit
Claims
1. an input vector scaler that generates a first scaled input vector based on the first input vector and a plurality of quantization scale factors; a first data type converter that generates a first fixed-point scaled input vector based on the first scaled input vector; a processing element array including a first processing element that generates first fixed point output elements based on the first fixed point scaled input vector and a first plurality of quantization code values, and a second processing element that generates second fixed point output elements based on the first fixed point scaled input vector and a second plurality of quantization code values; and a second data type converter that converts data types of the first and second fixed-point output elements to generate first and second output elements, respectively, and outputs a first output vector including the first and second output elements.
2. The input vector scaler comprises: further configured to generate a second scaled input vector based on the second input vector and the plurality of quantization scale factors; The first data type converter: further configured to generate a second fixed-point scaled input vector based on the second scaled input vector; The processing element array comprises: a third processing element generating a third fixed point output element based on the second fixed point scaled input vector and the first plurality of quantization code values, and a fourth processing element generating a fourth fixed point output element based on the second fixed point scaled input vector and the second plurality of quantization code values; The second data type converter:
2. The matrix multiplier of claim 1 , further configured to convert data types of the third and fourth fixed-point output elements to generate third and fourth output elements, respectively, and output a second output vector comprising the third and fourth output elements.
3. the first processing element and the second processing element are arranged in a first processing element row of the processing element array; 3. The matrix multiplier of claim 2, wherein the third and fourth processing elements are disposed in a second processing element row of the array of processing elements.
4. the first processing element and the third processing element are disposed in a first processing element column of the processing element array; 3. The matrix multiplier of claim 2, wherein the second processing element and the fourth processing element are disposed in a second processing element column of the array of processing elements.
5. The dimension of the first fixed-point scaled input vector is R times the dimension of the first input vector, where R is an integer equal to or greater than 2.
3. The matrix multiplier of claim 2, wherein the dimension of the second fixed-point scaled input vector is R times the dimension of the second input vector.
6. 6. The matrix multiplier of claim 5, wherein a dimension of the first input vector, a dimension of the second input vector, a dimension of the first output vector, and a dimension of the second output vector are identical to one another.
7. The first data type converter: a first exponent extraction circuit for extracting a first exponent that is greatest among the exponents of each of a first plurality of scaled input elements included in the first scaled input vector; a second exponent extraction circuit for extracting a second exponent that is largest among the exponents of each of a second plurality of scaled input elements included in the second scaled input vector; a first data type conversion circuit for converting a data type of each of the first plurality of scaled input elements to fixed point based on the first exponent to generate the first fixed-point scaled input vector; and 3. The matrix multiplier of claim 2, further comprising a second data type conversion circuit for converting a data type of each of the second plurality of scaled input elements to fixed point based on the second exponent to generate the second fixed-point scaled input vector.
8. The second data type converter: converting the data types of the first fixed-point output element and the second fixed-point output element to floating point based on the first exponent; 8. The matrix multiplier of claim 7, further configured to convert a data type of the third fixed-point output element and the fourth fixed-point output element to floating point based on the second exponent.
9. the exponent part of the first output element and the second output element corresponds to the first exponent; and 9. The matrix multiplier of claim 8, wherein the exponent portions of the third and fourth output elements correspond to the second exponent.
10. The first processing element comprises: configured to accumulate products of a first plurality of fixed-point scaled input elements included in the first fixed-point scaled input vector and each of the first plurality of quantization code values to generate the first fixed-point output element; The second processing element comprises:
3. The matrix multiplier of claim 2 configured to accumulate products of a second plurality of fixed-point scaled input elements included in the second fixed-point scaled input vector and each of the second plurality of quantization code values to generate the second fixed-point output element.
11. The first processing element comprises: an accumulation register; and an arithmetic and logic unit (ALU) including a first input terminal for sequentially receiving a first plurality of scaled input elements, a second input terminal for sequentially receiving the first plurality of quantization code values, and a third input terminal coupled to the accumulation register; The arithmetic logic unit:
11. The matrix multiplier of claim 10, configured to update the accumulation register via the third input terminal with a value obtained by adding a product of one scaled input element received via the first input terminal and one quantization code value received via the second input terminal.
12. an input vector scaler generating a first plurality of scaled input elements based on the first input elements and the first plurality of quantized scale factors and generating a second plurality of scaled input elements based on the second input elements and the second plurality of quantized scale factors; a first data type converter generating a first plurality of fixed-point scaled input elements based on the first plurality of scaled input elements and generating a second plurality of fixed-point scaled input elements based on the second plurality of scaled input elements; a first processing element that accumulates the first plurality of fixed-point scaled input elements and a second plurality of fixed-point scaled input elements based on a plurality of quantization code values to generate a first fixed-point output element; and a second data type converter for converting a data type of the first fixed-point output element to generate a first output element;
13. 13. The matrix multiplier of claim 12, wherein a number of the first plurality of quantized scale coefficients, a number of the second plurality of quantized scale coefficients, a number of the first plurality of scaled input elements, and a number of the second plurality of scaled input elements are the same as one another.
14. The first plurality of scaled input elements: each of the first plurality of quantization scale factors corresponds to a multiplication of the first input element; The second plurality of scaled input elements:
14. The matrix multiplier of claim 13, wherein each of the second plurality of quantization scale factors corresponds to a multiplication of the second input element.
15. a data type of each of the first input element and the second input element and the first plurality of quantization scale coefficients and the second plurality of quantization scale coefficients is floating point; 13. The matrix multiplier of claim 12, wherein the data types of the first plurality of fixed-point scaled input elements and the second plurality of scaled input elements are fixed-point corresponding to the same place value.
16. The first data type converter: an exponent extraction circuit for extracting a first exponent that is largest among the exponents of the first plurality of scaled input elements and the second plurality of scaled input elements; 16. The matrix multiplier of claim 15, further comprising a data type conversion circuit that converts a data type of each of the first plurality of scaled input elements and the second plurality of scaled input elements to fixed point based on the first exponent.
17. The second data type converter:
17. The matrix multiplier of claim 16 configured to determine an exponent portion of the first output element based on the first exponent.
18. 1. A method of operating a matrix multiplication apparatus, comprising: receiving first through N-th weights from an external device; (where N is an integer equal to or greater than 2); performing binary coding quantization (BCQ) on the first through N-th weights to generate first through (N×R)-th quantized code values and first through (N×R)-th quantized scale coefficients; where R is an integer equal to or greater than 2; receiving first through Nth input elements from the external device; scaling the first through Nth input elements based on the first through (N×R)th quantization scale factors to generate first through (N×R)th scaled input elements; and outputting a first output element generated by accumulating the first through (N×R)th scaled input elements based on the first through (N×R)th quantization code values.
19. The step of outputting a first output element comprises: converting data types of the first through (N×R)th scaled input elements into fixed point; accumulating the first through (N×R)th scaled input elements converted to fixed-point format based on the first through (N×R)th quantization code values to generate a first fixed-point output element; and 20. The method of claim 18, further comprising converting a data type of the first fixed-point output element to a floating-point data type to output the first output element.
20. The first output element:
20. The method of claim 18, wherein the first N×R scaled input elements are multiplied by first N×R quantization code values, respectively.