Matrix multiplier and operation method of matrix multiplication device including the same

By generating fixed point quantization input vectors in the matrix multiplier and using process element arrays for quantization encoding and quantization, the problem of matrix multiplication calculation speed and calculation amount in the prior art is solved, and fast calculation and efficient calculation are achieved.

JP2025073115APending Publication Date: 2025-05-12SAMSUNG ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024187872
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-08
Filing Date
2024-10-25
Publication Date
2025-05-12

AI Technical Summary

Technical Problem

The existing technology has a large calculation speed and calculation amount when performing matrix multiplication, which is difficult to meet the needs of fast computing of artificial intelligence models.

Method used

A matrix multiplier is used to generate input vectors quantized by fixed point quantization, quantization is performed using an array of processing elements for quantization, reducing the amount of calculation, and quantization is performed through a general quantizer to improve calculation efficiency.

Benefits of technology

The rapid calculation of matrix multiplication is realized, which reduces the calculation amount and increases the computing speed of matrix multiplier, which is suitable for the fast computing requirements of artificial intelligence models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025073115000001_ABST
    Figure 2025073115000001_ABST
Patent Text Reader

Abstract

To provide a matrix multiplier and a matrix multiplication device configured to perform matrix multiplication with a faster speed and with a smaller computation amount.SOLUTION: A matrix multiplier includes: an input vector scaler for generating a scaled input vector based on an input vector and a plurality of quantization scale coefficients; a first material type converter for generating a fixed point scaled input vector based on the scaled input vector; a processing element array comprising a processing element for generating first and second fixed point output elements based on the fixed point scaled input vector and a plurality of quantization sign values; and a second material type converter configured to convert material types of the first and second fixed point output elements to generate first and second output elements and output an output vector including the first and second output elements.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a semiconductor device, and more particularly to a matrix multiplier that performs matrix multiplication, and a matrix multiplication device including the same. [Background technology]

[0002] Recently, as artificial intelligence technology has developed, the amount of calculation required for an artificial intelligence model has increased dramatically. As a result, various techniques for shortening the operation time of an artificial intelligence model have been researched.

[0003] Generally, most of the operation time of an AI model is used for matrix multiplication. For example, an AI model uses most of its operation time for multiplying an input matrix and a weight matrix to calculate an output matrix. As a result, various algorithms such as BCQ (Binary Coding Quantization) are being researched to multiply an input matrix and a weight matrix with less computational effort. Summary of the Invention [Problem to be solved by the invention]

[0004] The present disclosure is intended to solve the above-mentioned technical problems. More specifically, an object of the present disclosure is to provide a matrix multiplier configured to perform matrix multiplication at a higher speed and with a smaller amount of calculations, and a matrix multiplication device including the same. [Means for solving the problem]

[0005] A matrix multiplier according to an embodiment of the present disclosure may include an input vector scaler that generates a first quantized scaled input vector based on a first input vector, a plurality of common scale factors, and first through Rth scale factor scale factors; a first data type converter that generates a first fixed-point quantized scaled input vector based on the first quantized scaled input vector; a processing element array including a first processing element that generates a first fixed-point output element based on the first fixed-point quantized scaled input vector and a first plurality of quantized code bits, and a second processing element that generates a second fixed-point output element based on the first fixed-point quantized scaled input vector and a second plurality of quantized code bits; and a second data type converter that converts data types of the first and second fixed-point output elements to generate first and second output elements, respectively, and outputs a first output vector including the first and second output elements.

[0006] A method of operating a matrix multiplication device according to an embodiment of the present disclosure may include the steps of receiving first to Nth weights from an external device; performing uniform binary coding quantization on the first to Nth weights to generate first to Nth common scale coefficients, first to Rth magnification scale coefficients, and first to (N×R)th quantization code bits; receiving first to Nth input elements from the external device; quantizing and scaling the first to Nth input elements based on the first to Nth common scale coefficients and the first to Rth magnification scale coefficients to generate first to (N×R)th quantized scaled input elements; and outputting a first output element generated based on the first to (N×R) quantized code bits and the first to (N×R)th quantized scaled input elements.

[0007] A matrix multiplier according to an embodiment of the present disclosure may include an input vector scaler that generates a first-scaled input vector based on a first input vector and first through Rth scale factors; a first data type converter that generates a first fixed-point scaled input vector based on the first-scaled input vector; a processing element array including a first processing element that generates a first fixed-point partial product based on the first fixed-point scaled input vector and a first plurality of quantization code bits; a second data type converter that converts a data type of the first fixed-point partial product to generate a first partial product; and a common scaler that generates a first output element based on a product of the first partial product and a first common scale factor, and outputs a first output vector including the first output elements. Effect of the Invention

[0008] According to the embodiments of the present disclosure, the amount of calculation of the matrix multiplier can be reduced. [Brief description of the drawings]

[0009] [Figure 1] 1 is a block diagram illustrating a matrix multiplication device according to an embodiment of the present disclosure. [Diagram 2] 2 illustrates the operation of a matrix multiplication device implemented to directly multiply the input matrix and weight matrix of FIG. 1; [Diagram 3] FIG. 2 is a diagram illustrating the operation of the uniform BCQ circuit of FIG. [Figure 4] FIG. 2 is a diagram illustrating the operation of the uniform BCQ circuit of FIG. [Diagram 5] FIG. 2 is a diagram illustrating the operation of the uniform BCQ circuit of FIG. 1, which performs binary coding quantization operations by columns of a weight matrix. [Figure 6] FIG. 6 illustrates the operation of a matrix multiplier according to the embodiment of FIG. 5. [Figure 7] 5 is a block diagram showing a configuration of the matrix multiplier of FIG. 1 according to the embodiment of FIG. 4. [Figure 8] FIG. 8 is a block diagram showing the configuration of an input vector scaler in FIG. [Figure 9] FIG. 9 illustrates in more detail the operation of the magnification scaling circuit of FIG. 8. [Figure 10] 9 illustrates in more detail how the magnification scaling circuit of FIG. 8 performs a magnification scaling operation. [Figure 11] FIG. 8 is a diagram showing the configuration of the first data-type converter of FIG. 7. [Figure 12] FIG. 12 is a diagram illustrating the operation of the exponent extraction circuit of FIG. [Figure 13] 12 is a diagram illustrating the operation of the material type conversion circuit of FIG. 11. [Figure 14] FIG. 8 is a block diagram showing in more detail the configuration of the processing element array of FIG. 7. [Figure 15] FIG. 8 is a block diagram showing in more detail a portion of the operation of the matrix multiplier of FIG. 7. [Figure 16] 16 is a block diagram showing in more detail the operation of the first processing element row of FIG. 15. [Figure 17] FIG. 17 illustrates a configuration of one of the processing elements of FIG. 16 implemented in accordance with one embodiment. [Figure 18] FIG. 8 is a diagram illustrating the operation of the second data-type converter of FIG. [Figure 19] 2 is a flowchart showing the operation of the matrix multiplication device of FIG. 1; [Figure 20] 20 is a flowchart showing step S150 of FIG. 19 in more detail. [Figure 21] 2 is a flowchart showing the operation of the matrix multiplication device of FIG. 1; [Figure 22] 22 is a flowchart showing step S250 of FIG. 21 in more detail. [Figure 23] FIG. 2 illustrates the operation of the BCQ circuit of FIG. 1 in accordance with one disclosed embodiment. [Figure 24] FIG. 24 illustrates a weight matrix approximated by the embodiment of FIG. 23. [Diagram 25]FIG. 25 is a diagram showing the quantization code bit matrix of FIG. 24. [Figure 26] 2 is a block diagram showing a configuration of a matrix multiplier of FIG. 1 implemented according to an embodiment of the present disclosure. [Figure 27] FIG. 27 is a block diagram illustrating the configuration of the input vector scaler of FIG. 26 in accordance with one embodiment. [Figure 28] FIG. 28 illustrates the operation of the common scaling circuit of FIG. 27 in more detail. [Figure 29] FIG. 27 is a block diagram illustrating the configuration of the input vector scaler of FIG. 26 in accordance with one embodiment. [Diagram 30] FIG. 30 illustrates the operation of the common scaling circuit of FIG. 29 in more detail. [Diagram 31] FIG. 30 illustrates in more detail the operation of the magnification scaling circuit of FIG. 29. [Diagram 32] FIG. 27 is a block diagram illustrating the configuration of the input vector scaler of FIG. 26 in accordance with one embodiment. [Diagram 33] FIG. 33 illustrates in more detail the operation of the magnification scaling circuit of FIG. 32. [Diagram 34] FIG. 33 illustrates the operation of the quantization scaling circuit of FIG. 32 in more detail. [Diagram 35] FIG. 27 is a block diagram showing the configuration of the first data-type converter of FIG. 26. [Diagram 36] FIG. 27 is a block diagram showing in more detail the configuration of the processing element array of FIG. 26. [Figure 37] FIG. 37 is a block diagram showing in more detail the operation of the processing element of FIG. 36. [Figure 38] FIG. 37 illustrates a configuration of one of the processing elements of FIG. 36 implemented in accordance with one embodiment. [Figure 39] FIG. 27 is a diagram illustrating the operation of the second data-type converter of FIG. 26. [Diagram 40] 2 is a flowchart showing the operation of the matrix multiplication device of FIG. 1; [Diagram 41] 41 is a flowchart showing step S350 of FIG. 40 in more detail. [Diagram 42] 2 is a flowchart showing the operation of the matrix multiplication device of FIG. 1; [Diagram 43] 43 is a flowchart showing in more detail step S440 of FIG. 42 implemented according to one embodiment. [Diagram 44] 43 is a flowchart showing in more detail step S440 of FIG. 42 implemented according to one embodiment. [Diagram 45] 43 is a flowchart showing in more detail step S440 of FIG. 42 implemented according to one embodiment. [Figure 46] 43 is a flowchart showing step S450 of FIG. 42 in more detail. [Figure 47] FIG. 2 illustrates the operation of the BCQ circuit of FIG. 1 according to one embodiment. [Figure 48] FIG. 2 is a block diagram showing a configuration of the matrix multiplier of FIG. 1 implemented according to one embodiment. [Figure 49] FIG. 49 is a block diagram illustrating the configuration of the input vector scaler of FIG. 48 in accordance with one embodiment. [Figure 50] FIG. 50 illustrates in more detail the operation of the magnification scaling circuit of FIG. 49. [Figure 51] FIG. 50 illustrates in more detail the operation of the common scaling circuit of FIG. 49. [Figure 52] FIG. 49 is a block diagram illustrating the configuration of the input vector scaler of FIG. 48 in accordance with one embodiment. [Diagram 53] FIG. 53 illustrates in more detail the operation of the common scaling circuit of FIG. 52. [Figure 54] FIG. 53 illustrates in more detail the operation of the magnification scaling circuit of FIG. 52 in accordance with one embodiment. [Figure 55] FIG. 49 is a block diagram illustrating the configuration of the input vector scaler of FIG. 48 in accordance with one embodiment. [Figure 56] FIG. 56 illustrates in more detail the operation of the magnification scaling circuit of FIG. 55. [Figure 57] FIG. 56 illustrates in more detail the operation of the quantization scaling circuit of FIG. 55. [Figure 58] FIG. 49 illustrates in more detail the operation of the processing element array of FIG. 48. [Figure 59] FIG. 27 is a block diagram showing the processing element array of FIG. 26 implemented in a systolic array format. [Figure 60] FIG. 60 is a diagram showing the configuration of the processing element in FIG. 59 in more detail. [Figure 61] 2 illustrates the operation of the matrix multiplication unit of FIG. 1 according to one embodiment. [Figure 62] FIG. 62 shows the full input matrix of FIG. 61. [Figure 63] FIG. 62 shows the full-weight matrix of FIG. 61. [Figure 64] FIG. 62 shows the full output matrix of FIG. 61. [Figure 65] FIG. 1 is a block diagram illustrating a neural processing system implemented in accordance with one embodiment. [Figure 66] FIG. 66 is a block diagram showing an artificial intelligence model driven by the neural processing system of FIG. 65. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] In the following, the embodiments of the present disclosure will be described clearly and in detail to the extent that a person having ordinary skill in the art of the present disclosure can easily carry out the present disclosure. Details such as detailed configurations and structures are provided simply to aid in the overall understanding of the embodiments of the present disclosure. Therefore, modifications of the embodiments described herein can be made by those skilled in the art without departing from the technical idea and scope of the present disclosure. Furthermore, descriptions of well-known functions and structures are omitted for clarity and conciseness. The configurations in the following drawings or detailed description may be connected to other components other than those shown in the drawings or described in the detailed description. The terms used in the present disclosure are defined in consideration of the functions of the present disclosure and are not limited to specific functions. The definitions of the terms are determined based on the matters described in the detailed description.

[0011] The components described with reference to terms such as driver or block used in the detailed description can be realized in the form of software, hardware, or a combination thereof. Exemplarily, the software may be machine code, firmware, embedded code, and application software. For example, the hardware may include electric circuits, electronic circuits, processors, computers, integrated circuit cores, pressure sensors, inertial sensors, MEMS (Micro Electro Mechanical System), manual elements, or a combination thereof.

[0012] In the following, for the sake of conciseness, matrices are referred to through square brackets “[”, “]” and sets are referred to through curly brackets “{”, “}”. However, the scope of the present disclosure is not limited to such notations.

[0013] 1 is a block diagram showing a matrix multiplication device according to an embodiment of the present disclosure. Referring to FIG. 1, the matrix multiplication device MMD may include a matrix multiplier 100 and a uniform BCQ circuit UBC.

[0014] The matrix multiplication device MMD can receive an input matrix XM. The input matrix XM can include a plurality of input vectors. Each of the plurality of input vectors can include a plurality of input elements. For example, the input matrix XM can be expressed by the following Equation 1:

number

number

number

number

number

[0015] For a simpler explanation, the following will representatively describe an embodiment in which the dimension of each of the input vectors included in the input matrix XM is 'n'. That is, the following will representatively describe an embodiment in which each of the input vectors includes 'n' input elements. In other words, the following will representatively describe an embodiment in which the input matrix XM includes 'n' columns.

[0016] In one embodiment, each of the input elements included in the input matrix XM may have a data type of FP16 (16-bits floating point) or FP32 (32-bits floating point), although the scope of this disclosure is not limited in this respect.

[0017] The matrix multiplication device MMD can receive a weight matrix WM. The weight matrix WM can include a plurality of weights. For example, the weight matrix WM can be expressed by the following Equation 2:

number

[0018] In one embodiment, each of the weights included in the weight matrix WM may have a data type of FP16 (16-bits floating point) or FP32 (32-bits floating point), although the scope of the present disclosure is not limited thereto.

[0019] The uniform BCQ circuit UBC can perform uniform binary coding quantization (BCQ) on the weight matrix WM. For example, the uniform BCQ circuit UBC can determine a plurality of common scale coefficients CSC, a plurality of multiplication scale coefficients MSC, and a plurality of quantization sign bits QSB based on the weight matrix WM.

[0020] More specifically, the uniform BCQ circuit UBC can convert each of the weights into a plurality of 'common scale factor CSC-multiplier MSC-quantization sign bit QSB combinations'. For example, the uniform BCQ circuit UBC can approximate each of the weights of the weight matrix WM based on a plurality of 'common scale factor CSC-multiplier MSC-quantization sign bit QSB combinations'. That is, the uniform BCQ circuit UBC can determine for each weight a 'common scale factor CSC-multiplier MSC-quantization sign bit QSB' indicating a quantized value of the weight. This will be explained in more detail below.

[0021] In one embodiment, each of the quantized code bits QSB can represent '0' or '1'.

[0022] In one embodiment, each of the common scale factors CSC can have the same data type as the weight of the weight matrix WM. For example, each of the common scale factors CSC can have a data type of FP16 or FP32. However, the scope of the present disclosure is not limited thereto. The operation of the uniform BCQ circuit UBC will be described in more detail with reference to FIG. 3 below.

[0023] The matrix multiplier 100 can receive a plurality of common scale coefficients CSC, a plurality of magnification scale coefficients MSC, and a plurality of quantization sign bits QSB. The matrix multiplier 100 can perform matrix multiplication on the input matrix XM and the weight matrix WM based on the plurality of common scale coefficients CSC, a plurality of magnification scale coefficients MSC, and a plurality of quantization sign bits QSB. For example, the matrix multiplier 100 can multiply the input matrix XM by an approximated weight matrix WM based on the plurality of common scale coefficients CSC, a plurality of magnification scale coefficients MSC, and a plurality of quantization sign bits QSB to generate an output matrix YM.

[0024] The output matrix YM may include multiple output vectors. Each of the multiple output vectors may include multiple output elements. For example, the output matrix YM may be expressed by the following Equation 3:

number

number

number

[0025] In one embodiment, 'n' and 'm' may be the same integers. For example, the weight matrix WM may be implemented as a square matrix. In this case, the dimension of the output vector included in the output matrix YM may be the same as the dimension of the input vector. However, the scope of the present disclosure is not limited thereto.

[0026] In one embodiment, when the matrix multiplication device MMD directly multiplies the input matrix XM and the weight matrix WM to calculate the output matrix YM, the matrix multiplication device MMD may have to process a very large amount of calculations. In this case, the operation speed of the matrix multiplication device MMD may be reduced. The operation of the matrix multiplication device MMD that directly multiplies the input matrix XM and the weight matrix WM to calculate the output matrix YM will be described in more detail with reference to FIG. 2 below.

[0027] On the other hand, when the matrix multiplication device MMD multiplies the input matrix XM by a weight matrix WM approximated based on a plurality of common scale coefficients CSC, a plurality of magnification scale coefficients MSC, and a plurality of quantization sign bits QSB to calculate the output matrix YM, the amount of calculation of the matrix multiplication device MMD is significantly reduced. The operation of the matrix multiplication device MMD that calculates the output matrix YM based on a plurality of common scale coefficients CSC, a plurality of magnification scale coefficients MSC, and a plurality of quantization sign bits QSB will be described in more detail with reference to the following drawings.

[0028] Fig. 2 is a diagram showing the operation of a matrix multiplication device realized to directly multiply the input matrix and the weight matrix in Fig. 1. Referring to Fig. 1 and Fig. 2, the matrix multiplication device MMD can directly multiply the input matrix XM and the weight matrix WM to calculate the output matrix YM.

[0029] The matrix multiplication device MMD may need to perform 'n' floating point multiplications and then 'n-1' floating point summations to calculate one output element included in the output matrix YM. For example, the matrix multiplication device MMD may need to perform 'n' floating point multiplications and then 'n-1' floating point summations to calculate one output element included in the output matrix YM in a manner similar to the following equation 4: 11) can be calculated.

number

number

number

[0030] 3 and 4 are diagrams showing the operation of the uniform BCQ circuit of FIG 1. First, referring to FIG 1 and FIG 3, the horizontal axis indicates the magnitude of the weights included in the weight matrix WM. Although FIG 3 is shown in a binary tree format, the scope of the present disclosure is not limited thereto.

[0031] The uniform BCQ circuit UBC can approximate one or more of the weights included in the weight matrix WM to a plurality of quantum levels QL (quantum levels).

[0032] The uniform BCQ circuit UBC can determine the number of multiple quantum levels QL based on a predetermined BCQ resolution. For example, the uniform BCQ circuit UBC can set each of the multiple weights to 2 R (where R is the BCQ resolution) quantum levels QL can be approximated. 2 R Each of the quantization levels QL can be determined based on a combination of the zero point value ZPV, the R quantization scale factors QSC, and the R quantization sign bits QSB.

[0033] However, in the following, for a simpler explanation than R, an embodiment in which R is '3' will be representatively described. For example, the uniform BCQ circuit UBC can approximate one or more of the weights included in the weight matrix WM to the first quantum level QL1 to the eighth quantum level QL8. In this case, each of the first quantum level QL1 to the eighth quantum level QL8 can be determined based on the following formula 5.

number

[0034] For example, referring to FIG. 4 again, the first quantum level QL1 may correspond to the case where the first quantized code bit QSB1 to the third quantized code bit QSB3 are all '1'. In this case, the first quantum level QL1 is "ZPV-α 1 -α 2 -α 3 ". Similarly, the eighth quantum level QL8 may correspond to the case where the first quantized code bit QSB1 to the third quantized code bit QSB3 are all '0'. In this case, the eighth quantum level QL8 is "ZPV+α 1 +α 2 +α 3 In this manner, the seventh quantum level QL7 may correspond to the case where the first quantized code bit QSB1 to the third quantized code bit QSB3 are '1', '0', and '0', respectively. In this case, the seventh quantum level QL7 may correspond to "ZPV-α 1 +α 2 +α3 ", however, the scope of the present disclosure is not limited in this respect.

[0035] Continuing with reference to FIG. 3, the intervals between the multiple quantum levels QL may be 'uniform'. That is, the intervals between the multiple quantum levels QL may be the same. In this case, the multiple quantization scale coefficients QSC may form a geometric progression with a common ratio of '2'. For example, the second quantization scale coefficient QSC2 may be twice the first quantization scale coefficient QSC1, and the third quantization scale coefficient QSC3 may be twice the second quantization scale coefficient QSC2.

[0036] Each of the quantization scale coefficients QSC can be expressed as a product of the same common scale coefficient CSC and a different magnification scale coefficient MSC, for example, the first quantization scale coefficient QSC1 can be expressed as a product of the common scale coefficient CSC and the first magnification scale coefficient MSC1, the second quantization scale coefficient QSC2 can be expressed as a product of the common scale coefficient CSC and the second magnification scale coefficient MSC2, and the third quantization scale coefficient QSC3 can be expressed as a product of the common scale coefficient CSC and the third magnification scale coefficient MSC3.

[0037] The multiplication factor scale coefficients MSC corresponding to the multiple quantization scale coefficients QSC may be different consecutive powers of 2. For example, the k-th multiplication factor scale coefficient MSCk is 2 k-1 In a more detailed example, the first magnification scale factor may be 2 0 The second magnification scale factor MSC2 may be 2 1 The third magnification scale factor MSC3 may be 2 2 may be also possible.

[0038] Therefore, 2 REach of the quantum levels QL is determined based on a combination of a zero point value ZPV, one common scale factor CSC, R magnification scale factors MSC, and R quantization code bits QSB. For example, the above-mentioned formula 5 can also be expressed as the following formula 6.

number

[0039] The uniform BCQ circuit UBC can approximate each weight included in the weight matrix WM to a quantum level having a closest value among multiple quantum levels QL. For example, the weight “w 11 When the magnitude of the weight “w 11 " can be approximated to the seventh quantum level QL7. In this case, the uniform BCQ circuit UBC has weight "w 11 " can be expressed as 'a combination of a zero point value ZPV, a common scale coefficient CSC, a plurality of magnification scale coefficients MSC, and a plurality of quantization sign bits QSB' corresponding to the seventh quantum level QL7. Similarly, the uniform BCQ circuit UBC can approximate the weights included in the weight matrix WM based on 'a zero point value ZPV, a plurality of common scale coefficients CSC, a plurality of magnification scale coefficients MSC, and a plurality of quantization sign bits QSB'.

[0040] However, for simpler explanation, an embodiment in which the zero point ZP is '0' will be representatively described below. That is, an embodiment in which the uniform BCQ circuit UBC symmetrically performs uniform binary coding quantization on each of the weights included in the weight matrix WM will be representatively described below. In other words, an embodiment in which the uniform BCQ circuit UBC approximates the weight matrix WM based on a plurality of common scale coefficients CSC, a plurality of magnification scale coefficients MSC, and a plurality of quantization code bits QSB will be described below. The operation of the matrix multiplication device MMD when the zero point ZP is not '0' (i.e., the operation of the matrix multiplication device MMD when the uniform BCQ circuit UBC asymmetrically performs uniform binary coding quantization) will be described in more detail with reference to the following Figures 47 to 58.

[0041] For the sake of simpler explanation, an embodiment in which the BCQ resolution is '3' is representatively described in Fig. 3, but the scope of the present disclosure is not limited thereto. For example, the BCQ resolution can be determined to be an integer equal to or greater than '2' depending on the type of artificial intelligence model driven based on the matrix multiplication device MMD.

[0042] In one embodiment, when the artificial intelligence model driven based on the matrix multiplication device MMD is a large language model (LLM), the BCQ resolution may be '3'. However, the scope of the present disclosure is not limited thereto.

[0043] In one embodiment, when the artificial intelligence model driven based on the matrix multiplication device MMD is an image object identification model, the BCQ resolution may be '2', but the scope of the present disclosure is not limited thereto.

[0044] In one embodiment, the uniform BCQ circuit UBC may perform a uniform binary coding quantization operation for each column of the weight matrix WM. For example, the uniform BCQ circuit UBC may determine a common scale factor CSC different from each other for each column of the weight matrix WM. In this case, the common scale factor CSC for the weights included in the first column of the weight matrix WM may be different from the common scale factor CSC for the weights included in the second column of the weight matrix WM. An embodiment in which the uniform BCQ circuit UBC performs a binary coding quantization operation for each column of the weight matrix WM will be described in more detail with reference to FIGS. 5 to 22 below. However, the scope of the present disclosure is not limited thereto.

[0045] In one embodiment, the uniform BCQ circuit UBC may perform a uniform binary coding quantization operation for each row of the weight matrix WM. For example, the uniform BCQ circuit UBC may determine a common scale factor CSC different from each other for each row of the weight matrix WM. In this case, the common scale factor CSC for the weights included in the first row of the weight matrix WM may be different from the common scale factor CSC for the weights included in the second row of the weight matrix WM. An embodiment in which the uniform BCQ circuit UBC performs a binary coding quantization operation for each row of the weight matrix WM will be described in more detail with reference to the following FIGS. 23 to 58. However, the scope of the present disclosure is not limited thereto.

[0046] 5 is a diagram showing the operation of the uniform BCQ circuit of FIG. 1, which performs a binary coding quantization operation for each column of a weight matrix. Referring to FIG. 1 to FIG. 5, the uniform BCQ circuit UBC can perform a uniform binary coding quantization operation for each column of a weight matrix WM.

[0047] First, the weight matrix WM can be expressed as the following Equation 7.

number

number

number

[0048] The uniform BCQ circuit UBC can perform a binary coding quantization operation for each column of the weight matrix WM based on the following Equation 8.

number

[0049]

number

number

number

number

number

[0050] In one embodiment, Equation 8 can be implemented with circuitry (e.g., an arithmetic logic unit (ALU)) that adds or subtracts values ​​based on the values ​​of the quantized sign bits. To this end, such circuitry can be implemented with a quantized sign bit vector (

number

number

number

number

number

[0051] That is, the uniform BCQ circuit UBC can approximate each weight of the weight matrix WM based on one common scale coefficient CSC (i.e., “s” value), multiple magnification scale coefficients MSC (i.e., powers of 2), and multiple quantization code bits QSB (i.e., “b” values). In this case, the same common scale coefficient CSC can be used for each column, and different common scale coefficients CSC can be used for different columns.

[0052] In this manner, the uniform BCQ circuit UBC can generate a plurality of common scale coefficients CSC, a plurality of magnification scale coefficients MSC, and a plurality of quantization sign bits QSB based on the weight matrix WM. The uniform BCQ circuit UBC can provide the plurality of common scale coefficients CSC, a plurality of magnification scale coefficients MSC, and a plurality of quantization sign bits QSB to the matrix multiplier 100. The following describes the operation of the matrix multiplier 100, which performs a matrix multiplication operation based on the plurality of common scale coefficients CSC, a plurality of magnification scale coefficients MSC, and a plurality of quantization sign bits QSB.

[0053] Fig. 6 is a diagram illustrating the operation of the matrix multiplier according to the embodiment of Fig. 5. Referring to Figs. 1 to 6, the matrix multiplier 100 can calculate an output matrix YM based on a plurality of common scale coefficients CSC, a plurality of magnification scale coefficients MSC, and a plurality of quantization code bits QSB.

[0054] In the following, for the sake of simplicity, the first input vector (i.e.

number

number

number

[0055] Referring to Equation 10, the matrix multiplier 100 can calculate y11 according to Equation 11 below.

number

number

[0056] The data type of the input elements included in the input vector may be floating point. In this case, the matrix multiplier 100 multiplies one output element (e.g., y 11 ), a total of (R × n) floating-point multiplication operations on products of powers of 2 must be performed.

[0057] In one embodiment, a floating-point multiplication operation for a product of a power of 2 can be performed with a very small amount of calculations. For example, the matrix multiplier 100 can perform the floating-point multiplication operation for a product of a power of 2 by changing the exponent part of the floating-point data type. The operation of the matrix multiplier 100 for performing the floating-point multiplication operation for a product of a power of 2 will be described in more detail with reference to FIG. 10 below.

[0058] Then, the matrix multiplier 100 multiplies each of the multiple scaled input vectors MSX by a different quantized code vector (e.g.,

number

number

[0059] In one embodiment, one scaled input vector MSX and one quantized code vector (e.g.,

number

number

[0060] In one embodiment, when the data type of each of the plurality of input elements is converted to a fixed point, the matrix multiplier 100 can calculate the partial product PSP with a smaller amount of calculations. For example, when the data types of each of the plurality of input elements are fixed points corresponding to the same exponent value, the matrix multiplier 100 must perform 'R×n' fixed point summation operations and 'R−1' fixed point summation operations to calculate the partial product PSP. A more detailed operation of the matrix multiplier 100 that converts the data type of each of the plurality of input elements to a fixed point will be described in more detail with reference to the following FIGS. 7 to 22.

[0061] The matrix multiplier 100 multiplies the partial product PSP11 and a common scale factor (i.e., s c1 ) to the output element (i.e., “y 11That is, the matrix multiplier 100 can compute the partial product PSP11 and the common scale factor (i.e., s c1 ) to generate the output element (i.e., “y 11 " can be calculated.

[0062] On the other hand, as mentioned above, R×N floating point multiplications for powers of 2 can be performed efficiently. As a result, the matrix multiplier 100 performs one floating point multiplication to produce one output element (e.g., y 11 ) can be calculated. In this case, unlike the one previously described with reference to Fig. 2, the number of floating-point multiplications performed by the matrix multiplier 100 is minimized. Therefore, according to the embodiment of Figs. 5 to 6, the operating speed of the matrix multiplication device MMD can be improved.

[0063] Fig. 7 is a block diagram showing a configuration of the matrix multiplier of Fig. 1 according to the embodiment of Fig. 4. Referring to Figs. 1 to 7, the matrix multiplier 100 includes a magnification scale coefficient buffer 110, an input vector scaler 120, a first data type converter 130, a quantization code bit buffer 140, a processing element array 150, a second data type converter 160, a common scale coefficient buffer 170, and a common scaler 180.

[0064] The multiplication factor scale factor buffer 110 can store a plurality of multiplication factor scale factors MSC provided from the uniform BCQ circuit UBC. The multiplication factor scale factor buffer 110 can provide a plurality of multiplication factor scale factors MSC to the input vector scaler 120.

[0065] The input vector scaler 120 may receive an input matrix XM. For example, the input vector scaler 120 may receive multiple input vectors that include multiple input elements (e.g.,

number

[0066] The input vector scaler 120 may perform multiplication scaling on the input matrix XM based on multiple multiplication scale factors MSC. For example, the input vector scaler 120 may generate multiple multiplication scaled input vectors MSX based on multiple input vectors. In this case, the multiple multiplication scaled input vectors MSX may be generated based on multiple input vectors (e.g.,

number

number

number

[0067] In one embodiment, a plurality of scaled input vectors MSX may be included in a scaled input matrix, where the row size of the scaled input matrix may be a constant multiple of the row size of the input matrix XM, and the column size of the scaled input matrix may be the same as the column size of the input matrix XM, although the scope of the present disclosure is not limited in this respect.

[0068] Each of the multiple scaled input vectors MSX is realized as a row vector having a dimension R times that of the corresponding input vector. For example, the first input vector (i.e.,

number

number

number

number

[0069] In this manner, the input vectors scaled by the second through (h)th factors (i.e.,

number

number

[0070] In one embodiment, the data type of each of the scaled input elements MSIE may be floating point, although the scope of the present disclosure is not limited in this respect.

[0071] The first data type converter 130 may receive a plurality of scaled input vectors MSX. For example, the first data type converter 130 may receive first through (h)-th scaled input vectors (i.e.,

number

[0072] The first data type converter 130 can extract the exponent EXP from each of the multiple scaled input vectors MSX. For example, the first data type converter 130 can extract the exponent EXP from the first scaled input vector (

number

number

[0073] The first data type converter 130 can convert the data type of the multiple scaled input vectors MSX into fixed point. For example, the first data type converter 130 can receive multiple scaled input vectors MSX and output multiple fixed point scaled input vectors MSXfxp. That is, the first data type converter 130 converts the first through (h)th fixed point scaled input vectors (i.e.,

number

[0074] More specifically, the first data type converter 130 may convert the data type of each of the multiple scaled input elements MSIE included in the scaled input vector MSX into a fixed point data type based on the extracted exponent. In this case, the first fixed point scaled input vector (i.e.,

number

[0075] The quantized sign bit buffer 140 can store a plurality of quantized sign bits QSB provided from the uniform BCQ circuit UBC. The quantized sign bit buffer 140 can provide a plurality of quantized sign bits QSB to the processing element array 150.

[0076] The processing element array 150 can receive the multiple quantization sign bits QSB and the multiple fixed-point scaled input vectors MSXfxp. The processing element array 150 can compute multiple fixed-point partial products PSPfxp based on the multiple fixed-point scaled input elements MSIEfxp and the multiple quantization sign bits QSB included in the multiple fixed-point scaled input vectors MSXfxp.

[0077] The processing element array 150 may include a plurality of processing elements arranged in a row direction and a column direction. The plurality of processing elements can respectively calculate different fixed-point partial products PSPfxp. The configuration and operation of each of the processing element arrays 150 will be described in more detail with reference to the following FIGS. 14 to 17.

[0078] The second data type converter 160 can receive a plurality of exponents EXP from the first data type converter 130. The second data type converter 160 can receive a plurality of fixed-point partial products PSPfxp from the processing element array 150. The second data type converter 160 can convert the data type of the plurality of fixed-point partial products PSPfxp to floating point based on the plurality of exponents EXP. That is, the second data type converter 160 can output a plurality of partial products PSP having a floating point format. A more detailed configuration and operation of the second data type converter 160 will be described in more detail with reference to FIG. 18 below.

[0079] The common scale factor buffer 170 can store a plurality of common scale factors CSC provided from the uniform BCQ circuit UBC. The common scale factor buffer 170 can provide a plurality of common scale factors CSC to the common scaler 180.

[0080] The common scaler 180 may receive a plurality of common scale factors CSC and a plurality of partial products PSP. The common scaler 180 may scale the plurality of partial products PSP based on the plurality of common scale factors CSC. For example, the common scaler 180 may multiply each of the plurality of partial products PSP by a corresponding common scale factor CSC to generate a plurality of output vectors included in the output matrix YM. A more detailed configuration and operation of the common scaler 180 will be described in more detail with reference to FIG. 15 below.

[0081] Fig. 8 is a block diagram showing the configuration of the input vector scaler of Fig. 7. Referring to Figs. 1 to 8, the input vector scaler 120 may include a first factor scaling circuit 121 to an (h)th factor scaling circuit 12h.

[0082] The first scaling circuit 121 to the (h)th scaling circuit 12h can receive input vectors different from each other. For example, the first scaling circuit 121 to the (h)th scaling circuit 12h can receive the first to (h)th input vectors (i.e.,

number

[0083] Each of the first scaling circuit 121 through the (h)th scaling circuit 12h can sequentially receive multiple input elements. For example, the first scaling circuit 121 receives x 11 ~x 1n The (h)th scaling circuit 12h receives x h1 ~x hn can be received in sequence.

[0084] Each of the first scaling circuit 121 to the (h)th scaling circuit 12h can sequentially receive a plurality of scaling coefficients MSC from the scaling coefficient buffer 110. For example, each of the first scaling circuit 121 to the (h)th scaling circuit 12h can sequentially receive a first scaling coefficient MSC1 to an Rth scaling coefficient MSCR (i.e., 2 0 ~2 R-1 ) can be received.

[0085] The first scaling circuit 121 to the (h)th scaling circuit 12h scale the input vectors scaled by the first to (h)th factors (i.e.,

number

number

[0086] Fig. 9 is a diagram showing in more detail the operation of the magnification scaling circuit of Fig. 8. For the sake of simpler explanation, the operation of the first magnification scaling circuit 121 will be representatively described below with reference to Figs. 1 to 9. However, the scope of the present disclosure is not limited thereto, and the second magnification scaling circuit 122 to the (h)th magnification scaling circuit 12h may operate in a similar manner.

[0087] The first factor scaling circuit 121 scales the first input vector (i.e.,

number

[0088] For example, the first factor scaling circuit 121 may be configured to scale the input element “x 11 " may be multiplied by the first through Rth scale coefficients MSCR to generate scaled input elements MSIE11_1 through MSIE11_R (illustrated by diagonal stripes).

[0089] Similarly, the first factor scaling circuit 121 scales the input element “x 12” may be multiplied by the first through Rth scale coefficients MSCR to generate scaled input elements MSIE12_1 through MSIE12_R (illustrated by a dotted pattern).

[0090] In this manner, the first scaling circuit 121 calculates x 13 ~x 1n A plurality of scaled input elements MSIE corresponding to:

[0091] The first scaling circuit 121 may sequentially output a plurality of scaled input elements MSIEs. For example, the first scaling circuit 121 may provide a plurality of scaled input elements MSIEs to the first data type converter 130.

[0092] That is, according to an embodiment of the present disclosure, the first factor scaling circuit 121 has one input element (e.g., x 11 ) based on multiple scaled input elements (e.g., scaled input elements illustrated with diagonal stripes (MSIE for x 11 )). In other words, the first factor scaling circuit 121 can generate a plurality of factor-scaled input elements MSIE by repeatedly using one input element. Therefore, according to the embodiment of the present disclosure, the input reuse of the matrix multiplier 100 is maximized, and thus the number of times the matrix multiplier 100 receives input elements from the outside is minimized. In this case, the number of times the matrix multiplier 100 accesses an external memory device that stores the input elements is minimized, and thus the operating efficiency and operating speed of the matrix multiplication device MMD can be improved.

[0093] Fig. 10 is a diagram illustrating in more detail how the scaler circuit of Fig. 8 performs a scaler operation. For the sake of simplicity, the following will representatively describe how the first scaler circuit 121 generates one scaled input element MSIE with reference to Figs. 1 to 10. However, the scope of the present disclosure is not limited thereto.

[0094] The first magnification scaling circuit 121 may receive an input element IE and a magnification scale factor MSC.

[0095] The input element IE may have a floating-point data type, for example, the input element IE may include a sign part SP, an exponent part EXPP, and a mantissa part MTSP.

[0096] The magnification scale factor MSC may be a power of 2. For example, the magnification scale factor MSC may be a power of 2. 0 2 of ~ R-1 It could be one of them.

[0097] The first factor scaling circuit 121 may generate a scaled input element MSIE by multiplying the input element IE by the factor scale coefficient MSC. In this case, the first factor scaling circuit 121 may generate the scaled input element MSIE by changing the value of the exponent part EXPP of the input element IE. That is, the first factor scaling circuit 121 may perform a floating-point multiplication operation for a power of 2 by increasing the value of the exponent part EXPP of the input element IE by an amount corresponding to the factor scale coefficient MSC.

[0098] For example, the first factor scaling circuit 121 may be a multiplier circuit having an input element IE and a k-th factor scale coefficient MSCk (i.e., 2 k-1When calculating the product of k and m, the first scaling circuit 121 may increase the value of the exponent part EXPP of the input element IE by 'k-1' to generate a scaled input element MSIE.

[0099] In a more detailed example, the three least significant bits of the exponent portion EXPP of the input element IE are '100', and the first magnification scaling circuit 121 scales the input element IE and the second magnification scale coefficient MSC2 (i.e., 2 1 When calculating the product of the first and second scaling circuits 121 and 122, the first and second scaling circuits 121 can generate a scaled input element MSIE by changing the three least significant bits of the exponent part EXPP of the input element IE to '101' (in other words, by increasing the value of the exponent part EXPP by '1').

[0100] Therefore, according to an embodiment of the present disclosure, the input vector scaler 120 can generate input vectors scaled by the first through (h)th factor (i.e.,

number

number

[0101] Figure 11 is a diagram showing the configuration of the first data type converter in Figure 7. Referring to Figures 1 to 11, the first data type converter 130 can include a first exponent extraction circuit 131_1 through an (h)th exponent extraction circuit 131_h, and a first data type conversion circuit 132_1 through an (h)th data type conversion circuit 132_h.

[0102] The first exponent extraction circuit 131_1 to the (h)th exponent extraction circuit 131_h can receive the input vector MSX scaled by different factors. For example, the first exponent extraction circuit 131_1 to the (h)th exponent extraction circuit 131_h can receive the input vector MSX scaled by the first to (h)th factors, respectively (i.e.,

number

number

[0103] Each of the first exponent extraction circuit 131_1 to the (h)th exponent extraction circuit 131_h can extract exponents from a plurality of scaled input elements MSIE received. For example, the first exponent extraction circuit 131_1 extracts an exponent from a first scaled input vector (i.e.,

number

number

[0104] The first exponent extraction circuit 131_1 to the (h)th exponent extraction circuit 131_h can provide the extracted exponents to the first data type conversion circuit 132_1 to the (h)th data type conversion circuit 132_h, respectively. Also, the first exponent extraction circuit 131_1 to the (h)th exponent extraction circuit 131_h can provide the extracted exponents to the second data type converter 160, respectively. The operation of the first exponent extraction circuit 131_1 to the (h)th exponent extraction circuit 131_h will be described in more detail with reference to FIG. 12 below.

[0105] The first data type conversion circuit 132_1 to the (h)th data type conversion circuit 132_h can receive the first exponent EXP1 to the (h)th exponent EXPh, respectively. The first data type conversion circuit 132_1 to the (h)th data type conversion circuit 132_h can convert the input vectors scaled by the first to (h)th scales (i.e.,

number

number

[0106] The first data type conversion circuit 132_1 to the (h)th data type conversion circuit 132_h can convert the data type of the received scaled input vector into a fixed point format based on the received exponent. The first data type conversion circuit 132_1 to the (h)th data type conversion circuit 132_h convert the first to the (h)th fixed point scaled input vectors (i.e.,

number

number

[0107] Fig. 12 is a diagram showing the operation of the exponent extraction circuit of Fig. 11. In the following, for the sake of simpler explanation, the operation of the first exponent extraction circuit 131_1 is representatively described. However, the scope of the present disclosure is not limited thereto.

[0108] 1 to 12, the first exponent extraction circuit 131_1 may receive a plurality of scaled input elements MSIE, for example, the first exponent extraction circuit 131_1 may receive the scaled input elements MSIE11_1 to MSIE1n_R described above with reference to FIG.

[0109] Each of the scaled input elements MSIE may have a data type of floating point. For example, each of the scaled input elements MSIE11_1 to MSIE1n_R may include a sign part SP, an exponent part EXPP, and a mantissa part MTSP.

[0110] The first exponent extraction circuit 131_1 can determine the largest value among the values ​​of the exponent parts EXPP of the received multiple scaled input elements MSIE. In this case, the first exponent extraction circuit 131_1 can determine the determined value as the first exponent EXP1. That is, the first exponent extraction circuit 131_1 can extract the largest exponent among the exponents of the scaled input elements MSIE11_1 to MSIE1n_R as the first exponent EXP1.

[0111] For the sake of simpler explanation, an embodiment in which the first exponent extraction circuit 131_1 extracts the largest value among the values ​​of the exponent part EXPP of the multiple scaled input elements MSIE is representatively described in Fig. 12, but the scope of the present disclosure is not limited thereto. For example, the first exponent extraction circuit 131_1 may be realized to extract the smallest value among the values ​​of the exponent part EXPP of the multiple scaled input elements MSIE.

[0112] Fig. 13 is a diagram showing the operation of the data type conversion circuit of Fig. 11. In the following, for the sake of simpler explanation, the operation of the first data type conversion circuit 132_1 is representatively described. However, the scope of the present disclosure is not limited thereto.

[0113] 1 to 13, the first data type conversion circuit 132_1 can receive a plurality of scaled input elements MSIE. For example, the first data type conversion circuit 132_1 can receive scaled input elements MSIE11_1 to MSIE1n_R.

[0114] The first data type conversion circuit 132_1 can receive a first exponent EXP1. The first data type conversion circuit 132_1 can convert the data type of the received multiple scaled input elements MSIE into fixed point according to the first exponent EXP1. That is, the first data type conversion circuit 132_1 can convert the scaled input elements MSIE11_1 to MSIE1n_R into fixed point scaled input elements MSIEfxp11_1 to MSIEfxp1n_R, respectively. However, for the sake of simpler explanation, the operation of the first data type conversion circuit 132_1 that converts the scaled input element MSIE11_1 into the fixed point scaled input element MSIEfxp11_1 will be representatively described below.

[0115] The first data type conversion circuit 132_1 can shift the mantissa of the scaled input element MSIE11_1 (hereinafter referred to as the first mantissa MTSPa) in the least significant bit (LSB) direction by the difference between the exponent of the scaled input element MSIE11_1 and the first exponent EXP1. For example, if the difference between the exponent of the scaled input element MSIE11_1 and the first exponent EXP1 is '4', the first data type conversion circuit 132_1 can insert '4' '0' bits in the most significant bit place of the first mantissa MTSPa.

[0116] The first data type conversion circuit 132_1 may determine a mantissa (hereinafter referred to as a second mantissa MTSPb) of the fixed-point scaled input element MSIEfxp11_1 based on the shifted first mantissa MTSPa. For example, the first data type conversion circuit 132_1 may cut off the lower bits of the shifted first mantissa MTSPa according to the code length of the second mantissa MTSPb. Alternatively, the first data type conversion circuit 132_1 may round the lower bits of the shifted first mantissa MTSPa based on various types of rounding algorithms such as "nearest even rounding" according to the code length of the second mantissa MTSPb. However, the scope of the present disclosure is not limited thereto.

[0117] That is, the maximum length of the mantissa of the fixed-point scaled input element MSIEfxp11_1 can be determined in advance.

[0118] In one embodiment, the code length of the first mantissa part MTSPa may be '10-bit' or '23-bit', but the scope of the present disclosure is not limited thereto.

[0119] In one embodiment, the code length of the second mantissa part MTSPb may be '7-bit', but the scope of the present disclosure is not limited thereto.

[0120] In one embodiment, the data type of each of the multiple fixed-point scaled input elements MSIEfxp may be INT8 (8-bit integer), although the scope of the present disclosure is not limited in this respect.

[0121] FIG. 14 is a block diagram showing the configuration of the processing element array of FIG. 7 in more detail. Referring to FIGS. 1 to 14, the processing element array 150 can include a plurality of processing elements PE arranged in row and column directions. In the following, for a simpler explanation, it is assumed that the plurality of processing elements PE are arranged along (h) rows and (m) columns. Also, a processing element arranged in the (i)th row and (j)th column of the processing element array 150 is referred to as "PEij". For example, a processing element arranged in the first row and second column of the processing element array 150 is referred to as "PE12".

[0122] The processing element array 150 can include a first processing element row PER1 to a (h)th processing element row PERh. Each of the first processing element row PER1 to the (h)th processing element row PERh can include a plurality of processing elements PE. For example, the first processing element row PER1 can include processing elements PE11 to PE1m.

[0123] The processing element array 150 can include a first processing element column PEC1 to an (m)th processing element column PECm. Each of the first processing element column PEC1 to the (m)th processing element column PECm can include a plurality of processing elements PE. For example, the first processing element column PEC1 can include processing elements PE11 to PEh1.

[0124] Each of the processing element rows PER can receive an input vector MSXfxp scaled by a different fixed-point magnification. For example, the first processing element row PER1 to the (h)th processing element row PERh receive the first to (h)th fixed-point magnification scaled input vectors MSXfxp, respectively (i.e.,

number

number

[0125] The processing elements included in the same processing element row can receive the same input vector scaled by a fixed-point factor. For example, each of the processing elements PE11 to PE1m receives an input vector scaled by a first fixed-point factor (i.e.,

number

[0126] Each of the processing element columns PEC receives a different quantized code bit vector (

number

number

number

[0127] The first to (m) quantized code bit vectors (i.e.,

number

number

number

number

number

number

number

number

number

number

number

[0128] The processing elements included in the same processing element column can receive the same quantized code bit vector. For example, each of the processing elements PE11 to PEh1 receives the first quantized code bit vector (i.e.,

number

number

[0129] Each of the multiple processing elements PE can calculate a different fixed-point partial product PSPfxp based on the received fixed-point scaled input vector MSXfxp and quantized code vector. That is, according to an embodiment of the present disclosure, one processing element PE can calculate one fixed-point partial product PSPfxp. For example, a processing element PEij can calculate a fixed-point partial product (PSPfxpij). In the following, a manner in which the fixed-point partial product PSPfxp is calculated in each processing element PE will be described in more detail.

[0130] 15 is a block diagram showing in more detail a portion of the operation of the matrix multiplier of FIG. 7. In the following, the first fixed-point scaled input vector (i.e.,

number

number

number

[0131] 1 to 15, the first processing element row PER1 receives a first fixed-point scaled input vector (i.e.,

number

[0132] Each of the processing elements PE11 to PE1m can receive a different quantized code bit vector. For example, the processing elements PE11 to PE1m receive the first to (m)th quantized code bit vectors (i.e.,

number

number

[0133] The processing elements PE11 to PE1m can respectively calculate fixed-point partial products PSPfxp different from one another. For example, the processing elements PE11 to PE1R can respectively calculate fixed-point partial products PSPfxp11 to PSPfxp1m.

[0134] The second data type converter 160 can receive the first exponent EXP1. The second data type converter 160 can receive the fixed-point partial products PSPfxp11-PSPfxp1m from the first processing element row PER1. The second data type converter 160 can convert the fixed-point partial products PSPfxp11-PSPfxp1m into a floating-point data type based on the first exponent EXP1. For example, the second data type converter 160 can convert the fixed-point partial products PSPfxp11-PSPfxp1m into the partial products PSP11-PSP1m, respectively.

[0135] The common scaler 180 may include a first multiplier circuit MUL1 through an m-th multiplier circuit MULm. The first multiplier circuit MUL1 through the m-th multiplier circuit MULm may receive the partial products PSP11 through PSP1m, respectively. The first multiplier circuit MUL1 through the m-th multiplier circuit MULm may receive the first common scale coefficient CSC1 through the m-th common scale coefficient CSCm, respectively, from the common scale coefficient buffer 170. In this case, the first common scale coefficient CSC1 through the m-th common scale coefficient CSCm may correspond to different column vectors of the weight matrix WM, respectively. For example, the first common scale coefficient CSC1 through the m-th common scale coefficient CSCm may be calculated by dividing the first common scale coefficient CSC1 through the m-th common scale coefficient CSCm by the s c1 ~s cm can correspond to each of the following:

[0136] Each of the first multiplier circuit MUL1 to the m-th multiplier circuit MULm can calculate an output element based on the received partial product and common scale factor. For example, the first multiplier circuit MUL1 multiplies the partial product PSP11 and the first common scale factor CSC1 to obtain y 11Similarly, the second multiplier MUL2 multiplies the partial product PSP12 and the second common scale factor CSC2 to obtain y 12 can be calculated.

[0137] In this manner, the matrix multiplier 100 uses the processing elements in the first processing element row PER1 to multiply the first input vector (i.e.,

number

number

[0138] In one embodiment, the first multiplier circuit MUL1 to the m-th multiplier circuit MULm can also multiply the partial products (e.g., partial products PSP21 to PSP2m) generated based on the second processing element row PER2 by the first common scale coefficient CSC1 to the m-th common scale coefficient CSCm, respectively, to generate a plurality of output elements. For example, the first multiplier circuit MUL1 multiplies the partial product PSP21 by the first common scale coefficient CSC1 to generate the output element “y 21 In other words, the first multiplier circuit MUL1 through the mth multiplier circuit MULm can multiply the partial products generated based on the first through mth processing element columns PEC1 through PECm by the first common scale coefficient CSC1 through the mth common scale coefficient CSCm, respectively. However, the scope of the present disclosure is not limited to a specific implementation of the common scaler 180.

[0139] In one embodiment, when the number of columns of the weight matrix WM is greater than the number of columns of the processing element array 150 or the number of rows of the weight matrix WM is greater than the number of rows of the processing element array 150, the matrix multiplier 100 can calculate the output elements based on various tiling techniques. Tiling techniques are described in more detail with reference to Figures 61-64 below.

[0140] Fig. 16 is a block diagram showing in more detail the operation of the first processing element row in Fig. 15. That is, hereinafter, the operation of the first processing element row PER1 will be representatively described with reference to Figs. 1 to 16. However, the scope of the present disclosure is not limited thereto, and the second processing element row PER2 to the (h)th processing element row PERh may also operate in a similar manner.

[0141] The first processing element row PER1 receives the first fixed-point scaled input vector (i.e.,

number

[0142] The processing elements PE11 to PE1m process the first to (m)-th quantized code bit vectors (i.e.,

number

number

number

number

[0143] To give a more detailed example, the first quantized code bit vector (i.e.,

number

[0144] The processing element PE11 can sequentially receive a plurality of corresponding 'quantization sign bit QSB-fixed point scaled input element MSIEfxp pairs'. For example, the processing element PE11 receives x 11 b corresponding to the scaled input element MSIEfxp 1_1_c1 ~b 1_R_c1 After receiving x 12 b corresponding to the scaled input element MSIEfxp 2_1_c1 ~b 2_R_c1 In this manner, the processing element PE11 receives b 1_1_c1 ~b n_R_c1 can be received in sequence.

[0145] The processing element PE11 can calculate the fixed-point partial product PSPfxp11 based on the order in which the quantization sign bit QSB and the multiple fixed-point scaled input elements MSIEfxp are received. For example, the processing element PE11 can calculate the fixed-point partial product PSPfxp11 by subtracting the sum of the fixed-point scaled input elements MSIEfxp11_1 to MSIEfxp1n_R corresponding to the quantization sign bit QSB indicating '1' from the sum of the fixed-point scaled input elements MSIEfxp11_1 to MSIEfxp1n_R corresponding to the quantization sign bit QSB indicating '0', thereby calculating the fixed-point partial product PSPfxp11.

[0146] Similarly, the processing element PE1j outputs the jth quantized code bit vector (i.e.,

number

number

number

number

[0147] To give a more detailed example, the processing element PE can calculate the fixed-point partial product PSPfxp according to the following equation 13.

number

[0148] Fig. 17 is a diagram showing a configuration of one of the processing elements of Fig. 16, which is realized according to an embodiment. Referring to Figs. 1 to 17, the processing element PE may include an arithmetic logic unit ALU (Arithmetic Logic Unit) and an accumulation register REG_ACC.

[0149] The arithmetic logic unit ALU may include a first input terminal TI1 to a third input terminal TI3 (input terminal) and an output terminal TO (output terminal). The first input terminal TI1 may sequentially receive a plurality of fixed-point scaled input elements MSIEfxp. The second input terminal TI2 may sequentially receive a plurality of quantization sign bits QSB. The third input terminal TI3 may be coupled to an accumulation register REG_ACC.

[0150] When the value received at the second input terminal TI2 is '0', the arithmetic logic unit ALU can store a value obtained by adding the values ​​received at the first input terminal TI1 and the third input terminal TI3 in the accumulation register REG_ACC through the output terminal TO. On the other hand, when the value received at the second input terminal TI2 is '1', the arithmetic logic unit ALU can store a value obtained by subtracting the value provided at the first input terminal TI1 from the value provided at the third input terminal TI3 in the accumulation register REG_ACC through the output terminal TO.

[0151] In this manner, the processing element PE can store the fixed-point partial product PSPfxp calculated according to the above-mentioned equation 13 in the accumulation register REG_ACC. In this case, the fixed-point partial product PSPfxp stored in the accumulation register REG_ACC is provided to the second data type converter 160. However, the scope of the present disclosure is not limited to a specific manner in which the processing element PE performs calculations and a specific configuration of the processing element PE.

[0152] In one embodiment, the processing element PE may not include a '1-bit adder' for performing a multiplication operation on the magnification scale factor MSC. That is, according to an embodiment of the present disclosure, since the magnification-scaled input element MSIE is provided to the processing element array 150, each processing element PE can be realized so as not to perform an operation on the magnification scale factor MSC. In this case, instead of each processing element PE including a circuit element for performing a multiplication operation on the magnification scale factor MSC, one magnification scaling circuit 121 is included for each processing element row PER, so that the size and production cost of the matrix multiplier 100 can be reduced.

[0153] FIG. 18 is a diagram illustrating the operation of the second data-based converter of FIG. 7. In the following, for the sake of simpler explanation, the first input vector (i.e.,

number

[0154] 1 to 18, the second data type converter 160 may receive a first exponent EXP1 from the first data type converter 130. The second data type converter 160 may receive a fixed-point partial product PSPfxp from one processing element PE. In this case, the fixed-point partial product PSPfxp may include a sign part SP and a mantissa part MTSP.

[0155] The second data type converter 160 may convert the data type of the fixed-point partial product PSPfxp to a floating-point data type to generate the partial product PSP. For example, the second data type converter 160 may add an exponent part EXPP to the fixed-point partial product PSPfxp. The second data type converter 160 may determine the exponent part EXPP of the partial product PSP to correspond to the first exponent EXP1.

[0156] The second data type converter 160 can add multiple '0' bits to the least significant bit place of the mantissa of the partial product PSP according to the difference in code length between the mantissa of the partial product PSP and the mantissa of the fixed-point partial product PSPfxp. However, the scope of the present disclosure is not limited thereto.

[0157] Fig. 19 is a flowchart showing the operation of the matrix multiplication device of Fig. 1. In the following, with reference to Figs. 1 to 19, one input vector (for example,

number

number

[0158] In step S110, the matrix multiplication device MMD may receive a weight matrix WM. For example, the uniform BCQ circuit UBC may receive a plurality of weights (e.g., w 11 ~w nm ) can be received.

[0159] In step S120, the matrix multiplication device MMD may perform uniform binary coding quantization on the weight matrix WM to generate a plurality of magnification scale coefficients MSC, a plurality of common scale coefficients CSC, and a plurality of quantization sign bits QSB. For example, the uniform BCQ circuit UBC may generate a plurality of magnification scale coefficients MSC, one common scale coefficient CSC, and a plurality of quantization sign bits QSB based on each weight. The uniform BCQ circuit UBC may provide the plurality of magnification scale coefficients MSC, one common scale coefficient CSC, and a plurality of quantization sign bits QSB to the matrix multiplier 100. In this case, the plurality of magnification scale coefficients MSC may be stored in the magnification scale coefficient buffer 110, the plurality of common scale coefficients CSC may be stored in the common scale coefficient buffer 170, and the plurality of quantization sign bits QSB may be stored in the quantization sign bit buffer 140. However, the scope of the present disclosure is not limited thereto.

[0160] In step S130, the matrix multiplication device MMD multiplies an input vector (e.g.,

number

[0161] In one embodiment, the matrix multiplication device MMD may perform step S130 regardless of the order of steps S110 to S120. For example, the matrix multiplication device MMD may perform step S130 before steps S110 to S120, or between steps S110 and S120.

[0162] In step S140, the matrix multiplication device MMD may perform scale on the input vector based on the multiple scale coefficients MSC to generate a scaled input vector MSX. For example, the input vector scaler 120 may scale each of the multiple input elements based on the multiple scale coefficients MSC provided from the scale coefficient buffer 110. More specifically, the input vector scaler 120 may multiply each of the multiple input elements by the multiple scale coefficients MSC to generate a multiple scaled input elements MSIE.

[0163] In step S150, the matrix multiplication device MMD can generate a plurality of partial products PSP based on the elements contained in the scaled input vector MSX (i.e., the plurality of scaled input elements MSIE) and the plurality of quantization sign bits QSB. Step S150 will be described in more detail with reference to FIG. 20 below.

[0164] In step S160, the matrix multiplication device MMD multiplies the plurality of partial products PSP and the plurality of common scale factors CSC, respectively, to generate an output vector (e.g., y 11 For example, the matrix multiplier 100 may multiply the partial product PSP11 and the first common scale factor CSC1 to generate one output element (e.g., y 11 Similarly, the matrix multiplier 100 can multiply the partial product PSP12 and a second common scale factor CSC2 to produce one output element (e.g., y 12 In this manner, the matrix multiplier 100 can generate multiple output elements that are included in the output vector.

[0165] Figure 20 is a flowchart showing in more detail step S150 of Figure 19. Referring to Figures 1 to 20, step S150 may include steps S151 to S153.

[0166] In step S151, the matrix multiplier 100 may convert the data type of the scaled input vector MSX to a fixed point to generate a fixed-point scaled input vector MSXfxp. For example, the first data type converter 130 may convert the data type of each of the plurality of scaled input elements MSIE to a fixed point to generate a plurality of fixed-point scaled input elements MSIEfxp.

[0167] In step S152, the matrix multiplier 100 may generate a plurality of fixed-point partial products PSPfxp based on the plurality of quantization sign bits QSB and the fixed-point scaled input vector MSXfxp. For example, the matrix multiplier 100 may generate the fixed-point partial products PSPfxp by sequentially adding or subtracting each of the fixed-point scaled input elements MSIEfxp based on the plurality of quantization sign bits QSB.

[0168] In step S153, the matrix multiplier 100 may convert the data type of the plurality of fixed-point partial products PSPfxp to floating point. For example, the second data type converter 160 may generate a plurality of partial products PSP based on the plurality of fixed-point partial products PSPfxp.

[0169] Fig. 21 is a flowchart showing the operation of the matrix multiplication device of Fig. 1. In the following, with reference to Figs. 1 to 18 and Fig. 21, one input vector (for example,

number

[0170] In step S210, the matrix multiplication device MMD may receive the first to (n)th weights. For example, the uniform BCQ circuit UBC receives a weight (e.g., w 11 ~w n1 ) can be received.

[0171] In step S220, the matrix multiplication device MMD may perform uniform binary coding quantization on the first through (n)th weights to generate first through (R)th magnification scale coefficients MSC, one common scale coefficient CSC, and first through (n×R)th quantization code bits QSB.

[0172] In step S230, the matrix multiplication device MMD may receive the first to n-th input elements. For example, the matrix multiplier 100 may multiply a plurality of input elements (e.g., x 11 ~x 1n ) can be received.

[0173] In one embodiment, the matrix multiplication device MMD may perform step S230 regardless of the order of steps S210 to S220. For example, the matrix multiplication device MMD may perform step S230 before steps S210 to S220, or between steps S210 and S220.

[0174] In step S240, the matrix multiplication device MMD may perform scale-up on the first through (n)th input elements based on the first through Rth scale coefficients MSC to generate first through (n×R)th scaled input elements. For example, the input vector scaler 120 may generate 'R' scaled input elements MSIE per input element based on the 'R' scale coefficients MSC.

[0175] In step S250, the matrix multiplication device MMD may generate one partial product PSP based on the first through (n×R)th scaled input elements and the first through (n×R)th quantization code bits QSB. For example, the processing element PE may calculate the partial product PSP by adding or subtracting the first through (n×R)th scaled input elements based on the first through (n×R)th quantization code bits QSB.

[0176] In step S260, the matrix multiplication device MMD multiplies the partial product PSP by a common scale factor CSC to generate an output element (e.g., y 11 For example, the common scaler 180 may multiply the partial product PSP and the common scale factor CSC to generate one output element. In this case, the generated output element may correspond to a value obtained by multiplying and adding the weights and input elements previously received through steps S210 and S230, respectively.

[0177] Figure 22 is a flowchart showing in more detail step S250 of Figure 21. Referring to Figures 1 to 18 and Figures 21 to 22, step S250 may include steps S251 to S253.

[0178] In operation S251, the matrix multiplier 100 may convert the data types of the first through (n×R)th scaled input elements MSIE to fixed point to generate the first through (n×R)th fixed-point scaled input elements MSIEfxp. For example, the first data type converter 130 may convert the first through (n×R)th scaled input elements MSIE to the first through (n×R)th fixed-point scaled input elements MSIEfxp, respectively.

[0179] In step S252, the matrix multiplier 100 may calculate the fixed-point partial product PSPfxp by subtracting the sum of the fixed-point scaled input elements MSIEfxp corresponding to the quantization sign bit QSB indicating '1' from the sum of the fixed-point scaled input elements MSIEfxp corresponding to the quantization sign bit QSB indicating '0'. For example, the processing element PE may calculate the fixed-point partial product PSPfxp in the manner previously described with reference to Equation 13.

[0180] In operation S253, the matrix multiplier 100 may convert the data type of the fixed-point partial product PSPfxp into a floating-point data type to generate the partial product PSP. For example, the second data type converter 160 may provide the partial product PSP to the common scaler 180.

[0181] 23 is a diagram illustrating an operation of the BCQ circuit of FIG. 1 according to an embodiment of the disclosure. Referring to FIG. 1 to FIG. 4 and FIG. 15, the uniform BCQ circuit UBC can perform a binary coding quantization operation for each row of the weight matrix WM.

[0182] First, the weight matrix WM can be expressed as the following Equation 14.

number

number

[0183] The uniform BCQ circuit UBC can perform a binary coding quantization operation for each row of the weight matrix WM based on the following Equation 15.

number

number

number

[0184] On the other hand, α k_ri can be expressed as a product of a common scale factor and a magnification scale factor. For example, α k_ri is 2 k-1 ×s cj Therefore, the above-mentioned formula 15 can be expressed as the following formula 16.

number

[0185] Figure 24 is a diagram showing a weight matrix approximated by the embodiment of Figure 23. Referring to Figures 1 to 4 and Figures 23 to 24, the uniform BCQ circuit UBC can perform uniform binary coding quantization on the weight matrix WM for each row.

[0186] Quantized Code Vector

number

number

number

[0187] As previously described with reference to Figures 5 and 6, the operation of the circuitry for multiplication based on the quantized sign bits QSB can be implemented by adding or subtracting values ​​based on the corresponding quantized sign bits QSB. Thus, such a circuitry can be implemented by multiplying the quantized sign bit vector (

number

number

number

number

[0188] To give a more detailed example, in the following, for the sake of simplicity, the output element in the first row and first column of the output matrix YM (i.e., y 11 ) is representatively described, however, the scope of the present disclosure is not limited thereto.

[0189] The uniform BCQ circuit UBC can provide a number of quantization sign bits QSB (i.e., “b” values), a number of common scale factors CSC (i.e., “s” values), and a number of magnification scale factors MSC (i.e., powers of 2) to the matrix multiplier 100.

[0190] The matrix multiplier 100 can calculate the output matrix YM based on the multiple quantization code bits QSB, multiple common scale coefficients CSC, and multiple magnification scale coefficients MSC. For example, the matrix multiplier 100 can calculate the output matrix YM by multiplying the input matrix XM by the approximated weight matrix WM in the manner previously described with reference to Equations 14 to 17.1 and FIG. 24.

[0191] In a more detailed example, the matrix multiplier 100 calculates y 11 can be calculated.

number

number

[0192] Meanwhile, referring to FIG. 24 and equations 14 to 17.1, the first quantization scale coefficient (i.e., α 1_r1 ~α 1_rn ), and the corresponding quantized code vector (e.g.,

number

number

number

number

number

number

number

number

number

[0193] In one embodiment, the first quantization code bit matrix QSBM_1 to the (R)th quantization code bit matrix QSBM_R can be realized with the same number of rows and columns as the weight matrix WM. For example, the number of rows of the first quantization code bit matrix QSBM_1 may be 'n' and the number of columns may be 'm'. The first quantization code bit matrix QSBM_1 to the (R)th quantization code bit matrix QSBM_R will be described in more detail with reference to FIG. 25 below.

[0194] Fig. 25 is a diagram showing the quantization code bit matrix of Fig. 24. Referring to Figs. 1 to 4 and Figs. 23 to 25, the uniform BCQ circuit UBC can perform uniform binary coding quantization on the weight matrix WM as a plurality of quantization code bit matrices QSBM, a plurality of common scale coefficients CSC, and a plurality of magnification scale coefficients MSC.

[0195] The number of quantized code bit matrices QSBM may be determined by the BCQ resolution (i.e., 'R'). For example, the uniform BCQ circuit UBC may perform binary coding quantization on the weight matrix WM to generate the first quantized code bit matrix QSBM_1 to the (R)th quantized code bit matrix QSBM_R.

[0196] Each of the first quantization code bit matrix QSBM_1 to the (R)th quantization code bit matrix QSBM_R can be realized with the same number of rows and columns as the weight matrix WM. For example, the number of rows of each of the first quantization code bit matrix QSBM_1 to the (R)th quantization code bit matrix QSBM_R may be 'n' and the number of columns may be 'm'. In this case, the quantization code bit QSB arranged in the (i)th row and the (j)th column of the (k)th quantization code bit matrix QSBM_k is expressed as "b j_k_ri " may also be used.

[0197] Each weight of the weight matrix WM is approximated based on the quantization code bit QSB arranged at the corresponding position of the first quantization code bit matrix QSBM_1 to the (R)th quantization code bit matrix QSBM_R. For example, the weight arranged in the (i)th row and (j)th column of the weight matrix WM (i.e., w ij ) is a quantized code bit (i.e., b j_1_ri ~b j_R_ri ) is approximated based on

[0198] Each of the first quantization code bit matrix QSBM_1 through the (R)th quantization code bit matrix QSBM_R may correspond to a plurality of common scale coefficients QSC. That is, as described above with reference to Figs. 23 and 24, when the uniform BCQ circuit UBC performs a uniform binary coding quantization operation for each row of the weight matrix WM, different rows of the weight matrix WM are approximated based on different common scale coefficients QSC, and weights included in the same row of the weight matrix WM are approximated based on the same common scale coefficient QSC.

[0199] As a result, the quantized code bit (i) included in the (k)th quantized code bit matrix QSBM_k (i.e., b j_k_ri ~b m_k_ri ) all have a common scale factor “s ri In this manner, the first to (n)-th rows of the first quantized code bit matrix QSBM_1 to the (k)-th quantized code bit matrix QSBM_k may correspond to the common scale coefficient “s r1 "~"s rn In a more detailed example, the first to (n)-th rows of the first quantized code bit matrix QSBM_1 may correspond to “s r1 "~"s rn " can be used.

[0200] The first quantization code bit matrix QSBM_1 to the (R)th quantization code bit matrix QSBM_R may correspond to different magnification scale coefficients MSC. For example, each of the multiple quantization code bits QSB included in the first quantization code bit matrix QSBM_1 corresponds to a first magnification scale coefficient MSC1 (i.e., 2 0 ), and each of the multiple quantization code bits QSB included in the second quantization code bit matrix QSBM_2 may correspond to the second magnification scale coefficient MSC2 (ie, 21).

[0201] In this manner, the weights (i.e., w ij ) is approximated based on the quantization code bits arranged in the (i)th row and the (j)th column of the first quantization code bit matrix QSBM_1 to the (R)th quantization code bit matrix QSBM_R, one common scale coefficient CSC corresponding to the (i)th row of the first quantization code bit matrix QSBM_1 to the (R)th quantization code bit matrix QSBM_R, and a plurality of magnification scale coefficients MSC corresponding to the first quantization code bit matrix QSBM_1 to the (R)th quantization code bit matrix QSBM_R, respectively.

[0202] More specifically, each of the weights included in the i-th row of the weight matrix is ​​w ij It is approximated based on R magnification scale factors MSC, one common scale factor CSC, and R quantization code bits QSB by the following Equation 20.

number

[0203] Fig. 26 is a block diagram showing a configuration of the matrix multiplier of Fig. 1 realized by an embodiment of the present disclosure. With reference to Figs. 1 to 4 and 23 to 26, the matrix multiplier 100 of Fig. 1 can be realized as a matrix multiplier 200 of Fig. 26.

[0204] The matrix multiplier 200 may include a magnification scale factor buffer 210, a common scale factor buffer 270, an input vector scaler 220, a first data type converter 230, a quantization sign bit buffer 240, a processing element array 250, and a second data type converter 260.

[0205] The magnification scale factor buffer 210 can store a plurality of magnification scale factors MSC provided from the uniform BCQ circuit UBC. The magnification scale factor buffer 210 can provide a plurality of magnification scale factors MSC to the input vector scaler 220.

[0206] The common scale factor buffer 270 can store a plurality of common scale factors CSC provided from the uniform BCQ circuit UBC. The common scale factor buffer 270 can provide a plurality of common scale factors CSC to the input vector scaler 220.

[0207] The input vector scaler 220 may receive an input matrix XM. For example, the input vector scaler 220 may receive multiple input vectors that include multiple input elements (e.g.,

number

number

[0208] The input vector scaler 220 may quantize scale the input matrix XM based on a plurality of common scale factors CSC and a plurality of magnification scale factors MSC. For example, the input vector scaler 220 may generate a plurality of quantized scaled input vectors QSX based on a plurality of input vectors. In this case, the plurality of quantized scaled input vectors QSX may be generated based on a plurality of input vectors (e.g.,

number

number

number

number

number

number

[0209] In one embodiment, a plurality of quantized scaled input vectors QSX may be included in a quantized scaled input matrix, where the row size of the quantized scaled input matrix may be a constant multiple of the row size of the input matrix XM, and the column size of the quantized scaled input matrix may be the same as the column size of the input matrix XM, although the scope of the present disclosure is not limited in this respect.

[0210] Each of the multiple quantized scaled input vectors QSX can be realized as a row vector having a dimension R times that of the corresponding input vector. For example, the first input vector (i.e.,

number

number

number

number

[0211] In this manner, the second through (h) quantized and scaled input vectors (i.e.,

number

number

number

number

[0212] In one embodiment, the data type of the plurality of input elements may be floating point, in which case the data type of each of the plurality of quantized scaled input elements QSIE may be floating point, although the scope of the present disclosure is not limited in this respect.

[0213] In one embodiment, the code length of each of the multiple input elements may be 16-bit or 32-bit, although the scope of the present disclosure is not limited in this respect.

[0214] In one embodiment, the code length of each of the multiple common scale factors CSC may be 16-bit or 32-bit, although the scope of the present disclosure is not limited in this respect.

[0215] The first data type converter 230 may receive a plurality of quantized scaled input vectors QSX. For example, the first data type converter 230 may receive first through (h)th quantized scaled input vectors (i.e.,

number

number

[0216] The first data type converter 230 can extract an exponent EXP from each of the multiple quantized scaled input vectors QSX. For example, the first data type converter 230 can extract an exponent EXP from each of the multiple quantized scaled input vectors QSX.

number

number

[0217] The first data type converter 230 can convert the data type of the plurality of quantized scaled input vectors QSX to fixed point. For example, the first data type converter 230 can receive the plurality of quantized scaled input vectors QSX and output the plurality of fixed point quantized scaled input vectors QSXfxp. That is, the first data type converter 230 converts the first through (h)th fixed point quantized scaled input vectors (i.e.

number

number

[0218] More specifically, the first data type converter 230 may convert the data type of each of the quantized scaled input elements QSIE included in the plurality of quantized scaled input vectors QSX into a fixed point data type based on the extracted exponents. In this case, the first fixed point quantized scaled input vector (i.e.,

number

[0219] The quantized sign bit buffer 240 can store a plurality of quantized sign bits QSB provided from the uniform BCQ circuit UBC. The quantized sign bit buffer 240 can provide a plurality of quantized sign bits QSB to the processing element array 250.

[0220] The processing element array 250 can receive the multiple quantization sign bits QSB and the multiple fixed-point quantization scaled input vectors QSXfxp. The processing element array 250 can generate a fixed-point output matrix YMfxp based on the multiple fixed-point quantization scaled input elements QSIEfxp and the multiple quantization sign bits QSB. The fixed-point output matrix YMfxp can be expressed as the following Equation 22:

number

number

number

[0221] The processing element array 250 may include a plurality of processing elements arranged in row and column directions. Each of the plurality of processing elements may calculate and output a different fixed-point output element of the above-mentioned Equation 22. The configuration and operation of each of the processing elements will be described in more detail with reference to the following Figures 36 to 38.

[0222] The second data type converter 260 may receive a plurality of exponents EXP from the first data type converter 230. The second data type converter 260 may receive a fixed-point output matrix YMfxp from the processing element array 250. For example, the second data type converter 260 may receive a plurality of fixed-point output elements from the processing element array 250.

[0223] The second data type converter 260 can convert the data type of the fixed-point output matrix YMfxp to floating point based on the multiple exponents EXP. That is, the second data type converter 260 can output the output matrix YM having a floating point data type. For example, the second data type converter 260 can convert the data type of each of the multiple received fixed-point output elements to floating point to generate multiple output elements. A more detailed configuration and operation of the second data type converter 160 will be described in more detail with reference to FIG. 39 below.

[0224] Fig. 27 is a block diagram showing a configuration of the input vector scaler of Fig. 26 according to an embodiment. Referring to Figs. 1 to 4 and 23 to 27, the input vector scaler 220 may include a factor scalar SCL_multiple and a common scalar SCL_common. The factor scalar SCL_multiple may include a first factor scaling circuit 221_1 to a (h)th factor scaling circuit 221_h. The common scalar SCL_common may include a first common scaling circuit 222_1 to a (h)th common scaling circuit 222_h.

[0225] The first scaling circuit 221_1 to the (h)th scaling circuit 221_h can receive different input vectors from each other. For example, the first scaling circuit 221_1 to the (h)th scaling circuit 221_h can receive the first to (h)th input vectors (i.e.,

number

number

[0226] That is, each of the first scaling circuit 221_1 to the (h)th scaling circuit 221_h can sequentially receive a plurality of input elements. For example, the first scaling circuit 221_1 can receive x 11 ~x 1n The (h)th scaling circuit 221_h can sequentially receive x h1 ~x hn can be received in sequence.

[0227] The magnification scaler SCL_multiple can sequentially receive a plurality of magnification scale coefficients MSC from the magnification scale coefficient buffer 210. For example, each of the first magnification scaling circuit 221_1 to the (h)th magnification scaling circuit 221_h can sequentially receive a plurality of magnification scale coefficients MSC from the magnification scale coefficient buffer 210. In a more detailed example, each of the first magnification scaling circuit 221_1 to the (h)th magnification scaling circuit 221_h can sequentially receive a first magnification scale coefficient MSC1 to an Rth magnification scale coefficient MSCR (i.e., 2 0 ~2 R-1 ) can be received.

[0228] Each of the first to (h)th scaling circuits 221_1 to 221_h can perform 'multiplication scaling' on the received input elements based on a plurality of multiplication scale coefficients MSC. That is, the first to (h)th scaling circuits 221_1 to 221_h can perform 'multiplication scaling' on the received input elements based on a plurality of multiplication scale coefficients MSC.

number

number

number

number

number

[0229] In one embodiment, the first scaling circuit 221_1 through the (h)th scaling circuit 221_h may operate on scaled input elements by increasing the exponent part EXPP of the input elements as previously described with reference to Fig. 10. In this case, the amount of operation of the first scaling circuit 221_1 through the (h)th scaling circuit 221_h may be minimized.

[0230] The first common scaling circuit 222_1 to the (h)th common scaling circuit 222_h may receive input vectors scaled by different factors. For example, the first common scaling circuit 222_1 to the (h)th common scaling circuit 222_h may receive input vectors scaled by first to (h)th factors (i.e.,

number

number

[0231] Each of the first common scaling circuit 222_1 through the (h)th common scaling circuit 222_h can sequentially receive a plurality of scaled input elements. For example, the first common scaling circuit 222_1 can sequentially receive the scaled input elements MSIE11_1 through MSIE1n_R described above with reference to FIG.

[0232] The common scaler SCL_common can sequentially receive the multiple common scale coefficients CSC from the common scale coefficient buffer 270. For example, each of the first common scaling circuit 222_1 through the (h)th common scaling circuit 222_h can sequentially receive the multiple common scale coefficients CSC from the common scale coefficient buffer 270.

[0233] The common scale coefficients CSC that each of the first common scaling circuit 222_1 to the (h)th common scaling circuit 222_h receives from the common scale coefficient buffer 270 may be identical to each other. For example, the common scale coefficients CSC that the first common scaling circuit 222_1 sequentially receives may be identical to the common scale coefficients CSC that the second common scaling circuit 222_2 sequentially receives.

[0234] The order in which each of the first common scaling circuit 222_1 to the (h)th common scaling circuit 222_h receives the common scale coefficient CSC may be the same as each other. For example, the common scale coefficient CSC that the first common scaling circuit 222_1 receives first may be the same as the common scale coefficient CSC that the second common scaling circuit 222_2 receives first. Similarly, the common scale coefficient CSC that the first common scaling circuit 222_1 receives second may be the same as the common scale coefficient CSC that the second common scaling circuit 222_2 receives second.

[0235] Each of the first common scaling circuit 222_1 through the (h)th common scaling circuit 222_h can perform 'common scaling' of the received scaled input elements MSIE based on a plurality of common scale factors CSC. That is, each of the first common scaling circuit 222_1 through the (h)th common scaling circuit 222_h can perform 'common scaling' of the first through (h)th quantized scaled input vectors (i.e.,

number

number

number

[0236] Fig. 28 is a diagram showing in more detail the operation of the common scaling circuit of Fig. 27. For a simpler explanation, the operation of the first common scaling circuit 222_1 will be representatively described below with reference to Figs. 1 to 4 and 23 to 28. However, the scope of the present disclosure is not limited thereto, and the second common scaling circuit 222_2 to the (h)th common scaling circuit 222_h may also operate in a similar manner.

[0237] The first common scaling circuit 222_1 scales the input vector by a first factor (i.e.,

number

[0238] The first common scaling circuit 222_1 may receive a plurality of common scale factors CSC in sequence. For example, the first common scaling circuit 222_1 may receive S r1 ~S rn In this case, S r1 ~S rn may correspond to different input elements. For example, S r1 ~S rn are x 11 ~x 1n It can correspond to.

[0239] The first common scaling circuit 222_1 scales the input vector by a first factor based on a plurality of common scale factors CSC (i.e.,

number

number

[0240] Similarly, the first common scaling circuit 222_1 scales each of the scaled input elements MSIE12_1 to MSIE12_R by a common scale factor “s r2 " and x 12 Multiple quantized scaled input elements QSIE for 1_r2× x 12 "~"α R_21× x 12 " (illustrated by the dot pattern).

[0241] In this manner, the first common scaling circuit 222_1 calculates x 13 ~x 1n A plurality of quantized scaled input elements QSIE corresponding to the respective quantized scaled input elements QSIE may be sequentially computed.

[0242] The first common scaling circuit 222_1 may sequentially output a plurality of quantized scaled input elements QSIE. For example, the first common scaling circuit 222_1 may provide a plurality of quantized scaled input elements QSIE to the first data type converter 230.

[0243] That is, according to the embodiment of Figs. 27-28, the input vector scaler 220 can perform magnification scaling on each of a plurality of input vectors and then common scaling on them to generate a quantization-scaled input vector.

[0244] Fig. 29 is a block diagram showing a configuration of the input vector scaler of Fig. 26 according to an embodiment. Referring to Figs. 1 to 4, 23 to 26, and 29, the input vector scaler 220 may include a common scalar SCL_common and a magnification scaler SCL_multiple. The common scalar SCL_common may include a first common scaling circuit 223_1 to a (h)th common scaling circuit 223_h. The magnification scaler SCL_multiple may include a first magnification scaling circuit 224_1 to a (h)th magnification scaling circuit 224_h.

[0245] The first common scaling circuit 223_1 to the (h)th common scaling circuit 223_h may receive different input vectors from each other. For example, the first common scaling circuit 223_1 to the (h)th common scaling circuit 223_h may receive the first to (h)th input vectors (i.e.,

number

number

[0246] Each of the first common scaling circuit 223_1 to the (h)th common scaling circuit 223_h can sequentially receive a plurality of input elements. For example, the first common scaling circuit 223_1 receives x 11 ~x 1n The (h)-th common scaling circuit 223_h can sequentially receive x h1 ~x hn can be received in sequence.

[0247] The common scaler SCL_common can sequentially receive the multiple common scale coefficients CSC from the common scale coefficient buffer 270. For example, each of the first common scaling circuit 223_1 through the (h)th common scaling circuit 223_h can sequentially receive the multiple common scale coefficients CSC from the common scale coefficient buffer 270.

[0248] The common scale coefficients CSC that each of the first common scaling circuit 223_1 to the (h)th common scaling circuit 223_h receives from the common scale coefficient buffer 270 may be the same as each other. The order in which each of the first common scaling circuit 223_1 to the (h)th common scaling circuit 223_h receives the common scale coefficients CSC may be the same as each other. The relationship between the common scale coefficients CSC that each of the first common scaling circuit 223_1 to the (h)th common scaling circuit 223_h receives is similar to that described above with reference to FIG. 27, and therefore detailed description thereof will be omitted.

[0249] Each of the first common scaling circuit 223_1 to the (h)th common scaling circuit 223_h can perform 'common scaling' on the received input vector based on a plurality of common scale factors CSC. That is, each of the first common scaling circuit 223_1 to the (h)th common scaling circuit 223_h can perform 'common scaling' on the received input vector based on a plurality of received input elements and a plurality of common scale factors CSC.

number

number

number

[0250] The first to (h)th scaling circuits 224_1 to 224_h may receive different common scaled input vectors. For example, the first to (h)th scaling circuits 224_1 to 224_h may receive first to (h)th common scaled input vectors (i.e.,

number

number

[0251] Each of the first scaling circuit 224_1 through the (h)th scaling circuit 224_h may sequentially receive a plurality of commonly scaled input elements. For example, the first scaling circuit 224_1 may receive a first commonly scaled input vector (i.e.,

number

[0252] The magnification scaler SCL_multiple can sequentially receive a plurality of magnification scale coefficients MSC from the magnification scale coefficient buffer 210. For example, each of the first magnification scaling circuit 224_1 through the (h)th magnification scaling circuit 224_h can sequentially receive a plurality of magnification scale coefficients MSC from the magnification scale coefficient buffer 210. In a more detailed example, each of the first magnification scaling circuit 224_1 through the (h)th magnification scaling circuit 224_h can sequentially receive a first magnification scale coefficient MSC1 through an Rth magnification scale coefficient MSCR (i.e., 2 0 ~2 R-1 ) can be received.

[0253] Each of the first to (h)th scaling circuits 224_1 to 224_h can perform 'multiplication scaling' on the received common-scaled input vector based on a plurality of multiplication scale coefficients MSC. That is, the first to (h)th scaling circuits 224_1 to 224_h can perform multiplication scaling on the first to (h)th common-scaled input vectors to generate the first to (h)th quantization-scaled input vectors (i.e.,

number

number

[0254] For example, the first factor scaling circuit 224_1 scales a first common scaled input vector (i.e.,

number

number

[0255] Fig. 30 is a diagram showing in more detail the operation of the common scaling circuit of Fig. 29. For a simpler explanation, the operation of the first common scaling circuit 223_1 will be representatively described below with reference to Figs. 1 to 4, 23 to 26, and 29 to 30. However, the scope of the present disclosure is not limited thereto, and the second common scaling circuit 223_2 to the (h)th common scaling circuit 223_h may also operate in a similar manner.

[0256] The first common scaling circuit 223_1 scales the first input vector (i.e.,

number

[0257] The first common scaling circuit 223_1 may sequentially receive a plurality of common scale factors CSC. For example, the first common scaling circuit 223_1 may receive S r1 ~S rn can be received in sequence.

[0258] The first common scaling circuit 223_1 scales a first input vector (i.e.,

number

number

[0259] Fig. 31 is a diagram showing in more detail the operation of the magnification scaling circuit of Fig. 29. For a simpler explanation, the operation of the first magnification scaling circuit 224_1 will be representatively described below with reference to Figs. 1 to 4, 23 to 26, and 29 to 31. However, the scope of the present disclosure is not limited thereto, and the second magnification scaling circuit 224_2 to the (h)th magnification scaling circuit 224_h may also operate in a similar manner.

[0260] The first factor scaling circuit 224_1 receives a first common scaled input vector (i.e.,

number

[0261] The first magnification scaling circuit 224_1 can sequentially receive the first magnification scale coefficient MSC1 through the Rth magnification scale coefficient MSCR.

[0262] The first factor scaling circuit 224_1 can multiply each of the received common scaled input elements CSIE11 to CSIE1n by a first factor scale coefficient MSC1 to an Rth factor scale coefficient MSCR to generate a plurality of quantized scaled input elements QSIE.

[0263] For example, the first scaling circuit 224_1 multiplies the common-scaled input element CSIE11 by the first to Rth scale coefficients MSCR, respectively, to obtain x 11 More specifically, the first factor scaling circuit 224_1 can generate a plurality of quantized scaled input elements QSIE for the common scaled input element CSIE11. 0 ~2 R-1 Multiplying each of these by α 1_r1×x11 "~"α R_r1×x11 3. A plurality of quantized scaled input elements QSIE may be generated (illustrated by diagonal stripes), each corresponding to a respective one of the quantized scaled input elements QSIE.

[0264] Similarly, the first scaling circuit 224_1 multiplies the common-scaled input element CSIE12 by the first to Rth scale coefficients MSCR, respectively, to obtain x 12 A number of quantized scaled input elements QSIE for x, y, and z can be generated (illustrated by the dotted pattern).

[0265] In this manner, the first factor scaling circuit 224_1 calculates x 13 ~x 1n A plurality of quantized scaled input elements QSIE corresponding to the respective quantized scaled input elements QSIE may be sequentially computed.

[0266] In one embodiment, the first through (h)th scaling circuits 224_1 through 224_h may operate the quantization-scaled input elements by increasing the exponent part EXPP of the scaled input elements, similar to that previously described with reference to Fig. 10. In this case, the amount of operation of the first through (h)th scaling circuits 224_1 through 224_h may be minimized.

[0267] The first factor scaling circuit 224_1 may sequentially output a plurality of quantized scaled input elements QSIE. For example, the first factor scaling circuit 224_1 may provide a plurality of quantized scaled input elements QSIE to the first source type converter 230.

[0268] That is, according to the embodiment of FIGS. 29 to 31, the input vector scaler 220 can perform common scaling on each of a plurality of input vectors and then scale them by a factor to generate a plurality of quantized scaled input vectors QSX.

[0269] Fig. 32 is a block diagram showing a configuration of the input vector scaler of Fig. 26 according to one embodiment. Referring to Figs. 1 to 4, 23 to 26, and 32, the input vector scaler 220 may include a factor scaling circuit 225 and a quantization scaler SCL_quantization. The quantization scaler SCL_quantization may include a first quantization scaling circuit 226_1 through a (h)th quantization scaling circuit 226_h.

[0270] The magnification scaling circuit 225 can sequentially receive a plurality of magnification scale coefficients MSC from the magnification scale coefficient buffer 210. For example, the magnification scaling circuit 225 can sequentially receive a first magnification scale coefficient MSC1 to an Rth magnification scale coefficient MSCR (i.e., 0 ~2 R-1 ) can be received sequentially.

[0271] The factor scaling circuit 225 may receive multiple common scale factors CSC in sequence from the common scale factor buffer 270. For example, the factor scaling circuit 225 may receive multiple common scale factors CSC in sequence from the common scale factor buffer 270. r1 ~s rn can be received in sequence.

[0272] The magnification scaling circuit 225 can generate multiple quantization scale factors QSC based on the multiple magnification scale factors MSC and the multiple common scale factors CSC. For example, the magnification scaling circuit 225 can multiply the multiple common scale factors CSC by each of the multiple magnification scale factors MSC to generate multiple quantization scale factors QSC. The operation of the magnification scaling circuit 225 will be described in more detail with reference to FIG. 33 below.

[0273] The quantization scaler SCL_quantization can receive a plurality of quantization scale factors QSC. For example, each of the first quantization scaling circuit 226_1 through the (h)th quantization scaling circuit 226_h can receive a plurality of quantization scale factors QSC from the factor scaling circuit 225. In this case, the quantization scale factors QSC received by each of the first quantization scaling circuit 226_1 through the (h)th quantization scaling circuit 226_h may be the same as each other.

[0274] The first quantization scaling circuit 226_1 through the (h)th quantization scaling circuit 226_h may receive different input vectors from each other. For example, the first quantization scaling circuit 226_1 through the (h)th quantization scaling circuit 226_h may receive the first through (h)th input vectors (i.e.,

number

number

[0275] Each of the first quantization scaling circuit 226_1 through the (h)th quantization scaling circuit 226_h can perform 'quantization scaling' on the received input vector based on a plurality of quantization scale coefficients QSC. For example, the first quantization scaling circuit 226_1 through the (h)th quantization scaling circuit 226_h can perform 'quantization scaling' on the received input vector based on a plurality of quantization scale coefficients QSC.

number

number

[0276] Fig. 33 is a diagram showing in more detail the operation of the magnification scaling circuit of Fig. 32. Referring to Figs. 1 to 4, 23 to 26, and 32 to 33, the magnification scaling circuit 225 generates a first magnification scale coefficient MSC1 to an Rth magnification scale coefficient MSCR (i.e., 2 0 ~2 R-1 ) can be received in sequence. The magnification scaling circuit 225 receives S r1 ~S rn can be received in sequence.

[0277] The magnification scaling circuit 225 is S r1 ~S rn Each of the multiple magnification scale factors MSC is multiplied to generate multiple quantization scale factors QSC (e.g., α 1_r1 ~α R_rn ) can be generated.

[0278] For example, the magnification scaling circuit 225 may be r1 The first row vector of the weight matrix WM (i.e.,

number

[0279] Similarly, the magnification scaling circuit 225 is r2 The first magnification scale coefficient MSC1 to the Rth magnification scale coefficient MSCR are multiplied by each other to obtain the second row vector of the weight matrix WM (i.e.,

number

[0280] In this manner, the first common scaling circuit 222_1 r3 ~S rn , a plurality of quantization scale factors QSC corresponding to the quantization scale factors QSC can be calculated in sequence.

[0281] In one embodiment, the factor scaling circuit 225 can calculate the quantization scale factor QSC in a manner that increases the exponent portion EXPP of the common scale factor CSC, similar to that described above with reference to FIG.

[0282] Fig. 34 is a diagram showing in more detail the operation of the quantization scaling circuit of Fig. 32. For a simpler explanation, the operation of the first quantization scaling circuit 226_1 will be representatively described below with reference to Figs. 1 to 4, 23 to 26, and 32 to 34. However, the scope of the present disclosure is not limited thereto, and the second quantization scaling circuit 226_2 to the (h)th quantization scaling circuit 226_h may also operate in a similar manner.

[0283] The first quantization and scaling circuit 226_1 converts the first input vector (i.e.,

number

[0284] The first quantization and scaling circuit 226_1 may receive a plurality of quantization scale factors QSC. For example, the first quantization and scaling circuit 226_1 may receive the α 1_r1 ~α R_r2n can be received in sequence.

[0285] The first quantization scaling circuit 226_1 may perform quantization scaling on the multiple input elements based on the multiple quantization scale factors QSC. For example, the first quantization scaling circuit 226_1 may multiply the multiple quantization scale factors QSC by corresponding input elements to generate multiple quantization scaled input elements QSIE.

[0286] In a more detailed example, the first quantization and scaling circuit 226_1 calculates x 11 Alpha 1_r1 ~α R_r1 Multiply each by x 11 A number of quantized scaled input elements QSIE for x, y, and z can be generated (illustrated by diagonal stripes).

[0287] Similarly, the first quantization and scaling circuit 226_1 calculates x 12 Alpha 1_r2 ~α R_r2 Multiply each by x 12 A number of quantized scaled input elements QSIE for x, y, and z can be generated (illustrated by the dotted pattern).

[0288] In this manner, the first quantization and scaling circuit 226_1 calculates x13 ~x 1n Based on the above, a plurality of quantized scaled input elements QSIE may be generated.

[0289] That is, according to the embodiment of Figures 32 to 34, the input vector scaler 220 generates a plurality of quantization scale coefficients, and then quantization scales each of a plurality of input vectors based on the generated plurality of quantization scale coefficients to generate a plurality of quantization scaled input vectors. In this case, unlike the embodiment previously described with reference to Figures 27 to 31, the input vector scaler 220 may include only one factor scaling circuit. However, the scope of the present disclosure is not limited thereto.

[0290] 27 to 34, the input vector scaler 220 scales one input element (e.g., x 11 ) based on the input vector scaler 220. In other words, the input vector scaler 220 can generate multiple quantized scaled input elements QSIE by repeatedly using one input element. Therefore, according to an embodiment of the present disclosure, the input reuse of the matrix multiplier 200 can be maximized, and the number of times the matrix multiplier 200 receives input elements from the outside can be minimized. In this case, the number of times the matrix multiplier 200 accesses an external memory device that stores the input elements can be minimized, and thus the operating efficiency and operating speed of the matrix multiplication device MMD can be improved.

[0291] Figure 35 is a block diagram showing the configuration of the first data type converter of Figure 26. Referring to Figures 1, 4, and 23 to 35, the first data type converter 230 can include a first exponent extraction circuit 231_1 through an (h)th exponent extraction circuit 231_h, and a first data type conversion circuit 232_1 through an (h)th data type conversion circuit 232_h.

[0292] The first exponent extraction circuit 231_1 to the (h)th exponent extraction circuit 231_h can receive the quantization-scaled input vector QSX different from each other. For example, the first exponent extraction circuit 231_1 to the (h)th exponent extraction circuit 231_h can receive the first to the (h)th quantization-scaled input vectors (i.e.,

number

number

[0293] Each of the first exponent extraction circuit 231_1 to the (h)th exponent extraction circuit 231_h can extract exponents from a plurality of quantized scaled input elements QSIE included in the received quantized scaled input vector QSX. For example, the first exponent extraction circuit 231_1 extracts an exponent from a first quantized scaled input vector (

number

number

[0294] The first exponent extraction circuit 231_1 to the (h)th exponent extraction circuit 231_h can provide the extracted exponents to the first data type conversion circuit 232_1 to the (h)th data type conversion circuit 232_h, respectively. Also, the first exponent extraction circuit 231_1 to the (h)th exponent extraction circuit 231_h can provide the extracted exponents to the second data type converter 260, respectively.

[0295] The specific manner in which each of the first exponent extraction circuit 231_1 to the (h)th exponent extraction circuit 231_h extracts exponents from the received elements is similar to the operation of the exponent extraction circuit previously described with reference to Figures 11 and 12, so a detailed description thereof will be omitted.

[0296] The first data type conversion circuit 232_1 to the (h)th data type conversion circuit 232_h can receive the first exponent EXP1 to the (h)th exponent EXPh, respectively. The first data type conversion circuit 232_1 to the (h)th data type conversion circuit 232_h can convert the first to the (h)th quantized scaled input vectors (i.e.,

number

number

[0297] Each of the first data type conversion circuit 232_1 to the (h)th data type conversion circuit 232_h can convert the data type of the received quantized scaled input vector into a fixed point based on the received exponent. That is, the first data type conversion circuit 232_1 to the (h)th data type conversion circuit 232_h convert the data type of the first to (h)th fixed point quantized scaled input vectors (i.e.,

number

number

number

[0298] The specific manner in which each of the first data type conversion circuit 232_1 to the (h)th data type conversion circuit 232_h converts the data type of each quantized scaled input element QSIE to a fixed point is similar to the operation of the data type conversion circuit previously described with reference to Figures 11 and 13, so a detailed description will be omitted.

[0299] FIG. 36 is a block diagram showing the configuration of the processing element array of FIG. 26 in more detail. Referring to FIGS. 1 to 4 and 23 to 36, the processing element array 250 can include a plurality of processing elements PE arranged in row and column directions. In the following, for a simpler explanation, it is assumed that the plurality of processing elements PE are arranged along (h) rows and (m) columns. Furthermore, a processing element arranged in the (i)th row and (j)th column of the processing element array 150 is referred to as "PEij". For example, a processing element arranged in the first row and second column of the processing element array 250 is referred to as "PE12".

[0300] The processing element array 250 can include a first processing element row PER1 to a (h)th processing element row PERh. Each of the first processing element row PER1 to the (h)th processing element row PERh can include a plurality of processing elements PE. For example, the first processing element row PER1 can include processing elements PE11 to PE1m.

[0301] The processing element array 150 can include a first processing element column PEC1 to an (m)th processing element column PECm. Each of the first processing element column PEC1 to the (m)th processing element column PECm can include a plurality of processing elements PE. For example, the first processing element column PEC1 can include processing elements PE11 to PEh1.

[0302] Different processing element rows may receive different fixed-point quantization scaled input vectors QSXfxp. For example, the first processing element row PER1 through the (h)th processing element row PERh may receive the first through (h)th fixed-point quantization scaled input vectors QSXfxp, respectively (i.e.,

number

number

[0303] The processing elements included in the same processing element row can receive the same fixed-point quantization scaled input vector QSXfxp. For example, each of the processing elements PE11 to PE1m receives a first fixed-point quantization scaled input vector (i.e.,

number

[0304] Different processing element columns can receive different quantized code bits QSBs. For example, the first processing element column PEC1 to the (m)th processing element column PECm can receive the first plurality of quantized code bits QSBs_1 to the (m)th plurality of quantized code bits QSBs_m, respectively.

[0305] The processing elements arranged in the same processing element column can receive the same plurality of quantized code bits QSBs. For example, each of the processing elements PE11 to PEh1 can receive a first plurality of quantized code bits QSBs_1, and each of the processing elements PE12 to PEh2 can receive a second plurality of quantized code bits QSBs_2.

[0306] Each of the processing elements PE can calculate a different fixed-point output element based on the received fixed-point quantized scaled input element QSIEfxp and the multiple quantized sign bits QSBs. That is, according to the embodiment of the present disclosure, one processing element PE can calculate one fixed-point output element. For example, a processing element PEij can calculate y' ij In the following, the fixed-point output element calculated in each processing element PE will be described in more detail.

[0307] The first processing element row PER1 to the (h)th processing element row PERh can respectively calculate different fixed-point output vectors. For example, the first processing element row PER1 to the (h)th processing element row PERh can respectively calculate the first to the (h)th fixed-point output vectors (i.e.,

number

number

[0308] Processing elements arranged in the same processing element row and different processing element columns can operate on different fixed-point output elements. For example, the processing elements PE11 to PE1m can operate on y' 11 ~y' 1m In a similar manner, the processing elements PE21 to PE2m can calculate y' 21 ~y' 2m The processing elements PEh1 to PEhm can respectively calculate y' h1 ~y' hm can be calculated respectively.

[0309] Fig. 37 is a block diagram showing in more detail the operation of the processing element in Fig. 36. Below, the operation of the first processing element row PER1 will be representatively described with reference to Figs. 1 to 4 and 23 to 37. However, the scope of the present disclosure is not limited thereto, and the second processing element row PER2 to the (h)th processing element row PERh may also operate in a similar manner.

[0310] The first processing element row PER1 receives the first fixed-point quantized scaled input vector (i.e.,

number

[0311] The processing elements PE11 to PE1m can receive the first plurality of quantized code bits QSBs_1 to the (m)th plurality of quantized code bits QSB_m, respectively. For example, the processing element PE11 can receive the first plurality of quantized code bits QSBs_1, and the processing element PE12 can receive the second plurality of quantized code bits QSBs_2.

[0312] The first plurality of quantized code bits QSBs_1 may include quantized code bits QSB arranged in the first columns of the first quantized code bit matrix QSBM_1 to the (R)th quantized code bit matrix QSBM_R, as previously described with reference to FIG. 25. For example, the first plurality of quantized code bits QSBs_1 may include quantized code bits (i.e., b 1_1_r1 ~b 1_R_rn ).

[0313] More specifically, the processing element PE11 selects the quantized code bits (i.e., b 1_1_r1 ~b 1_R_r1) are sequentially received, and then the quantized code bits (i.e., b 1_1_2n ~b 1_R_r2 In this manner, the processing element PE11 sequentially receives the quantized code bits (i.e., b 1_1_rn ~b 1_R_rn ) can be received in sequence.

[0314] That is, the processing element PE11 can sequentially receive a pair of corresponding quantized sign bits QSB and fixed-point quantized scaled input elements QSIEfxp.

[0315] The processing element PE11 selects a fixed-point output element (e.g., y') based on the order in which the quantized sign bit QSB and the multiple fixed-point quantized scaled input elements QSIEfxp are received. 11 For example, the processing element PE11 calculates a value obtained by subtracting a second value obtained by adding up fixed-point scaled input elements corresponding to the quantization sign bit QSB indicating '1' among the fixed-point quantization scaled input elements QSIEfxp from a first value obtained by adding up fixed-point quantization scaled input elements corresponding to the quantization sign bit QSB indicating '0' among the multiple fixed-point quantization scaled input elements QSIEfxp, to obtain a fixed-point output element "y' 11 " can be calculated.

[0316] In this manner, the processing element PE1j outputs the jth quantized code bit QSBs_j (i.e., b 1_1_cj ~b n_R_cjThe processing element PE1j can receive the first fixed-point quantized scaled input vector (i.e.,

number

number

[0317] To give a more detailed example, the processing element PE can calculate the fixed-point output element OEfxp according to the following equation 23.

number

[0318] Fig. 38 is a diagram showing a configuration of one of the processing elements of Fig. 36 realized according to an embodiment. Referring to Figs. 1 to 4 and 23 to 38, the processing element PE may include an arithmetic logic unit ALU and an accumulation register REG_ACC.

[0319] The arithmetic logic unit ALU may include a first input terminal TI1 to a third input terminal TI3 and an output terminal TO. The first input terminal TI1 may sequentially receive a plurality of fixed-point quantization scaled input elements QSIEfxp. The second input terminal TI2 may sequentially receive a plurality of quantization sign bits QSB. The third input terminal TI3 may be coupled to an accumulation register REG_ACC.

[0320] In one embodiment, each fixed-point quantized scaled input element QSIEfxp has a code length of 8-bits, and the arithmetic logic unit ALU may be configured to receive data in 8-bit units via a first input terminal TI1.

[0321] When the value received at the second input terminal TI2 is '0', the arithmetic logic unit ALU can add the values ​​received at the first input terminal TI1 and the third input terminal TI3 and store the result in the accumulation register REG_ACC through the output terminal TO. On the other hand, when the value received at the second input terminal TI2 is '1', the arithmetic logic unit ALU can subtract the value provided at the first input terminal TI1 from the value provided at the third input terminal TI3 and store the result in the accumulation register REG_ACC.

[0322] In this manner, the processing element PE can store the fixed-point output element OEfxp calculated according to the above-mentioned Equation 23 in the accumulation register REG_ACC. In this case, the fixed-point output element OEfxp stored in the accumulation register REG_ACC can be provided to the second data type converter 260. However, the scope of the present disclosure is not limited to a specific manner in which the processing element PE performs calculations and a specific configuration manner of the processing element PE.

[0323] In one embodiment, the fixed-point output element OEfxp may have a code length of 8-bits or more, for example, the fixed-point output element OEfxp may have a code length long enough to represent the accumulated magnitude of multiple fixed-point quantized scaled input elements QSIEfxp.

[0324] In one embodiment, the accumulation register REG_ACC may have a size of 8-bits or more, for example, the accumulation register REG_ACC may have a size large enough to store the fixed-point output element OEfxp.

[0325] In one embodiment, the fixed-point output element OEfxp may have a code length of 10-bits to 12-bits, although the scope of the present disclosure is not limited in this respect.

[0326] In one embodiment, the accumulation register REG_ACC may have a size of 10-bits to 12-bits, although the scope of the present disclosure is not limited in this respect.

[0327] In one embodiment, the multiple fixed-point quantization scaled input elements QSIEfxp received by the arithmetic logic unit ALU may correspond to the same exponent value. In this case, the arithmetic logic unit ALU can perform the operation of Equation 23 above without considering the place value of each of the multiple fixed-point quantization scaled input elements QSIEfxp. Therefore, the arithmetic logic unit ALU can operate the fixed-point output element OEfxp with a minimum amount of operations.

[0328] In one embodiment, the processing element PE may not include a '1-bit adder' for performing a multiplication operation on the magnification scale factor MSC. That is, according to an embodiment of the present disclosure, since the quantized scaled input element QSIE is provided to the processing element array 250, each processing element PE can be realized so as not to perform an operation on the magnification scale factor MSC. In this case, instead of including a circuit element for performing a multiplication operation on the magnification scale factor MSC in each processing element PE, one magnification scale circuit is included for each processing element row PER (e.g., when the input vector scaler 220 is realized by the embodiment previously described with reference to FIGS. 27 to 31) or one for the entire processing element array 250 (e.g., when the input vector scaler 220 is realized by the embodiment previously described with reference to FIGS. 32 to 34), thereby reducing the size and production cost of the matrix multiplier 200.

[0329] Figure 39 is a diagram showing the operation of the second data type converter of Figure 26. In the following, for the sake of simpler explanation, the operation of the second data type converter 260 for one fixed-point output element OEfxp is representatively described. However, the scope of the present disclosure is not limited thereto. For example, the second data type converter 260 may operate in a similar manner for any fixed-point output element OEfxp.

[0330] 1 to 4 and 23 to 39, the second data type converter 260 may receive a fixed-point output element OEfxp. For example, the second data type converter 260 may receive y' 11 ~y' hn One of the following can be received:

[0331] The second data type converter 260 can receive the first exponent EXP1 through the (h)th exponent EXPh from the first data type converter 230. In this case, the first exponent EXP1 through the (h)th exponent EXPh are respectively a first fixed-point output vector (i.e.,

number

number

[0332] The fixed-point output element OEfxp may include an exponent part SP and a mantissa part MTSP. The second data type converter 260 may convert the data type of the fixed-point output element OEfxp into a floating point to generate the output element OE.

[0333] The manner in which the second data type converter 260 converts the data type of the fixed-point output element OEfxp is similar to the manner in which the second data type converter 160 converts the data type of the fixed-point partial product PSPfxp, previously described with reference to FIG. 18, so a detailed description will be omitted.

[0334] Fig. 40 is a flowchart showing the operation of the matrix multiplication device of Fig. 1. In the following, with reference to Figs. 1 to 4 and Figs. 23 to 40, one input vector (for example,

number

number

[0335] In step S310, the matrix multiplication device MMD may receive a weight matrix WM. For example, the uniform BCQ circuit UBC may receive a plurality of weights (e.g., w 11 ~w nm ) can be received.

[0336] In step S320, the matrix multiplication device MMD may perform uniform binary coding quantization on the weight matrix WM to generate a plurality of magnification scale coefficients MSC, a plurality of common scale coefficients CSC, and a plurality of quantization sign bits QSB. For example, the uniform BCQ circuit UBC may approximate each weight based on a plurality of magnification scale coefficients MSC, one common scale coefficient CSC, and a plurality of quantization sign bits QSB. The uniform BCQ circuit UBC may provide the generated plurality of magnification scale coefficients MSC, a plurality of common scale coefficients CSC, and a plurality of quantization sign bits QSB to the matrix multiplier 200. In this case, the plurality of magnification scale coefficients MSC are stored in the magnification scale coefficient buffer 210, the plurality of common scale coefficients CSC are stored in the common scale coefficient buffer 270, and the plurality of quantization sign bits QSB are stored in the quantization sign bit buffer 240. However, the scope of the present disclosure is not limited thereto.

[0337] In step S330, the matrix multiplication device MMD multiplies the input vector (e.g.,

number

[0338] In one embodiment, the matrix multiplication device MMD may perform step S330 regardless of the order of steps S310 to S320. For example, the matrix multiplication device MMD may perform step S330 before steps S310 to S320, or between steps S310 and S320.

[0339] At step S340, the matrix multiplication device MMD may perform quantization scaling on the input vector based on the multiple magnification scale factors MSC and the multiple common scale factors CSC to generate a quantized scaled input vector QSX. For example, the input vector scaler 220 may perform quantization scaling on each of the multiple input elements based on the multiple magnification scale factors MSC and the multiple common scale factors CSC to generate a multiple quantized scaled input elements QSIE.

[0340] In step S350, the matrix multiplication device MMD generates an output vector (e.g.,

number

number

[0341] Figure 41 is a flowchart showing in more detail step S350 of Figure 40. Referring to Figures 1 to 4 and Figures 23 to 41, step S350 may include steps S351 to S353.

[0342] In operation S351, the matrix multiplier 200 may convert the data type of the quantization scaled input vector QSX to a fixed point to generate a fixed point quantization scaled input vector QSXfxp. That is, the matrix multiplier 200 may generate a fixed point quantization scaled input vector QSXfxp based on the quantization scaled input vector QSX. For example, the first data type converter 230 may convert the data type of each of a plurality of quantization scaled input elements QSIE to a fixed point to generate a plurality of fixed point quantization scaled input elements QSIEfxp.

[0343] In step S352, the matrix multiplier 100 generates a fixed-point output vector (e.g.,

number

[0344] More specifically, the processing element array 250 sequentially adds or subtracts a plurality of quantized scaled input elements QSIE based on a first plurality of quantized sign bits QSBs_1 to generate one fixed-point output element (e.g., y' 11 Similarly, the processing element array 250 may sequentially add or subtract the plurality of quantized scaled input elements QSIE based on the second plurality of quantized sign bits QSBs_2 to generate a single fixed-point output element (e.g., y' 12In this manner, the processing element array 250 can generate a fixed-point output vector (e.g.,

number

[0345] In step S353, the matrix multiplier 200 multiplies the fixed-point output vector (e.g.,

number

[0346] Fig. 42 is a flowchart showing the operation of the matrix multiplication device of Fig. 1. In the following, with reference to Figs. 1 to 4 and Figs. 23 to 42, one input vector (for example,

number

[0347] In step S410, the matrix multiplication device MMD may receive the first to (n)th weights. Step S410 is similar to step S210 described above with reference to FIG. 21, and therefore a detailed description thereof will be omitted.

[0348] In step S420, the matrix multiplication device MMD performs uniform binary coding quantization on the first through (n)th weights to generate the first through (R)th magnification scale coefficients MSC, the first through (n)th common scale coefficients CSC, and the first through (n×R)th quantization code bits QSB. That is, unlike S220 previously described with reference to FIG. 21, the uniform BCQ circuit UBC can generate the first through (n)th common scale coefficients CSC based on the first through (n)th weights.

[0349] In step S430, the matrix multiplication device MMD may receive the first to (n)th input elements. Step S430 is similar to step S230 described above with reference to FIG. 21, and therefore a detailed description thereof will be omitted.

[0350] In one embodiment, the matrix multiplication device MMD may perform step S430 regardless of the order of steps S410 to S420. For example, the matrix multiplication device MMD may perform step S430 before steps S410 to S420, or between steps S410 and S420.

[0351] In operation S440, the matrix multiplication device MMD may perform quantization scaling on the first through n-th input elements based on the first through (R)-th magnification scale coefficients MSC and the first through (n)-th common scale coefficients CSC to generate the first through (n×R)-th quantization scaled input elements QSIE. For example, the input vector scaler 220 may perform quantization scaling on the first through n-th input elements in various manners as previously described with reference to FIGS. 27 through 34 to generate the first through (n×R)-th quantization scaled input elements QSIE. The operation of the input vector scaler 220 will be described in more detail with reference to FIGS. 43 through 45 below.

[0352] In step S450, the matrix multiplication device MMD multiplies one output element (e.g., y 11 For example, the matrix multiplier 200 may calculate a value obtained by subtracting the sum of the quantized scaled input elements QSIE corresponding to the quantized sign bit QSB indicating '1' from the sum of the quantized scaled input elements QSIE corresponding to the quantized sign bit QSB indicating '0' among the first through (n×R) quantized scaled input elements QSIE, to generate one output element (e.g., y 11 ) can be generated.

[0353] Figures 43 to 45 are flowcharts showing in more detail step S440 of Figure 42 implemented according to an embodiment. Hereinafter, with reference to Figure 43, step S440 when input vector scaler 220 is implemented according to the embodiment previously described with reference to Figures 27 to 28 will be described, with reference to Figure 44, step S440 when input vector scaler 220 is implemented according to the embodiment previously described with reference to Figures 29 to 31 will be described, and with reference to Figure 45, step S440 when input vector scaler 220 is implemented according to the embodiment previously described with reference to Figures 32 to 34 will be described.

[0354] 1 to 4 and 23 to 43, step S440 can be realized as step S440a below. Step S440a may include steps S441a and S442a.

[0355] In step S441a, the input vector scaler 220 may multiply the first through n-th input elements by the first through (R)th scale coefficients MSC, respectively, to generate the first through (n×R)th scaled input elements MSIE. For example, the first scale scaling circuit 221_1 may multiply the first input element by the first scale coefficients MSC1 through (R)th scale coefficients MSCR to generate the first through Rth scaled input elements MSIE, and may multiply the second input element by the first scale coefficients MSC1 through (R)th scale coefficients MSCR to generate the (R+1)th through (2R)th scaled input elements MSIE.

[0356] In step S442a, the input vector scaler 220 may multiply each of the first through (n×R)th scaled input elements MSIE by a corresponding one of the first through (n)th common scale coefficients CSC to generate the first through (n×R)th quantization scaled input elements QSIE. For example, the first common scaling circuit 222_1 may multiply the first through Rth scaled input elements MSIE (e.g., scaled input elements MSIE11_1 through MSIE11_R) by a first common scale coefficient CSC (e.g., S r1 ) to obtain the first through (R) quantized and scaled input elements QSIE (e.g., x 11 Similarly, the first common scaling circuit 222_1 can generate quantized scaled input elements MSIE (e.g., scaled input elements MSIE12_1 to MSIE12_R) scaled by a second common scale factor CSC (e.g., S r2 ) to obtain the (R+1)th to (2R)th quantized and scaled input elements QSIE (e.g., x 12 quantized scaled input elements for

[0357] 1 to 4, 23 to 42, and 44, step S440 can be realized as step S440b below. Step S440b may include steps S441b and S442b.

[0358] In step S441b, the input vector scaler 220 may multiply the first through n-th input elements by the first through (n)-th common scale coefficients CSC, respectively, to generate the first through n-th commonly scaled input element CSIE. For example, the first common scaling circuit 223_1 may multiply the first through n-th input elements by the first through (n)-th common scale coefficients CSC, respectively, to generate the first through n-th commonly scaled input element CSIE.

[0359] In step S442b, the input vector scaler 220 may multiply the first through n-th common scaled input elements CSIE by the first through (R)-th magnification scale coefficients MSC, respectively, to generate the first through (n×R)-th quantization scaled input elements QSIE. For example, the first magnification scaling circuit 224_1 may multiply the first common scaled input element CSIE by the first through (R)-th magnification scale coefficients MSC to generate the first through (R)-th quantization scaled input elements QSIE, and may multiply the second common scaled input element CSIE by the first through (R)-th magnification scale coefficients MSC to generate the (R+1)-th through (2R)-th quantization scaled input elements QSIE.

[0360] 1 to 4, 23 to 42, and 45, step S440 can be realized as step S440c below. Step S440c may include steps S441c and S442c.

[0361] In step S441c, the input vector scaler 220 may multiply the first through (R)th magnification scale coefficients MSC by the first through (n)th common scale coefficients CSC, respectively, to generate the first through (n×R)th quantization scale coefficients QSC. For example, the magnification scaling circuit 225 may calculate products of different combinations of the first through (R)th magnification scale coefficients MSC and the first through (n)th common scale coefficients CSC to generate the first through (n×R)th quantization scale coefficients QSC (e.g., α 1_r1 ~α R_rn ) can be generated.

[0362] In step S442c, the input vector scaler 220 may multiply the first through (n×R) quantization scale coefficients QSC by corresponding ones of the first through n-th input elements to generate the first through (n×R) quantization scaled input elements QSIE. For example, the first quantization scaling circuit 226_1 multiplies the first input element (e.g., x 11 ) into the first to (R) quantization scale coefficients QSC (for example, α 1_r1 ~α R_r1 ) to generate the first through (R) quantization scaled input elements QSIE. Similarly, the first quantization scaling circuit 226_1 multiplies the second input element (e.g., x 12 ) to the (R+1)th to (2R)th quantization scale coefficients QSC (for example, α 1_r2 ~α R_r2 ) to generate the (R+1)-th to (2R)-th quantized and scaled input elements QSIE.

[0363] Figure 46 is a flowchart showing in more detail step S450 of Figure 42. Referring to Figures 1 to 4 and 23 to 46, step S450 may include steps S451 to S453.

[0364] In operation S451, the matrix multiplier 100 may convert the data type of the first through (n×R)th quantization scaled input elements QSIE to a fixed point to generate the first through (n×R)th fixed-point quantization scaled input elements QSIEfxp. For example, the first data type converter 230 may convert the first through (n×R)th quantization scaled input elements QSIE to the first through (n×R)th fixed-point quantization scaled input elements QSIEfxp, respectively.

[0365] In step S452, the matrix multiplier 200 may calculate one fixed-point output element OEfxp by subtracting the sum of the fixed-point quantization scaled input elements QSIEfxp corresponding to the quantization sign bit QSB indicating '1' among the first through (n×R) fixed-point quantization scaled input elements QSIEfxp from the sum of the fixed-point quantization scaled input elements QSIEfxp corresponding to the quantization sign bit QSB indicating '0' among the first through (n×R) fixed-point quantization scaled input elements QSIEfxp. For example, the processing element PE11 may calculate one fixed-point output element OEfxp (e.g., y' 11 ) can be calculated.

[0366] In step S453, the matrix multiplier 200 may convert the data type of the fixed-point output element OEfxp into a floating-point data type to generate one output element. For example, the second data type converter 260 converts the fixed-point output element OEfxp (e.g., y' 11 ) and receives an output element OE (e.g., y 11 ) can be output.

[0367] 47 is a diagram illustrating an operation of the BCQ circuit of FIG 1 according to an embodiment. Referring to FIG 1 to FIG 4 and FIG 47, the uniform BCQ circuit UBC can perform a binary coding quantization operation for each row of the weight matrix WM.

[0368] The uniform BCQ circuit UBC divides each row vector of the weight matrix WM into a plurality of quantization scale coefficients (e.g., “α”) and a plurality of quantization code vectors (e.g., “

number

number

number

number

number

number

number

number

number

[0369] Therefore, each of the multiple weights included in the weight matrix WM is approximated based on the multiple quantization scale factors QSC and the multiple quantization sign bits QSB according to the following Equation 26.

number

[0370] That is, one weight is approximated based on one zero point value ZPV, R magnification scale coefficients MSC, one common scale coefficient CSC, and R quantization code bits QSB.

[0371] In this manner, the uniform BCQ circuit UBC can approximate the weight matrix WM based on a plurality of zero point values ​​ZPV, a plurality of magnification scale coefficients MSC, a plurality of common scale coefficients CSC, and a plurality of quantization code bits QSB. In this case, unlike the one previously described with reference to FIG. 1, the uniform BCQ circuit UBC can further provide a plurality of zero point values ​​ZPV to the matrix multiplier 100.

[0372] The operation of the matrix multiplier 100 based on the weights approximated based on the zero point value ZPV, the multiplication factor scale factor MSC, the common scale factor CSC, and the quantization sign bit QSB will be described in more detail with reference to Figures 48 to 58 below.

[0373] Fig. 48 is a block diagram showing a configuration of the matrix multiplier of Fig. 1 realized according to one embodiment. With reference to Figs. 1 to 4 and 47 to 48, the matrix multiplier 100 of Fig. 1 is realized as a matrix multiplier 300 of Fig. 48.

[0374] The matrix multiplier 300 may include a magnification scale factor buffer 310, a common scale factor buffer 370, an input vector scaler 320, a first data type converter 330, a quantization sign bit buffer 340, a processing element array 350, and a second data type converter 360.

[0375] Each of the components of the matrix multiplier 300 may perform operations similar to those of the components of the matrix multiplier 200 previously described with reference to Figures 23 to 46. In the following, differences between the matrix multiplier 300 and the matrix multiplier 200 previously described with reference to Figures 23 to 46 will be mainly described.

[0376] The magnification scale coefficient buffer 310 can store the multiple magnification scale coefficients MSC and the multiple zero-point scale coefficients ZPSC provided from the uniform BCQ circuit UBC. The magnification scale coefficient buffer 210 can provide the multiple magnification scale coefficients MSC and the multiple zero-point scale coefficients ZPSC to the input vector scaler 220.

[0377] In one embodiment, each of the zero point scale coefficients ZPSC may be '1'. In the following, for easier explanation, an embodiment in which each of the zero point scale coefficients ZPSC is '1' will be representatively described. However, the scope of the present disclosure is not limited thereto.

[0378] The common scale factor buffer 370 can store the multiple common scale factors CSC and multiple zero point values ​​ZPV provided from the uniform BCQ circuit UBC. The common scale factor buffer 370 can provide the multiple common scale factors CSC and multiple zero point values ​​ZPV to the input vector scaler 320.

[0379] The input vector scaler 320 may receive an input matrix XM. The input vector scaler 320 may generate a plurality of quantized scaled input vectors QSX based on a plurality of common scale factors CSC, a plurality of zero point values ​​ZPV, a plurality of zero point scale factors ZPSC, and a plurality of magnification scale factors MSC. That is, the input vector scaler 320 may generate a plurality of quantized scaled input vectors QSX further based on a plurality of zero point values ​​ZPV and a plurality of zero point scale factors ZPSC.

[0380] Each of the multiple quantized scaled input vectors QSX can be realized as a row vector having a dimension R+1 times that of the corresponding input vector. For example, the first input vector (i.e.,

number

number

number

[0381] Similarly, the input vector scaler 320 outputs the second through (h) quantized scaled input vectors (i.e.,

number

number

number

number

[0382] The first data type converter 330 can receive a plurality of quantized scaled input vectors QSX. The first data type converter 330 can convert the data type of the plurality of quantized scaled input vectors QSX into fixed point. That is, the first data type converter 330 converts the first through (h)th fixed point quantized scaled input vectors (i.e.,

number

number

[0383] The quantization code bit buffer 340 can store the quantization code bits QSB and the zero point correction bits ZPCB provided from the uniform BCQ circuit UBC. The quantization code bit buffer 140 can provide the quantization code bits QSB and the zero point correction bits ZPCB to the processing element array 350.

[0384] In one embodiment, each of the zero point correction bits ZPCB may be '0'. In the following, for easier explanation, an embodiment in which each of the zero point correction bits ZPCB is '0' will be representatively described. However, the scope of the present disclosure is not limited thereto.

[0385] The processing element array 350 can receive the plurality of quantization sign bits QSB, the plurality of zero point correction bits ZPCB, and the plurality of fixed point quantization scaled input vectors QSXfxp. The processing element array 350 can generate a fixed point output matrix YMfxp based on the plurality of quantization sign bits QSB, the plurality of zero point correction bits ZPCB, and the plurality of fixed point quantization scaled input elements.

[0386] The processing element array 350 may include a plurality of processing elements arranged in a row direction and a column direction. Each of the plurality of processing elements may operate on a different fixed-point output element. A more detailed configuration and operation of the processing element array 350 will be described in more detail with reference to FIG. 58 below.

[0387] The second data type converter 360 can receive a plurality of fixed-point output elements from the processing element array 350. The second data type converter 360 can convert the data type of the fixed-point output matrix YMfxp to floating point according to a plurality of exponents EXP, that is, the second data type converter 360 can output an output matrix YM having a floating-point data type.

[0388] Fig. 49 is a block diagram showing a configuration of the input vector scaler of Fig. 48 according to one embodiment. Referring to Figs. 1 to 4 and 47 to 49, the input vector scaler 220 may include a factor scalar SCL_multiple and a common scalar SCL_common. The factor scalar SCL_multiple may include a first factor scaling circuit 321_1 to a (h)th factor scaling circuit 321_h. The common scalar SCL_common may include a first common scaling circuit 322_1 to a (h)th common scaling circuit 322_h.

[0389] The components of input vector scaler 320 may perform operations similar to those of input vector scaler 220 previously described with reference to Figures 27 to 28. In the following, differences between input vector scaler 320 and input vector scaler 220 previously described with reference to Figures 23 to 46 will be mainly described.

[0390] The magnification scaler SCL_multiple can sequentially receive the multiple magnification scale coefficients MSC and the multiple zero point scale coefficients ZPSC from the magnification scale coefficient buffer 310. For example, each of the first magnification scaling circuit 321_1 through the (h)th magnification scaling circuit 321_h can sequentially receive the multiple magnification scale coefficients MSC and the multiple zero point scale coefficients ZPSC from the magnification scale coefficient buffer 310.

[0391] Each of the first to (h)th scaling circuits 321_1 to 321_h can perform scaling on the received input vector based on a scaling factor coefficient MSC and a plurality of zero point scale coefficients ZPSC. For example, the first to (h)th scaling circuits 321_1 to 321_h can perform scaling on the received input vector based on the first to (h)th scaled input vectors (i.e.,

number

number

[0392] The common scaler SCL_common can sequentially receive the multiple common scale coefficients CSC and the multiple zero point values ​​ZPV from the common scale coefficient buffer 370. For example, each of the first common scaling circuit 322_1 through the (h)th common scaling circuit 322_h can sequentially receive the multiple common scale coefficients CSC and the multiple zero point values ​​ZPV from the common scale coefficient buffer 270.

[0393] Each of the first common scaling circuit 322_1 to the (h)th common scaling circuit 322_h can perform common scaling on the received scaled input vector MSX based on a plurality of common scale coefficients CSC and a plurality of zero point values ​​ZPV. For example, the first common scaling circuit 322_1 to the (h)th common scaling circuit 322_h can perform common scaling on the received scaled input vector MSX based on a plurality of common scale coefficients CSC and a plurality of zero point values ​​ZPV.

number

number

[0394] Fig. 50 is a diagram showing in more detail the operation of the magnification scaling circuit of Fig. 49. For a simpler explanation, the operation of the first magnification scaling circuit 321_1 will be representatively described below with reference to Figs. 1 to 4 and Figs. 47 to 50. However, the scope of the present disclosure is not limited thereto, and the second magnification scaling circuit 321_2 to the (h)th magnification scaling circuit 321_h may also operate in a similar manner.

[0395] The first factor scaling circuit 321_1 scales the first input vector (i.e.,

number

[0396] The first magnification scaling circuit 321_1 can sequentially receive the zero point scale coefficient ZPSC and the first magnification scale coefficient MSC1 through the Rth magnification scale coefficient MSCR.

[0397] The first magnification scaling circuit 321_1 can multiply each of the received multiple input elements by a zero-point scale coefficient ZPSC and a first magnification scale coefficient MSC1 to an Rth magnification scale coefficient MSCR to generate magnification-scaled input elements MSIE11_0 to MSIE1n_R.

[0398] For example, the first factor scaling circuit 321_1 scales the input element “x 11 " can be multiplied by the zero point scale factor ZPSC to generate the scaled input element MSIE11_0, "x11 ” may be multiplied by the first through Rth scale coefficients MSCR to generate scaled input elements MSIE11_1 through MSIE11_R (illustrated by diagonal stripes).

[0399] Similarly, the first factor scaling circuit 321_1 scales the input element “x 12 " can be multiplied by the zero point scale factor ZPSC to generate the scaled input element MSIE12_0, and "x 12 ” may be multiplied by the first through Rth scale coefficients MSCR to generate scaled input elements MSIE12_1 through MSIE12_R (illustrated by a dotted pattern).

[0400] In this manner, the first scaling circuit 321_1 calculates x 13 ~x 1n In this case, the scaled input elements (e.g., scaled input elements MSIE11_0 to MSIE1n_0) generated based on the zero-point scale factor ZPSC may be the same as different input elements. For example, the scaled input elements MSIE11_0 to MSIE1n_0 are each x 11 ~x 1n may be the same as

[0401] Fig. 51 is a diagram showing in more detail the operation of the common scaling circuit of Fig. 49. For the sake of simpler explanation, the operation of the first common scaling circuit 322_1 will be representatively described below with reference to Figs. 1 to 4 and Figs. 47 to 51. However, the scope of the present disclosure is not limited thereto, and the second common scaling circuit 322_2 to the (h)th common scaling circuit 322_h may also operate in a similar manner.

[0402] The first common scaling circuit 322_1 scales the input vector by a first factor (i.e.,

number

[0403] The first common scaling circuit 322_1 may sequentially receive a plurality of zero point values ​​ZPV and a plurality of common scale coefficients CSC. For example, the first common scaling circuit 322_1 may sequentially receive a first row vector of the weight matrix WM (i.e.,

number

number

[0404] The first common scaling circuit 322_1 generates a first quantized scaled input vector (i.e.,

number

[0405] For example, the first common scaling circuit 322_1 may generate one quantized scaled input element QSIE based on the product of the scaled input element MSIE11_0 and the first zero point value ZPV1, and may generate S of each of the scaled input elements MSIE11_1 to MSIE11_R. r1 Based on the product over , R quantized scaled input elements QSIE can be generated (illustrated by diagonal stripes).

[0406] Similarly, the first common scaling circuit 322_1 can generate one quantized scaled input element QSIE based on the product of the scaled input element MSIE12_0 and the second zero point value ZPV2, and the S of each of the scaled input elements MSIE12_1 to MSIE12_R can be calculated as follows: r2 Based on the product over R, R quantized scaled input elements QSIE can be generated (illustrated by the dotted pattern).

[0407] As a result, the first quantized scaled input vector (i.e.

number

[0408] Fig. 52 is a block diagram showing a configuration of the input vector scaler of Fig. 48 according to an embodiment. Referring to Figs. 1 to 4, 47 to 48, and 52, the input vector scaler 320 may include a common scalar SCL_common and a magnification scaler SCL_multiple. The common scalar SCL_common may include a first common scaling circuit 323_1 to a (h)th common scaling circuit 323_h. The magnification scaler SCL_multiple may include a first magnification scaling circuit 324_1 to a (h)th magnification scaling circuit 324_h.

[0409] The components of input vector scaler 320 may perform operations similar to those of input vector scaler 220 previously described with reference to Figures 29 to 31. The following mainly describes the differences between input vector scaler 320 and input vector scaler 220 previously described with reference to Figures 29 to 31.

[0410] The first common scaling circuit 323_1 to the (h)th common scaling circuit 323_h scale the first to (h)th input vectors (i.e.,

number

number

[0411] The common scaler SCL_common can sequentially receive the multiple common scale coefficients CSC and the multiple zero point values ​​ZPV from the common scale coefficient buffer 370. For example, each of the first common scaling circuit 323_1 to the (h)th common scaling circuit 323_h can sequentially receive the multiple common scale coefficients CSC and the multiple zero point values ​​ZPV from the common scale coefficient buffer 370.

[0412] Each of the first common scaling circuit 323_1 to the (h)th common scaling circuit 323_h can perform common scaling on the received input vector based on a plurality of common scale coefficients CSC and a plurality of zero point values ​​ZPV. For example, the first common scaling circuit 323_1 to the (h)th common scaling circuit 323_h can perform common scaling on the received input vector based on a plurality of common scale coefficients CSC and a plurality of zero point values ​​ZPV.

number

number

[0413] The magnification scaler SCL_multiple can sequentially receive the multiple magnification scale coefficients MSC and the multiple zero-point scale coefficients ZPSC from the magnification scale coefficient buffer 310. For example, each of the first magnification scaling circuit 324_1 through the (h)th magnification scaling circuit 324_h can sequentially receive the multiple magnification scale coefficients MSC and the multiple zero-point scale coefficients ZPSC from the magnification scale coefficient buffer 310.

[0414] Each of the first to (h)th factor scaling circuits 324_1 to 324_h can perform factor scaling on the received common scaled input vector CSX based on the factor scale coefficient MSC and the plurality of zero point scale coefficients ZPSC. For example, the first to (h)th factor scaling circuits 324_1 to 324_h can perform factor scaling on the received common scaled input vector CSX based on the factor scale coefficient MSC and the plurality of zero point scale coefficients ZPSC.

number

number

[0415] Fig. 53 is a diagram showing in more detail the operation of the common scaling circuit of Fig. 52. For a simpler explanation, the operation of the first common scaling circuit 323_1 will be representatively described below with reference to Figs. 1 to 4, 47 to 48, and 52 to 53. However, the scope of the present disclosure is not limited thereto, and the second common scaling circuit 323_2 to the (h)th common scaling circuit 323_h may also operate in a similar manner.

[0416] The first common scaling circuit 323_1 scales the first input vector (i.e.,

number

[0417] The first common scaling circuit 323_1 may sequentially receive a plurality of zero point values ​​ZPV and a plurality of common scale coefficients CSC. For example, the first common scaling circuit 323_1 may sequentially receive a first row vector of the weight matrix WM (i.e.,

number

number

[0418] The first common scaling circuit 323_1 performs, based on the order in which the plurality of input elements, the plurality of zero point values ​​ZPV, and the plurality of common scale coefficients CSC are received, The first commonly scaled input vector (i.e.,

number

[0419] For example, the first common scaling circuit 323_1 is x 11 and the first zero point value ZPV1, one common scaled input element CSIE11_0 may be generated based on the product of x 11 and R S r1 Based on the product over , R common scaled input elements CSIE11_1 to CSIE11_R can be generated (illustrated by diagonal stripes).

[0420] Similarly, the first common scaling circuit 323_1 calculates x12 and a second zero point value ZPV2, one common scaled input element CSIE12_0 may be generated based on the product of x 12 and R S r2 R common scaled input elements CSIE12_1 to CSIE12_R can be generated based on the product over (illustrated by the dotted pattern).

[0421] As a result, the first commonly scaled input vector (i.e.

number

[0422] Fig. 54 is a diagram illustrating in more detail the operation of the magnification scaling circuit of Fig. 52 according to one embodiment. For a simpler explanation, the operation of the first magnification scaling circuit 324_1 will be representatively described below with reference to Figs. 1 to 4, 47 to 48, and 52 to 54. However, the scope of the present disclosure is not limited thereto, and the second magnification scaling circuit 324_2 to the (h)th magnification scaling circuit 324_h may also operate in a similar manner.

[0423] The first factor scaling circuit 324_1 receives a first common scaled input vector (i.e.,

number

[0424] The first magnification scaling circuit 324_1 can repeatedly receive the zero point scale coefficient ZPSC and the first magnification scale coefficients MSC1 through Rth magnification scale coefficients MSCR.

[0425] The first factor scaling circuit 324_1 generates a first quantized scaled input vector (i.e.,

number

[0426] For example, the first factor scaling circuit 324_1 generates one quantized scaled input element QSIE (e.g., “ZPV1×x”) based on the product of the common scaled input element CSIE11_0 and the zero point scale factor ZPSC. 11 ”), which can be multiplied by scaled input elements CSIE11_1 to CSIE11_R and the first to Rth scale coefficients MSCR, respectively, to generate R quantized scaled input elements QSIE (illustrated by diagonal stripes).

[0427] Similarly, the first factor scaling circuit 324_1 generates a single quantized scaled input element QSIE (e.g., “ZPV2×x”) based on the product of the common scaled input element CSIE12_0 and the zero point scale factor ZPSC. 12 ”), which can be multiplied by scaled input elements CSIE12_1 through CSIE12_R and the first through Rth scale coefficients MSCR, respectively, to generate R quantized scaled input elements QSIE (illustrated by a dotted pattern).

[0428] As a result, the first quantized scaled input vector (i.e.

number

[0429] Figure 55 is a block diagram showing a configuration of the input vector scaler of Figure 48 according to one embodiment. With reference to Figures 1 to 4, 47 to 48, and 55, the input vector scaler 320 may include a factor scaling circuit 325 and a quantization scaler SCL_quantization. The quantization scaler SCL_quantization may include a first quantization scaling circuit 326_1 through a (h)th quantization scaling circuit 326_h.

[0430] The components of input vector scaler 320 may perform operations similar to those of input vector scaler 220 previously described with reference to Figures 32 to 34. The following mainly describes the differences between input vector scaler 320 and input vector scaler 220 previously described with reference to Figures 32 to 34.

[0431] The magnification scaling circuit 325 may sequentially receive a plurality of magnification scale coefficients MSC and a plurality of zero point scale coefficients ZPSC from the magnification scale coefficient buffer 310 .

[0432] The factor scaling circuit 325 may sequentially receive a plurality of common scale coefficients CSC and a plurality of zero point values ​​ZPV from a common scale coefficient buffer 370 .

[0433] The magnification scaling circuit 325 can output a plurality of quantization scale factors QSC based on a plurality of magnification scale factors MSC and a plurality of common scale factors CSC. The magnification scaling circuit 325 can output a plurality of zero point values ​​ZPV based on a plurality of zero point scale factors ZPSC and a plurality of zero point values ​​ZPV. The operation of the magnification scaling circuit 325 will be described in more detail with reference to FIG. 56 below.

[0434] The quantization scaler SCL_quantization can receive a plurality of quantization scale coefficients QSC and a plurality of zero point values ​​ZPV. For example, each of the first quantization scaling circuit 326_1 through the (h)th quantization scaling circuit 326_h can receive a plurality of quantization scale coefficients QSC and a plurality of zero point values ​​ZPV from the factor scaling circuit 325.

[0435] The first quantization scaling circuit 326_1 through the (h)th quantization scaling circuit 326_h may receive different input vectors from each other. For example, the first quantization scaling circuit 326_1 through the (h)th quantization scaling circuit 326_h may receive the first through (h)th input vectors (i.e.,

number

number

[0436] Each of the first quantization scaling circuit 326_1 through the (h)th quantization scaling circuit 326_h can perform 'quantization scaling' on the received input vector based on a plurality of quantization scale coefficients QSC and a plurality of zero point values ​​ZPV. For example, the first quantization scaling circuit 326_1 through the (h)th quantization scaling circuit 326_h can perform 'quantization scaling' on the received input vector based on a plurality of quantization scale coefficients QSC and a plurality of zero point values ​​ZPV.

number

number

[0437] Figure 56 is a diagram showing in more detail the operation of the magnification scaling circuit of Figure 55. With reference to Figures 1 to 4, 47 to 48, and 55 to 56, the magnification scaling circuit 325 can iteratively receive the zero point scale coefficient ZPSC and the first magnification scale coefficient MSC1 through the Rth magnification scale coefficient MSCR.

[0438] The magnification scaling circuit 325 may sequentially receive a plurality of zero point values ​​ZPV and a plurality of common scale coefficients CSC. For example, the magnification scaling circuit 325 may sequentially receive a first row vector of the weight matrix WM (i.e.,

number

number

[0439] The magnification scaling circuit 325 can sequentially output the multiple quantization scale coefficients QSC and the multiple zero point values ​​ZPV based on the order in which the multiple magnification scale coefficients MSC and the multiple zero point scale coefficients ZPSC were received, and the order in which the multiple common scale coefficients CSC and the multiple zero point values ​​ZPV were received.

[0440] For example, the magnification scaling circuit 325 may multiply the first zero point value ZPV1 by the zero point scale coefficient ZPSC to generate the first zero point value ZPV1, and may multiply the first magnification scale coefficients MSC1 to Rth magnification scale coefficients MSCR by R S r1 to obtain the first row vector of the weight matrix WM (i.e.,

number

[0441] Similarly, the magnification scaling circuit 325 can multiply the zero point scale coefficient ZPSC by the second zero point value ZPV2 to generate the second zero point value ZPV2, and can multiply the first magnification scale coefficient MSC1 to the Rth magnification scale coefficient MSCR by R S r2 to obtain the second row vector of the weight matrix WM (i.e.,

number

[0442] In this manner, the factor scaling circuit 325 can sequentially output a plurality of zero point values ​​ZPV and a plurality of quantization scale coefficients QSC.

[0443] FIG. 57 is a diagram showing in more detail the operation of the quantization scaling circuit of FIG. 55. For a simpler explanation, the operation of the first quantization scaling circuit 326_1 will be representatively described below with reference to FIGS. 1 to 4, 47 to 48, and 55 to 57. However, the scope of the present disclosure is not limited thereto, and the second factor scaling circuit 326_2 to the (h)th factor scaling circuit 326_h may also operate in a similar manner.

[0444] The first quantization and scaling circuit 326_1 converts the first input vector (i.e.,

number

[0445] The first quantization scaling circuit 326_1 can sequentially receive a plurality of zero point values ​​ZPV and a plurality of quantization scale coefficients QSC. For example, the first quantization scaling circuit 326_1 can sequentially receive a plurality of zero point values ​​ZPV and a plurality of quantization scale coefficients QSC described above with reference to FIG.

[0446] The first quantization and scaling circuit 326_1 converts the first input vector (i.e.,

number

number

[0447] For example, the first quantization and scaling circuit 326_1 calculates x 11 The first zero point value ZPV1 and α 1_r1 ~α R_r1 Multiply each by x 11 A number of quantized scaled input elements QSIE for x, y, and z can be generated (illustrated by diagonal stripes).

[0448] Similarly, the first quantization and scaling circuit 326_1 calculates x 12 The second zero point value ZPV2 and α 1_r2 ~α R_r2 Multiply each by x 12 A number of quantized scaled input elements QSIE for x, y, and z can be generated (illustrated by diagonal stripes).

[0449] As a result, the first quantized scaled input vector (i.e.

number

[0450] FIG. 58 is a diagram showing the operation of the processing element array of FIG. 48 in more detail. With reference to FIGS. 1 to 4 and 48 to 58, the processing element array 350 may include a plurality of processing element rows PER. Each of the plurality of processing element rows PER may include a plurality of processing elements. However, for a simpler explanation, the operation of the first processing element row PER1 will be representatively explained below. However, the scope of the present disclosure is not limited thereto.

[0451] The first processing element row PER1 includes processing elements PE11 to PE1m. The first processing element row PER1 outputs a first fixed-point quantized scaled input vector (i.e.,

number

[0452] Each of the processing elements PE11 to PE1m can receive a plurality of quantization code bits QSBs and a plurality of zero point correction bits ZPCBs. For example, the processing element PE11 can receive a first plurality of quantization code bits QSBs_1 and a plurality of zero point correction bits ZPCBs. The processing element PE12 can receive a second plurality of quantization code bits QSBs_2 and a plurality of zero point correction bits ZPCBs.

[0453] In a more detailed example, after receiving one zero point correction bit ZPCB, the processing element PE11 11 R quantized code bits QSB corresponding to 1_1_r1 ~b 1_R_r1 After that, the processing element PE11 receives one zero point correction bit ZPCB, and then receives x 12 R quantized code bits QSB corresponding to1_1_r2 ~b 1_R_r2 In this manner, the processing element PE11 can sequentially receive a plurality of quantization code bits QSB and a plurality of zero point correction bits ZPCB.

[0454] The processing element PE11 outputs a fixed-point output element (e.g., y' ) based on the order in which the quantization sign bit QSB and the plurality of zero point correction bits ZPCB are received and the order in which the plurality of fixed-point quantization scaled input elements QSIEfxp are received. 11 For example, the processing element PE11 can calculate a fixed-point output element (e.g., y') based on whether the quantization sign bit QSB or the zero point correction bit ZPCB corresponding to the received fixed-point quantized scaled input element QSIEfxp indicates '0' or '1', in the same manner as described above with reference to Figs. 37 and 38. 11 ) can be calculated.

[0455] Therefore, according to an embodiment of the present disclosure, the product of the zero point value ZPV and the input element is calculated in the matrix multiplier 300. In this case, the matrix multiplication device MMD can calculate the product of the input matrix XM and the weight matrix WM without including a separate multiplier for calculating the product of the zero point value ZPV and the input element.

[0456] In other words, according to the embodiment of the present disclosure, even if the uniform BCQ circuit UBC performs uniform binary coding quantization to make the weight matrix WM asymmetric, the matrix multiplier 300 can calculate the product of the input matrix XM and the weight matrix WM with a small amount of calculation. Therefore, according to the embodiment of the present disclosure, the versatility of the matrix multiplication device MMD can be increased.

[0457] Fig. 59 is a block diagram showing the processing element array of Fig. 26 realized in a systolic array manner. Referring to Figs. 1 to 4, 23 to 48, and 59, the processing element array 250 may include a plurality of processing elements PE arranged in row and column directions. The plurality of processing elements PE may operate in a systolic array manner.

[0458] The processing element array 250 is implemented to sequentially propagate a plurality of fixed-point quantization scaled input elements QSIEfxp in a row direction. For example, the first processing element row PER1 propagates a first fixed-point quantization scaled input vector (i.e.,

number

[0459] For example, the processing element PE11 may receive one fixed-point quantization scaled input element QSIEfxp at a first time point. The processing element PE11 may transmit the fixed-point quantization scaled input element QSIEfxp to the processing element PE12 arranged adjacent to the processing element PE11 in the row direction at a second time point after the first time point. In this manner, the processing element PE included in the first processing element row PER1 may sequentially transmit the multiple fixed-point quantization scaled input elements QSIEfxp provided from the first data type converter 230 to adjacent processing elements.

[0460] The processing element array 250 can be implemented to sequentially propagate the plurality of quantized code bits QSB in a column direction. For example, the first processing element column PEC1 can sequentially propagate a first plurality of quantized code bits QSBs_1 in a column direction.

[0461] For example, the processing element PE11 may receive one quantized code bit QSB at a first time point. The processing element PE11 may transmit the quantized code bit QSB to the processing element PE21 disposed adjacent to the processing element PE11 in the column direction at a second time point after the first time point. In this manner, the processing elements PE included in the first processing element column PEC1 may sequentially transmit the first plurality of quantized code bits QSBs_1 provided from the quantized code bit buffer 240 to adjacent processing elements.

[0462] The processing elements PE can generate different fixed-point output elements OEfxp. The processing element array 250 can be realized so as to sequentially propagate the fixed-point output elements OEfxp in the column direction. For example, the fixed-point output element OEfxp (i.e., y') calculated from the processing element PE11 can be realized as follows: 11 ) are sequentially propagated in the column direction. In this manner, the fixed-point output element OEfxp is transmitted to the second data type converter 260. The manner in which the fixed-point output element OEfxp is propagated is similar to the previously described quantization code bit QSB propagation method, and therefore a detailed description thereof will be omitted.

[0463] That is, each of the processing elements included in the processing element array 250 is realized to receive one or more of the fixed-point output element OEfxp, the quantization sign bit QSB, and the fixed-point quantization scaled input element QSIEfxp from the processing element arranged adjacently. Conversely, each of the processing elements included in the processing element array 250 is realized to transmit one or more of the fixed-point output element OEfxp, the quantization sign bit QSB, and the fixed-point quantization scaled input element QSIEfxp to the processing element arranged adjacently. A more detailed configuration of the processing element PE operating in a systolic array manner will be described in more detail with reference to FIG. 60 below.

[0464] For the sake of simpler explanation, Fig. 59 shows an embodiment in which the fixed-point output element OEfxp, the quantization sign bit QSB, and the fixed-point quantization scaled input element QSIEfxp are propagated in a systolic array manner, but the scope of the present disclosure is not limited thereto. For example, the processing element array 250 may be realized such that only one or two of the fixed-point output element OEfxp, the quantization sign bit QSB, and the fixed-point quantization scaled input element QSIEfxp are propagated in a systolic array manner.

[0465] In one embodiment, each of the processing elements included in the processing element array 250 may operate in response to the same control clock signal. In this case, each of the multiple processing elements may communicate the fixed-point output element OEfxp, the quantization sign bit QSB, and / or the fixed-point quantization scaled input element QSIEfxp to the other processing elements at the same time. However, the scope of the present disclosure is not limited in this respect.

[0466] For the sake of simpler explanation, Fig. 59 representatively illustrates an embodiment in which the processing element array 250 described with reference to Figs. 23 to 48 operates in a systolic array manner, but the scope of the present disclosure is not limited thereto. For example, according to an embodiment of the present disclosure, the processing element array 150 described with reference to Figs. 5 to 22 or the processing element array 350 described with reference to Figs. 49 to 58 may also operate in a systolic array manner.

[0467] Figure 60 is a diagram showing in more detail the configuration of the processing element of Figure 59. With reference to Figures 1 to 4, 23 to 48, and 59 to 60, the processing element PE may include an arithmetic logic unit ALU, an accumulation register REG_ACC, a quantization scaled input element register REG_QSIE, a quantization sign bit register REG_QSB, and an output register REG_OUT. For a simpler explanation, detailed explanations regarding the configuration and operation of the arithmetic logic unit ALU and the accumulation register REG_ACC described above with reference to Figure 38 will be omitted.

[0468] In the following, an embodiment in which the arithmetic logic unit ALU, the accumulation register REG_ACC, the quantization scaled input element register REG_QSIE, the quantization sign bit register REG_QSB, and the output register REG_OUT operate in response to the same control clock signal will be described as a representative example. However, the scope of the present disclosure is not limited thereto.

[0469] The quantized scaled input element register REG_QSIE can sequentially receive the multiple quantized scaled input elements QSIEfxp. The quantized scaled input element register REG_QSIE can receive the multiple quantized scaled input elements QSIEfxp and transmit them to the adjacent processing elements PE and the first input terminal TI1 after one period of the control clock signal has elapsed.

[0470] The quantized sign bit register REG_QSB can sequentially receive a plurality of quantized sign bits QSB. The quantized sign bit register REG_QSB can receive one quantized sign bit QSB and transmit it to the adjacent processing element PE and the second input terminal TI2 after one period of the control clock signal has elapsed.

[0471] The output register REG_OUT can receive the fixed-point output element OEfxp from the accumulation register REG_ACC. For example, the output register REG_OUT can transmit the fixed-point output element OEfxp to an adjacent processing element PE or to the second data type converter 260.

[0472] That is, the output register REG_OUT can provide the fixed-point output element OEfxp to the output register of another adjacent processing element, or can directly provide it to the second data type converter 260. For example, the fixed-point output element OEfxp (e.g., y') calculated by the processing element PE11 can be h1 ) is transmitted to the second data type converter 260 via the processing elements PE21 to PEh1 in sequence. However, the scope of the present disclosure is not limited to this.

[0473] For the sake of simplicity, FIG. 60 illustrates an embodiment in which each of the registers included in the processing element PE receives and outputs data every cycle of the control clock signal, but the scope of the present disclosure is not limited to the specific operation method of the registers in response to the control clock signal.

[0474] Figure 61 is a diagram illustrating the operation of the matrix multiplication device of Figure 1 according to one embodiment. With reference to Figures 1 to 48 and 61, the matrix multiplication device MMD can receive a full input matrix FXM. The full input matrix FXM may include the input matrix XM previously described with reference to Figures 1 to 34. For example, the full input matrix FXM may include a plurality of input matrices XM.

[0475] The matrix multiplication device MMD may receive a full-weight matrix FWM. The full-weight matrix FWM may include the weight matrix WM previously described with reference to Figures 1 to 48. For example, the full-weight matrix FWM may include a plurality of weight matrices WM.

[0476] The uniform BCQ circuit UBC can perform a uniform binary coding quantization operation on each of the multiple weight matrices WM included in the full-weight matrix FWM. For example, the uniform BCQ circuit UBC can generate multiple quantization sign bits QSB, multiple common scale coefficients CSC, and a magnification scale coefficient MSC from each of the multiple weight matrices WM.

[0477] The matrix multiplier 100 can receive a plurality of quantization code bits QSB, a plurality of common scale factors CSC, and a magnification factor MSC. The matrix multiplier 100 can perform matrix multiplication on the full-input matrix FXM and the full-weight matrix FWM based on the plurality of quantization code bits QSB, a plurality of common scale factors CSC, and a magnification factor MSC.

[0478] The matrix multiplier 100 may perform matrix multiplication on the full-input matrix FXM and the full-weight matrix FWM through one of various tiling techniques. For example, the matrix multiplier 100 may calculate the full-output matrix FYM by sequentially calculating the products of a plurality of input matrices XM and a plurality of weight matrices WM and then combining the calculated results.

[0479] FIG. 62 is a diagram showing the full-input matrix of FIG. 61. Referring to FIG. 1 to FIG. 48 and FIG. 61 to FIG. 62, the full-input matrix FXM may include a plurality of input matrices XM. In other words, the full-input matrix FXM is tiled with a plurality of input matrices XM arranged in the row direction and the column direction. Hereinafter, for a simpler explanation, the input matrix arranged in the (i)th row and the (j)th column of the full-input matrix FXM is referred to as "XM_ij".

[0480] In one embodiment, the input matrix XM, previously described with reference to FIGS. 1 to 48, may be one of a plurality of input matrices XM included in the full-input matrix FXM.

[0481] In one embodiment, each of the input matrices XM included in the full-input matrix FXM may have the same row size and column size as each other. For example, each of the input matrices XM may include 'n' input elements per row. Each of the input matrices XM may include 'h' input elements per column.

[0482] The row size of the full-input matrix FXM may be an integer multiple of the row size of each of the multiple input matrices XM. For example, one row of the full-input matrix FXM may include 'N' input elements, where 'N' may be an integer multiple of 'n'.

[0483] The column size of the full-input matrix FXM may be an integer multiple of the column size of each of the multiple input matrices XM. For example, one column of the full-input matrix FXM may include 'H' input elements, where 'H' may be an integer multiple of 'h'.

[0484] Figure 63 is a diagram showing the full-weight matrix of Figure 61. Referring to Figures 1 to 48 and Figures 61 to 63, the full-weight matrix FWM can include a plurality of weight matrices WM. In other words, the full-weight matrix FWM is tiled with a plurality of weight matrices WM arranged in the row direction and the column direction. Hereinafter, for a simpler explanation, the weight matrix arranged in the (i)th row and (j)th column of the full-weight matrix FWM is referred to as "WM_ij".

[0485] In one embodiment, the weight matrix WM described above with reference to FIGS. 1-48 may be one of a plurality of weight matrices WM included in the full-input matrix FXM.

[0486] In one embodiment, each of the weight matrices WM included in the full-weight matrix FWM may have the same row size and column size. For example, each of the weight matrices WM may include 'm' weights for each row. Each of the weight matrices WM may include 'n' weights for each column.

[0487] The row size of the full-weight matrix FWM may be an integer multiple of the row size of each of the weight matrices WM. For example, one row of the full-weight matrix FWM may include 'M' weights. In this case, 'M' may be an integer multiple of 'm'.

[0488] The column size of the full-weight matrix FWM may be an integer multiple of the row size of each of the weight matrices WM. For example, one column of the full-weight matrix FWM may include 'N' weights. In this case, 'N' may be an integer multiple of 'n'.

[0489] The uniform BCQ circuit UBC can perform a uniform binary coding quantization operation on each of a plurality of weight matrices WM included in the full-weight matrix FWM. In this case, a plurality of common scale coefficients CSC generated based on the weight matrix WM_11 may be different from a plurality of common scale coefficients CSC generated based on the weight matrix WM_12. Similarly, a plurality of quantization code bits QSB generated based on the weight matrix WM_11 may be different from a plurality of quantization code bits QSB generated based on the weight matrix WM_12. A specific manner in which the uniform BCQ circuit UBC performs a uniform binary coding quantization operation on each weight matrix WM is similar to that described above with reference to Figs. 1 to 48, and therefore a detailed description thereof will be omitted.

[0490] Figure 64 is a diagram showing the full-output matrix of Figure 61. Referring to Figures 1 to 48 and Figures 61 to 64, the full-output matrix FYM may correspond to the product of the full-input matrix FXM and the full-weight matrix FWM.

[0491] The full-output matrix FYM may include multiple sub-matrices FYM_sub arranged in row and column directions. In the following, for easier explanation, the sub-matrix arranged in the (i)th row and (j)th column of the full-output matrix FYM is referred to as “FYM_sub_ij”.

[0492] Each of the sub-matrices FYM_sub may have the same row size and column size as the others. The row size of each of the sub-matrices FYM_sub may be the same as the row size of the input matrix XM. The column size of each of the sub-matrices FYM_sub may be the same as the column size of the weight matrix WM. For example, each of the sub-matrices FYM_sub may include 'm' output elements per row. Each of the sub-matrices FYM_sub may include 'h' output elements per column.

[0493] The row size of the full-output matrix FYM may be the same as the row size of the full-weight matrix FWM, for example, the row size of the full-output matrix FYM may be 'M'.

[0494] The column size of the full-output matrix FYM may be the same as the column size of the full-input matrix FXM, for example, the column size of the full-output matrix FYM may be 'H'.

[0495] The matrix multiplication device MMD can calculate the full-output matrix FYM in units of sub-matrix FYM_sub. For example, the matrix multiplication device MMD can calculate one sub-matrix FYM_sub by adding the products of a plurality of tiled input matrices XM and a plurality of tiled weight matrices WM.

[0496] To give a more detailed example, when 'N' is three times 'n', the matrix multiplier 100 can calculate the sub-matrix FYM_sub_11 by sequentially calculating the product of the input matrix XM_11 and the weight matrix WM_11, the product of the input matrix XM_12 and the weight matrix WM_21, and the product of the input matrix XM_13 and the weight matrix WM_31, and then adding them. In this case, each product of the tiled input matrix and the tiled weight matrix may correspond to the output matrix YM previously described with reference to Figures 1 to 3 and Figures 15 to 33. That is, the matrix multiplication device MMD can calculate a first output matrix based on the product of the input matrix XM_11 and the weight matrix WM_11, calculate a second output matrix based on the product of the input matrix XM_12 and the weight matrix WM_21, and calculate a third output matrix based on the product of the input matrix XM_13 and the weight matrix WM_31. Thereafter, the matrix multiplication device MMD can add the above-mentioned first to third output matrices to calculate the sub-matrix FYM_sub_11, however, the scope of the present disclosure is not limited thereto.

[0497] That is, the matrix multiplication device MMD can be realized to accumulate a plurality of output matrices to calculate one sub-matrix (i.e., a portion of the full-output matrix FYM). For example, the matrix multiplication device MMD can be realized to temporarily store a plurality of output matrices in an external volatile memory device (e.g., an SRAM device) and then accumulate the output matrices to calculate one sub-matrix. In this manner, the matrix multiplication device MMD can calculate the full-output matrix FYM by sequentially calculating a plurality of sub-matrices FYM_sub.

[0498] Figure 65 is a block diagram showing a neural processing system implemented according to an embodiment. Referring to Figure 65, the neural processing system 2000 may include a central processing unit 2100, a neural processing unit 2200, a volatile memory device 2300, a nonvolatile memory device 2400, and a user interface 2500. The central processing unit 2100, the neural processing unit 2200, the volatile memory device 2300, the nonvolatile memory device 2400, and the user interface 2500 may be connected via a bus BUS.

[0499] The central processing unit 2100 can control the overall operation of the neural processing system 2000. For example, the central processing unit 2100 can control each component of the neural processing system 2000 to drive an artificial intelligence model.

[0500] In one embodiment, the artificial intelligence model implemented by the neural processing system 2000 may be any type of artificial intelligence model, such as a language model, an image identification model, an image generation model, a weather analysis model, etc. For example, the artificial intelligence model implemented by the neural processing system 2000 may be any type of artificial intelligence model, such as GPT-3, GPT-4, Pangu, GShard, Megatron-LM, etc. However, the scope of the present disclosure is not limited thereto.

[0501] In one embodiment, the artificial intelligence models executed by the neural processing system 2000 are capable of performing inference and / or training operations, although the scope of the present disclosure is not limited in this respect.

[0502] Each of the artificial intelligence models may include multiple processing layers. Each of the multiple processing layers may be implemented to receive layer input data and generate layer output data. In this case, the generated layer output data may be used as layer input data for other processing layers. For example, the generated layer output data from a first processing layer may be used as layer input data for a second processing layer. A more detailed description of the artificial intelligence models and processing layers is provided with reference to FIG. 66 below.

[0503] Each of the multiple processing layers may transform layer input data into layer output data based on a matrix multiplication operation. For example, each of the multiple processing layers may multiply an input matrix corresponding to the layer input data by a weight matrix to generate an output matrix corresponding to the layer output data. However, the scope of the present disclosure is not limited thereto, and each of the multiple processing layers may transform an input matrix corresponding to the layer input data in an arbitrary manner to generate output data. For example, each of the multiple processing layers may be realized to sequentially multiply an input matrix corresponding to the layer input data by a plurality of weight matrices to generate layer output data, or to transform an input matrix into layer output data based on an arbitrary transformation parameter. That is, the scope of the present disclosure is not limited to a specific manner in which each of the multiple processing layers transforms layer input data.

[0504] The neural processing unit 2200 may include a matrix multiplication device 2210. The matrix multiplication device 2210 may perform at least some of the operations included in the multiple processing layers. For example, the matrix multiplication device 2210 may perform matrix multiplication operations included in the multiple processing layers.

[0505] In one embodiment, matrix multiplication operations may account for a large portion of the processing load used by the neural processing system 2000 to execute each of the multiple processing layers.

[0506] In one embodiment, the matrix multiplication device 2210 can be realized by the matrix multiplication device MMD described above with reference to Figures 1 to 58. In this case, the neural processing unit 2200 can execute the operations included in multiple processing layers with a smaller amount of calculations. Therefore, the neural processing system 2000 including the matrix multiplication device MMD according to the embodiment of the present disclosure can run an artificial intelligence model at a faster speed.

[0507] The volatile memory device 2300 can be used as an operating memory for the neural processing unit 2200. For example, the volatile memory device 2300 can temporarily store data generated during the operation of the neural processing unit 2200.

[0508] In one embodiment, the neural processing unit 2200 can access the volatile memory device 2300 to perform operations included in the multiple processing layers. For example, the neural processing unit 2200 can be implemented to read parameters stored in the volatile memory device 2300 and perform operations on layer input data, or to temporarily store intermediate data generated during operations in the volatile memory device 2300.

[0509] In one embodiment, the calculation speed of the neural processing unit 2200 may be faster than the access speed of the neural processing unit 2200 to the volatile memory device 2300. As a result, a bottleneck phenomenon may occur in the operation speed of the artificial intelligence model due to the communication speed between the neural processing unit 2200 and the volatile memory device 2300.

[0510] In one embodiment, when the matrix multiplication device 2210 is realized by the matrix multiplication device MMD described above with reference to FIGS. 1 to 58, one output element can be calculated based on the processing element PE included in the matrix multiplier MMD. In this case, since one output element is generated without combining the output results of multiple processing elements PE, the number of accesses to the volatile memory device 2300 of the neural processing unit 2200 is minimized, and the artificial intelligence model is driven based on a smaller size of the processing element array (i.e., a smaller number of processing elements PE). Therefore, according to the embodiment of the present disclosure, the operating speed of the artificial intelligence model can be improved and the production cost of the neural processing system 2000 can be minimized.

[0511] In one embodiment, when the matrix multiplication unit 2210 is implemented by the matrix multiplication unit MMD described above with reference to Figures 1 to 58, each of the processing elements PE may not include a '1-bit adder' for performing a multiplication operation on the magnification scale factor MSC. That is, according to an embodiment of the present disclosure, the magnification scaling circuit provides the result of scaling the input elements to the processing element array, so that the size and production cost of each of the processing elements PE included in the matrix multiplication unit MMD can be reduced.

[0512] In one embodiment, the volatile memory device 2300 may be implemented with any type of volatile memory, such as dynamic random access memory (DRAM) or static random access memory (SRAM).

[0513] In one embodiment, the volatile memory device 2300 may be used as a buffer memory, working memory, or cache memory for the central processing unit 2100. However, the scope of the present disclosure is not limited in this respect.

[0514] The non-volatile memory device 2400 may store data for operation of the neural processing system 2000. For example, the non-volatile memory device 2400 may store various types of data, such as parameters for driving an operating system (OS) or an artificial intelligence model of the neural processing system 2000. However, the scope of the present disclosure is not limited thereto.

[0515] The central processing unit 2100 can communicate with a user through a user interface 2500. The central processing unit 2100 can provide model input data provided by a user through the user interface 2500 to the volatile memory device 2300 or the neural processing unit 2200. The central processing unit 2100 can return model output data generated by the artificial intelligence model based on the model input data to the user through the user interface 2500.

[0516] Figure 66 is a block diagram showing an artificial intelligence model driven by the neural processing system of Figure 65. Referring to Figure 65, the neural processing system 2000 can drive an artificial intelligence model AIM.

[0517] The artificial intelligence model AIM can receive the model input data MID. The artificial intelligence model AIM can include a first processing layer PL_1 to an L-th processing layer PL_L.

[0518] The artificial intelligence model AIM can generate model output data MOD by sequentially converting the model input data MID through the first processing layer PL_1 to the L-th processing layer PL_L. For example, the first processing layer PL_1 can receive the model input data MID and generate the second layer input data LID_2. The second processing layer PL_2 can receive the second layer input data LID_2 and generate the third layer input data LID_3. In this manner, the L-th processing layer PL_L can receive the L-th layer input data LID_L and generate the model output data MOD.

[0519] Each of the first processing layer PL_1 to the Lth processing layer PL_L may convert received data into output data through various types of operations. For example, the operations performed by the first processing layer PL_1 to convert the model input data MID into the second layer input data LID_2 include a matrix multiplication operation. Similarly, each of the first processing layer PL_1 to the Lth processing layer PL_L may need to perform a matrix multiplication operation to convert the received layer input data. However, the scope of the present disclosure is not limited thereto, and some of the first processing layer PL_1 to the Lth processing layer PL_L may not need to perform a matrix multiplication operation.

[0520] In one embodiment, the matrix multiplication operations performed by each of the first processing layer PL_1 to the L-th processing layer PL_L may be performed through a matrix multiplication device 2210.

[0521] In one embodiment, when the matrix multiplication device 2210 is realized by the matrix multiplication device MMD previously described with reference to Figures 1 to 48, the matrix multiplication device 2210 can output the result of the matrix multiplication operation at a faster speed. Therefore, according to the embodiment of the present disclosure, the operating speed of the artificial intelligence model AIM can be improved.

[0522] For the sake of simpler explanation, FIG. 66 representatively describes an embodiment in which the artificial intelligence model AIM is composed of a plurality of processing layers that operate in series, but the scope of the present disclosure is not limited to this. For example, the artificial intelligence model AIM may further include a processing layer that operates in parallel with at least a part of the first processing layer PL_1 to the Lth processing layer PL_L described above. That is, the scope of the present disclosure is not limited to a specific implementation method of the artificial intelligence model AIM.

[0523] The above is a specific embodiment for carrying out the present disclosure. The present disclosure includes not only the above-mentioned embodiment, but also an embodiment that can be simply modified or easily modified. The present disclosure also includes a technique that can be easily modified and implemented using the embodiment. Therefore, the scope of the present disclosure should not be limited to the above-mentioned embodiment, but should be determined not only by the claims described below, but also by equivalents to the claims of the present disclosure. [Explanation of symbols]

[0524] MMD: Matrix Multiplication Device XM: Input Matrix WM: Weight matrix YM: Output matrix 100: Matrix multiplier UBC: Uniform BCQ circuit

Claims

1. 1. A matrix multiplier comprising: an input vector scaler for generating a first quantized scaled input vector based on a first input vector, a plurality of common scale factors, and first through Rth scale factor scale factors, where R is an integer equal to or greater than 2; a first data type converter that generates a first fixed-point quantized scaled input vector based on the first quantized scaled input vector; a processing element array including a first processing element that generates a first fixed point output element based on the first fixed point quantized scaled input vector and a first plurality of quantized code bits, and a second processing element that generates a second fixed point output element based on the first fixed point quantized scaled input vector and a second plurality of quantized code bits; a second data type converter that converts data types of the first fixed-point output elements and the second fixed-point output elements, respectively, to generate first output elements and second output elements, and outputs a first output vector including the first output elements and the second output elements; A matrix multiplier, including:

2. the first input vector includes a first input element and a second input element; The first quantized scaled input vector is a first plurality of quantized scaled input elements generated based on the first input element, a first common scale factor that is one of the plurality of common scale factors, and the first through Rth scale factor scale factors; a second plurality of quantized scaled input elements generated based on the second input element, a second common scale factor that is one of the plurality of common scale factors, and the first through Rth scale factor scale factors; 2. The matrix multiplier of claim 1, comprising:

3. The input vector scaler comprises: a factor scaling circuit that generates input elements scaled by a first factor through an Rth factor based on the first input element and the first through Rth factor scale factors, and generates input elements scaled by (R+1)th through (2R)th factors based on the second input element and the first through Rth factor scale factors; a common scaling circuit for generating the first plurality of quantized scaled input elements based on a product of each of the first through Rth scaled input elements relative to the first common scale factor, and for generating the second plurality of quantized scaled input elements based on a product of each of the (R+1)th through (2R)th scaled input elements relative to the second common scale factor; 3. The matrix multiplier of claim 2, comprising:

4. The input vector scaler comprises: a common scaling circuit that generates a first commonly scaled input element based on a product of the first input element and the first common scale factor, and that generates a second commonly scaled input element based on a product of the second input element and the second common scale factor; a factor scaling circuit for generating the first plurality of quantized scaled input elements based on a product of each of the first through R factor scale coefficients to the first common scaled input element, and for generating the second plurality of quantized scaled input elements based on a product of each of the first through R factor scale coefficients to the second common scaled input element; 3. The matrix multiplier of claim 2, comprising:

5. The input vector scaler comprises: a magnification scaling circuit that generates first to Rth quantization scale coefficients based on a product of each of the first to Rth magnification scale coefficients and the first common scale coefficient, and generates (R+1)th to (2Rth) quantization scale coefficients based on a product of each of the first to Rth magnification scale coefficients and the second common scale coefficient; a first quantization scaling circuit that generates the first plurality of quantized scaled input elements based on the first input element and the first through Rth quantization scale factors, and generates the second plurality of quantized scaled input elements based on the second input element and the (R+1)th through (2R)th quantization scale factors; 3. The matrix multiplier of claim 2, comprising:

6. The input vector scaler comprises: generating a second quantized scaled input vector based on a second input vector, the plurality of common scale factors, and the first through Rth scale factor scale factors; The first data type converter includes: further configured to generate a second fixed-point quantized scaled input vector based on the second quantized scaled input vector; The processing element array comprises: a third processing element for generating a third fixed point output element based on the second fixed point quantization scaled input vector and the first plurality of quantization code bits, and a fourth processing element for generating a fourth fixed point output element based on the second fixed point quantization scaled input vector and the second plurality of quantization code bits; The second data type converter comprises: and further configured to convert the data types of the third fixed-point output element and the fourth fixed-point output element, respectively, to generate a third output element and a fourth output element, and output a second output vector including the third output element and the fourth output element.

6. The matrix multiplier of claim 5.

7. the second input vector includes a third input element and a fourth input element; The input vector scaler comprises: a second quantization scaling circuit for generating a third plurality of quantized scaled input elements included in the second quantization scaled input vector based on the third input element and the first through Rth quantization scale factors, and for generating a fourth plurality of quantized scaled input elements included in the second quantization scaled input vector based on the fourth input element and the (R+1)th through (2R)th quantization scale factors; The matrix multiplier of claim 6 further comprising:

8. the first processing element is arranged in a first processing element row and a first processing element column of the processing element array; the second processing element is arranged in the first processing element row and a second processing element column of the processing element array; the third processing element is arranged in a second processing element row and in the first processing element column of the processing element array; the fourth processing element is arranged in the second processing element row and the second processing element column of the processing element array; 7. The matrix multiplier of claim 6.

9. the first common scale factor has a floating point data type; The magnification scaling circuit includes: and generating the first to Rth quantization scale coefficients by changing a value of an exponent part of the first common scale coefficient based on the first to Rth magnification scale coefficients.

6. The matrix multiplier of claim 5.

10. the first to Rth magnification scale coefficients form a geometric progression having a common ratio of 2; 2. The matrix multiplier of claim 1.

11. the input vector scaler is configured to generate the first quantized scaled input vector further based on a plurality of zero points; the first processing element is configured to generate the first fixed-point output element further based on a plurality of zero point correction bits respectively corresponding to the plurality of zero points; the second processing element configured to generate the second fixed-point output element further based on the plurality of zero point correction bits.

2. The matrix multiplier of claim 1.

12. 1. A method of operating a matrix multiplication apparatus, comprising: receiving first through N-th weights from an external device (where N is an integer equal to or greater than 2); performing uniform binary coding quantization on the first through N-th weights to generate first through N-th common scale coefficients, first through R-th scale factor scale coefficients, and first through (N×R)-th quantized code bits, where N and R are integers equal to or greater than 2; receiving first through Nth input elements from the external device; quantizing and scaling the first through Nth input elements based on the first through Nth common scale coefficients and first through Rth scale factor scale coefficients to generate first through (N×R)th quantized and scaled input elements; outputting first output elements generated based on the first through (N×R) quantized code bits and the first through (N×R) quantized scaled input elements; 4. A method of operation comprising:

13. The step of generating the first through (N×R) quantized scaled input elements comprises: multiplying the first through Rth magnification scale coefficients by the first through Nth common scale coefficients, respectively, to generate first through (N×R)th quantization scale coefficients; multiplying a corresponding one of the first through Nth input elements by each of first through (N×R)th quantization scale coefficients to generate the first through (N×R)th quantization scaled input elements; The method of claim 12, comprising:

14. The first to Nth common scale coefficients have a floating-point data type; the first through Rth magnification scale factors are different powers of 2; 14. The method of claim 13.

15. The step of outputting the first output element comprises: converting a data type of the first through (N×R)th quantization scaled input elements into a fixed point to generate first through (N×R)th fixed point quantization scaled input elements; computing a first fixed-point output element based on the first through (N×R)th quantized code bits and the first through (N×R)th fixed-point quantized scaled input elements; converting a data type of the first fixed-point output element into a floating-point data type and outputting the first output element; The method of claim 12, comprising:

16. The first fixed-point output element is a first value obtained by adding a first plurality of quantized scaled input elements corresponding to a first plurality of quantized code bits indicating '0' among the first to (N×R) quantized code bits among the first to (N×R) quantized scaled input elements, a second value obtained by adding a second plurality of quantized scaled input elements corresponding to a second plurality of quantized code bits indicating '1' among the first to (N×R) quantized code bits among the first to (N×R) quantized scaled input elements, minus 16. The method of claim 15.

17. 1. A matrix multiplier comprising: an input vector scaler that generates a first scaled input vector based on a first input vector and first through Rth scale factors, where R is an integer equal to or greater than 2; a first data type converter that generates a first fixed-point scaled input vector based on the first scaled input vector; a processing element array including a first processing element that generates a first fixed-point partial product based on the first fixed-point scaled input vector and a first plurality of quantization code bits; a second data type converter for converting a data type of the first fixed-point partial product to generate a first partial product; a common scaler that generates a first output element based on a product of the first partial product and a first common scale factor, and outputs a first output vector including the first output element; A matrix multiplier, including:

18. The first input vector includes first to Nth input elements (where N is an integer equal to or greater than 2); The input vector scaler comprises: a factor scaling circuit for multiplying the first through Nth input elements, respectively, by the first through Rth scale factors to generate first through (N×R)th factor scaled input elements included in the first scaled input vector; 20. The matrix multiplier of claim 17, comprising:

19. The processing element array comprises: a second processing element that generates a second fixed-point partial product based on the first fixed-point scaled input vector and a second plurality of quantization sign bits; The second data type converter comprises: further configured to convert the data type of the second fixed-point partial products to generate second partial products, respectively; The common scaler comprises: further configured to generate second output elements, the second output elements being included in the first output vector, based on the products of the second partial products and a second common scale factor.

20. The matrix multiplier of claim 17.

20. the first fixed-point scaled input vector includes a first plurality of fixed-point scaled input elements and a second plurality of fixed-point scaled input elements; The first processing element comprises: from a first value corresponding to a sum of the first plurality of fixed-point scaled input elements, each of which corresponds to a quantization code bit indicating '0' among the first plurality of quantization code bits; calculating a value obtained by subtracting a second value corresponding to a sum of the second plurality of fixed-point scaled input elements, the second value corresponding to a quantization code bit indicating a '1' among the first plurality of quantization code bits; generating the first fixed-point partial product.

20. The matrix multiplier of claim 17.