Matrix multiplication device, matrix multiplication system including same, and operation method of matrix multiplication device

By dividing the weight matrix into multiple trimming submatrices and using the processing element array to perform matrix multiplication operations, the problem of large amount and slow calculation of matrix multiplication in the prior art is solved, and the rapid execution of matrix multiplication and resource saving is achieved.

CN120123627APending Publication Date: 2025-06-10SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410846033.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-07
Filing Date
2024-06-27
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art has large calculations and slow speeds when performing matrix multiplication, making it difficult to meet the fast computing needs of artificial intelligence models.

Method used

By dividing the weight matrix into multiple trimming submatrices and using the processing element array to perform matrix multiplication operations, the calculation amount and resource requirements are reduced.

Benefits of technology

It realizes the rapid execution of matrix multiplication, reduces the computational volume and resource requirements, and improves the operation speed of artificial intelligence models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123627A_ABST
    Figure CN120123627A_ABST
Patent Text Reader

Abstract

The invention provides a matrix multiplication device, a matrix multiplication system including the same, and an operation method of the matrix multiplication device. The matrix multiplication apparatus includes a weight memory circuit, an input matrix buffer, a first array of processing elements, and a second array of processing elements. A weight memory circuit stores a first pruned sub-matrix and a second pruned sub-matrix. The input matrix buffer receives an input matrix including a plurality of input elements, outputs a first plurality of input elements corresponding to a first plurality of remaining weights of the first pruned sub-matrix, and outputs a second plurality of input elements corresponding to a second plurality of remaining weights of the second pruned sub-matrix. A first processing element array receives the first plurality of remaining weights and the first plurality of input elements and outputs a first sub-output matrix. A second processing element array receives the second plurality of remaining weights and the second plurality of input elements and outputs a second sub-output matrix.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is based on and claims priority to Korean Patent Application No. 10-2023-0177028, filed with the Korean Intellectual Property Office on December 7, 2023, the entire contents of which are incorporated herein by reference. Technical Field

[0002] The disclosed aspects relate to semiconductor devices, and more particularly, to a matrix multiplication device, a matrix multiplication system, and a matrix multiplication method for performing matrix multiplication. Background Art

[0003] As artificial intelligence technology has recently developed, the amount of computation in artificial intelligence models is rapidly increasing. Accordingly, various techniques are being studied to shorten the operation time of artificial intelligence models.

[0004] Generally, most of the operation time of an artificial intelligence model is spent on matrix multiplication. For example, an artificial intelligence model spends most of its operation time calculating an output matrix by performing multiplication of an input matrix and a weight matrix. Accordingly, various algorithms (such as pruning, etc.) are being studied to perform multiplication of an input matrix and a weight matrix with less computation. Summary of the Invention

[0005] The disclosed aspects may solve the above technical problems and / or other technical problems. For example, one or more aspects of the disclosure provide a matrix multiplier and a matrix multiplication device including the matrix multiplier, the matrix multiplier and the matrix multiplication device including the matrix multiplier being configured to perform matrix multiplication at a faster speed, with less computation, and / or with reduced resources.

[0006] According to the disclosed aspect, there is provided a matrix multiplication device including: a weight memory circuit configured to store a first pruned sub-matrix including a first plurality of remaining weights and a second pruned sub-matrix including a second plurality of remaining weights; an input matrix buffer configured to receive an input matrix including a plurality of input elements, output a first plurality of input elements corresponding to the first plurality of remaining weights among the plurality of input elements, and output a second plurality of input elements corresponding to the second plurality of remaining weights among the plurality of input elements; a first processing element array configured to receive the first plurality of remaining weights and the first plurality of input elements and output a first sub-output matrix; and a second processing element array configured to receive the second plurality of remaining weights and the second plurality of input elements and output a second sub-output matrix.

[0007] According to another aspect of the disclosure, there is provided a matrix multiplication system, the matrix multiplication system comprising: a matrix pruning device configured to generate a first plurality of pruned sub-matrices and a second plurality of pruned sub-matrices based on a weight matrix; and a matrix multiplication device comprising: a first processing element array configured to generate a first plurality of sub-output matrices based on an input matrix and the first plurality of pruned sub-matrices, a second processing element array configured to generate a second plurality of sub-output matrices based on the input matrix and the second plurality of pruned sub-matrices, and an output merging circuit configured to generate an output matrix based on the first plurality of sub-output matrices and the second plurality of sub-output matrices.

[0008] According to another aspect of the disclosure, there is provided an operating method of a matrix multiplication device, the operating method comprising: generating a plurality of pruning groups by partitioning a weight matrix; classifying the plurality of pruning groups into a first plurality of pruning groups and a second plurality of pruning groups based on the group importance of each of the plurality of pruning groups; generating a first plurality of sub-matrices based on the first plurality of pruning groups; generating a second plurality of sub-matrices based on the second plurality of pruning groups; generating a first plurality of pruned sub-matrices by pruning each of the first plurality of sub-matrices; generating a second plurality of pruned sub-matrices by pruning each of the second plurality of sub-matrices; and generating an output matrix based on the product of the input matrix and each of the first plurality of pruned sub-matrices and the second plurality of pruned sub-matrices. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other aspects, features and advantages of the disclosure will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings, in which: Figure 1 is a block diagram of a matrix multiplication system according to an embodiment of the disclosure.

[0010] Figure 2 shows the operation of a matrix multiplication system implemented to directly multiply a Figure 1 weight matrix with an input matrix.

[0011] Figure 3 , Figure 4 and Figure 5 are diagrams showing the operation of a matrix pruning device of Figure 1 .

[0012] Figure 6 is a block diagram showing in detail the Figure 1 matrix multiplication device.

[0013] Figure 7 is a table showing the remaining weight indices and remaining weights stored in a Figure 6 weight memory circuit according to an embodiment of the disclosure.

[0014] Figure 8 , Figure 9 , Figure 10 , Figure 11 and Figure 12 are block diagrams showing the operations of some configurations of Figure 6 .

[0015] Figure 13 Shows Figure 6 the operation of the output combining circuit of

[0016] Figure 14 is a flowchart of a method for operating a matrix multiplication system according to the disclosed embodiments.

[0017] Figure 15 is a flowchart showing more details of Figure 14 operation S200.

[0018] Figure 16 is a flowchart showing more details of Figure 14 operation S300.

[0019] Figure 17 is a flowchart showing more details of Figure 16 operation S310.

[0020] Figure 18 is a flowchart showing more details of Figure 16 operation S320.

[0021] Figure 19 Shows Figure 6 the iterative operation of the processing element array of

[0022] Figure 20 is a block diagram showing the Figure 19 processing element array using the systolic array method.

[0023] Figure 21 Shows more details of Figure 20 the configuration of the processing element.

[0024] Figure 22 is a timing diagram showing the operation of each of the Figure 20 processing elements.

[0025] Figure 23 Shows a method for adjusting the operation time of the Figure 19 processing element array according to an embodiment.

[0026] Figure 24 Shows a method for adjusting the operation time of the Figure 19 processing element array according to an embodiment.

[0027] Figure 25 is a flowchart of a method for synchronizing Figure 19 the operation times of an array of processing elements.

[0028] Figure 26 and Figure 27 shows the operation of a matrix pruning device according to another embodiment.

[0029] Figure 28 is a block diagram of a matrix multiplication system according to an embodiment.

[0030] Figure 29 shows Figure 28 the operation of the matrix pruning device.

[0031] Figure 30 is a more detailed block diagram showing Figure 28 the matrix multiplication device.

[0032] Figure 31 shows according to an embodiment Figure 1 the matrix multiplication system.

[0033] Figure 32 shows Figure 31 the full input matrix.

[0034] Figure 33 shows Figure 31 the full weight matrix.

[0035] Figure 34 shows Figure 31 the full output matrix.

[0036] Figure 35 is a block diagram of a neural processing system implemented according to an embodiment.

[0037] Figure 36 is a block diagram of an artificial intelligence model run by Figure 35 the neural processing system. DETAILED DESCRIPTION

[0038] Hereinafter, embodiments of the disclosure will be described clearly and in detail to the extent that an ordinary person skilled in the art of the disclosed technology can easily implement the disclosure. Details (such as detailed configurations and structures) are provided only to facilitate a comprehensive understanding of the disclosed embodiments. Therefore, without departing from the spirit and scope of the disclosed technology, those of ordinary skill in the art can implement modifications to the embodiments described in the disclosure. In addition, descriptions of well-known functions and structures are omitted for clarity and conciseness. Configurations in the disclosed drawings or detailed descriptions may be connected to elements other than those shown in the drawings or described in the detailed descriptions. Terms used in the disclosure are defined in consideration of the functions of the disclosure and are not limited to specific functions. The definitions of the terms can be determined based on the details described in the detailed descriptions.

[0039] Elements described with reference to terms used in the detailed description (such as, drivers, blocks, etc.) can be implemented in the form of software, hardware, or a combination thereof. For example, software can be machine code, firmware, embedded code, and application software. For example, hardware can include circuits, electronic circuits, processors, computers, integrated circuit cores, pressure sensors, inertial sensors, microelectromechanical systems (MEMS), passive components, or a combination thereof.

[0040] The features described herein can be implemented in different forms and should not be construed as limited to the examples described herein. Instead, the examples described herein have been provided only to illustrate some of the many feasible ways of implementing the methods, apparatuses, and / or systems described herein, and many feasible ways will be apparent after understanding the disclosure of the present application.

[0041] The terms used herein are only used to describe various examples and will not be used to limit the disclosure. Unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. The terms "comprising", "including", and "having" indicate the presence of the recited features, numbers, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, numbers, operations, components, elements, and / or combinations thereof.

[0042] Unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as commonly understood by those of ordinary skill in the art to which the present disclosure pertains and based on an understanding of the disclosure of the present application. Unless explicitly defined as such herein, terms (such as those defined in common dictionaries) should be interpreted as having a meaning consistent with their meaning in the relevant art and the context of the disclosure of the present application, and should not be interpreted in an idealized or overly formalized sense. The use of the term "may" herein with respect to an example or embodiment (e.g., with respect to what an example or embodiment may include or implement) means that there is at least one example or embodiment in which such a feature is included or implemented, while not all example embodiments are limited thereto.

[0043] Hereinafter, for more concise description, matrices are referred to by the square brackets "[" and "]". However, the scope of the disclosure is not limited to the symbolic method.

[0044] Figure 1 is a block diagram of a matrix multiplication system according to an embodiment of the disclosure. Referring to Figure 1 , the matrix multiplication system MMS may include a matrix pruning device 100 and a matrix multiplication device 200.

[0045] In one embodiment, the matrix multiplication system MMS can be used to drive an artificial intelligence model. For example, the matrix multiplication system MMS can be used to perform the matrix multiplication operations required to execute an artificial intelligence model. However, the disclosed scope is not limited to the specific embodiments using the matrix multiplication system MMS, and thus, according to another embodiment, the matrix multiplication system MMS can be used to implement or execute other operations, processes, and functions. In one example embodiment, the artificial intelligence model can receive and process multimedia data. For example, the multimedia data can include one or more of text data, voice data, and image data. The artificial intelligence model can perform corresponding inference operations based on the input multimedia data.

[0046] The matrix pruning device 100 can receive the weight matrix WM. The weight matrix WM can include a plurality of weights. For example, the weight matrix WM can be represented as given in Equation 1.

[0047] (Equation 1) Referring to Equation 1, WM can represent the weight matrix WM, to can respectively represent different weights. For example, can be the weight set in the i-th row and j-th column of the weight matrix WM.

[0048] Hereinafter, an embodiment in which the weight matrix WM includes "n" weight rows and "m" weight columns will be representatively described, where n and m are positive integers.

[0049] In one embodiment, each of the weights included in the weight matrix WM can have a floating-point data type. However, the disclosed scope is not limited thereto, and thus, according to another embodiment, the weights included in the weight matrix WM can have other data types.

[0050] The matrix pruning device 100 can generate a plurality of pruned small sub-matrices PSM_S and a plurality of pruned large sub-matrices PSM_L based on the weight matrix WM. For example, the matrix pruning device 100 can generate a plurality of pruned small sub-matrices PSM_S by pruning a small sub-matrix (hereinafter referred to as SM_S) corresponding to a part of the weight matrix WM, and can generate a plurality of pruned large sub-matrices PSM_L by pruning a large sub-matrix (hereinafter referred to as SM_L) corresponding to another part of the weight matrix WM. For example, the small sub-matrix corresponds to the first part of the weight matrix WM, and the large sub-matrix corresponds to the second part of the weight matrix WM.

[0051] According to an embodiment, the matrix pruning device 100 may generate a plurality of pruned small submatrices PSM_S by changing some of the weights included in the small submatrix SM_S to zero elements (e.g., "0"), and may generate a plurality of pruned large submatrices PSM_L by changing some of the weights included in the large submatrix SM_L to zero elements. For example, the matrix pruning device 100 may generate a plurality of pruned small submatrices PSM_S by changing one or more of the weights included in the small submatrix SM_S to zero elements (i.e., "0"), and may generate a plurality of pruned large submatrices PSM_L by changing one or more of the weights included in the large submatrix SM_L to zero elements.

[0052] According to an embodiment, the matrix pruning device 100 may approximate the weight matrix WM based on the plurality of pruned small submatrices PSM_S and the plurality of pruned large submatrices PSM_L. In this case, a plurality of zero elements may be included in the plurality of pruned small submatrices PSM_S and the plurality of pruned large submatrices PSM_L. Therefore, in an example case where a matrix multiplication operation of the weight matrix WM is performed based on the plurality of pruned small submatrices PSM_S and the plurality of pruned large submatrices PSM_L, the amount of computation required for the matrix multiplication operation may be reduced. The operation of the matrix pruning device 100 will be described in more detail with reference to Figure 3 、 Figure 4 and Figure 5 The operation of the matrix pruning device 100 will be described in more detail.

[0053] In one embodiment, the matrix pruning device 100 may be used in training an artificial intelligence (AI) model executed based on a matrix multiplication system MMS. For example, while the AI model is being trained, the matrix pruning device 100 may be implemented to generate a plurality of pruned small submatrices PSM_S and a plurality of pruned large submatrices PSM_L. However, the disclosed scope is not limited to the detailed embodiments of using the matrix pruning device 100. Therefore, according to another embodiment, the matrix pruning device 100 may be implemented to perform other operations, processes, and / or functions related to various other types of tasks.

[0054] The matrix multiplication device 200 may receive an input matrix XM. The input matrix XM may include a first input element row to an m-th input element row (hereinafter referred to as "IER") and a first input element column to a b-th input element column (hereinafter referred to as "IEC"), where b is a positive integer. Each of the first input element row to the m-th input element row may include a plurality of input elements. For example, the input matrix XM may be represented as given in Equation 2.

[0055] (Equation 2) Referring to Equation 2, XM may represent the input matrix XM, and to can represent different input elements. For example, to can represent the input elements included in the first input element row of the input matrix XM, and to can represent the input elements included in the m-th input element row.

[0056] Hereinafter, an embodiment in which the input matrix XM includes m input element rows and b input element columns will be representatively described.

[0057] In one embodiment, each input element included in the input matrix XM may have a floating-point data type. However, the disclosed scope is not limited thereto, and thus, according to another embodiment, each input element included in the input matrix XM may have another data type.

[0058] The matrix multiplication device 200 may receive a plurality of trimmed small submatrices PSM_S and a plurality of trimmed large submatrices PSM_L from the matrix pruning device 100. The matrix multiplication device 200 may generate an output matrix YM based on the multiplication of the input matrix XM with each of the plurality of trimmed small submatrices PSM_S and the plurality of trimmed large submatrices PSM_L. For example, the matrix multiplication device 200 may generate the output matrix YM by combining the multiplication results into each of the plurality of trimmed small submatrices PSM_S and the plurality of trimmed large submatrices PSM_L. The more detailed configuration and operation of the matrix multiplication device 200 will be described in more detail with reference to Figures 6 to 13 The more detailed configuration and operation of the matrix multiplication device 200 will be described in more detail.

[0059] In one embodiment, the matrix multiplication device 200 may be used for inference of an AI model executed based on the matrix multiplication system MMS. For example, when the AI performs an inference operation, the matrix multiplication device 200 may be implemented to generate the output matrix YM based on the input matrix XM, the plurality of trimmed small submatrices PSM_S, and the plurality of trimmed large submatrices PSM_L. However, the disclosed scope is not limited to the specific embodiments using the matrix multiplication device 200.

[0060] The output matrix YM may include a plurality of output elements. The output matrix YM may include the first output element row to the n-th output element row (hereinafter referred to as OER). Each of the first output element row to the n-th output element row may include a plurality of output elements. For example, the output matrix YM may be represented as given in Equation 3.

[0061] (Equation 3) In this case, YM may represent the output matrix YM, and to may respectively represent different output elements. For example, to can represent output elements included in the first output element row of the output matrix YM, and to can represent output elements included in the n-th output element row.

[0062] In one embodiment, "n" and "m" can be the same positive integer. For example, the weight matrix WM can be implemented as a square matrix. In this case, the input matrix XM and the output matrix YM can have the same dimensions. However, the disclosed scope is not limited thereto, and thus, according to another embodiment, "n" and "m" can be different.

[0063] In one embodiment, in the example case where the matrix multiplication system MMS calculates the output matrix YM by directly multiplying the weight matrix WM with the input matrix XM, the matrix multiplication system MMS may have to handle a very large amount of computation. In this case, the operation speed of the matrix multiplication system MMS may deteriorate. The operation of the matrix multiplication system MMS for obtaining the output matrix YM by directly multiplying the weight matrix WM with the input matrix XM will be described in more detail with reference to Figure 2 more specifically describes the operation of the matrix multiplication system MMS for obtaining the output matrix YM by directly multiplying the weight matrix WM with the input matrix XM.

[0064] In another example case where the matrix multiplication system MMS calculates the output matrix YM by multiplying the input matrix XM with a weight matrix WM approximated based on a plurality of trimmed small sub-matrices PSM_S and a plurality of trimmed large sub-matrices PSM_L, compared with the case of obtaining the output matrix YM by directly multiplying the input matrix XM with the weight matrix WM, the amount of computation of the matrix multiplication system MMS can be greatly reduced. For example, since a plurality of zero elements can be included in each of the plurality of trimmed small sub-matrices PSM_S and the plurality of trimmed large sub-matrices PSM_L, the amount of computation of the matrix multiplication system MMS can be correspondingly greatly reduced. The operation of the matrix multiplication system MMS for calculating the output matrix YM based on the plurality of trimmed small sub-matrices PSM_S and the plurality of trimmed large sub-matrices PSM_L will be described in more detail with reference to the following drawings.

[0065] Figure 2 illustrates the operation of a matrix multiplication system implemented to directly multiply Figure 1 the weight matrix with the input matrix. Referring to Figure 1 and Figure 2 , the matrix multiplication system MMS can obtain the output matrix YM by directly multiplying the weight matrix WM with the input matrix XM.

[0066] For example, to obtain an output element included in the output matrix YM, the matrix multiplication system MMS may perform m multiplication operations and then perform (m - 1) summation operations. For example, the matrix multiplication system MMS may calculate one of the output elements included in the output matrix YM in a manner similar to the following Equation 4 (e.g., y 11 ).

[0067] (Equation 4) In this way, the matrix multiplication system MMS may need to perform (n × b × m) multiplication operations and (n × b × (m - 1)) summation operations to obtain the output matrix YM. Therefore, due to the excessive amount of calculations processed by the matrix multiplication system MMS, the operation speed of the AI model driven by the matrix multiplication system MMS may deteriorate.

[0068] Figure 3 、 Figure 4 and Figure 5 show Figure 1 the operation of the matrix pruning device. Referring to Figures 1 to 3 , the matrix pruning device 100 may generate a plurality of pruning groups PG based on the weight matrix WM. For example, the matrix pruning device 100 may divide the weight matrix WM into a plurality of pruning groups PG. For example, the matrix pruning device 100 may generate one pruning group PG for every two adjacent rows of the weight matrix WM. In this case, the first weight row and the second weight row of the weight matrix WM may form the first pruning group PG1, the third weight row and the fourth weight row may form the second pruning group PG2, the fifth weight row and the sixth weight row may form the third pruning group PG3, and the seventh weight row and the eighth weight row may form the fourth pruning group PG4. However, the disclosed scope is not limited thereto, and the matrix pruning device 100 may divide the weight matrix WM into a plurality of pruning groups PG in any manner. For example, the matrix pruning device 100 may divide the weight matrix WM such that each of the plurality of pruning groups PG includes a plurality of weight rows spaced apart from each other. For example, the first weight row and the third weight row of the weight matrix WM may form the first pruning group, and the second weight row and the fourth weight row of the weight matrix WM may form the second pruning group. In another embodiment, more than two rows may form a pruning group.

[0069] For a more concise description, in Figure 3 , an embodiment is representatively described in which the matrix pruning device 100 divides the weight matrix WM such that each of the plurality of pruning groups PG includes two adjacent weight rows, but the disclosed scope is not limited thereto, and thus, according to other embodiments, more than two rows may form a pruning group. For example, the matrix pruning device 100 may divide the weight matrix WM such that each of the plurality of pruning groups PG includes 2 i(where i is an integer greater than 0) adjacent weight rows.

[0070] In one embodiment, each of the plurality of pruning groups PG may include the same number of weight rows. However, the disclosed scope is not limited thereto.

[0071] In one embodiment, each of the plurality of pruning groups PG may include different weight rows. That is, each of the plurality of pruning groups PG may include only weight rows. For example, the weight rows included in the first pruning group PG1 may not be included in other pruning groups. However, the disclosed scope is not limited thereto, and thus, according to another embodiment, the weight rows included in the first pruning group PG1 may be included in another pruning group.

[0072] The matrix pruning device 100 may calculate the group importance of each of the plurality of pruning groups PG. For example, the matrix pruning device 100 may calculate the group importance of each of the first pruning group PG1 to the fourth pruning group PG4.

[0073] In one embodiment, the group importance of each of the plurality of pruning groups PG may be determined based on the weights included in the pruning group PG. For example, the group importance of each of the plurality of pruning groups PG may be determined based on the sum of the weights included in the pruning group PG, the sum of the absolute values of the weights included in the pruning group PG, or the sum of the squares of the weights included in the pruning group PG. However, the disclosed scope is not limited thereto, and the group importance of each of the plurality of pruning groups PG may be any kind of function value calculated based on the weights included in the pruning group PG.

[0074] In one embodiment, the group importance of each of the plurality of pruning groups PG may be determined based on the deviation of the weights included in each pruning group PG. For example, a pruning group PG containing a relatively large number of outlier weights may have a relatively high group importance. For example, a first group having a larger number of outlier weights than a second group may be determined to have a higher group importance. However, the disclosed scope is not limited thereto.

[0075] The matrix pruning device 100 may classify the plurality of pruning groups PG into important pruning groups IPG and unimportant pruning groups UIPG based on the group importance. For example, the matrix pruning device 100 may classify a pruning group PG having a group importance higher than a group importance threshold (hereinafter referred to as GIPTH) as an important pruning group IPG, and classify a pruning group PG having a group importance lower than the group importance threshold GIPTH as an unimportant pruning group UIPG.

[0076] For a more detailed example, in the case where the group importance of the first pruning group PG1 and the fourth pruning group PG4 is higher than the group importance threshold GIPTH, the matrix pruning device 100 may classify the first pruning group PG1 and the fourth pruning group PG4 as important pruning groups IPG. In another example case where the group importance of the second pruning group PG2 and the third pruning group PG3 is lower than the group importance threshold GIPTH, the matrix pruning device 100 may classify the second pruning group PG2 and the third pruning group PG3 as unimportant pruning groups UIPG.

[0077] The matrix pruning device 100 may generate a plurality of small sub - matrices SM_S based on the plurality of important pruning groups IPG. For example, the matrix pruning device 100 may generate one small sub - matrix SM_S based on one important pruning group IPG.

[0078] The matrix pruning device 100 may generate a plurality of large sub - matrices SM_L based on the plurality of unimportant pruning groups UIPG. For example, the matrix pruning device 100 may generate one large sub - matrix SM_L based on two or more unimportant pruning groups UIPG.

[0079] For example, the large sub - matrix SM_L may be generated based on a larger number of pruning groups PG than the small sub - matrix SM_S. For example, the matrix pruning device 100 may combine a plurality of unimportant pruning groups UIPG to generate one large sub - matrix SM_L, and may convert one important pruning group IPG into one small sub - matrix SM_S. For example, the matrix pruning device 100 may convert one or more important pruning groups IPG into one or more small sub - matrices SM_S respectively.

[0080] In other words, the number of pruning groups PG included in one large sub - matrix SM_L may be greater than the number of pruning groups PG included in one small sub - matrix SM_S. In this case, the number of weight rows included in one large sub - matrix SM_L may be greater than the number of weight rows included in one small sub - matrix SM_S.

[0081] In one embodiment, the number of pruning groups PG included in one large sub - matrix SM_L may be an integer multiple of the number of pruning groups PG included in one small sub - matrix SM_S.

[0082] In one embodiment, the number of pruning groups PG included in one large sub - matrix SM_L may be 2 i times (where i is any integer greater than or equal to 1) the number of pruning groups PG included in one small sub - matrix SM_S.

[0083] In the following, for a more concise description, an embodiment will be representatively described in which the matrix trimming device 100 converts an important trimming group IPG into a small sub-matrix SM_S and merges two unimportant trimming groups UIPG to generate a large sub-matrix SM_L. However, the disclosed scope is not limited thereto. For example, the matrix trimming device 100 may be implemented to merge multiple important trimming groups IPG to generate a small sub-matrix SM_S, or merge three or more unimportant trimming groups UIPG to generate a large sub-matrix SM_L.

[0084] The matrix trimming device 100 may convert the first trimming group PG1 into the first small sub-matrix SM_S1, and convert the fourth trimming group PG4 into the second small sub-matrix SM_S2. The matrix trimming device 100 may generate the first large sub-matrix SM_L1 by merging the second trimming group PG2 and the third trimming group PG3. In this way, the matrix trimming device 100 may generate multiple small sub-matrices SM_S and multiple large sub-matrices SM_L based on multiple trimming groups PG.

[0085] For example, the matrix trimming device 100 may divide the weight matrix WM into multiple small sub-matrices SM_S and multiple large sub-matrices SM_L. In this case, the number of rows of the weight matrix WM (i.e., "n") may be equal to the sum of the number of rows of multiple small sub-matrices SM_S and the number of rows of multiple large sub-matrices SM_L.

[0086] In one embodiment, the matrix trimming device 100 may generate a large sub-matrix SM_L by merging adjacent unimportant trimming groups UIPG. However, the disclosure is not limited thereto, and the matrix trimming device 100 may generate a large sub-matrix SM_L by merging separated unimportant trimming groups UIPG.

[0087] In one embodiment, the matrix trimming device 100 may be implemented to generate multiple small sub-matrices SM_S by dividing an important trimming group IPG and converting an important trimming group IPG into a large sub-matrix SM_L. Reference will be made to Figures 26 to 27 An embodiment in which the matrix trimming device 100 generates multiple small sub-matrices by dividing an important trimming group IPG will be described in more detail.

[0088] In one embodiment, the matrix trimming device 100 may also generate multiple medium sub-matrices SM_M based on the group importance of each of the multiple trimming groups PG. In this case, the sizes of the small sub-matrix SM_S, the medium sub-matrix SM_M, and the large sub-matrix SM_L may be different. For example, the number of weighted rows of the small sub-matrix SM_S, the medium sub-matrix SM_M, and the large sub-matrix SM_L may be different. Reference will be made to Figures 28 to 30Describe in more detail how the matrix pruning device 100 classifies the weight matrix WM into three or more sub-matrix types with different sizes.

[0089] Referring Figures 1 to 4 , the matrix pruning device 100 can generate a plurality of pruned small sub-matrices (PSM_S) by pruning a plurality of small sub-matrices SM_S respectively. For example, the matrix pruning device 100 can generate a first pruned small sub-matrix PSM_S1 by pruning the first small sub-matrix SM_S1. Hereinafter, for a more concise description, in Figure 4 , the operation of the matrix pruning device 100 that prunes the first small sub-matrix SM_S1 is representatively described, but the disclosed scope is not limited thereto. For example, the matrix pruning device 100 can also prune the second small sub-matrix SM_S2 in a similar manner.

[0090] The matrix pruning device 100 can prune the first small sub-matrix SM_S1 column by column. For example, and can form a pruning unit PU, and and can form another pruning unit PU. However, the disclosed scope is not limited thereto, and each of the plurality of columns included in the first small sub-matrix SM_S1 can form a different pruning unit PU.

[0091] The matrix pruning device 100 can calculate the column importance (or group importance) of each of the plurality of pruning units PU.

[0092] In one embodiment, the column importance of each of the plurality of pruning units PU can be determined based on the values of the weights included in the pruning unit PU. For example, the column importance of each of the plurality of pruning units PU can be determined based on the sum of the values of the weights included in the pruning unit PU, the column importance of each of the plurality of pruning units PU can be determined based on the sum of the absolute values of the weights included in the pruning unit PU, or the column importance of each of the plurality of pruning units PU can be determined based on the sum of the squares of the values of the weights included in the pruning unit PU. However, the disclosed scope is not limited thereto, and the column importance of each of the plurality of pruning units PU can be any type of function value calculated based on the weights included in the pruning unit PU.

[0093] In one embodiment, the column importance of the pruning unit PU and the group importance of the pruning group PG can be calculated based on the same function. However, the disclosure is not limited thereto, and thus, the column importance of the pruning unit PU and the group importance of the pruning group PG can be calculated based on one or more different functions.

[0094] The matrix pruning device 100 may prune each of the multiple pruning units PU based on the column importance of each of the multiple pruning units PU. For example, the matrix pruning device 100 may keep the top four pruning units among the multiple pruning units PU included in the first small sub-matrix SM_S1 that have relatively high column importance, and change the weights included in the remaining pruning units to zero elements (i.e., "0"). For example, the matrix pruning device 100 may obtain four pruning units among the multiple pruning units PU included in the first small sub-matrix SM_S1 that have the four highest column importances. For a more detailed example, the matrix pruning device 100 may keep the weights included in the second, fifth, sixth, and m-th columns corresponding to the pruning units PU with relatively high column importance, and may change the weights included in the other columns to zero elements. In this way, the matrix pruning device 100 may determine the elements of the first pruned small sub-matrix PSM_S1.

[0095] Hereinafter, for more concise description, the weights (e.g., non-zero elements) included in the pruned small sub-matrix PSM_S will be referred to as residual weights RW. For example, the residual weights RW included in the first pruned small sub-matrix PSM_S1 may be referred to as residual weights RW_PSM_S1.

[0096] In one embodiment, the number of columns of the pruned small sub-matrix PSM_S including the residual weights may be determined in advance. That is, in the example case of the pruned small sub-matrix SM_S, the number of pruning units for which the weights will be kept may be determined in advance. For example, the matrix pruning device 100 may be configured to keep a predetermined number of pruning units among the multiple pruning units PU included in the first small sub-matrix SM_S1 that have relatively high column importance.

[0097] For more concise description, hereinafter, an embodiment in which the matrix pruning device 100 performs a pruning operation to keep the four pruning units with the highest column importance among the multiple pruned small sub-matrices PSM_S and converts the remaining pruning units to zero elements will be representatively described. However, the disclosed scope is not limited thereto. For example, the number of pruning units with relatively high column importance obtained by the matrix pruning device 100 may be different from four. According to another embodiment, the matrix pruning device 100 performs a pruning operation by keeping the weights included in the pruning units with column importance higher than the column importance threshold and changing the weights included in the pruning units PU with column importance lower than the column importance threshold to zero elements.

[0098] Similarly, referring to Figures 1 to 5, the matrix trimming device 100 can generate a plurality of trimmed large sub - matrices PSM_L by trimming each of the plurality of large sub - matrices SM_L. For example, the matrix trimming device 100 can generate a first trimmed large sub - matrix PSM_L1 by trimming the first large sub - matrix SM_L1. Hereinafter, for a more concise description, the operation of the matrix trimming device 100 that trims the first large sub - matrix SM_L1 will be representatively described in Figure 5 , but the disclosed scope is not limited thereto. For example, the matrix trimming device 100 can trim each of the plurality of large sub - matrices SM_L described previously with reference to Figure 3 in a similar manner.

[0099] The matrix trimming device 100 can trim the first large sub - matrix SM_L1 in units of columns. For example, to can form a trimming unit PU, and to can form another trimming unit PU. However, the disclosed scope is not limited thereto, and the weights of each of the plurality of columns included in the first large sub - matrix SM_L1 can form different trimming units PU.

[0100] The matrix trimming device 100 can trim each of the plurality of trimming units PU based on the column importance of each of the plurality of trimming units PU. For example, the matrix trimming device 100 can change the weights included in the second to fifth columns and the m - th column of the first large sub - matrix SM_L1 to zero elements. In another example, the matrix trimming device 100 can not change the weights included in the first column and the sixth column of the first large sub - matrix SM_L1. The method by which the matrix trimming device 100 trims each of the plurality of trimming units PU included in the first large sub - matrix SM_L1 is similar to the method described previously with reference to Figure 4 , so its detailed description will not be provided.

[0101] According to the disclosed embodiment, the matrix trimming device 100 can trim the small sub - matrix SM_S and the large sub - matrix SM_L based on trimming units of different sizes. According to an embodiment, the matrix trimming device 100 can trim the small sub - matrix SM_S based on a trimming unit having a smaller size than the trimming unit of the large sub - matrix SM_L. In this case, the important pruning groups IPG included in the small sub - matrix SM_S can be trimmed relatively finely, so the error caused by trimming the weight matrix WM (e.g., the operation error of an artificial intelligence model including the matrix multiplication system MMS) can be minimized.

[0102] Hereinafter, for a more concise description, the weights (e.g., non-zero elements) included in the pruned large sub-matrix PSM_L may be referred to as remaining weights RW. For example, the remaining weights RW included in the first pruned large sub-matrix PSM_L1 may be referred to as remaining weights RW_PSM_L1.

[0103] In one embodiment, the matrix pruning device 100 may prune a plurality of small sub-matrices SM_S such that the number of columns of the remaining weights including each of the plurality of pruned small sub-matrices PSM_S is the same. In other words, the matrix pruning device 100 may prune the plurality of small sub-matrices SM_S at the same pruning ratio.

[0104] In one embodiment, the matrix pruning device 100 may prune a plurality of large sub-matrices SM_L such that the number of columns of the remaining weights including each of the plurality of pruned large sub-matrices PSM_L is the same. That is, the matrix pruning device 100 may prune the plurality of large sub-matrices SM_L at the same pruning ratio.

[0105] In one embodiment, the number of columns of the remaining weights including the plurality of pruned small sub-matrices PSM_S and the number of columns of the remaining weights including the plurality of pruned large sub-matrices PSM_L may be the same. In another embodiment, the number of columns of the remaining weights including the plurality of pruned small sub-matrices PSM_S and the number of columns of the remaining weights including the plurality of pruned large sub-matrices PSM_L may be different. For example, the matrix pruning device 100 may adjust the pruning ratios of the plurality of small sub-matrices SM_S and the plurality of large sub-matrices SM_L to be the same or different from each other. Embodiments in which the matrix pruning device 100 adjusts the pruning ratios of the plurality of small sub-matrices SM_S and the plurality of large sub-matrices SM_L will be described in more detail with reference to Figure 24 more detailed embodiments of the matrix pruning device 100 adjusting the pruning ratios of the plurality of small sub-matrices SM_S and the plurality of large sub-matrices SM_L.

[0106] Hereinafter, for a more concise explanation, the column numbers of the remaining weights including the pruned small sub-matrix PSM_S and the pruned large sub-matrix PSM_L are referred to as remaining weight indices RWI. For example, the remaining weight index RWI of the pruned small sub-matrix PSM_S is referred to as "RWI_PSM_S", and the remaining weight index RWI of the pruned large sub-matrix PSM_L is referred to as "RWI_PSM_L". For example, the remaining weight index RWI_PSM_L may indicate a plurality of column numbers of the remaining weights RW included in the plurality of pruned large sub-matrices PSM_L, and the remaining weight index RWI_PSM_S may indicate a plurality of column numbers of the remaining weights RW included in the plurality of pruned small sub-matrices PSM_S. For a more detailed example, the remaining weight index RWI of the first pruned small sub-matrix PSM_S1 is referred to as RWI_PSM_S1, and the remaining weight index RWI of the first pruned large sub-matrix PSM_L1 is referred to as RWI_PSM_L1. However, the disclosed scope is not limited thereto.

[0107] Figure 6 is a block diagram of a matrix multiplication device that shows in detail Figure 1 . Referring to Figures 1 to 6 , the matrix multiplication device 200 may include a weight memory circuit 210, a control logic circuit 220, an input matrix buffer 230, a first processing element array 240_S, a second processing element array 240_L, and an output merging circuit 250.

[0108] Hereinafter, for a more concise description, embodiments will be representatively described in which the first processing element array 240_S performs matrix multiplication operations on each of a plurality of pruned small sub-matrices PSM_S and an input matrix XM, and the second processing element array 240_L performs matrix multiplication operations on each of a plurality of pruned large sub-matrices PSM_L and the input matrix XM. However, the disclosed scope is not limited thereto, and thus, according to another embodiment, the matrix multiplication device 200 may include a first plurality of processing element arrays and a second plurality of processing element arrays, the first plurality of processing element arrays respectively perform matrix multiplication operations on each of a plurality of pruned small sub-matrices PSM_S and the input matrix XM, and the second plurality of processing element arrays respectively perform matrix multiplication operations on each of a plurality of pruned large sub-matrices PSM_L and the input matrix XM. In addition, according to yet another embodiment, the matrix multiplication device 200 may include only a single processing element array, and the single processing element array performs matrix multiplication operations on both a plurality of pruned small sub-matrices PSM_S and a plurality of pruned large sub-matrices PSM_L and the input matrix XM.

[0109] The weight memory circuit 210 may receive (e.g., and store) a plurality of pruned large sub-matrices PSM_L and a plurality of pruned small sub-matrices PSM_S. That is, the weight memory circuit 210 may receive a weight matrix WM approximated in a format of a plurality of pruned sub-matrices having different sizes.

[0110] In one embodiment, the number of columns of each of the plurality of pruned large sub-matrices PSM_L and the plurality of pruned small sub-matrices PSM_S may be the same.

[0111] In one embodiment, the number of rows of the plurality of pruned large sub-matrices PSM_L may be a positive integer multiple of the number of rows of the plurality of pruned small sub-matrices PSM_S.

[0112] The weight memory circuit 210 may store a plurality of pruned large sub-matrices PSM_L and a plurality of pruned small sub-matrices PSM_S in the form of remaining weights RW and remaining weight indices RWI, respectively. For example, the weight memory circuit 210 may store the remaining weights RW and the column numbers arranging the remaining weights RW. In this case, the weight memory circuit 210 may not store all elements (e.g., zero elements and remaining weights) included in the plurality of pruned large sub-matrices PSM_L and the plurality of pruned small sub-matrices PSM_S. However, the disclosure is not limited thereto, and thus, according to another embodiment, the weight memory circuit 210 or another memory circuit may store some or all elements (e.g., zero elements and remaining weights) included in the plurality of pruned large sub-matrices PSM_L and the plurality of pruned small sub-matrices PSM_S. The method by which the weight memory circuit 210 stores the plurality of pruned large sub-matrices PSM_L and the plurality of pruned small sub-matrices PSM_S will be described in more detail below with reference to Figure 7 how the weight memory circuit 210 stores the plurality of pruned large sub-matrices PSM_L and the plurality of pruned small sub-matrices PSM_S will be described in more detail.

[0113] In one embodiment, the weight memory circuit 210 may be a dynamic random access memory (DRAM). However, the scope of the disclosure is not limited thereto, and the weight memory circuit 210 may be any type of memory circuit.

[0114] The control logic circuit 220 may receive the remaining weight index RWI of each of the plurality of pruned large sub-matrices PSM_L and the plurality of pruned small sub-matrices PSM_S from the weight memory circuit 210. The control logic circuit 220 may control the operation of the input matrix buffer 230 based on the remaining weight index RWI.

[0115] The input matrix buffer 230 may receive the input matrix XM. For example, the input matrix buffer 230 may receive a plurality of input elements included in the input matrix XM.

[0116] The control logic circuit 220 may control the input matrix buffer 230 to output the input elements to be multiplied with the plurality of pruned small sub-matrices PSM_S based on the remaining weight index RWI of the plurality of pruned small sub-matrices PSM_S. Similarly, the control logic circuit 220 may control the input matrix buffer 230 to output the input elements to be multiplied with the plurality of pruned large sub-matrices PSM_L based on the remaining weight index RWI of the plurality of pruned large sub-matrices PSM_L.

[0117] The input matrix buffer 230 can output input elements corresponding to the remaining weight index RWI based on the control of the control logic circuit 220. For example, the input matrix buffer 230 can provide the input elements to be multiplied by the multiple pruned small sub-matrices PSM_S to the first processing element array 240_S based on the remaining weight index RWI_PSM_S provided by the control logic circuit 220. Similarly, the input matrix buffer 230 can provide the input elements to be multiplied by the multiple pruned large sub-matrices PSM_L to the second processing element array 240_L based on the remaining weight index RWI_PSM_L provided by the control logic circuit 220.

[0118] For a more detailed example, in the case where the first processing element array 240_S obtains the product of the first pruned small sub-matrix PSM_S1 and the input matrix XM, the control logic circuit 220 can provide the remaining weight index RWI (e.g., 2, 5, 6, m) to the input matrix buffer 230. In this case, the input matrix buffer 230 can provide the input elements included in the second input element row, the fifth input element row, the sixth input element row, and the m-th input element row of the input matrix XM to the first processing element array 240_S. In this way, according to the remaining weight index RWI provided by the control logic circuit 220, the input matrix buffer 230 can provide the input elements corresponding to each of the multiple pruned small sub-matrices PSM_S to the first processing element array 240_S, and can provide the input elements corresponding to each of the multiple pruned large sub-matrices PSM_L to the second processing element array 240_L respectively.

[0119] The first processing element array 240_S can receive the remaining weights RW_PSM_S included in the multiple pruned small sub-matrices PSM_S from the weight memory circuit 210. For example, the first processing element array 240_S can receive the remaining weight RW_PSM_S1 included in the first pruned small sub-matrix PSM_S1, and can receive the remaining weight RW_PSM_S2 included in the second pruned small sub-matrix PSM_S2.

[0120] The second processing element array 240_L can receive the remaining weights RW_PSM_L included in the multiple pruned large sub-matrices PSM_L from the weight memory circuit 210. Since the method by which the second processing element array 240_L receives the remaining weights is similar to the method by which the first processing element array 240_S receives the remaining weights, a further detailed description thereof will not be provided.

[0121] Each of the first processing element array 240_S and the second processing element array 240_L may generate sub-output matrices having different sizes from each other based on the received remaining weight RW and input element IE. For example, the first processing element array 240_S may generate a plurality of small sub-output matrices SYM_S corresponding to the product of the input matrix XM and each of the plurality of trimmed small sub-matrices PSM_S, and the second processing element array 240_L may generate a plurality of large sub-output matrices SYM_L corresponding to the product of the input matrix XM and each of the plurality of trimmed large sub-matrices PSM_L.

[0122] For example, the first processing element array 240_S may calculate a plurality of output elements based on the remaining weight RW_PSM_S included in the plurality of trimmed small sub-matrices PSM_S and the input element IE corresponding to the plurality of trimmed small sub-matrices PSM_S. Similarly, the second processing element array 240_L may obtain a plurality of output elements based on the remaining weight RW_PSM_L included in the plurality of trimmed large sub-matrices PSM_L and the input element IE corresponding to the plurality of trimmed large sub-matrices PSM_L. Reference will be made to Figure 8 、 Figure 9 、 Figure 10 、 Figure 11 and Figure 12 describe the configuration and operation of the first processing element array 240_S and the second processing element array 240_L in more detail.

[0123] In one embodiment, the number of rows of the plurality of trimmed small sub-matrices PSM_S may be the same as the number of rows of the plurality of small sub-output matrices SYM_S. Similarly, the number of rows of the plurality of trimmed large sub-matrices PSM_L may be the same as the number of rows of the plurality of large sub-output matrices SYM_L.

[0124] In one embodiment, the number of columns of each of the plurality of small sub-output matrices SYM_S and each of the plurality of large sub-output matrices SYM_L may be the same as the number of columns of the output matrix YM.

[0125] The output merging circuit 250 may receive the plurality of small sub-output matrices SYM_S and the plurality of large sub-output matrices SYM_L. The output merging circuit 250 may generate the output matrix YM by merging the plurality of small sub-output matrices SYM_S and the plurality of large sub-output matrices SYM_L. Reference will be made to Figure 13 describe the operation of the output merging circuit 250 in more detail.

[0126] The control logic circuit 220 can control the overall operation of the matrix multiplication device 200. For example, the control logic circuit 220 can provide commands or control signals for controlling the operation timings of each of the weight memory circuit 210, the control logic circuit 220, the input matrix buffer 230, the first processing element array 240_S, the second processing element array 240_L, and the output merging circuit 250.

[0127] Figure 7 is a table showing the remaining weight indices and remaining weights stored in Figure 6 the weight memory circuit. Referring to Figures 1 to 7 , the weight memory circuit 210 can store the remaining weights RW and the remaining weight indices RWI of each of the plurality of pruned large sub-matrices PSM_L and the plurality of pruned small sub-matrices PSM_S. Hereinafter, for a more concise description, the remaining weights RW and the remaining weight indices RWI of the first pruned small sub-matrix PSM_S1 and the first pruned large sub-matrix PSM_L1 are shown in Figure 7 , but the disclosed scope is not limited thereto. For example, the weight memory circuit 210 can store the remaining weights RW and the remaining weight indices RWI of each of the plurality of pruned large sub-matrices PSM_L and the plurality of pruned small sub-matrices PSM_S in a manner similar to the remaining weights RW and the remaining weight indices RWI stored for the first pruned small sub-matrix PSM_S1 and the first pruned large sub-matrix PSM_L1.

[0128] The weight memory circuit 210 can store the remaining weight indices 2, 5, 6, and m of the first pruned small sub-matrix PSM_S1. That is, the weight memory circuit 210 can store the column numbers including the remaining weights in the first pruned small sub-matrix PSM_S1 as the remaining weight indices RWI.

[0129] In one embodiment, the weight memory circuit 210 can store the remaining weights RW of the first pruned small sub-matrix PSM_S1 for each row. For example, the weight memory circuit 210 can store the remaining weights w 12 , w 15 , w 16 , and w 1m included in the first row of the first pruned small sub-matrix PSM_S1 in consecutive addresses (e.g., storing w 12 , w 15 , w 16 , and w 1m sequentially), and can store the remaining weights w 22 , w 25 , w 26 , w 2m included in the second row of the first pruned small sub-matrix PSM_S1 in consecutive addresses.

[0130] In the example case where a plurality of residual weights RW included in a row of the pruned small submatrix PSM_S are stored in consecutive addresses, the plurality of residual weights RW required to calculate each output element can be read consecutively from the weight memory circuit 210. In this case, the time required to provide the plurality of residual weights RW stored in the weight memory circuit 210 to the processing element array can be minimized, and therefore, the operating speed of the matrix multiplication device 200 can be improved. However, the scope of the disclosure is not limited thereto. For example, the weight memory circuit 210 may divide and store a plurality of residual weights RW included in a single row of the first pruned small submatrix PSM_S1 into non-continuous addresses (e.g., storing a plurality of residual weights RW included in a single row of the first pruned small submatrix PSM_S1 in an interleaved manner).

[0131] The weight memory circuit 210 may store the remaining weight indexes RWI “1” and “6” of the first pruned large submatrix PSM_L1 . That is, the weight memory circuit 210 may store the remaining remaining weight indexes RWI as the column numbers including the remaining weights in the first pruned large submatrix PSM_L1 .

[0132] In one embodiment, the weight memory circuit 210 may store the remaining weights RW of the first pruned large submatrix PSM_L1 for each row. For example, similar to the method of storing the remaining weights RW of the first pruned small submatrix PSM_S1, the weight memory circuit 210 may store the remaining weights RW of the first pruned large submatrix PSM_L1.

[0133] Figures 8 to 12 It is shown in detail Figure 6 A block diagram of the operation of some configurations of the present invention is provided. Hereinafter, for a more concise description, a method in which the matrix multiplication device 200 generates a first small sub-output matrix SYM_S1 corresponding to the multiplication of the first pruned small sub-matrix PSM_S1 and the input matrix XM, and generates a first large sub-output matrix SYM_L1 corresponding to the multiplication of the first pruned large sub-matrix PSM_L1 and the input matrix XM will be representatively described. However, the scope of the disclosure is not limited thereto, and similar to the above method, the matrix multiplication device 200 may generate a plurality of small sub-output matrices SYM_S corresponding to the multiplication of each of a plurality of pruned small sub-matrices PSM_S and the input matrix XM, and may generate a plurality of large sub-output matrices SYM_L corresponding to the multiplication of each of a plurality of pruned large sub-matrices PSM_L and the input matrix XM.

[0134] Reference Figures 1 to 8, the input matrix buffer 230 may receive the remaining weight index RWI_PSM_S1 of the first pruned small sub-matrix PSM_S1. The input matrix buffer 230 may provide the input elements included in the input element row IER corresponding to the remaining weight index RWI_PSM_S1 to the first processing element array 240_S.

[0135] For example, the input matrix buffer 230 may output the input elements included in the second row, fifth row, sixth row, and m-th row of the input matrix XM to the first processing element array 240_S. For a more detailed example, refer to Figure 4 and Figure 8 , the input matrix buffer 230 may to (e.g., the input elements included in the second input element row), to (e.g., the input elements included in the fifth input element row), to (e.g., the input elements included in the sixth input element row) and to (e.g., the input elements included in the m-th input element row) to the first processing element array 240_S.

[0136] Similarly, the input matrix buffer 230 may receive the remaining weight index RWI_PSM_L1 of the first pruned large sub-matrix PSM_L1. The input matrix buffer 230 may provide the input elements included in the input element row IER corresponding to the remaining weight index RWI_PSM_L1 to the second processing element array 240_L.

[0137] In other words, the input matrix buffer 230 may output the input elements included in the first row and sixth row of the input matrix XM to the second processing element array 240_L. For a more detailed example, refer to Figure 5 and Figure 8 , the input matrix buffer 230 may to (e.g., the input elements included in the first input element row) and to (e.g., the input elements included in the sixth input element row) to the second processing element array 240_L.

[0138] The first processing element array 240_S may receive the remaining weight RW_PSM_S1 of the first pruned small sub-matrix PSM_S1. The first processing element array 240_S may generate the first small output matrix SYM_S1 based on the input elements corresponding to the remaining weight RW_PSM_S1 and the remaining weight index RWI_PSM_S1. It will be referred toFigures 9 to 10 A method for the first processing element array 240_S to generate the first small sub-output matrix SYM_S1 is described in more detail.

[0139] Similarly, the second processing element array 240_L may receive the remaining weights RW_PSM_L1 of the first pruned large sub-matrix PSM_L1. The second processing element array 240_L may generate the first large sub-output matrix SYM_L1 based on the input elements corresponding to the remaining weights RW_PSM_L1 and the remaining weight index RWI_PSM_L1. Reference will be made to Figures 11 to 12 A method for the second processing element array 240_L to generate the first large sub-output matrix SYM_L1 is described in more detail.

[0140] In one embodiment, the first processing element array 240_S may be larger in size than the second processing element array 240_L. For example, the first processing element array 240_S may include more processing elements than the second processing element array 240_L. Reference will be made to Figures 9 to 12 The configurations of the first processing element array 240_S and the second processing element array 240_L are described in more detail.

[0141] In one embodiment, the first processing element array 240_S and the second processing element array 240_L may be different parts of a processing element array. For example, the first processing element array 240_S and the second processing element array 240_L may be different regions included in a processing element array that are assigned to perform matrix multiplication operations on matrices of different sizes and have different sizes from each other. However, the disclosed scope is not limited thereto.

[0142] Figure 9 is a block diagram showing in detail Figure 8 the configuration and operation of the first processing element array. Referring to Figures 1 to 9 , the first processing element array 240_S may include a plurality of processing elements PE arranged in row and column directions.

[0143] The plurality of processing elements PE included in the first processing element array 240_S may be arranged in two rows and b columns. That is, the number of rows of the first processing element array 240_S may be the same as the number of rows of each of the plurality of small sub-matrices SM_S, and the number of columns of the first processing element array 240_S may be the same as the number of columns of the input matrix XM.

[0144] Hereinafter, for a more concise description, the processing element arranged in the (i)-th row and the (j)-th column of the first processing element array 240_S will be referred to as PEij_S. For example, the processing element arranged in the first row and the second column of the first processing element array 240_S will be referred to as PE12_S.

[0145] The first processing element array 240_S may include a first processing element row PER1_S and a second processing element row PER2_S. Each of the first processing element row PER1_S and the second processing element row PER2_S may include a plurality of processing elements PE that are different from each other. For example, the first processing element row PER1_S may include processing elements PE11_S to PE1b_S, and the second processing element row PER2_S may include processing elements PE21_S to PE2b_S.

[0146] The first processing element row PER1_S to the second processing element row PER2_S may receive the remaining weights RW arranged in different rows of the first pruned sub-matrix PSM_S1. For example, the (i)-th processing element row PERi_S may receive the remaining weights RW arranged in the (i)-th row of the first pruned sub-matrix PSM_S1.

[0147] More specifically, each of the processing elements included in the first processing element row PER1_S may receive , , and , and each of the processing elements included in the second processing element row PER2_S may receive , , and .

[0148] In one embodiment, the first processing element column PEC1_S to the b-th processing element column PECb_S may receive the input elements IE arranged in different columns of the input matrix XM. For example, the first processing element column PEC1_S may receive the input elements IE arranged in the input element row corresponding to the remaining weight index RWI_PSM_S1 and the first input element column, and the second processing element column PEC2_S may receive the input elements IE arranged in the input element row corresponding to the remaining weight index RWI_PSM_S1 and the second input element column. More specifically, each processing element included in the first processing element column PEC1_S may receive , , and , and each processing element included in the second processing element column PEC2_S may receive , , and . Similarly, the third processing element column PEC3_S to the b-th processing element column PECb_S may receive the input elements arranged in different columns of the input matrix XM.

[0149] Each of the processing elements PE included in the first processing element array 240_S can output different output elements based on the received multiple residual weights RW and the received multiple input elements IE. For example, the processing element PE11_S can output y 11 , and the processing element PE12_S can output y 12 . In this way, the first processing element row PER1_S can output y 11 to y 1b , and the second processing element row PER2_S can output y 21 to y 2b . For example, each of the different processing element rows in the first processing element array 240_S can calculate the output elements included in different rows of the output matrix YM, and each of the different processing element columns in the first processing element array 240_S can calculate the output elements included in different columns of the output matrix YM. A specific method for each processing element PE included in the first processing element array 240_S to obtain output elements will be described in more detail with reference to Figure 10 .

[0150] Figure 10 More specifically shown Figure 9 of the operation of the processing element row. That is, hereinafter, the operation of the first processing element row PER1_S will be described representatively with reference to Figures 1 to 10 . However, the scope of the disclosure is not limited thereto, and other processing element rows included in the first processing element array 240_S can also operate in a similar manner.

[0151] The first processing element row PER1_S can include processing elements PE11_S to PE1b_S. Each of the processing elements PE11_S to PE1b_S can sequentially receive the residual weights RW_PSM_S1_R1 included in the first row of the first pruned sub-matrix PSM_S1. That is, each of the processing elements PE11_S to PE1b_S can sequentially receive , , and .

[0152] Each of the processing elements PE11_S to PE1b_S can sequentially receive the input elements IE included in different columns of the input matrix XM. For example, PE11_S can sequentially receive the input element row corresponding to the residual weight index RWI_PSM_S1 and the input elements included in the first input element column. That is, the processing element PE11_S can sequentially receive , , and . Similarly, the processing element PE12_S can sequentially receive , , and , and the processing element PE1b_S can sequentially receive , , and . For the sake of more concise illustration, the detailed description of the input elements received by the processing elements in the first processing element row PER1_S is omitted.

[0153] Each of the processing elements PE11_S to PE1b_S can calculate different output elements based on the order of receiving multiple residual weights RW and input elements IE. For example, the processing elements PE11_S to PE1b_S can calculate y 11 to y 1b .

[0154] For a more detailed example, the processing element PE11_S can output y 12 by adding the partial sums obtained by multiplying w 15 , w 16 , w 1m and x 21 , x 51 , x 61 , x m1 respectively. That is, the processing element PE11_S can obtain the output element y11 according to Equation 5 below. 11 That is, the processing element PE11_S can obtain the output element y11 according to Equation 5 below.

[0155] (Equation 5) In this way, each of the processing elements PE11_S to PE1b_S can calculate y11 to y1b respectively based on the order of receiving multiple residual weights RW and input elements IE.

[0156] Figure 11 is a block diagram showing the configuration and operation of the second processing element array in more detail. Referring to Figure 8 and Figures 1 to 8 and Figure 11 , the second processing element array 240_L can include a plurality of processing elements PE arranged in row and column directions.

[0157] The plurality of processing elements PE included in the second processing element array 240_L can be arranged in four rows and b columns. That is, the number of rows of the second processing element array 240_L can be the same as the number of rows of each of the plurality of large sub-matrices SM_L, and the number of columns of the second processing element array 240_L can be the same as the number of columns of the input matrix XM.

[0158] According to the disclosed embodiments, the size of the second processing element array 240_L may be larger than the size of the first processing element array 240_S. For example, the number of rows in the second processing element array 240_L may be greater than the number of rows in the first processing element array 240_S.

[0159] Hereinafter, the processing element arranged in the (i)-th row and (j)-th column of the second processing element array 240_L will be referred to as "PEij_L".

[0160] The second processing element array 240_L may include a first processing element row PER1_L to a fourth processing element row PER4_L. Each of the first processing element row PER1_L to the fourth processing element row PER4_L may include a plurality of different processing elements PE. For example, the first processing element row PER1_L may include processing elements PE11_L to PE1b_L.

[0161] The first processing element row PER1_L to the fourth processing element row PER4_L may receive the remaining weights RW arranged in different rows of the first trimmed large sub-matrix PSM_L1. For example, the (i)-th processing element row PERi_L may receive the remaining weights RW arranged in the (i)-th row of the first trimmed large sub-matrix PSM_L1.

[0162] More specifically, each of the processing elements included in the first processing element row PER1_L may receive w 31 and w 36 (e.g., the remaining weights included in the first row of the first trimmed large sub-matrix PSM_L1). Similarly, each of the processing elements included in the second processing element row PER2_L may receive w 41 and w 46 ; each of the processing elements included in the third processing element row PER3_L may receive w 51 and w 56 ; and each of the processing elements included in the fourth processing element row PER4_L may receive w 61 and w 66 .

[0163] The first processing element column PEC1_L to the b-th processing element column PECb_L can receive input elements IE arranged in different columns. For example, the first processing element column PEC1_L can receive the input elements IE arranged in the input element row corresponding to the remaining weight index RWI_PSM_L1 and the first input element column; and the second processing element column PEC2_L can receive the input elements IE arranged in the input element row corresponding to the remaining weight index RWI_PSM_L1 and the second input element column. More specifically, each of the processing elements included in the first processing element column PEC1_L can receive and , and each of the processing elements included in the second processing element column PEC2_L can receive and . In a similar manner, the third processing element column PEC3_L to the b-th processing element column PECb_L can be capable of receiving the input elements IE arranged in different columns of the input matrix XM.

[0164] Each of the processing elements PE included in the second processing element array 240_L can output different output elements based on the received multiple remaining weights RW and the received multiple input elements IE. For example, the processing element PE11_L can output , and the processing element PE12_L can output . In this way, the first processing element row PER1_L can output to , and the second processing element row PER2_L can output to . That is to say, different processing element rows of the second processing element array 240_L can calculate the output elements included in different rows of the output matrix; and different processing element columns of the second processing element array 240_L can calculate the output elements included in different columns of the output matrix. The detailed method for each processing element PE included in the second processing element array 240_L to obtain the output elements will be described in more detail with reference to Figure 12 .

[0165] Figure 12 More specifically shows Figure 11 the operation of the processing element row. That is to say, hereinafter, the operation of the first processing element row PER1_L will be described representatively with reference to Figures 1 to 8 and Figures 11 to 12 . However, the disclosed scope is not limited thereto, and other processing element rows included in the second processing element array 240_L can also operate in a similar manner.

[0166] The first row of processing elements PER1_L may include processing elements PE11_L to PE1b_L. Each of the processing elements PE11_L to PE1b_L may sequentially receive the remaining weights RW_PSM_L1_R1 included in the first row of the first trimmed large sub-matrix PSM_L1. That is, each of the processing elements PE11_L to PE1b_L may sequentially receive w 31 and w 36 .

[0167] Each of the processing elements PE11_L to PE1b_L may sequentially receive input elements IE included in different columns of the input matrix XM. For example, the processing element PE11_L may sequentially receive the input elements included in the input element row corresponding to the remaining weight index RWI_PSM_L1 and the first input element column. In other words, the processing element PE11_L may sequentially receive x 11 and x 61 . Similarly, the processing element PE12_1 may sequentially receive x 12 and x 62 , and the processing element PE1b_L may sequentially receive x 1b and x 6b .

[0168] Each of the processing elements PE11_L to PE1b_L may calculate different output elements based on the order of receiving the multiple remaining weights RW and the input elements IE. For example, the processing elements PE11_L to PE1b_L may calculate y 31 to y 3b respectively.

[0169] For a more detailed example, the processing element PE11_L may output y 31 by adding the partial sums obtained by multiplying w 36 and w 11 respectively, and x 61 and x 31 . In this way, each of the processing elements PE11_L to PE1b_L may calculate y 31 to y 3b respectively based on the order of receiving the multiple remaining weights RW and the input elements IE.

[0170] In this way, the first processing element array 240_S can generate a plurality of small sub-output matrices SYM_S by multiplying each of the plurality of trimmed small sub-matrices PSM_S by the input matrix XM. Similarly, the second processing element array 240_L can generate a plurality of large sub-output matrices SYM_L by multiplying each of the plurality of trimmed large sub-matrices PSM_L by the input matrix XM. In this case, the plurality of small sub-output matrices SYM_S and the plurality of large sub-output matrices SYM_L can correspond to different rows of the output matrix YM. In other words, a part of the weight matrix WM on which the "second processing element array 240_L will perform a matrix multiplication operation with the input matrix XM" can be different from a part of the weight matrix WM on which the "first processing element array 240_S will perform a matrix multiplication operation with the input matrix XM".

[0171] Figure 13 illustrates Figure 6 the operation of the output merging circuit. Referring to Figures 1 to 13 , the output merging circuit 250 can receive a plurality of small sub-output matrices SYM_S and a plurality of large sub-output matrices SYM_L. For example, the output merging circuit 250 can sequentially receive the plurality of small sub-output matrices SYM_S from the first processing element array 240_S, and can sequentially receive the plurality of large sub-output matrices SYM_L from the second processing element array 240_L.

[0172] The output merging circuit 250 can generate the output matrix YM by merging the plurality of small sub-output matrices SYM_S and the plurality of large sub-output matrices SYM_L. In this case, the row numbers in the output matrix YM corresponding to the plurality of small sub-output matrices SYM_S can be the same as the row numbers in the weight matrix WM of the small sub-matrices SM_S respectively corresponding to the plurality of small sub-output matrices SYM_S. For example, the first small sub-matrix SM_S1 for generating the first small sub-output matrix SYM_S1 can correspond to the first and second rows in the weight matrix WM. In this case, the first small sub-output matrix SYM_S1 can correspond to the first and second rows in the output matrix YM.

[0173] Similarly, the row numbers in the output matrix YM corresponding to the plurality of large sub-output matrices SYM_L can be the same as the row numbers in the weight matrix WM of the large sub-matrices SM_L respectively corresponding to the plurality of large sub-output matrices SYM_L.

[0174] In this way, the arrangement order of the plurality of small sub-output matrices SYM_S and the plurality of large sub-output matrices SYM_L in the output matrix YM can correspond to the order of the plurality of small sub-matrices SM_S and the plurality of large sub-matrices SM_L described above with reference to Figure 3 description.

[0175] Thus, according to the disclosed embodiments, each of the plurality of small sub-output matrices SYM_S and the plurality of large sub-output matrices SYM_L may correspond to different rows of the output matrix YM. For example, the first small sub-output matrix SYM_S1 may correspond to the first to second rows of the output matrix YM, the first large sub-output matrix SYM_L1 may correspond to the third to sixth rows of the output matrix YM, and the second small sub-output matrix SYM_S2 may correspond to the third to sixth rows of the output matrix YM.

[0176] According to the disclosed embodiments, the first processing element array 240_S and the second processing element array 240_L may sequentially calculate a part of the output matrix YM, and the output merging circuit 250 may generate the output matrix YM by merging the calculation results from the first processing element array 240_S and the second processing element array 240_L.

[0177] Figure 14 is a flowchart of an operation method of a matrix multiplication system according to the disclosed embodiments. Refer to Figures 1 to 14 , in operation S100, the method may include receiving an input matrix XM and a weight matrix WM. For example, the matrix multiplication system MMS may receive the input matrix XM and the weight matrix WM. For example, the matrix multiplication device 200 may receive the input matrix XM, and the matrix pruning device 100 may receive the weight matrix WM.

[0178] In operation S200, the method may include generating one or more pruned small sub-matrices PSM_S and one or more pruned large sub-matrices PSM_L based on the weight matrix WM. For example, the matrix multiplication system MMS may generate a plurality of pruned small sub-matrices PSM_S and a plurality of pruned large sub-matrices PSM_L based on the weight matrix WM. For example, the matrix pruning device 100 may generate a plurality of pruned small sub-matrices PSM_S and a plurality of pruned large sub-matrices PSM_L by pruning the weight matrix WM divided into a plurality of pruning groups PG. Refer to Figure 15 The operation of the matrix pruning device 100 in operation S200 will be described in more detail.

[0179] In operation S300, the method may include generating the output matrix YM based on the multiplication of each of the plurality of pruned small sub-matrices PSM_S and the plurality of pruned large sub-matrices PSM_L with the input matrix XM. For example, the matrix multiplication system MMS may generate the output matrix YM based on the multiplication of each of the plurality of pruned small sub-matrices PSM_S and the plurality of pruned large sub-matrices PSM_L with the input matrix XM. For example, the matrix multiplication device 200 may generate the output matrix YM by merging the results of multiplying the input matrix XM by the plurality of pruned small sub-matrices PSM_S and the results of multiplying the input matrix XM by the plurality of pruned large sub-matrices PSM_L. Refer to Figure 16The operation of the matrix multiplication device 200 in operation S300 will be described in more detail.

[0180] Figure 15 is a flowchart showing Figure 14 the operation S200 in more detail. Referring to Figures 1 to 15 , operation S200 may include operations S210 to S260.

[0181] In operation S210, the method may include generating a plurality of pruning groups PG based on the weight matrix WM. For example, the matrix pruning device 100 may generate a plurality of pruning groups PG by dividing the weight matrix WM. For example, the matrix pruning device 100 may generate a plurality of pruning groups PG by dividing the weight matrix WM into two weight row units. However, the disclosed scope is not limited thereto, and the matrix pruning device 100 may generate a plurality of pruning groups PG by dividing the weight matrix WM into any number of weight row units.

[0182] In operation S220, the method may include classifying each of the plurality of pruning groups PG as an important pruning group IPG or an unimportant pruning group UIPG. For example, the matrix pruning device 100 may classify each of the plurality of pruning groups PG as an important pruning group IPG or an unimportant pruning group UIPG by determining the group importance of each of the plurality of pruning groups PG. For example, the matrix pruning device 100 may obtain the group importance of each of the plurality of pruning groups PG. After obtaining the group importance of each of the plurality of pruning groups PG, the matrix pruning device 100 may classify the pruning group PG having a group importance higher than the group importance threshold GIPTH as an important pruning group IPG, and classify the pruning group PG having a group importance lower than the group importance threshold as an unimportant pruning group UIPG.

[0183] In operation S230, the method may include generating a plurality of small submatrices SM_S based on the plurality of important pruning groups IPG. For example, the matrix pruning device 100 may generate a plurality of small submatrices SM_S based on the plurality of important pruning groups IPG. For example, the matrix pruning device 100 may convert one important pruning group IPG into one small submatrix SM_S. However, the disclosed scope is not limited thereto, and the matrix pruning device 100 may generate one small submatrix SM_S by combining a plurality of important pruning groups IPG.

[0184] In one embodiment, each of the plurality of small submatrices SM_S may be a subset of the weight matrix WM. For example, each of the plurality of small submatrices SM_S may correspond to a different important pruning group IPG in the weight matrix WM.

[0185] In operation S240, the method may include generating a plurality of large sub-matrices SM_L based on a plurality of unimportant pruning groups UIPG. For example, the matrix pruning device 100 may generate a plurality of large sub-matrices SM_L based on a plurality of unimportant pruning groups UIPG. For example, the matrix pruning device 100 may generate a plurality of large sub-matrices SM_L by generating one large sub-matrix SM_L through merging two unimportant pruning groups UIPG. However, the disclosed scope is not limited thereto, and the matrix pruning device 100 may generate one large sub-matrix SM_L by merging a random number of important pruning groups IPG.

[0186] In one embodiment, each of the plurality of large sub-matrices SM_L may be a subset of the weight matrix WM. For example, each of the plurality of large sub-matrices SM_L may correspond to two or more unimportant pruning groups UIPG in the weight matrix WM.

[0187] In operation S250, the method may include generating a plurality of pruned small sub-matrices PSM_S based on a plurality of small sub-matrices SM_S. For example, the matrix pruning device 100 may generate a plurality of pruned small sub-matrices PSM_S by pruning each of the plurality of small sub-matrices SM_S. For example, the matrix pruning device 100 may prune each of the plurality of small sub-matrices SM_S in units of columns.

[0188] In one embodiment, after operation S250, the matrix pruning device 100 may store the plurality of pruned small sub-matrices PSM_S in the weight memory circuit 210.

[0189] In operation S260, the method may include generating a plurality of pruned large sub-matrices PSM_L based on a plurality of large sub-matrices SM_L. For example, the matrix pruning device 100 may generate a plurality of pruned large sub-matrices PSM_L by pruning each of the plurality of large sub-matrices SM_L respectively. For example, the matrix pruning device 100 may prune each of the plurality of large sub-matrices SM_L in units of columns.

[0190] In one embodiment, after operation S260, the matrix pruning device 100 may store the plurality of pruned large sub-matrices PSM_L in the weight memory circuit 210.

[0191] For a more concise description, in Figure 15 , an embodiment of sequentially performing operation S230 to operation S260 is representatively described, but the disclosed scope is not limited thereto. For example, operation S230 and operation S240 may be performed in parallel, and operation S250 and operation S260 may be performed in parallel.

[0192] Figure 16 is a flowchart of operation S300 that shows more details of Figure 14 . Refer to Figures 1 to 16, operation S300 may include operations S310 to S330.

[0193] In operation S310, the method may include obtaining a plurality of small sub-output matrices SYM_S based on a plurality of trimmed small sub-matrices PSM_S and an input matrix XM. For example, matrix multiplication device 200 may calculate a plurality of small sub-output matrices SYM_S by multiplying the plurality of trimmed small sub-matrices PSM_S by the input matrix XM respectively. For example, by using the first processing element array 240_S, matrix multiplication device 200 may calculate a first small sub-output matrix SYM_S1 corresponding to the product of the first trimmed small sub-matrix PSM_S1 and the input matrix XM, and may calculate a second small sub-output matrix SYM_S2 corresponding to the product of the second trimmed small sub-matrix PSM_S2 and the input matrix XM.

[0194] In operation S320, the method may include obtaining a plurality of large sub-output matrices SYM_L based on the input matrix XM and a plurality of trimmed large sub-matrices PSM_L. For example, matrix multiplication device 200 may calculate a plurality of large sub-output matrices SYM_L by multiplying the input matrix XM and the plurality of trimmed large sub-matrices PSM_L respectively. For example, by using the second processing element array 240_L, matrix multiplication device 200 may calculate a first large sub-output matrix SYM_L1 corresponding to the product of the first trimmed large sub-matrix PSM_L1 and the input matrix XM, and may calculate a second large sub-output matrix SYM_L2 corresponding to the product of the second trimmed large sub-matrix PSM_L2 and the input matrix XM.

[0195] For a more concise description, in Figure 16 , an embodiment of sequentially performing operations S310 to S320 is representatively described, but the disclosed scope is not limited thereto. For example, operation S310 may be performed in parallel with operation S320.

[0196] In operation S330, the method may include generating an output matrix YM based on the plurality of small sub-output matrices SYM_S and the plurality of large sub-output matrices SYM_L. For example, matrix multiplication device 200 may generate the output matrix YM by combining the plurality of small sub-output matrices SYM_S and the plurality of large sub-output matrices SYM_L. For example, output merging circuit 250 may generate the output matrix YM by combining the plurality of small sub-output matrices SYM_S and the plurality of large sub-output matrices SYM_L.

[0197] Figure 17 is a flowchart showing operation S310 in more detail. Referring to Figure 16 , operation S310 may include operations S311 to S314. Figures 1 to 17 , operation S310 may include operations S311 to S314.

[0198] In operation S311, the method may include identifying a remaining weight RW and a remaining weight index RWI of a pruned small sub-matrix PSM_S among a plurality of pruned small sub-matrices PSM_S. For example, the matrix multiplication device 200 may identify a remaining weight RW and a remaining weight index RWI of a pruned small sub-matrix PSM_S among a plurality of pruned small sub-matrices PSM_S. For example, the matrix multiplication device 200 may identify a remaining weight RW and a remaining weight index RWI of one of the plurality of pruned small sub-matrices PSM_S stored in the weight memory circuit 210 in the form of the remaining weight RW and the remaining weight index RWI.

[0199] In operation S312, the method may include identifying an input element IE corresponding to the remaining weight index RWI from the input matrix XM. For example, the matrix multiplication device 200 may identify an input element IE corresponding to the remaining weight index RWI from the input matrix XM. For example, the control logic circuit 220 may identify an input element IE included in the input element row corresponding to the remaining weight index RWI identified in the above operation S311.

[0200] In operation S313, the method may include obtaining a small output matrix SYM_S based on the identified input element IE and the remaining weight RW. For example, the matrix multiplication device 200 may calculate the small output matrix SYM_S based on the identified input element IE and the remaining weight RW. For example, the first processing element array 240_S may calculate a plurality of output elements included in a small output matrix SYM_S based on the remaining weight RW identified in operation S311 and the input element IE identified in operation S312.

[0201] In operation S314, the method may include determining whether the operations for all the pruned small sub-matrices PSM_S have been completed. For example, the matrix multiplication device 200 may determine whether the operations for all the pruned small sub-matrices PSM_S have been completed. For example, the matrix multiplication device 200 may determine whether the small output matrix SYM_S has been calculated based on each of all the pruned small sub-matrices PSM_S generated in the above operation S250.

[0202] In operation S314, based on determining that the calculation for any of the pruned small sub-matrices PSM_S has not been completed, the above operation S311 may be iteratively executed. In this way, the first processing element array 240_S may be able to sequentially generate a plurality of small output matrices SYM_S corresponding to the product of the input matrix XM and each of the plurality of pruned small sub-matrices PSM_S.

[0203] In operation S314, based on determining that the calculations for all the pruned small sub-matrices PSM_S have been completed, operation S310 may be terminated.

[0204] Figure 18 is a flowchart that shows more details of Figure 16 operation S320. Referring to Figures 1 to 18 , operation S320 may include operations S321 to S324.

[0205] In operation S321, the method may include identifying a remaining weight RW and a remaining weight index RWI of a pruned large sub-matrix PSM_L among a plurality of pruned large sub-matrices PSM_L. For example, matrix multiplication device 200 may identify a remaining weight RW and a remaining weight index RWI of a pruned large sub-matrix PSM_L among a plurality of pruned large sub-matrices PSM_L. For example, matrix multiplication device 200 may identify a remaining weight RW and a remaining weight index RWI of one of the plurality of pruned large sub-matrices PSM_L stored in weight memory circuit 210 in the form of the remaining weight RW and the remaining weight index RWI.

[0206] In operation S322, the method may include identifying an input element IE corresponding to the remaining weight index RWI from input matrix XM. For example, matrix multiplication device 200 may identify an input element IE corresponding to the remaining weight index RWI from input matrix XM. For example, control logic circuit 220 may identify an input element IE included in the input element row corresponding to the remaining weight index RWI identified in the above operation S321.

[0207] In operation S323, the method may include obtaining a large sub-output matrix SYM_L based on the identified input element IE and the remaining weight RW. For example, matrix multiplication device 200 may calculate a large sub-output matrix SYM_L based on the identified input element IE and the remaining weight RW. For example, second processing element array 240_L may calculate a plurality of output elements included in one large sub-output matrix SYM_L based on the remaining weight RW identified in operation S321 and the input element IE identified in operation S322.

[0208] In operation S324, the method may include determining whether the calculation for all pruned large sub-matrices PSM_L has been completed. For example, matrix multiplication device 200 may determine whether the calculation for all pruned large sub-matrices PSM_L has been completed. For example, matrix multiplication device 200 may determine whether a large sub-output matrix SYM_L has been calculated based on each of all the pruned large sub-matrices PSM_L generated in the above operation S260.

[0209] In operation S324, based on determining that the calculation for any pruned large sub-matrix PSM_L is not completed, the above operation S321 may be repeatedly performed. In this way, the second processing element array 240_L may be able to sequentially generate a plurality of large sub-output matrices SYM_L corresponding to the product of each of the input matrix XM and the plurality of pruned large sub-matrices PSM_L.

[0210] In operation S324 , based on determining that operations on all pruned large sub-matrices PSM_L have been completed, operation S320 may be terminated.

[0211] According to the disclosed embodiment, the first processing element array 240_S may sequentially generate a plurality of small output sub-matrices SYM_S based on a plurality of pruned small sub-matrices PSM_S, and the second processing element array 240_L may sequentially generate a plurality of large output sub-matrices SYM_L based on a plurality of pruned large sub-matrices PSM_L. For example, the first processing element array 240_S and the second processing element array 240_L may each operate in parallel and may obtain different parts of the output matrix YM.

[0212] Figure 19 Show Figure 6 Iteration operation on an array of processing elements. Figures 1 to 19 , the first processing element array 240_S may sequentially receive "P" pruned small sub-matrices PSM_S. The second processing element array 240_L may sequentially receive "Q" pruned large sub-matrices PSM_L. Here, "P" and "Q" are positive integers.

[0213] In one embodiment, the value obtained by adding P times the number of rows of the pruned small submatrix PSM_S and Q times the number of rows of the pruned large submatrix PSM_L may be the same as the number of rows (i.e., n) of the weight matrix WM. For example, "P" pruned small submatrices PSM_S and "Q" pruned large submatrices PSM_L may be generated by pruning "P" small submatrices SM_S and "Q" large submatrices SM_L generated based on one weight matrix WM, respectively.

[0214] The first processing element array 240_S may perform matrix multiplication on each of the "P" pruning sub-matrices PSM_S. For example, the first processing element array 240_S may generate "P" sub-output matrices SYM_S corresponding to the product of the input matrix XM and each of the "P" pruning sub-matrices PSM_S. For example, the first processing element array 240_S may iteratively perform matrix multiplication (e.g., referring to Figure 9 and Figure 10 The matrix multiplication described in (a) is performed “P” times to generate a small output sub-matrix SYM_S corresponding to the product of a pruned small sub-matrix PSM_S and the input matrix XM.

[0215] Similarly, the second processing element array 240_L may perform a matrix multiplication operation on each of the "Q" pruned large sub-matrices PSM_L. For example, the second processing element array 240_L may generate "Q" large sub-output matrices SYM_L corresponding to the product of the input matrix XM and each of the "Q" pruned large sub-matrices PSM_L. For example, the second processing element array 240_L may iteratively perform matrix multiplication (e.g., referring to Figure 11 and Figure 12 The matrix multiplication described in (a) is performed “Q” times to generate a large sub-output matrix SYM_L corresponding to the product of a pruned large sub-matrix PSM_L and the input matrix XM.

[0216] The first processing element array 240_S and the second processing element array 240_L may operate in parallel with each other. For example, the first processing element array 240_S and the second processing element array 240_L may independently perform "P" times of matrix multiplication operations and "Q" times of matrix multiplication operations, respectively. However, the disclosure is not limited thereto, and therefore, according to another embodiment, the first processing element array 240_S and the second processing element array 240_L may operate in a sequential order.

[0217] Hereinafter, for a more concise description, the total time taken by the first processing element array 240_S to perform the matrix multiplication operation “P” times may be referred to as a first operation time, and the total time taken by the second processing element array 240_L to perform the matrix multiplication operation “Q” times may be referred to as a second operation time.

[0218] In the example case where the operation times of the first processing element array 240_S and the second processing element array 240_L are synchronized, the matrix multiplication operations of the first processing element array 240_S and the second processing element array 240_L can be started and ended at the same time. That is, in the example case where the first operation time and the second operation time are synchronized, the operation efficiency of the matrix multiplication system MMS can be maximized.

[0219] For a more detailed example, when "P" and "Q" are the same, the first operation time and the second operation time may be synchronized, and the time taken by the first processing element array 240_S to perform one matrix multiplication operation is the same as the time taken by the second processing element array 240_L to perform one matrix multiplication operation. However, the scope of the disclosure is not limited thereto.

[0220] In the following, reference will be made to Figures 20 to 25 A method for synchronizing a first operating time and a second operating time is described in detail.

[0221] Figure 20is shown using a systolic array scheme Figure 19 A block diagram of an array of processing elements of FIG. Figure 20 , we will representatively describe the implementation of the systolic array scheme. Figure 19 The configuration and operation of the first processing element array 240_S are described in detail, but the scope of the disclosure is not limited thereto. For example, the second processing element array 240_L may also be implemented using a contraction array solution. According to other embodiments, the first processing element array 240_S and the second processing element array 240_L may be implemented using another solution.

[0222] Reference Figures 1 to 20 , the first processing element array 240_S may include a plurality of processing elements PE arranged in row and column directions. For example, the processing elements PE may include but are not limited to PE11_S, PE12_S to PE1b_S and PE21_S, PE22_S to PE2b_S. The plurality of processing elements PE may operate in a systolic array method.

[0223] The first processing element array 240_S may sequentially receive the residual weights RW_PSM_S1 included in the first pruning sub-matrix PSM_S1. The first processing element array 240_S may be implemented to sequentially propagate the residual weights RW_PSM_S1 in a row direction. For example, the first processing element row PER1_S may sequentially propagate the residual weights RW_PSM_S1_R1 (e.g., w 12 、w 15 、w 16 、w 1m ).

[0224] For a more detailed example, the processing element PE11_S may receive one residual weight RW at time point 0. The processing element PE11_S may receive a different residual weight RW at a first time point after time point 0. The processing element PE11_S may send the residual weight RW received at time point 0 to the processing element PE12_S located adjacent to the processing element PE11_S in the row direction at the first time point.

[0225] In this manner, the processing elements PE included in the first processing element row PER1_S may sequentially transmit a plurality of residual weights RW (eg, residual weights RW_PSM_S1_R1 included in the first row of the first pruning submatrix PSM_S1 ) provided from the weight memory circuit 210 to adjacent processing elements.

[0226] Similarly, the second processing element row PER2_S may sequentially propagate a plurality of residual weights RW (eg, residual weights RW_PSM_S1_R2 included in the second row of the first pruning small sub-matrix PSM_S1 ) provided from the weight memory circuit 210 in the row direction.

[0227] Furthermore, the time point at which the second processing element row PER2_S initially receives the residual weight RW may be later than the time point at which the first processing element row PER1_S initially receives the residual weight RW. For example, the processing element PE21_S may initially receive the residual weight RW at a first time point after the 0th time point.

[0228] The first processing element array 240_S may be implemented to sequentially propagate a plurality of input elements IE in a column direction. For example, the first processing element column PEC1_S may sequentially propagate input elements (eg, x 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52 21 、x 51 、x 61 and x m1 ) and the first pruned small sub-matrix PSM_S1.

[0229] For a more detailed example, the processing element PE11_S may receive one input element IE at time point 0, and may also receive another input element IE at a first time point after time point 0. The processing element PE11_S may send the input element received at time point 0 to the processing element PE21_S disposed adjacent to the processing element PE11_S in the column direction at the first time point.

[0230] In this manner, the processing elements PE included in the first processing element column PEC1_S may sequentially transmit the plurality of input elements IE provided from the input matrix buffer 230 to adjacent processing elements. Similarly, each of the second processing element column PEC2_S to the b-th processing element column PECb_S may sequentially propagate the plurality of input elements IE provided from the input matrix buffer 230 in the column direction.

[0231] In addition, the first processing element column PEC1_S to the b-th processing element column PECb_S may initially receive the input element IE at different time points. For example, the time point at which the second processing element column PEC2_S initially receives the input element IE may be later than the time point at which the first processing element column PEC1_S initially receives the input element IE. For example, the processing element PE21_S may initially receive the input element IE at a first time point after the 0th time point, and the processing element PE31_S may initially receive the input element IE at a second time point after the first time point.

[0232] Each of the plurality of processing elements PE may be configured as described above. Figures 9 to 10The described method is similar to generating different output elements OE. For example, each of the plurality of processing elements PE may generate different output elements OE by accumulating the product of the input elements IE received at the same time point and the residual weights RW.

[0233] The first processing element array 240_S may be implemented to sequentially propagate the obtained output elements OE in the column direction. For example, the output element OE (eg, y 11 ) can be propagated sequentially in the column direction. In this way, the output element OE calculated by each processing element PE can be sent to the output merging circuit 250. The method for propagating the output element OE is similar to the method for propagating the input element IE, and therefore a further detailed description thereof will not be provided.

[0234] For example, each of the processing elements included in the first processing element array 240_S may be implemented to receive one or more of the input element IE, the remaining weight RW, and the output element OE from the processing element located adjacent to the first processing element array 240_S. In addition, each of the processing elements included in the first processing element array 240_S may be implemented to send one or more of the input element IE, the remaining weight RW, and the output element OE to the processing elements located adjacent to each other. Figure 21 A more detailed configuration and operation of each of the plurality of processing elements PE is described in more detail.

[0235] For a more concise explanation, Figure 20 200_S is a representative example of an embodiment in which the input element IE, the remaining weight RW, and the output element OE are all propagated in a contraction array scheme, but the scope of the disclosure is not limited thereto. For example, the first processing element array 240_S may be implemented to propagate only a portion of the input element IE, the remaining weight RW, and the output element OE in a contraction array scheme.

[0236] In one embodiment, each of the processing elements included in the first processing element array 240_S may perform calculations based on the same control clock signal. In this case, each of the plurality of processing elements may send input elements IE, residual weights RW, and / or output elements OE to other processing elements at the same time point. However, the scope of the disclosure is not limited thereto.

[0237] Figure 21 Shown in more detail Figure 20 Configuration of the processing elements. Figures 1 to 21The processing element PE may include an arithmetic logic unit ALU, an accumulation register REG_ACC, a residual weight register REG_RW and an input element register REG_IE. The arithmetic logic unit ALU may include a first input terminal TI1, a second input terminal TI2, a third input terminal TI3 and an output terminal TO.

[0238] The remaining weight register REG_RW may receive a plurality of remaining weights RW. For example, the remaining weight register REG_RW may sequentially receive a plurality of remaining weights RW. The remaining weight register REG_RW may receive the remaining weight RW and transmit the remaining weight RW to the adjacent processing element PE and the first input terminal TI1 after one cycle of the control clock signal CLK has passed.

[0239] The input element register REG_IE may receive a plurality of input elements IE. For example, the input element register REG_IE may sequentially receive a plurality of input elements IE. The input element register REG_IE may receive one input element IE and send the one input element IE to the adjacent processing element PE and the second input terminal TI2 after one cycle of the control clock signal CLK has passed.

[0240] The arithmetic logic unit ALU may receive a plurality of residual weights RW through the first input terminal TI1 and a plurality of input elements IE through the second input terminal TI2. For example, the arithmetic logic unit ALU may sequentially receive a plurality of residual weights RW through the first input terminal TI1 and may sequentially receive a plurality of input elements IE through the second input terminal TI2.

[0241] The third input terminal TI3 may be connected to the accumulation register REG_ACC. The arithmetic logic unit ALU may receive data stored in the accumulation register REG_ACC through the third input terminal TI3.

[0242] The arithmetic logic unit ALU can calculate a value by adding the product of the residual weight RW received via the first input terminal TI1 and the input element IE received via the second input terminal TI2 to the data received via the third input terminal TI3. ​​The arithmetic logic unit ALU can update the value stored in the accumulation register REG_ACC by providing the calculated value to the accumulation register REG_ACC via the output terminal TO. In other words, the arithmetic logic unit ALU can update the calculated value in the accumulation register REG_ACC by the following equation 6.

[0243] (Equation 6) here, may represent an input element IE received by the arithmetic logic unit ALU via the second input terminal TI2, may represent the residual weight RW received by the arithmetic logic unit ALU via the first input terminal TI1, may represent the accumulated value received by the arithmetic logic unit ALU from the accumulation register REG_ACC through the third input terminal TI3, and It can represent the value updated to the accumulation register REG_ACC through the calculation of the arithmetic logic unit ALU.

[0244] In one embodiment, the product of IEin and RWin may be referred to as a partial sum (hereinafter referred to as PSUM).

[0245] That is, the arithmetic logic unit ALU can sequentially accumulate partial sums corresponding to the products of the residual weight RW received through the first input terminal TI1 and the input element IE received through the second input terminal TI2, respectively. In this way, the arithmetic logic unit ALU can obtain the output element OE by accumulating multiple partial sums in the accumulation register REG_ACC.

[0246] The accumulation register REG_ACC may store the output element OE calculated by the arithmetic logic unit ALU. The accumulation register REG_ACC may provide the output element OE to the output merging circuit 250 in a systolic array scheme. In other words, the accumulation register REG_ACC may transfer the output element OE to the adjacent processing element PE or to the output merging circuit 250.

[0247] For a more detailed example, the output element OE (e.g., y) calculated by the processing element PE11_S 11 ) may be sent to the output merging circuit 250 through the processing element PE21_S. However, the scope of the disclosure is not limited thereto.

[0248] In one embodiment, the accumulator register REG_ACC may send the output element OE to the accumulator register of the adjacent processing element PE. However, the scope of the disclosure is not limited thereto, and the accumulator register REG_ACC may also transfer the output element OE to a separate output element register included in the adjacent processing element PE.

[0249] The arithmetic logic unit ALU, the accumulation register REG_ACC, the residual weight register REG_RW and the input element register REG_IE can each operate based on the control clock signal CLK. In this case, the arithmetic logic unit ALU can accumulate the partial sum in the accumulation register REG_ACC for each cycle of the control clock signal CLK, and the accumulation register REG_ACC, the residual weight register REG_RW and the input element register REG_IE can each update the stored data in each cycle of the control clock signal CLK.

[0250] Figure 22 It is shown Figure 20 A timing diagram of the operation of each of the processing elements. Figures 1 to 22 , each of the plurality of processing elements PE included in the first processing element array 240_S may operate based on the control clock signal CLK.

[0251] In the following, for a more concise description, it is assumed that the interval between each time point from the 0th time point t0 to the (b+4)th time point tb+4 is the same as the period of the control clock signal CLK. However, the scope of the disclosure is not limited thereto, and the interval between the 0th time point t0 and the (b+4)th time point tb+4 may also be a positive integer multiple of the period of the control clock signal CLK.

[0252] For example, at time point t0, the processing element PE11_S may and The partial sum PSUM11_1 corresponding to the product of is accumulated to the accumulation register REG_ACC. Then, at the first time point t1, the processing element PE11_S can also and The partial sum PSUM11_2 corresponding to the product of is accumulated into the accumulation register REG_ACC. In this way, at the second time point t2 and the third time point t3, the processing element PE11_S can sequentially accumulate the partial sum PSUM11_3 and the partial sum PSUM11_4 in the accumulation register REG_ACC. In this case, the output element y 11 After the third time point t3, it will be stored in the accumulation register REG_ACC of the processing element PE11_S.

[0253] In addition, the processing element PE12_S may receive the residual weight RW and the input element IE one cycle later than the processing element PE11_S. Therefore, the processing element PE12_S may accumulate the residual weight RW and the input element IE in the accumulation register REG_ACC at a first time point t1 later than the 0th time point t0. 22 and x 21In this way, the processing element PE12_S may be able to sequentially receive a plurality of residual weights RW and input elements IE and generate an output element y at a fourth time point t4 later than the third time point t3. 12 .

[0254] Similarly, the processing element PE1b_S may receive the remaining weight RW and the input element IE at a period (b-1) times later than the processing element PE11_S of the control clock signal CLK. Therefore, the processing element PE1b_S may generate the output element y at the (b+2)th time point tb+2. 1b .

[0255] Furthermore, the time points at which the processing elements PE21_S to PE2b_S included in the second processing element row PER2_S initially receive the residual weight RW and the input element IE may be delayed by one cycle of the control clock signal CLK compared to the time points at which the processing elements PE11_S to PE1b_S included in the first processing element row PER1_S initially receive the residual weight RW and the input element IE. For example, the processing element PE21_S may receive the residual weight RW and the input element IE one cycle of the control clock signal CLK later than the processing element PE11_S, and the processing element PE2b_S may receive the residual weight RW and the input element IE one cycle of the control clock signal CLK later than the processing element PE1b_S. Therefore, the processing element PE2b_S may generate the output element y at the (b+3)th time point tb+3. 2b .

[0256] That is, the time taken for the first processing element array 240_S to perform a matrix multiplication operation to generate a small sub-output matrix SYM_S corresponding to the product of one pruned small sub-matrix PSM_S and the input matrix XM can be determined based on the sum of: i) the number of columns including the residual weights RW of the pruned small sub-matrix PSM_S, ii) the number of columns of the input matrix XM (e.g., the number of columns of the first processing element array 240_S), and iii) the number of rows of the pruned small sub-matrix PSM_S (e.g., the number of rows of the first processing element array 240_S).

[0257] Similarly, the time taken for the second processing element array 240_L to perform a matrix multiplication operation to generate a large sub-output matrix SYM_L corresponding to the product of one pruned large sub-matrix PSM_L and the input matrix XM can be determined based on the sum of: i) the number of columns including the residual weights RW of the pruned large sub-matrix PSM_L, ii) the number of columns of the input matrix XM (e.g., the number of columns of the second processing element array 240_L), and iii) the number of rows of the pruned large sub-matrix PSM_L (e.g., the number of rows of the second processing element array 240_L).

[0258] Therefore, refer to Figure 19 The first operation time described can be determined by adjusting the size of "P" and / or the number of columns comprising the remaining weights RW of the pruned sub-matrix PSM_S. Figure 19 The described second operation time may be determined by adjusting the size of "Q" and / or the number of columns including the remaining weights RW of the pruned large sub-matrix PSM_L.

[0259] In reference Figure 19 In the described example case where "P" and "Q" are the same, when the sum of the number of columns of the residual weights RW of the pruned small submatrix PSM_S and the number of rows of the pruned small submatrix PSM_S is the same as the sum of the number of columns of the residual weights RW of the pruned large submatrix PSM_L and the number of rows of the pruned large submatrix PSM_L, the first operation time and the second operation time can be synchronized.

[0260] Figure 23 The method for adjusting the Figure 19 A method of processing element array operations at a time. Figure 1 , Figure 3 , Figure 19 and Figure 23 , the matrix pruning device 100 may divide the weight matrix WM into the first pruning group PG1 to the hth pruning group PGh. The matrix pruning device 100 may calculate the group importance of each of the first pruning group PG1 to the hth pruning group PGh. In this case, the group importance information of the first pruning group PG1 to the hth pruning group PGh may be GIP1 to GIPh, respectively. The method in which the matrix pruning device 100 calculates the group importance of each of the first pruning group PG1 to the hth pruning group PGh has been previously described with reference to Figure 3 is described, and thus a detailed description thereof will not be provided.

[0261] In one embodiment, the matrix pruning device 100 may classify the group types of the first pruning group PG1 to the hth pruning group PGh based on the first group importance threshold GIPTH1. In this case, the pruning group having a group importance higher than the first group importance threshold GIPTH1 may be classified as an important pruning group IPG, and the pruning group having a group importance lower than the first group importance threshold GIPTH1 may be classified as an unimportant pruning group UIPG. For a more detailed example, in the case where GIP1, GIP4, and GIPh are higher than the first group importance threshold GIPTH1, the first pruning group PG1, the fourth pruning group PG4, and the hth pruning group PGh may be classified as an important pruning group IPG. In the example case where GIP2 and GIP3 are lower than the first group importance threshold GIPTH1, the second pruning group PG2 and the third pruning group PG3 may be classified as unimportant pruning groups UIPG.

[0262] In one embodiment, the matrix pruning apparatus 100 may classify the group type of the first pruning group PG1 to the hth pruning group PGh based on the second group importance threshold GIPTH2. In this case, the number of pruning groups classified as the important pruning group IPG may be changed compared to the case where the group type of each group in the first pruning group PG1 to the hth pruning group PGh is classified based on the first group importance threshold GIPTH1. For example, in the case where GIP4 and GIPh are higher than the first group importance threshold GIPTH1 and lower than the second group importance threshold GIPTH2, the fourth pruning group PG4 and the hth pruning group PGh may be classified as the unimportant pruning group UIPG based on the case where the matrix pruning apparatus 100 classifies the group type of each of the first pruning group PG1 to the hth pruning group PGh based on the second group importance threshold GIPTH2.

[0263] According to the disclosed embodiment, the number of pruning groups classified into the important pruning group IPG and the unimportant pruning group UIPG may be changed based on what group importance threshold the matrix pruning apparatus 100 classifies each group type from the first pruning group PG1 to the h-th pruning group PGh. In other words, since the number of the plurality of small sub-matrices SM_S and the number of the plurality of large sub-matrices SM_L may be changed by adjusting the group importance threshold, the number of the plurality of pruned small sub-matrices PSM_S and the number of the plurality of pruned large sub-matrices PSM_L may be adjusted. In this case, the number of times the first processing element array 240_S performs a matrix multiplication operation (e.g., as shown in FIG. 1 ) may be changed. Figure 19 The number of times the second processing element array 240_L performs matrix multiplication operations (for example, as shown in FIG. Figure 19 The "Q") can be adjusted, and thus the first operation time and the second operation time can be synchronized. Therefore, the operation efficiency of the matrix multiplication system MMS can be optimized by adjusting the group importance threshold.

[0264] Figure 24 The method for adjusting the Figure 19 A method of processing element array operations at a time. Figure 1 , Figure 3 , Figure 4 , Figure 5 , Figure 19 and Figure 24 , the matrix pruning device 100 can generate multiple small sub-matrices SM_S and multiple large sub-matrices SM_L based on the weight matrix WM. The matrix pruning device 100 can generate multiple pruned small sub-matrices PSM_S by pruning each of the multiple small sub-matrices SM_S, and can generate multiple pruned large sub-matrices PSM_L by pruning each of the multiple large sub-matrices SM_L.

[0265] The matrix pruning apparatus 100 may prune each of the plurality of small sub-matrices SM_S at the same pruning ratio. In this case, the number of columns including the remaining weight of each of the plurality of pruned small sub-matrices PSM_S may be the same.

[0266] The matrix pruning apparatus 100 may prune each of the plurality of large sub-matrices SM_L at the same pruning ratio. In this case, the number of columns including the remaining weights of each of the plurality of pruned large sub-matrices PSM_L may be the same.

[0267] The matrix pruning device 100 may adjust the pruning ratio of multiple small sub-matrices SM_S or the pruning ratio of multiple large sub-matrices SM_L. Hereinafter, for a more concise description, an embodiment in which the matrix pruning device 100 adjusts the pruning ratio of multiple small sub-matrices SM_S will be representatively described. However, the scope of the disclosure is not limited thereto, and the matrix pruning device 100 may be able to adjust the pruning ratio of multiple small sub-matrices SM_S and / or multiple large sub-matrices SM_L in a similar manner.

[0268] The matrix pruning apparatus 100 may calculate the column importance of each of the plurality of pruning units PU included in the first small sub-matrix SM_S1. In this case, the column importances of the pruning units PU corresponding to the first column to the mth column of the first small sub-matrix SM_S1 may be CIP1 to CIPh, respectively. The method of calculating the column importance of each of the plurality of pruning units PU by the matrix pruning apparatus 100 is previously described with reference to Figures 4 to 5 are described, and thus a further detailed description thereof will not be provided.

[0269] In one embodiment, the matrix pruning device 100 may prune a plurality of pruning units PU included in the first small sub-matrix SM_S1 based on a first pruning ratio PRR1. For example, the matrix pruning device 100 may maintain the first four pruning units having the highest column importance among the plurality of pruning units PU included in the first small sub-matrix SM_S1, and may convert the weights of the remaining pruning units into zero elements. For a more detailed example, in the case where the column importance information of the first four pruning units among CIP1 to CIPh are CIP2, CIP5, CIP6, and CIPm, respectively, the matrix pruning device 100 may maintain the weights included in the second column, the fifth column, the sixth column, and the mth column of the first small sub-matrix SM_S1, and convert the weights included in the remaining columns into zero elements.

[0270] In one embodiment, the matrix pruning device 100 may prune a plurality of pruning units PU included in the first small sub-matrix SM_S1 based on the second pruning ratio PRR2. For example, the matrix pruning device 100 may retain the first three pruning units with the highest column importance among the plurality of pruning units PU included in the first small sub-matrix SM_S1, and may convert the weights of the remaining pruning units into zero elements. For a more detailed example, in a case where the column importances of the first three pruning units among CIP1 to CIPh are CIP2, CIP5, and CIP6, respectively, the matrix pruning device 100 may retain the weights included in the second column, the fifth column, and the sixth column of the first small sub-matrix SM_S1, and convert the weights included in the remaining columns into zero elements. In other words, in a case where the matrix pruning device 100 prunes the plurality of pruning units PU included in the first small sub-matrix SM_S1 based on the second pruning ratio PRR2, the number of columns including the remaining weights of the first pruning small sub-matrix PSM_S1 may be reduced.

[0271] According to the disclosed embodiment, the number of remaining weights included in the plurality of pruned small sub-matrices PSM_S may vary depending on which pruning ratio the matrix pruning apparatus 100 uses to prune the plurality of small sub-matrices SM_S. In other words, the number of partial sums to be calculated by each processing element included in the first processing element array 240_S may be changed by adjusting the pruning ratio of the plurality of small sub-matrices SM_S. Therefore, the time taken by the first processing element array 240_S to perform a matrix multiplication operation on a small sub-matrix SM_S and the input matrix XM may be adjusted, and thus, the previously referenced Figure 19 The first operation time described may be adjusted.

[0272] Similarly, according to what pruning ratio the matrix pruning apparatus 100 uses to prune the plurality of large sub-matrices SM_L, the time taken by the second processing element array 240_L to perform a matrix multiplication operation on one large sub-matrix SM_L and the input matrix XM can be adjusted, and thus, the above reference Figure 19 The second operation time described may be adjusted.

[0273] Figure 25 Is used for synchronization Figure 19 A flowchart of a method for processing operation time of an element array. Figures 1 to 25 In S1000, the method may include generating one or more pruned small sub-matrices PSM_S and one or more pruned large sub-matrices PSM_L based on the weight matrix WM. For example, the matrix multiplication system MMS may generate a plurality of pruned small sub-matrices PSM_S and a plurality of pruned large sub-matrices PSM_L based on the weight matrix WM. Operation S1000 is related to Figure 14 The described operation S200 is similar, and thus a further detailed description thereof will not be provided.

[0274] In operation S2000, the method may include performing a matrix multiplication operation based on a pruned small submatrix PSM_S and a pruned large submatrix PSM_L. For example, the matrix multiplication system MMS may perform a matrix multiplication operation based on a plurality of pruned small submatrices PSM_S and a plurality of pruned large submatrices PSM_L. For example, the matrix multiplication system MMS may calculate an output matrix YM corresponding to the product of the weight matrix WM and the input matrix XM by calculating the product of a plurality of pruned small submatrices PSM_S and a plurality of pruned large submatrices PSM_L with the input matrix XM. In more detail, the matrix multiplication system MMS may use a first processing element array 240_S to calculate the product of each of the plurality of pruned small submatrices PSM_S and the input matrix XM, and may use a second processing element array 240_L to calculate the product of each of the plurality of pruned large submatrices PSM_L and the input matrix XM.

[0275] In operation S3000, the method may include determining whether the operation times of the first processing element array 240_S and the second processing element array 240_L are synchronized. For example, the matrix multiplication system MMS may determine whether the operation times are synchronized. For example, the matrix multiplication system MMS may determine whether the first operation time spent by the first processing element array 240_S to generate a plurality of small sub-output matrices SYM_S is the same as the second operation time spent by the second processing element array 240_L to generate a plurality of large sub-output matrices SYM_L. In other words, the matrix multiplication system MMS may determine whether the time consumed by the first processing element array 240_S to complete outputting a plurality of small sub-output matrices SYM_S is the same as the time consumed by the second processing element array 240_L to complete outputting a plurality of large sub-output matrices SYM_L.

[0276] In operation S3000, based on determining the operation time synchronization, the operation of the matrix multiplication system MMS may be terminated.

[0277] In operation S3000, based on determining that the operation times are not synchronized, the matrix multiplication system MMS may perform operation S4000.

[0278] In operation S4000, the method may include adjusting a group importance threshold GIPTH and / or a pruning ratio PRR. For example, the matrix multiplication system MMS may adjust the group importance threshold GIPTH and / or the pruning ratio PRR based on determining that the operation time is not synchronized. For example, the matrix multiplication system MMS may adjust the number of multiple pruned small sub-matrices PSM_S and the number of multiple pruned large sub-matrices PSM_L by adjusting the group importance threshold GIPTH. Additionally or optionally, the matrix multiplication system MMS may adjust the number of partial sums to be obtained by a processing element by adjusting the pruning ratios of multiple small sub-matrices SM_S or multiple large sub-matrices SM_L. Then, the above operation S1000 may be iteratively performed.

[0279] In this way, the matrix multiplication system MMS can synchronize the first operation time with the second operation time. In this case, the operation times of the first processing element array 240_S and the second processing element array 240_L are synchronized, and thus the operation efficiency of the matrix multiplication system MMS can be maximized.

[0280] In the example case where the sum of the number of rows of the first processing element array 240_S and the number of columns including the remaining weights in the plurality of pruned small sub-matrices PSM_S is the same as the sum of the number of rows of the second processing element array 240_L and the number of columns including the remaining weights in the plurality of pruned large sub-matrices PSM_L, the first operation time and the second operation time may be synchronized. That is, the matrix multiplication system MMS may adjust the sum of the number of rows of the first processing element array 240_S and the number of columns including the remaining weights in the plurality of pruned small sub-matrices PSM_S to be the same as the sum of the number of rows of the plurality of second processing element arrays 240_L and the number of columns including the remaining weights in the pruned large sub-matrix PSM_L through the above-mentioned operations S1000 to S4000. However, the scope of the disclosure is not limited thereto.

[0281] Figure 26 and Figure 27 The operation of the matrix pruning device according to another embodiment is shown. First, referring to Figure 1 , Figure 3 and Figure 26, the matrix pruning device 100 may generate a plurality of pruning groups PG. The matrix pruning device 100 may calculate the group importance of each of the plurality of pruning groups PG. For example, the matrix pruning device 100 may calculate the group importance of each of the first pruning group PG1 to the fourth pruning group PG4. The matrix pruning device 100 may classify each of the plurality of pruning groups PG as an important pruning group IPG or an unimportant pruning group UIPG based on the group importance of each of the plurality of pruning groups PG. The matrix pruning device 100 classifies each of the plurality of pruning groups PG as an important pruning group IPG or an unimportant pruning group UIPG in the same manner as the reference Figure 3 The manner of description is similar, and therefore a further detailed description thereof will not be provided.

[0282] The matrix pruning device 100 may generate a plurality of small sub-matrices SM_S based on a plurality of important pruning groups IPG. For example, the matrix pruning device 100 may generate a plurality of small sub-matrices SM_S by dividing each of a plurality of important pruning groups IPG into two or more small sub-matrices SM_S. For a more detailed example, the matrix pruning device 100 may divide the first pruning group PG1 by rows to generate a first small sub-matrix SM_1 and a second small sub-matrix SM_2. Similarly, the matrix pruning device 100 may generate a third small sub-matrix SM_3 and a fourth small sub-matrix SM_4 by dividing the fourth pruning group PG4 by rows.

[0283] In other words, the matrix pruning apparatus 100 may generate a small sub-matrix SM_S by dividing one pruning group PG.

[0284] The matrix pruning device 100 may generate a plurality of large sub-matrices SM_L based on a plurality of unimportant pruning groups UIPG. For example, the matrix pruning device 100 may generate a large sub-matrix SM_L based on one unimportant pruning group UIPG. Figure 27 , the second pruning group PG2 may be converted into the first largest sub-matrix SM_L1, and the third pruning group PG3 may be converted into the second largest sub-matrix SM_L2.

[0285] For example, refer to Figure 26 and Figure 27 Examples, with reference to Figure 3 Similar to the described embodiment, the important pruning group IPG can be converted into a sub-matrix of a relatively small size, and the unimportant pruning group UIPG can be converted into a sub-matrix of a relatively large size. In this case, since the important pruning group IPG can be pruned relatively finely, the error caused by the pruning weight matrix WM (for example, the operation error of the artificial intelligence model including the matrix multiplication system MMS) can be minimized.

[0286] Figure 28 is a block diagram of a matrix multiplication system according to an embodiment.Figures 1 to 28 , the matrix multiplication system MMS may include a matrix pruning device 300 and a matrix multiplication device 400.

[0287] The matrix pruning device 300 may receive a weight matrix WM. The matrix pruning device 300 may generate a plurality of pruned small sub-matrices PSM_S, a plurality of pruned medium sub-matrices PSM_M, and a plurality of pruned large sub-matrices PSM_L based on the weight matrix WM.

[0288] The number of rows included in each of the plurality of pruned medium sub-matrices PSM_M may be greater than the number of rows included in each of the plurality of pruned small sub-matrices PSM_S, and less than the number of rows included in each of the plurality of pruned large sub-matrices PSM_L.

[0289] The total number of rows included in the plurality of pruned small sub-matrices PSM_S, the plurality of pruned medium sub-matrices PSM_M, and the plurality of pruned large sub-matrices PSM_L may be equal to the number of rows of the weight matrix WM.

[0290] For example, the matrix pruning device 300 may generate three types of pruned sub-matrices based on the weight matrix WM. However, the disclosure is not limited thereto, and thus, according to another embodiment, the matrix pruning device 300 may generate more than three types of pruned sub-matrices based on the weight matrix WM. The operation of the matrix pruning device 300 will be described in more detail with reference to Figure 29 The operation of the matrix pruning device 300 will be described in more detail.

[0291] The matrix multiplication device 400 may receive an input matrix XM. The matrix multiplication device 400 may generate an output matrix YM based on the products of the plurality of pruned small sub-matrices PSM_S, the plurality of pruned medium sub-matrices PSM_M, and the plurality of pruned large sub-matrices PSM_L and the input matrix XM.

[0292] Figure 29 Shown Figure 28 The operation of the matrix pruning device is shown. Referring to Figures 1 to 28 , the matrix pruning device 300 may generate a first pruning group PG1 to an nth pruning group PGn based on the weight matrix WM. Hereinafter, for a more concise description, an embodiment in which the matrix pruning device 300 generates one pruning group for each row of the weight matrix WM is described, but the scope of the disclosure is not limited thereto. For example, in a manner similar to that described with reference to Figure 3 The matrix pruning device 300 may generate one pruning group for every two rows of the weight matrix WM.

[0293] The matrix pruning device 300 may calculate the group importance of each of the first pruning group PG1 to the nth pruning group PGn. The matrix pruning device 300 may classify each of the first pruning group PG1 to the nth pruning group PGn as an important pruning group IPG or an unimportant pruning group UIPG based on the group importance of each of the first pruning group PG1 to the nth pruning group PGn. For example, the matrix pruning device 300 may classify the fifth pruning group PG5 as an important pruning group IPG, and classify the first pruning group PG1 to the fourth pruning group PG4 and the sixth pruning group PG6 to the seventh pruning group PG7 as unimportant pruning groups UIPG.

[0294] The matrix pruning device 300 may generate a plurality of small sub-matrices SM_S based on a plurality of pruning groups classified as important pruning groups IPG. Figure 3 The methods described are similar and therefore further description thereof will not be provided.

[0295] The matrix pruning device 300 may generate a plurality of medium sub-matrices SM_M and a plurality of large sub-matrices SM_L based on a plurality of pruning groups classified as unimportant pruning groups UIPG. In this case, the number of pruning groups included in each of the plurality of medium sub-matrices SM_M may be greater than the number of pruning groups included in each of the plurality of small sub-matrices SM_S, and less than the number of pruning groups included in each of the plurality of large sub-matrices SM_L. For example, the number of pruning groups included in each of the plurality of small sub-matrices SM_S may be 1, the number of pruning groups included in each of the plurality of medium sub-matrices SM_M may be 2, and the number of pruning groups included in each of the plurality of large sub-matrices SM_L may be 4.

[0296] Next, the matrix pruning device 300 may generate a plurality of pruned small sub-matrices PSM_S by pruning a plurality of small sub-matrices SM_S, may generate a plurality of pruned medium sub-matrices PSM_M by pruning a plurality of medium sub-matrices SM_M, and may generate a plurality of pruned large sub-matrices PSM_L by pruning a plurality of large sub-matrices SM_L. Figure 4 and Figure 5 The methods described are similar, and thus a detailed description thereof will be omitted.

[0297] In one embodiment, the matrix trimming device 300 may generate multiple medium sub - matrices SM_M by combining two adjacent unimportant trimming groups UIPG, and may generate multiple large sub - matrices SM_L by combining four adjacent unimportant trimming groups UIPG. However, the disclosed scope is not limited thereto. For example, the matrix trimming device 300 may classify the first trimming group PG1 to the nth trimming group PGn into important trimming groups, medium - important trimming groups, and unimportant trimming groups respectively based on the group importance information of the first trimming group PG1 to the nth trimming group PGn. In this case, the matrix trimming device 300 may be implemented to generate multiple small sub - matrices based on multiple important trimming groups, generate multiple medium sub - matrices based on multiple medium - important trimming groups, and generate multiple large sub - matrices based on multiple unimportant trimming groups. For example, the matrix trimming device 300 may generate multiple small sub - matrices based on multiple important trimming groups having a group importance higher than a first group importance threshold, generate multiple medium sub - matrices based on multiple medium - important trimming groups having a group importance higher than a second group importance threshold and lower than the first group importance threshold, and generate multiple large sub - matrices based on multiple unimportant trimming groups having a group importance lower than the second group importance threshold. In other words, the disclosed scope is not limited to the specific manner in which the matrix trimming device 300 generates sub - matrices of different sizes. Although the terms "higher than" or "lower than" are used to classify the trimming groups, the disclosure is not limited thereto, and thus, according to another embodiment, other combinations such as "equal to or higher than", "equal to or lower than", etc. may be used to classify the trimming groups.

[0298] Figure 30 is shown in more detail Figure 28 block diagram of the matrix multiplication device. Referring to Figures 1 to 30 , the matrix multiplication device 400 may include a weight memory circuit 410, a control logic circuit 420, an input matrix buffer 430, a first processing element array 440_S, a second processing element array 440_M, a third processing element array 440_L, and an output merging circuit 450. The configurations and functions of the weight memory circuit 410, the control logic circuit 420, the input matrix buffer 430, and the output merging circuit 450 are similar to those described previously with respect to Figure 6 and thus further detailed descriptions thereof will not be provided.

[0299] The first processing element array 440_S may receive the remaining weights RW_PSM_S included in multiple trimmed small sub - matrices PSM_S. The first processing element array 440_S may receive multiple input elements IE corresponding to the remaining weights RW_PSM_S. The first processing element array 440_S may generate multiple small output matrices SYM_S based on the remaining weights RW_PSM_S and the received input elements IE.

[0300] The second processing element array 440_M may receive the remaining weights RW_PSM_M included in a plurality of pruned medium sub-matrices PSM_M. The second processing element array 440_M may receive a plurality of input elements IE corresponding to the remaining weights RW_PSM_M. The second processing element array 440_M may generate a plurality of medium sub-output matrices SYM_M based on the remaining weights RW_PSM_M and the received input elements IE.

[0301] The third processing element array 440_L may receive the remaining weights RW_PSM_L included in a plurality of pruned large sub-matrices PSM_L. The third processing element array 440_L may receive a plurality of input elements IE corresponding to the remaining weights RW_PSM_L. The third processing element array 440_L may generate a plurality of large sub-output matrices SYM_L based on the remaining weights RW_PSM_L and the received input elements IE.

[0302] In one embodiment, the number of columns of each of the first processing element array 440_S, the second processing element array 440_M, and the third processing element array 440_L may be the same.

[0303] In one embodiment, the number of rows of the first processing element array 440_S may be the same as the number of rows of the plurality of pruned small sub-matrices PSM_S. The number of rows of the second processing element array 440_M may be the same as the number of rows of the plurality of pruned medium sub-matrices PSM_M. The number of rows of the third processing element array 440_L may be the same as the number of rows of the plurality of pruned large sub-matrices PSM_L.

[0304] The output merging circuit 450 may generate an output matrix YM by merging a plurality of small sub-output matrices SYM_S, a plurality of medium sub-output matrices SYM_M, and a plurality of large sub-output matrices SYM_L.

[0305] Figure 31 Showing a Figure 1 matrix multiplication system according to an embodiment. Referring to Figures 1 to 31 , the matrix multiplication system MMS may include a matrix pruning device 100, a matrix multiplication device 200, and an output matrix tiling device OMTD.

[0306] The matrix pruning device 100 may receive a full weight matrix FWM. The full weight matrix FWM may include the weight matrix WM previously described with reference to Figures 1 to 30 . For example, the full weight matrix FWM may include a plurality of weight matrices WM.

[0307] The matrix pruning device 100 can generate a plurality of pruned small sub-matrices PSM_S and a plurality of pruned large sub-matrices PSM_L based on each of a plurality of weight matrices WM included in the full weight matrix FWM.

[0308] The matrix multiplication device 200 can receive the full input matrix FXM. Referring to Figures 1 to 30 , the full input matrix FXM can include the input matrix XM described previously with reference to Figures 1 to 30 .

[0309] The matrix multiplication device 200 can generate a plurality of output matrices YM based on the plurality of pruned small sub-matrices PSM_S and the plurality of pruned large sub-matrices PSM_L. For example, the matrix multiplication device 200 can generate the plurality of output matrices YM as the result of multiplying each of the plurality of pruned small sub-matrices PSM_S and the plurality of pruned large sub-matrices PSM_L by a corresponding input matrix XM.

[0310] The output matrix tiling device OMTD can receive the plurality of output matrices YM. The output matrix tiling device OMTD can generate a full output matrix FYM corresponding to the product of the full weight matrix FWM and the full input matrix FXM based on the plurality of output matrices YM.

[0311] That is, the matrix multiplication system MMS can perform matrix multiplication on the full input matrix FXM and the full weight matrix FWM by one of various tiling techniques. In other words, the matrix multiplication system MMS can generate the full output matrix FYM by sequentially obtaining the products of the plurality of input matrices XM and the plurality of weight matrices WM and then combining the obtained results.

[0312] Figure 32 Shown Figure 31 is the full input matrix. Referring to Figures 1 to 32 , the full input matrix FXM can include a plurality of input matrices XM. In other words, the full input matrix FXM can be tiled with a plurality of input matrices XM arranged in the row direction and the column direction. Hereinafter, for a more concise description, the input matrix arranged in the (i)-th row and the (j)-th column of the full input matrix FXM may be referred to as "XM_ij".

[0313] In one embodiment, the input matrix XM described with reference to Figures 1 to 30 can be one of the plurality of input matrices XM included in the full input matrix FXM.

[0314] In one embodiment, each of the plurality of input matrices XM included in the full input matrix FXM may have the same row size and column size. For example, each of the plurality of input matrices XM may include b input elements per row. Each of the plurality of input matrices XM may include m input elements per column.

[0315] The row size of the full input matrix FXM can be a positive integer multiple of the row sizes of each input matrix XM. For example, one row of the full input matrix FXM can contain B input elements. In this case, B can be a positive integer multiple of b.

[0316] The column size of the full input matrix FXM can be an integer multiple of the column sizes of each of the multiple input matrices XM. For example, one column of the full input matrix FXM can include M input elements. In this case, M can be a positive integer multiple of m.

[0317] Figure 33 shown Figure 31 of the full weight matrix. Referring to Figures 1 to 33 , the full weight matrix FWM can include multiple weight matrices WM. In other words, the full weight matrix FWM can be tiled into multiple weight matrices WM arranged in the row direction and the column direction. Hereinafter, for a more concise description, the weight matrix arranged in the (i)-th row and the (j)-th column of the full weight matrix FWM may be referred to as "WM_ij".

[0318] In one embodiment, the weight matrix WM described previously with reference to Figures 1 to 30 can be one of the multiple weight matrices WM included in the full weight matrix FWM.

[0319] In one embodiment, each of the multiple weight matrices WM included in the full weight matrix FWM can have the same row size and column size. For example, each of the multiple weight matrices WM can include m weights per row. Each of the multiple weight matrices WM can include n weights per column.

[0320] The row size of the full weight matrix FWM can be a positive integer multiple of the row sizes of each of the multiple weight matrices WM. For example, M weights can be included in one row of the full weight matrix FWM. In this case, M can be a positive integer multiple of m.

[0321] The column size of the full weight matrix FWM can be a positive integer multiple of the row sizes of each of the multiple weight matrices WM. For example, N weights can be included in one column of the full weight matrix FWM. In this case, N can be an integer multiple of n.

[0322] The matrix pruning device 100 may generate a plurality of pruned small sub-matrices PSM_S and a plurality of pruned large sub-matrices PSM_L based on each of a plurality of weight matrices WM included in the full weight matrix FWM. In this case, the plurality of pruned small sub-matrices PSM_S and the plurality of pruned large sub-matrices PSM_L generated based on the weight matrix WM_11 may be different from the plurality of pruned small sub-matrices PSM_S and the plurality of pruned large sub-matrices PSM_L generated based on the weight matrix WM_12. The method by which the matrix pruning device 100 generates a plurality of pruned small sub-matrices PSM_S and a plurality of pruned large sub-matrices PSM_L based on each of the weight matrices WM is similar to the method described with reference to Figures 1 to 30 and thus a further detailed description thereof will not be provided.

[0323] Figure 34 shown Figure 31 of the full output matrix. With reference to Figures 1 to 34 , the full output matrix FYM may correspond to the product of the full weight matrix FWM and the full input matrix FXM.

[0324] The full output matrix FYM may include a plurality of sub-matrices FYM_sub arranged in a row direction and a column direction. Hereinafter, for a more concise description, the sub-matrix arranged in the (i)-th row and the (j)-th column of the full output matrix FYM may be referred to as FYM_sub_ij.

[0325] Each of the plurality of sub-matrices FYM_sub may have the same row size and column size. The column size of each of the plurality of sub-matrices FYM_sub may be the same as the column size of the weight matrix WM. The row size of each of the plurality of sub-matrices FYM_sub may be the same as the row size of the input matrix XM. For example, each of the plurality of sub-matrices FYM_sub may have b output elements per row. Each of the plurality of sub-matrices FYM_sub may include n output elements per column.

[0326] The row size of the full output matrix FYM may be the same as the row size of the full input matrix FXM. For example, the row size of the full output matrix FYM may be B.

[0327] The column size of the full output matrix FYM may be the same as the column size of the full weight matrix FWM. For example, the column size of the full output matrix FYM may be N.

[0328] The matrix multiplication system MMS may calculate the full output matrix FYM in units of the sub-matrix FYM_sub. For example, the matrix multiplication system MMS may calculate one sub-matrix FYM_sub by adding the products of a plurality of input matrices XM and a plurality of weight matrices WM.

[0329] For a more detailed example, in the case where M is three times m, the matrix multiplication device 200 may sequentially calculate a first output matrix corresponding to the product of the weight matrix WM_11 and the input matrix XM_11, a second output matrix corresponding to the product of the weight matrix WM_12 and the input matrix XM_21, and a third output matrix corresponding to the product of the weight matrix WM_13 and the input matrix XM_31. In this case, the output matrix tiling device OMTD may calculate the sub-matrix FYM_sub_11 by adding the first to third output matrices. In this way, the output matrix tiling device OMTD may be able to generate the full output matrix FYM by sequentially obtaining a plurality of sub-matrices FYM_sub.

[0330] Figure 35 is a block diagram of a neural processing system implemented according to an embodiment. Referring to Figure 35 , the neural processing system 2000 may include a central processing unit (CPU) 2100, a neural processing unit (NPU) 2200, a volatile memory device 2300, a non-volatile memory device 2400, and a user interface 2500. The CPU 2100, NPU 2200, volatile memory device 2300, non-volatile memory device 2400, and user interface 2500 may be connected to each other via a bus BUS.

[0331] The CPU 2100 may control the overall operation of the neural processing system 2000. For example, the CPU 2100 may control each component of the neural processing system 2000 to run an artificial intelligence model.

[0332] In one embodiment, the artificial intelligence model executed by the neural processing system 2000 may be one of any type of artificial intelligence model (such as, a language model, an image recognition model, an image generation model, a weather analysis model, etc.). For example, the artificial intelligence model executed by the neural processing system 2000 may be one of any type of artificial intelligence model (such as, GPT-3, GPT-4, Pangu, GShard, Megatron-LM, etc.). However, the disclosed scope is not limited thereto.

[0333] In one embodiment, the artificial intelligence model executed by the neural processing system 2000 may perform inference operations and / or training operations. However, the disclosed scope is not limited thereto.

[0334] Each artificial intelligence model may include a plurality of processing layers. Each of the plurality of processing layers may be implemented to generate "layer output data" by receiving "layer input data". In this case, the generated layer output data may be used as the layer input data for other processing layers. For example, the layer output data generated from the first processing layer may be used as the layer input data for the second processing layer. Referring to Figure 36Provide a more detailed description of the artificial intelligence model and the processing layer.

[0335] Each of the multiple processing layers may convert the layer input data into layer output data based on matrix multiplication operations. For example, each of the multiple processing layers may generate an output matrix corresponding to the layer output data by multiplying an input matrix corresponding to the layer input data by a weight matrix. However, the disclosed scope is not limited thereto, and each of the multiple processing layers may be capable of generating the layer output data by converting the input matrix corresponding to the layer input data in any scheme. For example, each of the multiple processing layers may be implemented to convert the input matrix into the layer output data based on any transformation parameters. In other words, the disclosed scope is not limited to the specific manner in which each of the multiple processing layers converts the layer input data.

[0336] The NPU 2200 may include a matrix multiplication system 2210. The matrix multiplication system 2210 may execute at least a part of the operations included in the multiple processing layers. For example, the matrix multiplication system 2210 may execute matrix multiplication operations included in the multiple processing layers.

[0337] In one embodiment, the matrix multiplication operation may account for most of the processing load required for the neural processing system 2000 to execute each of the multiple processing layers.

[0338] In one embodiment, the matrix multiplication system 2210 may be implemented as the matrix multiplication system MMS described above with reference to Figures 1 to 30 For example, the matrix multiplication system 2210 may include a matrix pruning device 100. In this case, the matrix multiplication system 2210 may generate a plurality of pruned small submatrices PSM_S and a plurality of pruned large submatrices PSM_L based on a weight matrix WM divided into a plurality of pruning groups PG. In this case, since the weight matrix WM may be pruned in consideration of the importance of each of the plurality of pruning groups PG, the error of the output matrix generated by the matrix multiplication system 2210 may be minimized. Therefore, according to the disclosed embodiment, since a high pruning ratio may be applied to low-importance weights and a low pruning ratio may be applied to high-importance weights, the operation speed of the artificial intelligence model may be increased, and thus the reduction in the operation accuracy of the intelligent model may be minimized.

[0339] In one embodiment, the matrix multiplication system 2210 may be implemented as the one described above with reference to Figures 1 to 30The described matrix multiplication system MMS. For example, the matrix multiplication system 2210 may include a matrix multiplication device 200. In this case, the matrix multiplication system 2210 may use the first processing element array 240_S to calculate the product of each of the plurality of pruned small sub-matrices PSM_S and the input matrix, and may use the second processing element array 240_L to calculate the product of each of the plurality of pruned large sub-matrices PSM_L and the input matrix. In this case, product operations may be performed on matrices of different sizes according to the sizes of the first processing element array 240_S and the second processing element array 240_L, and thus the operation efficiency of the matrix multiplication system 2210 may be maximized.

[0340] The volatile memory device 2300 may be used as the operation memory of the NPU 2200. For example, the volatile memory device 2300 may temporarily store data generated during the operation of the NPU 2200.

[0341] In one embodiment, the NPU 2200 may access the volatile memory device 2300 to perform operations included in a plurality of processing layers. For example, the NPU 2200 may be implemented to read parameters stored in the volatile memory device 2300 and perform operations on layer input data, or may be implemented to temporarily store intermediate data generated during the operation in the volatile memory device 2300.

[0342] In one embodiment, the volatile memory device 2300 may be implemented using any type of volatile memory (such as dynamic random access memory (DRAM) or static random access memory (SRAM)).

[0343] In one embodiment, the volatile memory device 2300 may be used as a buffer memory, an operation memory, or a cache memory of the CPU 2100. However, the disclosed scope is not limited thereto.

[0344] The non-volatile memory device 2400 may store data for the operation of the neural processing system 2000. For example, the non-volatile memory device 2400 may store various types of data (such as the operating system (OS) of the neural processing system 2000 or parameters for running an artificial intelligence model). However, the disclosed scope is not limited thereto.

[0345] The CPU 2100 may communicate with the user through the user interface 2500. The CPU 2100 may provide model input data provided by the user to the volatile memory device 2300 or the NPU 2200 through the user interface 2500. The CPU 2100 may return model output data generated by the artificial intelligence model based on the model input data to the user through the user interface 2500.

[0346] Figure 36 is a block diagram of an artificial intelligence model run by a neural processing system. Refer to Figure 35 and Figure 35 and Figure 36 , the neural processing system 2000 can run the artificial intelligence model AIM.

[0347] The artificial intelligence model AIM can receive model input data MID. The artificial intelligence model AIM can include a plurality of processing layers PL_1 to PL_L.

[0348] The artificial intelligence model AIM can generate model output data MOD by sequentially transforming the model input data MID through the first processing layer PL_1 to the Lth processing layer PL_L. For example, the first processing layer PL_1 can receive the model input data MID and generate second layer input data LID_2. The second processing layer PL_2 can receive the second layer input data LID_2 and generate third layer input data LID_3. In this way, the Lth processing layer PL_L can receive the Lth layer input data LID_L and generate the model output data MOD.

[0349] Each of the first processing layer PL_1 to the Lth processing layer PL_L can transform the received data into the data to be output through various types of operations. For example, a matrix multiplication operation can be included in the operations performed by the first processing layer PL_1 to transform the model input data MID into the second layer input data LID_2. Similarly, each of the first processing layer PL_1 to the Lth processing layer PL_L may need to perform a matrix multiplication operation to transform the received layer input data. However, the disclosed scope is not limited thereto, and some of the first processing layer PL_1 to the Lth processing layer PL_L may not perform a matrix multiplication operation.

[0350] In one embodiment, the matrix multiplication operations performed by each of the first processing layer PL_1 to the Lth processing layer PL_L can be executed by a matrix multiplication system 2210.

[0351] In one embodiment, in the case where the matrix multiplication system 2210 is implemented as the matrix multiplication system MMS described above with reference to Figures 1 to 30 , the matrix multiplication system 2210 may be able to output the result of the matrix multiplication operation with higher precision and faster speed. Therefore, according to the disclosed embodiment, the operation speed of the artificial intelligence model AIM can be improved.

[0352] For a more concise description, in Figure 36Among them, an embodiment composed of multiple processing layers serially operated by the artificial intelligence model AIM is representatively described, but the disclosed scope is not limited thereto. For example, the artificial intelligence model AIM may also include a processing layer that operates in parallel with at least some of the above-mentioned first processing layer PL_1 to the L-th processing layer PL_L. In other words, the disclosed scope is not limited to the specific implementation method of the artificial intelligence model AIM.

[0353] The above content is for implementing the disclosed specific embodiments. The disclosure will include not only the above embodiments, but also embodiments that have been simply redesigned or can be easily modified. In addition, the disclosure will also include technologies that can be easily modified and implemented using the embodiments. Therefore, the scope of the disclosure should not be limited to the above embodiments, but should be determined by the appended patent claims and those equivalent to the disclosed patent claims.

Claims

1. A matrix multiplication device, comprising: a weight memory circuit configured to: store a first pruned sub-matrix including a first plurality of residual weights and a second pruned sub-matrix including a second plurality of residual weights; Input matrix buffer, configured as: receives an input matrix comprising a plurality of input elements, outputting a first plurality of input elements corresponding to the first plurality of residual weights among the plurality of input elements, and outputting a second plurality of input elements corresponding to the second plurality of residual weights among the plurality of input elements; The first processing element array is configured as follows: receiving the first plurality of residual weights and the first plurality of input elements, and outputting a first sub-output matrix; and The second processing element array is configured as follows: receiving the second plurality of residual weights and the second plurality of input elements, and Output the second sub-output matrix.

2. The matrix multiplication device according to claim 1, wherein: The number of columns of the first pruned sub-matrix is ​​the same as the number of columns of the second pruned sub-matrix, and The number of rows of the second pruning sub-matrix is ​​n times the number of rows of the first pruning sub-matrix, where n is a positive integer.

3. The matrix multiplication device according to claim 2, wherein: The first processing element array includes a first plurality of processing elements arranged in a row direction and a column direction, The second processing element array includes a second plurality of processing elements arranged in a row direction and a column direction, The number of rows of the first processing element array is the same as the number of rows of the first pruning sub-matrix, The number of rows of the second processing element array is the same as the number of rows of the second pruning sub-matrix, and The number of columns of the first processing element array is the same as the number of columns of the second processing element array.

4. The matrix multiplication device according to claim 1, wherein: The first pruning sub-matrix further includes a first plurality of zero elements, and the first plurality of residual weights and the first plurality of zero elements are arranged in different columns of the first pruning sub-matrix, and The second pruned sub-matrix also includes a second plurality of zero elements, and the second plurality of residual weights and the second plurality of zero elements are arranged in different columns of the second pruned sub-matrix.

5. The matrix multiplication device according to claim 4, wherein: The input matrix buffer is configured as: outputting the first plurality of input elements based on a first residual weight index, the first residual weight index indicating a first plurality of column numbers including the first plurality of residual weights in a first pruning sub-matrix, and The second plurality of input elements are output based on second residual weight indices indicating a second plurality of column numbers including the second plurality of residual weights in a second pruning sub-matrix.

6. The matrix multiplication device according to claim 5, wherein: The first plurality of input elements are arranged in a first plurality of input element rows in an input matrix, The first plurality of residual weights are arranged in a first plurality of columns in a first pruning sub-matrix, The row numbers of the first plurality of input element rows in the input matrix correspond to the column numbers of the first plurality of columns in the first pruning sub-matrix, The second plurality of input elements are arranged in a second plurality of input element rows in the input matrix, The second plurality of residual weights are arranged in a second plurality of columns in a second pruning sub-matrix, and The row numbers of the second plurality of input element rows in the input matrix correspond to the column numbers of the second plurality of columns in the second pruning sub-matrix.

7. The matrix multiplication device according to claim 1, wherein: The weight memory circuit is also configured to: storing a first plurality of pruned sub-matrices including a first pruned sub-matrix, and storing a second plurality of pruned sub-matrices including a second pruned sub-matrix, The first processing element array is further configured to: generate a first plurality of sub-output matrices, each of the first plurality of sub-output matrices corresponding to a product of an input matrix and one of the first plurality of pruned sub-matrices, and The second processing element array is further configured to generate a second plurality of sub-output matrices, each of which corresponds to a product of the input matrix and one of the second plurality of pruned sub-matrices.

8. The matrix multiplication device according to claim 7, wherein: A first operation time for generating the first plurality of sub-output matrices by the first processing element array corresponds to a second operation time for generating the second plurality of sub-output matrices by the second processing element array.

9. The matrix multiplication device according to claim 8, wherein: Based on the same number of the first plurality of pruned sub-matrices and the same number of the second plurality of pruned sub-matrices, the time for the first processing element array to generate the first plurality of sub-output matrices is the same as the time for the second processing element array to generate the second plurality of sub-output matrices.

10. The matrix multiplication device according to claim 8, wherein: Among the first plurality of residual weights, two or more residual weights arranged in the same row of the first pruning sub-matrix are stored in consecutive addresses in the weight memory circuit, and Among the second plurality of residual weights, two or more residual weights arranged in the same row of the second pruning sub-matrix are stored in consecutive addresses in the weight memory circuit.

11. The matrix multiplication device according to any one of claims 1 to 10, further comprising: The output merging circuit is configured as: An output matrix is ​​generated based on the first sub-output matrix and the second sub-output matrix.

12. A matrix multiplication system comprising: A matrix pruning device generates a first plurality of pruned sub-matrices and a second plurality of pruned sub-matrices based on the weight matrix; as well as A matrix multiplication device comprising: a first processing element array configured to generate a first plurality of sub-output matrices based on an input matrix and the first plurality of pruned sub-matrices, a second processing element array configured to generate a second plurality of sub-output matrices based on the input matrix and the second plurality of pruned sub-matrices, and The output merging circuit is configured to generate an output matrix based on the first plurality of sub-output matrices and the second plurality of sub-output matrices.

13. The matrix multiplication system of claim 12, wherein: the number of columns of each of the first plurality of pruning sub-matrices and the second plurality of pruning sub-matrices is a first number, The number of rows of each of the first plurality of pruning sub-matrices is a second number, and The number of rows of each of the second plurality of pruning sub-matrices is a third number, the third number is n times the second number, and n is a positive integer.

14. The matrix multiplication system of claim 13, wherein: the number of columns of the first processing element array and the number of columns of the second processing element array are a first number, The number of rows of the first processing element array is a second number, and The number of rows of the second processing element array is a third number.

15. The matrix multiplication system of claim 14, wherein: The matrix pruning device is further configured to: generate a third plurality of pruned sub-matrices based on the weight matrix, wherein the third plurality of pruned sub-matrices include a fourth number of rows and a first number of columns, the fourth number being m times the second number, wherein m is a positive integer, The matrix multiplication apparatus further comprises: a third processing element array configured to generate a third plurality of sub-output matrices based on the input matrix and the third plurality of pruned sub-matrices, The output merging circuit is further configured to: generate an output matrix further based on the third plurality of sub-output matrices, and The fourth number is smaller than the third number.

16. The matrix multiplication system of claim 12, wherein: A first operation time for generating the first plurality of sub-output matrices by the first processing element array corresponds to a second operation time for generating the second plurality of sub-output matrices by the second processing element array.

17. The matrix multiplication system of claim 16, wherein: Based on the first operation time and the second operation time, the matrix pruning apparatus is further configured to determine the number of columns including the remaining weights of each of the first plurality of pruned sub-matrices and the number of columns including the remaining weights of each of the second plurality of pruned sub-matrices.

18. The matrix multiplication system of claim 16, wherein: Based on the first operation time and the second operation time, the matrix pruning apparatus is further configured to: determine the number of the first plurality of pruned sub-matrices and the number of the second plurality of pruned sub-matrices.

19. A method for operating a matrix multiplication device, comprising: Generate multiple pruning groups by partitioning the weight matrix; classifying the plurality of prune groups into a first plurality of prune groups and a second plurality of prune groups based on a group importance of each of the plurality of prune groups; generating a first plurality of sub-matrices based on the first plurality of pruning groups; generating a second plurality of sub-matrices based on the second plurality of pruning groups; generating a first plurality of pruned sub-matrices by pruning each of the first plurality of sub-matrices; generating a second plurality of pruned sub-matrices by pruning each of the second plurality of sub-matrices; as well as An output matrix is ​​generated based on a product of an input matrix and each of the first plurality of pruning sub-matrices and the second plurality of pruning sub-matrices.

20. The operating method according to claim 19, wherein: Each of the first plurality of sub-matrices corresponds to one of the first plurality of pruning groups, and Each of the second plurality of sub-matrices corresponds to two or more of the second plurality of pruning groups.

Citation Information

Cited By

  • Matrix multiplication apparatus

    TWI914280B