Method and device for generating matrix index information, and method and device for processing matrix using matrix index information
By generating matrix index information that includes quantization and position data for non-zero elements, the method addresses the high memory overhead in neural networks, enhancing efficiency and reducing power consumption in matrix operations.
Patent Information
- Application Number
- PCT/KR2024/021207
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-20
- Filing Date
- 2024-12-27
- Publication Date
- 2025-07-03
AI Technical Summary
The increasing depth of layers in neural network models leads to a high memory overhead due to the growing number of parameters, particularly in sparse matrices, which is exacerbated by inefficient memory access and power consumption.
Generating matrix index information that includes quantization and position information of non-zero elements, allowing for efficient memory access and reduced memory transactions by standardizing data sizes to multiples of unit data blocks, and selectively loading only necessary elements for operations.
Improves memory access efficiency, reduces power consumption, and enhances the speed and accuracy of matrix operations by minimizing unnecessary data transfers and optimizing data loading based on matrix index information.
Smart Images

Figure KR2024021207_03072025_PF_FP_ABST
Abstract
Description
Method and device for generating matrix index information, method and device for processing matrix using matrix index information
[0001] The present invention relates to a method and device for generating index information of a matrix, and a method and device for processing a matrix using the index information of the matrix, and more particularly, to a method and device for generating index information for a target matrix including a sparse matrix, and a method and device for processing a matrix using the index information of the matrix.
[0002]
[0003] With the recent advancement of neural network models, such as the Convolutional Neural Network (CNN) model, utilized in service fields such as image recognition, the depth of layers that neural network models must process is increasing. These factors have led to an increase in the number of parameters, such as the weight matrix, in neural network models, and high memory overhead has emerged as a significant issue.
[0004] To address this, research has been conducted on matrix indexing methods that can efficiently perform operations on sparse matrices, by utilizing the fact that the pruning technique performed to solve the overfitting problem of neural network models turns the weight matrix into a sparse matrix. CSR (Compressed Sparse Row) is widely used as an indexing method for sparse matrices. Processing elements of processors, such as ALU (Arithmetic Logic Unit), load matrix data from memory using the index information for the matrix generated through the matrix indexing method, and perform operations on the matrix using the loaded matrix data.
[0005] Furthermore, to efficiently perform matrix operations in neural network models, quantization techniques are utilized to reduce the data size of matrix elements. For example, matrix elements typically expressed in 32 bits can be quantized and operated on in smaller bit sizes, such as 16 or 8 bits.
[0006] Recently, mixed precision quantization, which adjusts the quantization bits of matrix elements according to the importance of the element or whether it is to be operated on with outliers, has been attracting attention.
[0007]
[0008] The present invention provides a method for generating matrix index information and a method for processing matrix information, which can improve memory access efficiency and reduce the number of memory accesses.
[0009] In addition, the present invention provides a method for transmitting and processing matrix data for efficient matrix operations.
[0010] In addition, the present invention provides a method and device for generating index information for an operation area of an operand matrix, and a method and device for processing a matrix using the index information for an operation area of an operand matrix.
[0011] In addition, the present invention provides a method and device for generating matrix index information, and a method and device for processing matrix, which can reduce the number of memory accesses for matrix operations.
[0012] In addition, the present invention provides a method and device for generating matrix index information including characteristic information of a matrix element, and a method and device for processing a matrix using matrix index information including characteristic information of a matrix element.
[0013]
[0014] According to one embodiment of the present invention for achieving the above object, a method for generating matrix index information is provided, including: a step of identifying a non-zero element in a target matrix; and a step of generating matrix index information including quantization information for the target matrix, data size information of the target matrix, and position information of the non-zero element.
[0015] In addition, according to another embodiment of the present invention for achieving the above-mentioned object, a method for generating matrix index information is provided, including: a step of identifying a non-zero element in a target matrix; and a step of generating matrix index information including data size information of the target matrix and position information of each of the non-zero elements, wherein a bit string representing the position information includes a bit corresponding to each position of an element in the target matrix, and in the bit string, a bit value corresponding to the position of the zero element of the target matrix and a bit value corresponding to the position of the non-zero element are different from each other.
[0016] In addition, according to another embodiment of the present invention for achieving the above-mentioned purpose, a method for generating matrix index information is provided, including: a step of identifying a non-zero element in a target matrix; and a step of generating matrix index information including quantization information for the target matrix, information on the number of the non-zero elements, and information on the position of the non-zero elements.
[0017] In addition, according to another embodiment of the present invention for achieving the above-mentioned object, a method for generating matrix index information is provided, including: a step of identifying a non-zero element in a target matrix; and a step of generating matrix index information including quantization information for the target matrix and position information of each of the non-zero elements, wherein a bit string representing the position information includes a bit corresponding to each position of an element in the target matrix, and in the bit string, a bit value corresponding to the position of the zero element of the target matrix and a bit value corresponding to the position of the non-zero element are different from each other.
[0018] In addition, according to another embodiment of the present invention for achieving the above-mentioned object, a matrix processing method is provided, including a step of loading a non-zero element of a first target matrix from a memory using matrix index information for the first target matrix; and a step of transferring the loaded data to a calculator, wherein the matrix index information includes quantization information for the first target matrix, data size information of the first target matrix, and position information of the non-zero element.
[0019] In addition, according to another embodiment of the present invention for achieving the above-mentioned object, a matrix processing method is provided, including a step of loading a non-zero element of a first target matrix from a memory using matrix index information for the first target matrix; and a step of transferring the loaded data to a calculator, wherein the matrix index information includes quantization information for the first target matrix, information on the number of the non-zero elements, and information on the position of the non-zero elements.
[0020] In addition, according to another embodiment of the present invention for achieving the above-mentioned purpose, a matrix data transmission method is provided, including a step of selecting an element necessary for an operation of the first and second matrices from among the elements of the first and second matrices, respectively, using matrix index information for each of the first and second matrices; and a step of transmitting the selected element to an upper memory or an operator.
[0021] In addition, according to another embodiment of the present invention for achieving the above-mentioned purpose, a matrix data processing method is provided, including the steps of: selectively loading elements necessary for operations on the first and second matrices from among the elements of the first and second matrices stored in a lower memory into an upper memory or an operator using matrix index information for each of the first and second matrices; and performing operations on the first and second matrices using the loaded elements.
[0022] In addition, according to another embodiment of the present invention for achieving the above-mentioned purpose, a matrix data transmission method is provided, including a step of identifying a non-zero element among elements of each of the first and second matrices using matrix index information for each of the first and second matrices; and a step of transmitting a memory address value for the non-zero element generated using the matrix index information and the non-zero element to an upper memory or an operator.
[0023] In addition, according to another embodiment of the present invention for achieving the above-mentioned object, a matrix data processing method is provided, including the steps of loading a non-zero element from among the elements of each of the first and second matrices stored in a lower memory using matrix index information for each of the first and second matrices, and loading a memory address value for the non-zero element into an upper memory; and performing an operation on the first and second matrices using a non-zero element selected according to the memory address value from among the loaded non-zero elements.
[0024] In addition, according to another embodiment of the present invention for achieving the above-mentioned purpose, a method for generating matrix index information is provided, including the steps of: determining an operation area in which an operation is performed with a non-zero element of a second operand matrix in a first operand matrix; and generating partial matrix index information for the operation area using a non-zero element included in the operation area.
[0025] In addition, according to another embodiment of the present invention for achieving the above-mentioned object, a method for processing a matrix is provided, including the steps of loading elements of the first and second operand matrices using matrix index information of the first and second operand matrices; and the step of performing an operation on the loaded elements, wherein the matrix index information of the first operand matrix is partial matrix index information generated using a non-zero element included in an operation area of the first operand matrix, with which an operation is performed with a non-zero element of the second operand matrix.
[0026] In addition, according to another embodiment of the present invention for achieving the above-mentioned object, a matrix index information generating device is provided, which includes a memory; and a processor electrically connected to the memory, wherein the processor determines an operation area in which an operation is performed with a non-zero element of a second operand matrix in a first operand matrix, and generates partial matrix index information for the operation area using a non-zero element included in the operation area.
[0027] In addition, according to another embodiment of the present invention for achieving the above object, there is provided a matrix processing device including a memory; and a processor electrically connected to the memory, wherein the processor loads elements of the first and second operand matrices using matrix index information of the first and second operand matrices and performs an operation on the loaded elements, and the matrix index information of the first operand matrix uses matrix index information, which is partial matrix index information generated using a non-zero element included in an operation area of the first operand matrix, where an operation is performed with a non-zero element of the second operand matrix.
[0028] In addition, according to another embodiment of the present invention for achieving the above-mentioned purpose, a method for generating matrix index information is provided, including: a step of identifying a non-zero element in a target matrix for a deep learning model; and a step of generating matrix index information including characteristic information on the non-zero element, wherein the characteristic information includes at least one of information indicating whether an operation with an outlier element of an operand matrix for the target matrix is performed and weight information on the deep learning model.
[0029] In addition, according to another embodiment of the present invention for achieving the above-mentioned object, a method for generating matrix index information is provided, including the steps of loading a target matrix for a deep learning model; and the step of generating matrix index information including characteristic information for elements of the target matrix, wherein the characteristic information includes at least one of information indicating whether an operation with an outlier element of an operand matrix for the target matrix is performed and lightweight information for the deep learning model.
[0030] In addition, according to another embodiment of the present invention for achieving the above-mentioned object, a method for processing a matrix using matrix index information is provided, including: loading matrix index information for a target matrix for a deep learning model; and performing weight reduction for the deep learning model using the matrix index information, wherein the matrix index information includes characteristic information for a non-zero element of the target matrix, and the characteristic information includes at least one of information indicating whether an operation with an outlier of an operand matrix for the target matrix is performed and weight reduction information for the deep learning model.
[0031] In addition, according to another embodiment of the present invention for achieving the above-mentioned object, a device for generating matrix index information is provided, comprising: a memory for storing elements of a target matrix for a deep learning model; and a processor electrically connected to the memory, wherein the processor identifies a non-zero element in the target matrix and generates matrix index information including characteristic information about the non-zero element, and the characteristic information includes at least one of information indicating whether an operation with an outlier of an operand matrix for the target matrix is performed and weight information about the deep learning model.
[0032] In addition, according to another embodiment of the present invention for achieving the above object, a matrix processing device is provided, including a memory for storing matrix index information for a target matrix for a deep learning model; and a processor electrically connected to the memory, wherein the processor performs weight reduction for the deep learning model using the matrix index information, the matrix index information includes characteristic information for a non-zero element of the target matrix, and the characteristic information includes at least one of information indicating whether an operation with an outlier of an operand matrix for the target matrix is performed and weight reduction information for the deep learning model.
[0033]
[0034] According to one embodiment of the present invention, data is loaded through memory access according to the data size of the matrix, so memory access efficiency is improved, the number of memory accesses is reduced, and power consumption can also be reduced.
[0035] In addition, according to one embodiment of the present invention, since the data size of a matrix is standardized to correspond to a multiple or divisor of the size of a unit data block loaded through memory access, data of multiple matrices can be loaded simultaneously, and thus memory access efficiency can be further improved.
[0036] In addition, according to one embodiment of the present invention, the speed and efficiency of matrix operations can be improved by selectively transmitting data required for matrix operations from lower memory to upper memory or an operator without transmitting matrix index information from lower memory to upper memory or an operator.
[0037] In addition, according to one embodiment of the present invention, elements included in the operation area of the first operand matrix and the second operand matrix on which the operation is performed can be selectively loaded and used in the operation.
[0038] In addition, according to one embodiment of the present invention, the number of memory accesses can be reduced by selectively loading elements included in the operation area of the first operand matrix and the second operand matrix on which the operation is performed according to matrix index information.
[0039] In addition, according to one embodiment of the present invention, lightweighting of an efficient deep learning model can be performed by checking matrix index information including characteristic information of a thesis element.
[0040] In addition, according to one embodiment of the present invention, by checking matrix index information including characteristic information of a nonzero element, a problem in which a nonzero element that is operated with an outlier element is not quantized or a low level of quantization is applied to a nonzero element that is operated with an outlier element, thereby lowering the accuracy of a multiplication result, can be prevented.
[0041] In addition, according to one embodiment of the present invention, by confirming matrix index information including characteristic information of a non-zero element, pruning is performed according to the pruning priority, thereby alleviating the problem of accuracy degradation that may occur due to pruning.
[0042]
[0043] FIG. 1 is a drawing for explaining a method for generating matrix index information according to one embodiment of the present invention.
[0044] FIG. 2 is a diagram showing matrix index information according to one embodiment of the present invention.
[0045] FIG. 3 is a diagram for explaining a method for generating matrix index information according to another embodiment of the present invention.
[0046] FIG. 4 is a diagram showing matrix index information according to another embodiment of the present invention.
[0047] FIG. 5 is a drawing for explaining a matrix processing device using matrix index information according to one embodiment of the present invention.
[0048] FIG. 6 is a diagram for explaining a matrix processing method using matrix index information according to one embodiment of the present invention.
[0049] FIG. 7 is a diagram showing memory storage data according to one embodiment of the present invention.
[0050] FIG. 8 is a drawing for explaining a matrix data processing device according to one embodiment of the present invention.
[0051] FIG. 9 is a drawing for explaining a matrix data transmission method according to one embodiment of the present invention.
[0052] Figure 10 is a drawing for explaining a method of selecting elements required for operation.
[0053] FIG. 11 is a drawing for explaining a matrix data transmission method according to another embodiment of the present invention.
[0054] Figure 12 is a diagram for explaining the memory address value for the nonzero element.
[0055] FIG. 13 is a drawing for explaining a matrix data processing method according to one embodiment of the present invention.
[0056] FIG. 14 is a drawing for explaining a matrix data processing method according to another embodiment of the present invention.
[0057] Figure 15 is a diagram for explaining the convolution operation of a matrix.
[0058] FIG. 16 is a diagram for explaining a method for generating matrix index information according to another embodiment of the present invention.
[0059] FIG. 17 is a diagram for explaining a method for generating partial matrix index information according to one embodiment of the present invention.
[0060] FIG. 18 is a diagram for explaining a method for generating partial matrix index information according to another embodiment of the present invention.
[0061] FIG. 19 is a diagram for explaining a matrix processing method using matrix index information according to another embodiment of the present invention.
[0062] FIG. 20 is a diagram for explaining a method for generating matrix index information according to another embodiment of the present invention.
[0063] FIG. 21 is a diagram showing matrix index information according to another embodiment of the present invention.
[0064] FIG. 22 is a drawing for explaining a matrix processing method according to another embodiment of the present invention.
[0065]
[0066] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated and described in detail in the drawings. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention. Throughout the description of each drawing, similar reference numerals have been used to designate similar components.
[0067]
[0068] The present invention relates to a method for generating matrix index information and a method for processing a matrix or transmitting matrix data using the matrix index information. One embodiment of the present invention proposes a method for generating matrix index information and a method for processing a matrix, which can improve memory access efficiency and reduce the number of memory accesses. In addition, another embodiment of the present invention proposes a method for transmitting and processing matrix data to enable efficient matrix operations to be performed using the matrix index information. In addition, another embodiment of the present invention proposes a method and a device for generating index information for an operation area of an operand matrix, and a method and a device for processing a matrix using the index information for the operation area of an operand matrix. In addition, another embodiment of the present invention proposes a method and a device for generating matrix index information including characteristic information of a matrix element, and a method and a device for processing a matrix using the matrix index information including characteristic information of a matrix element.
[0069] Hereinafter, embodiments according to the present invention will be described in detail with reference to the attached drawings.
[0070]
[0071] First embodiment
[0072] As mentioned above, quantization techniques are used for efficient matrix operations, and quantization can be applied to each matrix according to different quantization bits. Here, the quantization bits correspond to the number of bits in which the elements of the matrix are quantized. For example, a 16-bit quantization bit indicates that the elements of the matrix are quantized into 16-bit values.
[0073] And, since the data size of the matrix is determined by the product of the quantization bit and the number of non-zero elements, if the quantization bit of the matrix is set differently for each matrix, even if the number of non-zero elements is the same for matrices, the size of the matrix data loaded through memory access will be different for each matrix. However, since the matrix index information such as the currently used CSR does not include information about the quantization bit of the matrix, memory access that reflects the data size of the matrix cannot be performed, and therefore, memory access efficiency is lowered, and power consumption is bound to increase due to the increase in the number of memory accesses.
[0074] Accordingly, the present invention proposes a method of generating matrix index information including quantization information of a matrix and processing a matrix using such matrix index information in order to increase the efficiency of memory access for loading matrix data and reduce power consumed according to memory access.
[0075] In particular, one embodiment of the present invention standardizes the data size of a matrix and includes information about the matrix data size in matrix index information, so that, according to one embodiment of the present invention, the efficiency of memory access for loading matrix data can be further improved.
[0076] A method for generating matrix index information and a method for processing a matrix according to an embodiment of the present invention can be performed in a computing device including a memory and a processor, and the processor can be a general-purpose processor or a separate processor for artificial intelligence operations such as a deep learning accelerator.
[0077] Hereinafter, embodiments according to the present invention will be described in detail with reference to the attached drawings.
[0078]
[0079] FIG. 1 is a drawing for explaining a method for generating matrix index information according to an embodiment of the present invention, and FIG. 2 is a drawing showing matrix index information according to an embodiment of the present invention.
[0080] Referring to FIG. 1, a computing device according to an embodiment of the present invention checks for non-zero elements in a target matrix (S110). In one embodiment, the target matrix may be a weight matrix including weight values of an artificial neural network, and the computing device can check the number and positions of non-zero elements in the target matrix by checking for non-zero elements in the target matrix.
[0081] A computing device generates matrix index information including quantization information for a target matrix, data size information of the target matrix, and position information of a nonzero element (S120). The matrix index information may be in the form of a bit string, and may include a first bit string indicating quantization information, a second bit string indicating data size information, and a third bit string indicating position information. Here, the bit string means a sequence including at least one bit.
[0082] As described above, the quantization information is information on quantization bits indicating the number of bits in which the non-zero elements of the target matrix are quantized, and may be set in various ways depending on the embodiment. In addition, the data size of the target matrix is determined according to the number of quantization bits and non-zero elements, and may be one of the preset candidate data sizes. The candidate data size may be determined according to the size of the unit data block loaded through memory access, and in one embodiment, the data size of the target matrix may be determined as a multiple or a divisor of the size of the unit data block loaded through memory access. In addition, the unit data block size is a size of data that can efficiently load data from the memory, and may be a unit data size with high transmission efficiency, for example, determined by the number of bits of a data bus connected to the memory to which data is to be loaded or the transmission operation characteristics of the memory to which data is to be loaded.
[0083] For example, referring to FIG. 2, the quantization bits can be set to 64 bits, 32 bits, 16 bits, or 8 bits, as in the first table (210), and a bit string can be mapped for each quantization bit. The candidate data size of the matrix can be set according to the size of the matrix, as in the second table (220), and a bit string can be mapped for each candidate data size. The candidate data sizes of a 3x3 matrix can be set to 0 bits and 64 bits, the candidate data sizes of a 5x5 matrix can be set to 0 bits, 64 bits, 128 bits, and 192 bits, and the candidate data sizes of a 7x7 matrix can be set to 0 bits, 64 bits, 128 bits, 192 bits, 256 bits, 320 bits, and 384 bits. Here, 64 bits corresponds to the unit data block size, and 128, 192, 256, 320, and 384 are all multiples of 64. If the unit data block size changes, the candidate data sizes of the matrix may also change.
[0084] If the quantization bit of the target matrix (230) is 16 bits, the bit value of the first bit string (241) is determined as 10 by the first table (210). In addition, the data size of the target matrix (230) corresponds to the product of the number of non-zero elements and the quantization bits, and since the number of non-zero elements (a, b, c, d) of the 3x3 target matrix (230) is 4, the bit value of the second bit string (242) is determined as 1.
[0085] The third bit string (243) includes, as an example, a bit corresponding to each position of an element in the target matrix, and in the third bit string (243), a bit value corresponding to the position of a zero element of the target matrix and a bit value corresponding to the position of a non-zero element may be assigned different values. For example, a bit value corresponding to the position of a zero element may be 0, and a bit value corresponding to the position of a non-zero element may be 1. Accordingly, in the example of FIG. 2, the third bit string (243) may be determined as 100011100. There is no limitation on the format of the third bit string indicating position information, and a bit string used to indicate position information in CSR may be used as the third bit string.
[0086] Meanwhile, as described above, the data size of the target matrix corresponds to the product of the number of non-zero elements and the quantization bits, and the computing device can adjust the quantization bits or the number of non-zero elements of the target matrix to standardize the data size of the target matrix to a preset size.
[0087] As an example, the computing device can adjust the number of non-zero elements according to the quantization bits so that the data size of the target matrix becomes a multiple or divisor of the unit data block size. That is, when the quantization bits are set, the computing device can adjust the number of non-zero elements according to the set quantization bits. The computing device can adjust the number of non-zero elements by reducing the number of non-zero elements through pruning.
[0088] In the example of Fig. 2, if the number of non-zero elements of the target matrix (230) is 5, the computing device can change one of the 5 non-zero elements into a zero element through pruning so that the data size of the target matrix becomes 64 bits. At this time, the non-zero element to be changed into a zero element can be determined according to the pruning algorithm.
[0089] Alternatively, the computing device may adjust the number of quantization bits according to the number of non-zero elements so that the data size of the target matrix becomes a multiple or divisor of the unit data block size. That is, when the number of non-zero elements is determined by pruning or the like, the computing device may adjust the quantization bits according to the determined number of non-zero elements.
[0090] In the example of Fig. 2, if the number of non-zero elements of the target matrix (230) is reduced to two through pruning, the quantization bits can be adjusted to 32 bits. In this case, the computing device determines the first bit string (241) as 01.
[0091] Through this matrix index information, the computing device can load data of the target matrix from memory in a size that is a multiple or divisor of the unit data block size, distinguish each non-zero element in the loaded data, and also confirm the location of the non-zero element.
[0092] Meanwhile, according to an embodiment, the computing device may identify a non-zero element in the target matrix and generate matrix index information including data size information of the target matrix and position information of each non-zero element. That is, the computing device may generate matrix index information that does not include quantization information for the target matrix.
[0093] The third bit string (243) described above includes information on the number of non-zero elements along with information on the position of the non-zero elements. When this bit string is used as position information, since quantization information can be estimated from the matrix index information, quantization information can be omitted from the matrix index information. In the example of Fig. 2, the number of 1s in the third bit string (243) can be estimated to be 4 for the non-zero elements of the target matrix, and since the data size of the target matrix is 64 bits, it can be predicted that the quantization bit of the non-zero elements is 16 bits.
[0094]
[0095] FIG. 3 is a drawing for explaining a method for generating matrix index information according to another embodiment of the present invention, and FIG. 4 is a drawing showing matrix index information according to another embodiment of the present invention.
[0096] Referring to FIG. 3, a computing device according to an embodiment of the present invention identifies a non-zero element in a target matrix (S310), and generates matrix index information including quantization information for the target matrix, information on the number of non-zero elements, and information on the position of the non-zero elements (S320).
[0097] The matrix index information may include a first bit string representing quantization information, a second bit string representing count information, and a third bit string representing position information. The second bit string is a bit string in which count information is expressed in binary.
[0098] As shown in Fig. 4, when the target matrix (410) includes four non-zero elements and the quantization bits of the non-zero elements are 16 bits, the first bit string (421) can be set to 10, the second bit string (422) to 0100, and the third bit string (423) to 100011100.
[0099] The computing device can calculate the data size of the target matrix from the matrix index information, and load data for the target matrix from the memory according to the calculated data size. In the example of Fig. 4, since the quantization bit is 16 bits and the number of non-zero elements is 4, the data size of the target matrix (410) is 64 bits, and therefore the computing device can load data of 64 bits in size from the memory.
[0100] Meanwhile, according to an embodiment, the computing device can identify non-zero elements in the target matrix and generate matrix index information including quantization information for the target matrix and position information of each non-zero element. That is, the computing device can generate matrix index information that does not separately include information on the number of non-zero elements when the position information of the non-zero elements is allocated for each non-zero element position.
[0101] Since the third bit string (423) described above includes information on the number of non-zero elements along with the position information of the non-zero elements, when this bit string is used as position information, information on the number of non-zero elements can be omitted from the matrix index information. In the example of Fig. 4, it can be estimated that the number of non-zero elements of the target matrix is 4 through the number of 1s in the third bit string (423).
[0102]
[0103] FIG. 5 is a drawing for explaining a matrix processing device using matrix index information according to one embodiment of the present invention.
[0104] Referring to FIG. 5, a matrix processing device according to an embodiment of the present invention includes a bitstream generation unit (510), a data loading unit (520), and a calculation unit (530). Depending on the embodiment, a memory may be further included. The matrix processing device according to an embodiment of the present invention may be an example of the computing device described above.
[0105] The bit string generation unit (510) generates matrix index information. The bit string generation unit (510) can generate matrix index information for the first target matrix, as in the embodiment described above.
[0106] The data loading unit (520) can use matrix index information for the first target matrix to check the data size of the first target matrix, access the first memory (540), and load non-zero element data equivalent to the data size of the first target matrix from the first memory (540). The data loading unit (520) can load the non-zero element data of the first target matrix at once using burst mode or the like.
[0107] If the matrix index information includes data size information of the target matrix, the data loading unit (520) can check the data size of the first target matrix from the data size information. Alternatively, if the matrix index information includes quantization information for the target matrix and information on the number of non-zero elements, the data loading unit (520) can check the data size of the first target matrix from the quantization bits and the number information.
[0108] The data loading unit (520) can use the second table (220) in which the data size and bit value are mapped to check the data size of the first target matrix from the matrix index information.
[0109] In addition, the data loading unit (520) can load the non-zero element data of the first target matrix using the memory address values for the non-zero elements stored in the first memory (540). In one embodiment, the memory address values assigned to the non-zero element data may be in a continuous form according to a preset rule, and the memory address values for the non-zero element data of the plurality of target matrices may be assigned in a continuous pattern so as to correspond to the order of the indexes assigned to the target matrices. Accordingly, the data loading unit (520) can determine the address values of the non-zero element data of the first target matrix using the number of non-zero elements previously loaded from the memory, and can load the non-zero element data of the first target matrix from the first memory (540) using the determined memory address values.
[0110] The operation unit (530) performs an operation on the first target matrix using the loaded data. The operation unit (530) can perform an operation on the elements of another second target matrix loaded by the data loading unit (520) and the non-zero elements of the first target matrix. The elements of the second target matrix can be stored in the second memory (550), and the second target matrix can also be stored in the second memory (550) together with matrix index information in the same manner as the first target matrix.
[0111] The operation unit (530) may include a plurality of processing elements for parallel operation, and each processing element may be assigned a non-zero element of a first target matrix and a non-zero element of a second target matrix. Each processing element may perform an operation on the assigned non-zero element of the first target matrix and the non-zero element of the second target matrix.
[0112] The data loading unit (520) transmits the loaded data to the calculation unit (530), and at this time, the loaded data can be divided into each non-zero element data using matrix index information, and the divided non-zero elements can be transferred to each calculation unit. The data loading unit (520) can identify each non-zero element data by dividing the loaded data into quantization bit units using quantization information, and in order to identify the quantization bit, the first table (210) in which quantization bits and bit values are mapped can be used.
[0113] In addition, the data loading unit (520) can use the position information to check the position of the nonzero element in the first target matrix of the nonzero element, and can also check the position of the nonzero element in the second target matrix. Accordingly, the data loading unit (520) can transfer a pair of nonzero elements in an operation relationship among the nonzero elements of the first and second target matrices to each operator.
[0114]
[0115] FIG. 6 is a diagram for explaining a matrix processing method using matrix index information according to an embodiment of the present invention, and FIG. 7 is a diagram showing memory storage data according to an embodiment of the present invention.
[0116] Referring to FIG. 6, a matrix processing device according to an embodiment of the present invention loads a non-zero element of a first target matrix from a memory using matrix index information for the first target matrix (S610). The matrix index information may include quantization information for the first target matrix, data size information of the first target matrix, and position information of the non-zero element, or may include data size information and position information without quantization information.
[0117] The size of the loaded data is determined based on the data size information of the first target matrix, and the data size of the first target matrix may be a size corresponding to a multiple or divisor of the size of the unit data block loaded through memory access. In other words, the matrix processing device can load data of a size corresponding to the data size of the first target matrix from memory.
[0118] As illustrated in FIG. 7, matrix index information (710, 720) for target matrices and non-zero elements (730) can be stored in the memory. FIG. 7 illustrates an example in which index information for a 3X3 target matrix is stored in the memory. In the matrix index information (710) of the first target matrix, if the first bit column is 11, the second bit column is 1, and the third bit column is 100011100, the quantization bit is 16 bits and the data size is 64 bits by the first and second tables (210, 220), so if the unit data block size is 64 bits or more, the matrix processing device can load 64 bits of data from the memory. And the 64-bit loaded data can be divided into 16-bit units, and each of the divided 16-bit units of data can correspond to the non-zero elements (730) 0.1, 0.25, -0.5, 0.2.
[0119] If the data size of the first target matrix is smaller than the unit data block size, the matrix processing device can load the non-zero elements of the second target matrix together. For example, if the unit data block size is 128 bits, since the unit data block size is larger than the data size of the first target matrix, which is 64 bits, the matrix processing device can load all or part of the non-zero elements of the second target matrix together with the non-zero elements of the first target matrix using the matrix index information (720) for the second target matrix. In the example of Fig. 7, since the data sizes of the first and second target matrices are both 64 bits, if the unit data block size is 128 bits, all of the non-zero elements of the first and second target matrices can be loaded from the memory at once through memory access.
[0120] Returning to FIG. 6 again, in step S610, the matrix processing device can load elements of a third target matrix, which is a target of operation with the first target matrix, from memory, and at this time, using positional information about the first target matrix, only elements that are a target of multiplication with the non-zero element of the first target matrix, among the elements of the third target matrix, can be loaded from memory.
[0121] And the matrix processing device transfers the loaded data in step S610 to the operator (S620). As described above, since the loaded data can be distinguished into each non-zero element by quantization information, each distinguished non-zero element can be transmitted to each of a plurality of operators. The operator that performs matrix operations can perform operations on the non-zero elements of the first and third target matrices.
[0122] According to one embodiment of the present invention, data is loaded through memory access according to the data size of the matrix, so memory access efficiency is improved, the number of memory accesses is reduced, and power consumption can also be reduced.
[0123] In addition, according to one embodiment of the present invention, since the data size of a matrix is standardized to correspond to a multiple or divisor of the size of a unit data block loaded through memory access, data of multiple matrices can be loaded simultaneously, and thus memory access efficiency can be further improved.
[0124] Meanwhile, according to an embodiment, in step S610, the matrix processing device can load the non-zero elements of the first target matrix from the memory using matrix index information for the first target matrix, which includes quantization information for the first target matrix, information on the number of non-zero elements, and position information of the non-zero elements.
[0125] At this time, the size of the loaded data can be determined by the quantization information and number information of the first target matrix, and the loaded data can be divided into each non-zero element by the quantization information or number information.
[0126] For example, if the quantization bit of the first target matrix is 16 bits and the number of non-zero elements is 3, 48 bits of data can be loaded from memory at once through memory access. Furthermore, the 48 bits of loaded data can be divided into 16-bit units, and each of the divided 16-bit units of data can correspond to a non-zero element.
[0127] In addition, in step S610, the matrix processing device can load elements of a second target matrix, which is a target of operation with the first target matrix, from memory, and at this time, using position information about the first target matrix, only elements that are a target of multiplication with the non-zero element of the first target matrix, among the elements of the second target matrix, can be loaded from the memory.
[0128] And the matrix processing device can transmit the loaded data in step S610 to the operator. As described above, since the loaded data can be distinguished into each non-zero element by quantization information, each distinguished non-zero element can be transmitted to each of the multiple operators. The operator performing the matrix operation can perform operations on the non-zero elements of the first and second target matrices.
[0129]
[0130] Second embodiment
[0131] As described above, an operator that generally performs matrix operations loads matrix data from memory using matrix index information for the matrix, which is the operation target, i.e., the operand. Matrix operations are mostly performed through multiplication and accumulation operations between elements of matrices. However, since a sparse matrix contains more zero elements than non-zero elements and the result of multiplication of a zero element is zero, it is inefficient to load all elements of the matrix, which is the operand.
[0132] Accordingly, the present invention proposes a method for performing matrix operations by selectively loading elements required for matrix operations. Matrix data is stored in a lower memory, such as a main memory, and the stored matrix data is transferred to an operator via an upper memory, such as a cache memory or a register. In one embodiment of the present invention, among the matrix data stored in the lower memory, data required for operations are selectively transferred to an upper memory or an operator, and the matrix data is processed.
[0133] In particular, as the number of zero elements in a matrix increases, the size of matrix index information may become larger than the matrix data. According to one embodiment of the present invention, matrix data required for an operation may be selectively transmitted to an upper memory or an operator without the matrix index information being transmitted from a lower memory to an upper memory or an operator, so that the speed and efficiency of matrix operations may be improved.
[0134] Hereinafter, embodiments according to the present invention will be described in detail with reference to the attached drawings.
[0135]
[0136] FIG. 8 is a drawing for explaining a matrix data processing device according to one embodiment of the present invention. The matrix data processing device may be a computing device including a general-purpose processor or an artificial intelligence accelerator.
[0137] Referring to FIG. 8, a matrix data processing device according to an embodiment of the present invention includes a calculation unit (810) and a lower memory (820). The calculation unit (810) includes an upper memory and may include a plurality of calculation units for parallel calculation. The upper memory may be a cache memory (811), a register (812), a buffer memory, etc., and the lower memory (820) may be a main memory, a memory including a controller equipped with a calculation function, or a Processing-In-Memory (PIM).
[0138] The operation unit (810) generates index information for a matrix, and may generate the matrix index information using various matrix indexing methods such as the aforementioned CSR. Such matrix index information may include, as an example, position information for a non-zero element of the matrix, information on the number of non-zero elements, information on the size of the matrix, etc. The matrix index information may be stored in the lower memory (820) together with matrix data, i.e., elements of the matrix.
[0139] The matrix data processing device performs an operation by loading elements for a plurality of matrices stored in a lower memory (820), wherein the operation may be a multiplication operation for elements of the matrix.
[0140] As an example, when the matrix data processing device performs an operation on the first and second matrices, elements of each of the first and second matrices stored in the lower memory (820) are loaded, and the operation unit (810) can perform the operation on the first and second matrices using the loaded elements. At this time, the matrix data processing device may not load all elements of each of the first and second matrices, but may selectively load elements necessary for the operation on the first and second matrices among the elements of each of the first and second matrices. In addition, matrix index information may be used to selectively load elements necessary for the operation on the first and second matrices.
[0141] As described above, since the elements for which the result of the multiplication operation for accumulation is zero do not require the operation, the elements required for the operation may be the elements of each of the first and second matrices for which the result of the multiplication operation is nonzero.
[0142] Alternatively, as an example, when performing an operation on the first and second matrices, the matrix data processing device may load a non-zero element from among the elements of the first and second matrices stored in the lower memory by using matrix index information for each of the first and second matrices, and load a memory address value for the non-zero element, thereby performing an operation on the first and second matrices.
[0143] Among the loaded nonzero elements, a nonzero element selected according to a memory address value can be provided to the operation unit (810).
[0144] In this way, according to one embodiment of the present invention, even if the calculation unit does not request individual matrix data required using matrix index information, matrix data required for calculation is selectively transmitted from the lower memory to the upper memory or the calculation unit, so that matrix calculation delay due to transmission and reception of matrix index information does not occur, and thus matrix calculation speed and efficiency can be improved.
[0145]
[0146] FIG. 9 is a diagram for explaining a matrix data transmission method according to an embodiment of the present invention, and FIG. 10 is a diagram for explaining a method for selecting elements required for an operation. In FIGS. 9 and 10, a matrix data transmission method performed in a lower memory is explained as an embodiment.
[0147] Referring to FIG. 9, a lower memory according to an embodiment of the present invention selects elements necessary for the operation of the first and second matrices from among the elements of the first and second matrices, respectively, using matrix index information for each of the first and second matrices (S910). As described above, the elements necessary for the operation are elements of the first and second matrices, respectively, for which the multiplication operation result is nonzero, and the matrix index information may include position information for the nonzero elements of the first and second matrices.
[0148] The position information for the non-zero element can be expressed as a bit string, and when the first and second matrices are 3X3 matrices consisting of 9 elements, the bit string can be composed of 9 bit values, as illustrated in FIG. 10. The first bit string (1010) represents position information for the first matrix, and the second bit string (1020) represents position information for the second matrix. The bit string includes a bit value corresponding to each position of an element in the matrix, and 0 can be assigned to the bit value corresponding to the position of the zero element of the matrix, and 1 can be assigned to the bit value corresponding to the position of the non-zero element. Through the first bit string (1010), it can be confirmed that the number of non-zero elements among the elements of the first matrix is 4, and that the first, fourth, fifth, and ninth elements of the first matrix are non-zero elements.
[0149] In an example such as FIG. 10, when a multiplication operation is performed between elements of the first and second matrices, the fifth element (1005) and the ninth element (1009) of the first and second matrices are common non-zero elements, and since only the result of the multiplication operation of the fifth element (1005) and the ninth element (1009) of the first and second matrices is non-zero, the lower memory can select the fifth element (1005) and the ninth element (1009) of the first and second matrices as elements required for the operation. That is, among the non-zero elements of the first and second matrices, elements whose element values at the same corresponding positions are all non-zero can be selected as elements required for the operation.
[0150] Returning to FIG. 9, the lower memory transmits the element selected in step S910 to the upper memory or the operator (S920). At this time, the lower memory can create a third matrix including the selected element and transmit the elements of the third matrix to the upper memory or the operator. In other words, the lower memory can create a new matrix composed of the selected element and transmit the selected element to the upper memory or the operator.
[0151] Meanwhile, according to an embodiment, the lower memory can use the matrix index information to determine the number of times that operations are required for the elements of the first and second matrices, and in step S920, the data selected in step S910 can be continuously transmitted to the upper memory or the operator through a burst mode or the like, depending on the number of times that operations are required. The number of times that operations are required can be determined according to the number of elements required for the operations, and in the example described above, the lower memory can continuously transmit the fifth element and the ninth element of the first and second matrices to the upper memory or the operator.
[0152]
[0153] FIG. 11 is a diagram for explaining a matrix data transmission method according to another embodiment of the present invention, and FIG. 12 is a diagram for explaining a memory address value for a nonzero element. In FIGS. 11 and 12, a matrix data transmission method performed in a lower memory is explained as an example.
[0154] Referring to FIG. 11, a lower memory according to an embodiment of the present invention uses matrix index information for each of the first and second matrices to identify a non-zero element among the elements of each of the first and second matrices (S1110).
[0155] The lower memory can store a non-zero element among the elements of each of a plurality of matrices including the first and second matrices, and the lower memory can identify the non-zero element using matrix index information including at least one of the number information and position information for the non-zero element of the plurality of matrices.
[0156] And, according to one embodiment of the present invention, the lower memory transmits the memory address value for the nonzero element generated through matrix index information and the nonzero element to the upper memory or the operator (S1120).
[0157] Here, the memory address value may be an address value for the entire nonzero element transmitted to the upper memory or the operator. Or, depending on the embodiment, it may be an address value for a nonzero element required for the operation of the first and second matrices among the nonzero elements of the first and second matrices, respectively. And, as described above, the nonzero element required for the operation corresponds to the nonzero element of the first and second matrices, respectively, whose multiplication operation result is nonzero.
[0158] Among the nonzero elements transmitted to the upper memory, a nonzero element selected according to the memory address value can be transmitted to the operator, and the selected nonzero elements can be transmitted to the operator continuously. As in the example described above, when the memory address value is an address value for a nonzero element required for the operation of the first and second matrices, among the nonzero elements transmitted to the upper memory, a nonzero element required for the operation of the first and second matrices can be selectively transmitted to the operator.
[0159] The lower memory generates a memory address value for a non-zero element, and for example, as shown in FIG. 12, first and second bit strings (1010, 1020) indicating the location information of the first and second matrices are given, and when the non-zero elements (1210, 1220) of the first and second matrices are a, b, c, d and A, B, C, respectively, the memory address value can be assigned to each non-zero element in a pattern that increases by 1 from the first element to the last element. When the memory address value of the first non-zero element a of the first matrix is 1, the memory address value of the second non-zero element b can be 2, the memory address value of the third non-zero element c can be 3, and the memory address value of the fourth non-zero element d can be 4. And if the memory address value of the first nonzero element A of the second matrix is 100, the memory address value of the second nonzero element B can be 101, and the memory address value of the third nonzero element C can be 102.
[0160] In this example, among the non-zero elements a, b, c, d and A, B, C of the first and second matrices transferred to the upper memory, the fifth elements (1005) of the first and second matrices, b and B, and the ninth elements (1009) of the first and second matrices, d and C, can be selected and transferred to the operator according to the memory address value.
[0161]
[0162] FIG. 13 is a drawing for explaining a matrix data processing method according to one embodiment of the present invention.
[0163] Referring to FIG. 13, a matrix data processing device according to an embodiment of the present invention selectively loads elements necessary for operations on the first and second matrices from among the elements of the first and second matrices stored in the lower memory using matrix index information for each of the first and second matrices (S1310). The elements necessary for operations can be loaded into the upper memory or the operator.
[0164] And, using the loaded elements, operations are performed on the first and second matrices (S1320).
[0165]
[0166] FIG. 14 is a diagram for explaining a matrix data processing method according to another embodiment of the present invention.
[0167] Referring to FIG. 14, a matrix data processing device according to an embodiment of the present invention uses matrix index information for each of the first and second matrices to load a non-zero element from among the elements of each of the first and second matrices stored in a lower memory, and loads a memory address value for the non-zero element (S1410). The non-zero element and the memory address value can be loaded into an upper memory.
[0168] And, among the loaded nonzero elements, the nonzero elements selected according to the memory address values are used to perform operations on the first and second matrices (S1420). Here, as described above, the memory address values are address values for nonzero elements required for the operations on the first and second matrices among the loaded nonzero elements, and the nonzero elements required for the operations can correspond to the nonzero elements of each of the first and second matrices, whose multiplication operation results in nonzero.
[0169]
[0170] Third embodiment
[0171] Figure 15 is a diagram for explaining the convolution operation of a matrix.
[0172] As illustrated in FIG. 15, when a convolution operation is performed between a first operand matrix (1510) and a second operand matrix (1520), the second operand matrix (1520) shifts according to a stride value, and an operation area in the first operand matrix (1510) for each element of the second operand matrix (1520) is determined. Here, the operation area refers to an area in the first operand matrix (1510) that includes elements in which an operation is performed with each element of the second operand matrix (1520). The operation area may vary for each element of the second operand matrix (1520). In FIG. 15, the grids of the first and second operand matrices (1510, 1520) represent elements of the first and second operand matrices (1510, 1520).
[0173] When a convolution operation is performed between a first operand matrix (1510) of 7x7 size as in FIG. 15 and a second operand matrix (1520) of 3x3 size, for example, the operation area in the first operand matrix (1510) for the first element (1521) of the second operand matrix (1520) corresponds to the hatched area in the first operand matrix (1510). That is, the operation is performed on some elements included in the operation area, not all elements included in the first operand matrix (1510).
[0174] In this case, for the operation on the first element (1521) of the second operand matrix (1520), it is unnecessary to load all elements included in the first operand matrix (1510), and it is efficient to selectively load some elements included in the operation area of the first operand matrix (1510) and use them for the operation.
[0175] Accordingly, the present invention proposes a matrix processing method that generates matrix index information for elements included in an operation area of a first operand matrix (1510), and uses this matrix index information to selectively load elements included in an operation area of the first operand matrix (1510) so that they can be used for operations.
[0176] In particular, one embodiment of the present invention proposes a method for generating matrix index information for an operation area by using non-zero elements included in the operation area, considering that zero elements among the elements included in the operation area are also unnecessary for the operation. The matrix index information for the operation area can be generated according to various matrix index information generation algorithms, such as the existing CSR.
[0177] A method for generating matrix index information and a method for processing a matrix according to one embodiment of the present invention can be performed in a computing device including a memory and a processor electrically connected to the memory, and the processor can perform a series of processes for generating matrix index information and processing a matrix.
[0178]
[0179] FIG. 16 is a drawing for explaining a method for generating matrix index information according to another embodiment of the present invention.
[0180] Referring to FIG. 16, a computing device according to an embodiment of the present invention determines an operation area in which an operation is performed with a non-zero element of a second operand matrix in a first operand matrix (S1610). As described above, the operation area refers to an area in the first operand matrix that includes elements in which an operation is performed with each element of the second operand matrix, and may be determined according to an operation method such as a convolution operation or a transformer operation, the sizes of the first and second operand matrices, etc.
[0181] And the computing device generates partial matrix index information for the operation area using the nonzero elements included in the operation area (S1620). The partial matrix index information includes information on the number of nonzero elements included in the operation area and information on the location of the nonzero elements included in the operation area in the operation area.
[0182] As an example, the computing device can generate new matrix index information from a non-zero element included in an operation area, or can generate partial matrix index information by updating the entire matrix index information generated from the non-zero element included in the first operand matrix according to the non-zero element included in the operation area. That is, the computing device can generate partial matrix index information from the entire matrix index information, which is matrix index information for the first operand matrix that has already been generated. The computing device can generate partial matrix index information by updating the entire matrix index information that has already been generated according to the operation area of the first operand matrix that varies depending on the second operand matrix that is operated with the first operand matrix, and the operation method, etc.
[0183] The full matrix index information includes information on the number of non-zero elements included in the first operand matrix and information on the position of the non-zero elements included in the first operand matrix in the first operand matrix, and can be generated according to various index information generation methods. For example, the full matrix index information can be generated according to the CSR method, or can be generated in the form of a bit string including a first bit string indicating information on the number of non-zero elements and a second bit string indicating information on the position of the non-zero elements. In addition, the partial matrix index information is expressed in the same type as the full matrix index information.
[0184] Since the number of non-zero elements included in the first operand matrix and the number of non-zero elements included in the operation area are different, the number information included in the entire matrix index information is updated according to the number of non-zero elements included in the operation area. In addition, since the position information of the non-zero elements included in the operation area is determined based on the operation area, the position information included in the entire matrix index information is also updated according to the position of the non-zero elements in the operation area.
[0185] Meanwhile, in step S1620, the computing device can generate a memory address value for a nonzero element included in the operation area using the partial matrix index information, and as an example, can generate a continuous memory address value. The computing device can identify the location of the nonzero element included in the operation area using the partial matrix index information, and generate the memory address value of the nonzero element included in the operation area as a continuous value so that the nonzero element included in the operation area can be loaded continuously according to a burst mode or the like.
[0186] According to one embodiment of the present invention, the number of memory accesses can be reduced by selectively loading elements included in the operation area of the first operand matrix and the second operand matrix on which the operation is performed according to matrix index information.
[0187]
[0188] FIG. 17 is a diagram for explaining a method for generating partial matrix index information according to one embodiment of the present invention.
[0189] Referring to FIG. 17, the entire matrix index information for the first operand matrix (1710) may be, as an example, in the form of a bit string including a first bit string (1721) indicating the number information of non-zero elements and a second bit string (1722) indicating the position information of the non-zero elements. In addition, the first and second bit strings (1721, 1722) include the number information and the position information of the non-zero elements in each row of the first operand matrix (1710). That is, the first and second bit strings (1721, 1722) may be generated in units of rows of the first operand matrix (1710). According to an example, the entire matrix index information may include the first and second bit strings including the number information and the position information of the non-zero elements in each column of the first operand matrix (1710).
[0190] Among the grids of the first operand matrix (1710), grids with alphabets displayed correspond to non-zero elements, and grids without alphabets displayed correspond to zero elements. Since the number of non-zero elements (A, B, C) in the first row of the first operand matrix (1710) is 3, the bit value of the first bit string (1721) becomes '11', and the bit value of the second bit string (1722) becomes '0011001'. The second bit string (1722) may be composed of bits corresponding to each element of the first operand matrix (1710), and since the first operand matrix (1710) is a 7X7 matrix, 7 bits are allocated to each of the second bit strings (1722). In addition, depending on whether the element is a zero element or a non-zero element, the bit values corresponding to the elements may be allocated differently.
[0191] And, when the operation area is determined as a hatched part in the first operand matrix (1710) by an operation with the second operand, the computing device can remove the number information and position information of non-zero elements in rows or columns not included in the operation area from each of the first and second bit strings (1721, 1722) of the entire matrix index information (1720), and update the first bit string (1721) according to the number of non-zero elements included in the operation area, thereby generating partial matrix index information (1730).
[0192] In the second bit string (1722) of the full matrix index information (1720), the bits corresponding to the operation area are the bits included within the dotted line (1723), and the bits included within the dotted line (1723) correspond to the second bit string (1732) of the partial matrix index information (1730). In addition, since the size of the operation area is 5X5, the partial matrix index information (1730) includes the first and second bit strings (1731, 1732) corresponding to five rows. The first bit string (1731) of the partial matrix index information (1730) includes information on the number of non-zero elements included in each row of the operation area.
[0193] In the first and second bit strings (1721, 1722) of the full matrix index information (1720), bits corresponding to the first and last rows and the first and last columns of the first operand matrix (1710) are deleted, and the first bit string (1721) of the full matrix index information (1720) is updated according to the number of non-zero elements included in the operation area, thereby generating partial matrix index information (1730).
[0194] The computing device can use the partial matrix index information (1730) to generate memory address values for nonzero elements included in the operation area by row or column units of the operation area, and as an example, can generate continuous memory address values. For example, the memory address values for nonzero elements (I, J, K, L) included in the third row of the operation area are assigned as continuous values. That is, the computing device generates continuous memory address values for nonzero elements in the order of upper bits to lower bits in the second bit column (1732) of the partial matrix index information (1730).
[0195] At this time, the starting memory address value in each row or column of the operation area corresponds to the memory address value for the first non-zero element in each row or column of the operation area, and corresponds to the memory address value for the first bit in each of the second bit columns (1732) of the partial matrix index information (1730). Here, the first bit corresponds to the most significant bit indicating the non-zero element in each of the second bit columns (1732) of the partial matrix index information (1730).
[0196] And, in each of the second bit columns (1722) of the entire matrix index information, if there is a second bit that is a higher bit than the first bit, that is, if there is a second bit that is a higher bit than the first bit in the second bit column (1722) before the update, the computing device can increase the starting memory address value according to the number of second bits. For example, since there is one second bit (1742) that is a higher bit than the first bit (1741) of the third second bit column of the partial matrix index information (1730) in the fourth second bit column of the entire matrix index information (1722), the starting memory address value of the non-zero element corresponding to the first bit (1741) of the third second bit column of the partial matrix index information can be increased by 1. Accordingly, the starting memory address value for the non-zero element of the third row of the operation area can be 1 greater than the starting memory address values for the non-zero elements of the remaining rows of the operation area.
[0197]
[0198] FIG. 18 is a diagram for explaining a method for generating partial matrix index information according to another embodiment of the present invention.
[0199] FIG. 18 illustrates an embodiment in which partial matrix index information is generated from full matrix index information for a first operand matrix (1710) generated according to CSR.
[0200] According to the CSR algorithm, matrix index information including row index information and column index information is generated. The row index information includes accumulated information on the number of nonzero elements in each row, and the column index information includes information on the location of nonzero elements in each row.
[0201] Since the number of non-zero elements in the first row of the first operand matrix (1710) is 3, the number of non-zero elements in the second row is 1, and the number of non-zero elements in the third row is 3, the index values included in the row index information (1811) of the entire matrix index information show an increasing pattern of 3, 4, and 7. In addition, the non-zero elements in the first row are located in the third, fourth, and seventh columns, and when the index values correspond to the first to seventh columns in the order of 0 to 6, the index values included in the column index information (1812) of the entire matrix index information show a pattern such as 2, 3, and 6.
[0202] Afterwards, when the operation area is determined as the hatched part in the first operand matrix (1710) by operation with the second operand, the computing device can generate updated row index information (1821) and column index information (1822) according to the nonzero element included in the operation area.
[0203]
[0204] FIG. 19 is a diagram for explaining a matrix processing method using matrix index information according to another embodiment of the present invention.
[0205] Referring to FIG. 19, a computing device according to an embodiment of the present invention loads elements of the first and second operand matrices using matrix index information of the first and second operand matrices (S1910). Here, the matrix index information of the first operand matrix corresponds to partial matrix index information generated using non-zero elements included in the operation area of the first operand matrix, and the operation area of the first operand matrix corresponds to an area including elements on which operations are performed with non-zero elements of the second operand matrix in the first operand matrix. The computing device can load elements of the first and second operand matrices stored in the memory from the memory.
[0206] As described above, the partial matrix index information may correspond to the entire matrix index information generated from the nonzero elements included in the first operand matrix, updated according to the nonzero elements included in the operation area. In addition, the entire matrix index information may include information on the number of nonzero elements included in the first operand matrix and information on the position of the nonzero elements included in the first operand matrix in the first operand matrix, and the partial matrix index information may include information on the number of nonzero elements included in the operation area and information on the position of the nonzero elements included in the operation area in the operation area.
[0207] And the computing device performs an operation on the loaded element in step S1910 (S1920). As an example, the computing device may perform a convolution operation or a transformer operation on the first and second operand matrices using the loaded element.
[0208]
[0209] Fourth embodiment
[0210] As mentioned above, recently, methods for reducing the weight of deep learning models, such as mixed-precision quantization, have been used, and in order for this type of quantization to be performed efficiently, matrix index information needs to include characteristic information of matrix elements, such as quantization information. Accordingly, the present invention proposes a method for generating matrix index information including characteristic information of matrix elements, and a matrix processing method using such matrix index information.
[0211] One embodiment of the present invention can generate matrix index information including at least one characteristic for a matrix element. As an example, the matrix index information can include characteristic information such as information on quantization bits of the matrix element. Here, the quantization bits correspond to the number of bits by which an element of the matrix is to be quantized. A computing device can quantize an element of the matrix using the matrix index information. For example, a matrix element having 16 bits of quantization bits can be quantized into a 16-bit value, and a matrix element having 8 bits of quantization bits can be quantized into an 8-bit value.
[0212] A method for generating matrix index information and a method for processing a matrix according to an embodiment of the present invention can be performed in a computing device including a memory and a processor electrically connected to the memory, and the processor can be a general-purpose processor or a separate processor for artificial intelligence operations, such as a deep learning accelerator. The memory stores elements of a matrix or matrix index information, and the processor can load elements or matrix index information stored in the memory to generate matrix index information or perform a series of processes for processing a matrix.
[0213] In this way, a computing device according to an embodiment of the present invention can load a target matrix for a deep learning model and generate matrix index information containing characteristic information about elements of the target matrix. Furthermore, this characteristic information can include information indicating whether an operation with an outlier element of an operand matrix for the target matrix is performed, and lightweight information for the deep learning model.
[0214] Hereinafter, embodiments according to the present invention will be described in detail with reference to the attached drawings.
[0215]
[0216] FIG. 20 is a drawing for explaining a method for generating matrix index information according to another embodiment of the present invention, and FIG. 21 is a drawing showing matrix index information according to another embodiment of the present invention.
[0217] Referring to FIG. 20, a computing device according to an embodiment of the present invention checks for non-zero elements in a target matrix for a deep learning model (S2010). In one embodiment, the target matrix may be a weight matrix including weight values of an artificial neural network. Since the characteristic information described below is significant for non-zero elements, the non-zero elements may be checked in step S2010. However, as described above, the matrix index information may include characteristic information for all elements of the target matrix, and step S2010 may be selectively performed depending on the embodiment.
[0218] The computing device generates matrix index information including characteristic information on a non-zero element (S2020). Here, the characteristic information includes at least one of information indicating whether an operation with an outlier element of an operand matrix for a target matrix is performed, as described above, and lightweight information for a deep learning model. Here, the lightweight information represents information that can be utilized when performing lightweighting of a deep learning model.
[0219] When an operation is performed between a target matrix and an operand matrix, if a non-zero element multiplied with an outlier element of the operand matrix is quantized or quantized to a low number of bits, the accuracy of the multiplication result may be very low. Therefore, the non-zero element multiplied with an outlier element of the operand matrix needs to be either not quantized or, if quantized, quantized to a high number of bits. If information indicating whether to operate with an outlier element of the operand matrix for the target matrix is included in the matrix index information, the problem of low accuracy of the multiplication result can be prevented because the non-zero element that is operated with an outlier element may not be quantized or may be quantized to a high number of bits by checking the matrix index information. In this respect, the information indicating whether to operate with an outlier element can also be said to be information indicating whether quantization is possible.
[0220] In step S2020, the computing device can identify the location of an outlier element in the operand matrix, and, using the location of the outlier element, generate information indicating whether an operation with the outlier element is performed. That is, by identifying the location of the outlier element, the non-zero element that is to be operated with the outlier element can be identified, and information indicating whether an operation with the outlier element is performed can be generated.
[0221] The lightweight information for the deep learning model may include at least one of quantization information and pruning information for the nonzero element. Here, the quantization information represents information for quantization used when the nonzero element is quantized, and may include at least one of whether quantization is applied to the nonzero element, the quantization priority of the nonzero element, and the quantization bit of the nonzero element. That is, the information representing whether quantization is applied to the nonzero element is information representing whether the nonzero element is an element to be quantized, and the quantization priority of the nonzero element represents the priority when the nonzero element is quantized. In addition, the quantization bit of the nonzero element represents the quantization bit when the nonzero element is quantized.
[0222] And the pruning information is information for pruning that indicates whether pruning is applied to the nonzero element, and is information that indicates the nonzero element to be pruned.
[0223] The characteristic information can be set by the user or determined by preset rules or mixed-precision quantization algorithms. For example, as described above, the characteristic information can be generated by identifying the location of an outlier element or can be set according to the accuracy of the operation result, the importance of the element, etc. For non-zero elements with low importance, a high quantization priority can be assigned or the quantization bits can be set low. Alternatively, when the user increases the weight reduction level, the quantization bits of non-zero elements can be set low or the number of non-zero elements to be pruned can increase.
[0224] The matrix index information may be expressed as a bit string including bit values, as an example, and the bit values may be assigned to each element of the target matrix. For example, a non-zero element assigned with bit value 0 may correspond to an element that is not quantized, not pruned, or not operated on with an outlier element, and a non-zero element assigned with bit value 1 may correspond to an element that is quantized, pruned, or operated on with an outlier element. Alternatively, a non-zero element assigned with bit value 01 may correspond to an element that is quantized to 4 bits, a non-zero element assigned with bit value 10 may correspond to an element that is quantized to 8 bits, and a non-zero element assigned with bit value 11 may correspond to an element that is quantized to 16 bits.
[0225] In addition, the matrix index information may further include at least one of position information and number information of the non-zero element and size information of the target matrix, and the above-described bit value may be expressed in a form in which at least one of the position information, number information, and size information and characteristic information are integrated. The characteristic information may be provided in a form integrated with various matrix index information expression methods such as CSR.
[0226] As an example, as illustrated in FIG. 21, matrix index information for a 3X3 sized target matrix (2100) including three non-zero elements may include first and second bit strings (2110, 2120). Depending on the embodiment, one of the first and second bit strings (2110, 2120) may be selectively used as matrix index information.
[0227] The first bit string (2110) contains information on the number of non-zero elements of the target matrix (2100). Since the number of non-zero elements (a, b, c) is 3, the bit value of the first bit string (2110) becomes '0011', which corresponds to 3.
[0228] The second bit string (2120) includes characteristic information and position information of the non-zero elements of the target matrix. The second bit string (2120) includes bit values assigned to each element of the target matrix, and the bit values indicate characteristic information and whether the element is non-zero. In addition, since the bit values are assigned to each element of the target matrix, the bit values indicate the positions of the non-zero elements. The bit values for zero elements are assigned as 0, and the bit values for non-zero elements can be assigned as non-zero bit values.
[0229] 2 bits can be allocated to each element of the target matrix, and since the number of elements of the target matrix (2100) is 9, the second bit string (2120) can be composed of a total of 18 bits. The bit value allocated to the non-zero element a is 01, the bit value allocated to the non-zero element b is 10, and the bit value allocated to the non-zero element c is 11, so according to the example described above, the non-zero elements a, b, and c correspond to elements that are quantized into 4 bits, 8 bits, and 16 bits, respectively.
[0230] According to one embodiment of the present invention, not only can an efficient deep learning model be made lightweight by checking matrix index information including characteristic information of a nonzero element, but also a problem in which the accuracy of a multiplication result is lowered due to a nonzero element that is operated with an outlier element not being quantized or being quantized with a high number of bits can be prevented.
[0231]
[0232] FIG. 22 is a drawing for explaining a matrix processing method according to another embodiment of the present invention.
[0233] Referring to FIG. 22, a computing device according to an embodiment of the present invention loads matrix index information for a target matrix for a deep learning model (S2210). The computing device can load the matrix index information from a memory storing the matrix index information.
[0234] And the computing device performs weight reduction on the deep learning model using the loaded matrix index information (S2220). The computing device can perform weight reduction on the deep learning model using the characteristic information on the non-zero elements of the target matrix included in the matrix index information, and can perform weight reduction by quantizing or pruning the elements of the target matrix.
[0235] For example, the computing device may not perform quantization on non-zero elements that operate with outliers of the operand matrix for the target matrix, or may perform quantization with a high bit count. Alternatively, the computing device may perform quantization on non-zero elements according to the quantization bits assigned to each non-zero element. Alternatively, the computing device may use pruning information assigned to each non-zero element to change non-zero elements to zero elements.
[0236]
[0237] The technical contents described above may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination. The program commands recorded on the medium may be those specially designed and configured for the embodiments, or may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. The hardware devices may be configured to operate as one or more software modules to perform the operations of the embodiments, and vice versa.
[0238]
[0239] Although the present invention has been described with reference to specific details such as specific components and limited embodiments and drawings, these have been provided only to help a more general understanding of the present invention, and the present invention is not limited to the above embodiments, and those with ordinary skill in the art to which the present invention pertains can make various modifications and variations based on these descriptions. Therefore, the spirit of the present invention should not be limited to the described embodiments, and all things that are equivalent or equivalent to the following claims as well as the claims are considered to fall within the scope of the spirit of the present invention.
Claims
1. A step of checking non-zero elements in the target matrix; and A step of generating matrix index information including quantization information for the target matrix, data size information of the target matrix, and position information of the nonzero element. A method for generating matrix index information including .
2. In paragraph 1, The above matrix index information is A first bit string representing the above quantization information; A second bit string representing the above data size information; and Contains a third bit string representing the above location information, The above quantization information is The above topic is information about the quantization bits of the element, The data size of the above target matrix is One of the preset candidate sizes How to generate matrix index information.
3. In paragraph 2, The data size of the above target matrix is A multiple or divisor of the size of the unit data block loaded through memory access. How to generate matrix index information.
4. In paragraph 3, The step to check the element with the above topic is The number of elements in the above-mentioned nonzero is adjusted according to the quantization bits so that the data size of the above-mentioned target matrix becomes a multiple or divisor of the size of the above-mentioned unit data block. How to generate matrix index information.
5. In paragraph 3, The step to check the element with the above topic is The quantization bits are adjusted according to the number of elements in the thesis so that the data size of the target matrix is a multiple or divisor of the unit data block size. How to generate matrix index information.
6. In the target matrix, a step of checking non-zero elements; and A step of generating matrix index information including data size information of the target matrix and position information of each of the above-mentioned zero elements, The bit string representing the above location information is Contains a bit corresponding to each position of an element in the above target matrix, In the above bit string, the bit value corresponding to the position of the zero element of the target matrix and the bit value corresponding to the position of the non-zero element are different from each other. How to generate matrix index information.
7. In the target matrix, a step of checking non-zero elements; and A step of generating matrix index information including quantization information for the target matrix, information on the number of nonzero elements, and information on the position of the nonzero elements. A method for generating matrix index information including .
8. In paragraph 7, The above matrix index information is A first bit string representing the above quantization information; A second bit string representing the above quantity information; and Contains a third bit string representing the above location information, The above quantization information is Information about the quantization bits of the element in the above topic How to generate matrix index information.
9. In the target matrix, a step of checking non-zero elements; and A step of generating matrix index information including quantization information for the target matrix and position information for each of the nonzero elements, The bit string representing the above location information is Contains a bit corresponding to each position of an element in the above target matrix, In the above bit string, the bit value corresponding to the position of the zero element of the target matrix and the bit value corresponding to the position of the non-zero element are different from each other. How to generate matrix index information.
10. A step of loading the nonzero element of the first target matrix from memory using matrix index information for the first target matrix; and A step of transmitting the loaded data to a calculator is included. The above matrix index information is Including quantization information for the first target matrix, data size information of the first target matrix, and position information of the nonzero element. A matrix processing method that uses matrix index information.
11. In paragraph 10, The size of the loaded data above is It is determined according to the data size information of the first target matrix above, The data size of the first target matrix above is A size that corresponds to a multiple or divisor of the unit data block size loaded through memory access. A matrix processing method that uses matrix index information.
12. In paragraph 11, The step of loading the element from memory as above topic is If the data size of the first target matrix is smaller than the unit data block size, the nonzero elements of the second target matrix are loaded together. A matrix processing method that uses matrix index information.
13. In paragraph 11, The above loaded data is By the above quantization information, each of the above-mentioned elements is separated into the above-mentioned elements and transmitted to the operator. The step of loading the element from memory as above topic is Using the above location information, among the elements of the third target matrix, only the elements that are the target of multiplication with the nonzero element of the first target matrix are loaded from the memory. A matrix processing method that uses matrix index information.
14. A step of loading the nonzero element of the first target matrix from memory using matrix index information for the first target matrix; and A step of transmitting the loaded data to a calculator is included. The above matrix index information is Including quantization information for the first target matrix, information on the number of nonzero elements, and information on the position of the nonzero elements. A matrix processing method that uses matrix index information.
15. In paragraph 14, The size of the loaded data above is Determined by the above quantization information and the above number information, A matrix processing method that uses matrix index information.
16. In paragraph 15, The above loaded data is By the above quantization information or the above number information, each of the above-mentioned elements is separated into the above-mentioned elements and transmitted to the operator. The step of loading the element from memory as above topic is Using the above location information, among the elements of the second target matrix, only the elements that are the target of multiplication with the nonzero element of the first target matrix are loaded from the memory. A matrix processing method that uses matrix index information.
Citation Information
Patent Citations
Hierarchical sparse coding method of pruned deep neural network with extremely high compression ratio
CN112418424A
Image sensor
KR1020210046102A
Self-status reporting method for programmable logic controller
KR1020250058209A
Sparse Matrix Storage in a Database
US20150242484A1
Apparatus and Method of Using Dual Indexing in Input Neurons and Corresponding Weights of Sparse Neural Network
US20180330235A1