Decompression device and control method thereof
By designing a decompression device including memory, decoder and processor, using multiple logic circuits and representative value matrix and other technical means, the problems of accuracy and decompression speed in the compression process in the prior art are solved, and efficient parallel processing and high-accuracy compression and decompression effects are achieved.
Patent Information
- Application Number
- CN202010435215.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-06
- Filing Date
- 2020-05-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-05-21
AI Technical Summary
The prior art is difficult to maintain accuracy when compressing artificial intelligence models, and at the same time, the operation speed is slow during the decompression process and cannot be effectively processed in parallel.
A decompression device is designed, including a memory, a decoder and a processor, decompress the compressed data through multiple logic circuits, and obtain data in the form that a neural network can process in the processor. The device also contains representative value matrix, prune index matrix, and patch information to update and improve decompressed data.
Through parallel processing and the collaborative work of multiple logic circuits, the speed and efficiency of understanding the compression process is improved, ensuring that high accuracy is maintained in the artificial intelligence model while increasing the compression rate.
Smart Images

Figure CN111985632B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims and is based on the priority of Korean Patent Application No. 10-2019-0060991 filed in the Korean Intellectual Property Office on May 24, 2019, Korean Patent Application No. 10-2019-0117081 filed in the Korean Intellectual Property Office on September 23, 2019, and Korean Patent Application No. 10-2019-0140720 filed in the Korean Intellectual Property Office on November 6, 2019, the entire disclosure of each of these Korean patent applications is incorporated herein by reference. Technical Field
[0003] The present disclosure relates to a decompression device and a control method thereof for decompressing a compressed artificial intelligence model in an artificial intelligence (AI) system, which uses a machine learning algorithm such as deep learning and its application to simulate the cognitive and judgment functions of the human brain. Background Art
[0004] In recent years, pruning and quantization have been used to increase compression rates while minimizing performance degradation of deep learning models. For example, a weight matrix in which weights equal to a certain value or less are pruned to zero can be divided into a first data set representing non-zero values, a second data accumulating the number of non-zero weights per row, and a third data storing the column index corresponding to each non-zero value. Thereafter, the first to third data can be quantized. On the other hand, the weight matrix can represent the weight parameters of the deep learning model in matrix form.
[0005] However, in order to restore the original weight matrix from the quantized data, it is necessary to release the quantization and restore the original weight matrix from the first to third data. That is, before restoring the original weight matrix, it is impossible to divide the quantized data into multiple groups and process each group in parallel.
[0006] Therefore, research is actively being conducted to maintain accuracy while increasing the compression rate during compression, and to ensure operation speed through parallel processing during decompression.
[0007] The above information is presented as background information only to assist with an understanding of the present disclosure. No determination has been made, and no assertion is made, as to whether any of the above might be applicable as prior art with respect to the present disclosure. Summary of the invention
[0008] Aspects of the present disclosure are intended to address at least the above-mentioned problems and / or disadvantages and to provide at least the advantages described below. Therefore, one aspect of the present disclosure provides a decompression device and a method for the decompression device.
[0009] Additional aspects will be set forth in part in the description which follows and, in part, will be obvious from the description, or may be learned by practice of the presented embodiments.
[0010] According to one aspect of the present disclosure, a decompression device is provided. The decompression device includes a memory, a decoder, and a processor, wherein the memory is configured to store compressed data to be decompressed and used in neural network processing of an artificial intelligence model, the decoder is configured to include a plurality of logic circuits related to a compression method of compressed data, decompress the compressed data through the plurality of logic circuits based on the input of the compressed data, and output the decompressed data, and the processor is configured to obtain data in a form processable by the neural network from the data output by the decoder.
[0011] The memory is further configured to store a representative value matrix corresponding to the compressed data, wherein the processor is further configured to: obtain data in a form processable by a neural network based on the decompressed data and the representative value matrix, and perform neural network processing using the data in the form processable by the neural network, and wherein the decompressed data and the representative value matrix include a matrix obtained by quantizing an original matrix included in the artificial intelligence model.
[0012] The memory is further configured to store a pruning index matrix corresponding to the compressed data, wherein the processor is further configured to update the decompressed data based on the pruning index matrix, wherein the pruning index matrix includes a matrix obtained in a pruning process of the original matrix, and wherein the pruning index matrix is used in a process of obtaining the compressed data.
[0013] The memory is configured to further store patch information corresponding to the compressed data, wherein the processor is further configured to change the binary data value of at least one element of the multiple elements included in the decompressed data based on the patch information, and wherein the patch information includes error information generated in the process of obtaining the compressed data.
[0014] The memory is further configured to: store a first pruning index matrix corresponding to the compressed data and store a second pruning index matrix corresponding to the compressed data, wherein the processor is further configured to: obtain a pruning index matrix based on the first pruning index matrix and the second pruning index matrix, and update the decompressed data based on the pruning index matrix, wherein the pruning index matrix includes a matrix obtained in a pruning process of an original matrix, wherein the pruning index matrix is used in a process of obtaining the compressed data, and wherein the first pruning index matrix and the second pruning index matrix are respectively obtained based on each of a first sub-matrix and a second sub-matrix obtained by factorizing the original matrix.
[0015] The decompressed data includes a matrix obtained by interleaving an original matrix and then quantizing the interleaved matrix, and wherein the processor is further configured to: deinterleave the data in a neural network processable form according to a manner corresponding to the interleaving, and perform neural network processing using the deinterleaved data.
[0016] The processor includes a plurality of processing elements arranged in a matrix, and wherein the processor is further configured to perform neural network processing using the plurality of processing elements.
[0017] The decompressed data includes a matrix obtained by dividing an original matrix into a plurality of matrices having the same number of columns and rows and quantizing one of the divided plurality of matrices.
[0018] The memory is further configured to store other compressed data that is decompressed and used in the neural network processing of the artificial intelligence model, wherein the decompression device also includes another decoder, which is configured to: include multiple other logic circuits related to the compression method of other compressed data, decompress the other compressed data through multiple other logic circuits based on the input of other compressed data, and output the decompressed other data, and wherein the processor is further configured to: obtain other data in a form that can be processed by the neural network from the decompressed other data output by the other decoder, and obtain a matrix in which each element includes multiple binary data by coupling the neural network processable data and the other data in a form that can be processed by the neural network.
[0019] The decompression device is implemented as a chip.
[0020] According to another aspect of the present disclosure, a control method for a decompression device is provided, the decompression device including a plurality of logic circuits related to a compression method of compressed data. The control method includes receiving compressed data to be decompressed and used in neural network processing of an artificial intelligence model by the plurality of logic circuits, decompressing the compressed data by the plurality of logic circuits and outputting the decompressed data, and obtaining data in a form processable by the neural network from the data output by the plurality of logic circuits.
[0021] The control method may include obtaining data in a form processable by a neural network based on decompressed data and a representative value matrix corresponding to the compressed data; and performing neural network processing using the data in a form processable by a neural network, wherein the decompressed data and the representative value matrix include a matrix obtained by quantizing an original matrix included in an artificial intelligence model.
[0022] The control method may include updating the decompressed data based on a pruning index matrix corresponding to the compressed data, wherein the pruning index matrix includes a matrix obtained in a pruning process of the original matrix, and wherein the pruning index matrix is used in a process of obtaining the compressed data.
[0023] The control method may include changing a binary data value of at least one element of a plurality of elements included in the decompressed data based on patch information corresponding to the compressed data, wherein the patch information includes error information generated in a process of obtaining the compressed data.
[0024] The control method may include obtaining a pruning index matrix based on a first pruning index matrix corresponding to the compressed data and a second pruning index matrix corresponding to the compressed data; and updating the decompressed data based on the pruning index matrix, wherein the pruning index matrix includes a matrix obtained in a pruning process of an original matrix, wherein the pruning index matrix is used in a process of obtaining the compressed data, and wherein the first pruning index matrix and the second pruning index matrix are respectively obtained based on each of a first submatrix and a second submatrix obtained by factorizing the original matrix.
[0025] The decompressed data includes a matrix obtained by interleaving an original matrix and then quantizing the interleaved matrix, wherein the control method also includes deinterleaving the data in a neural network processable form according to a manner corresponding to the interleaving, and wherein during the process of performing the neural network processing, the deinterleaved data is used to perform the neural network processing.
[0026] In performing the neural network processing, a plurality of processing elements arranged in a matrix form are used to perform the neural network processing.
[0027] The decompressed data includes a matrix obtained by dividing an original matrix into a plurality of matrices having the same number of columns and rows and quantizing one of the divided plurality of matrices.
[0028] The decompression device also includes multiple other logic circuits related to the compression method of other compressed data, and the control method also includes: receiving other compressed data to be decompressed and used in the neural network processing of the artificial intelligence model, decompressing the other compressed data by multiple other logic circuits, outputting the decompressed other data, obtaining other data in a neural network processable form output from the multiple logic circuits, and obtaining a matrix including multiple binary data in each element by coupling the neural network processable data and the other data in a neural network processable form.
[0029] The decompression device is implemented as a chip.
[0030] Other aspects, advantages, and salient features of the present disclosure will become apparent to those skilled in the art from the following detailed description, which, in conjunction with the accompanying drawings, discloses various embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The above and other aspects, features and advantages of certain embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings, in which:
[0032] Figure 1 is a block diagram showing a configuration of a decompression device according to an embodiment of the present disclosure;
[0033] Figure 2 is a diagram describing an electronic system according to an embodiment of the present disclosure;
[0034] Figure 3A , Figure 3B , Figure 3C and Figure 3D is a diagram describing a method for obtaining a first matrix, a first pruning index matrix, and a second pruning index matrix according to various embodiments of the present disclosure;
[0035] Figure 4 is a diagram describing a method for obtaining a first matrix according to an embodiment of the present disclosure;
[0036] Figure 5A and Figure 5B is a diagram describing a decompression device according to various embodiments of the present disclosure;
[0037] Figure 6 is a diagram depicting an original matrix according to an embodiment of the present disclosure;
[0038] Fig. 7A , Figure 7B and Figure 7C is a diagram describing a method for decompressing a first matrix according to various embodiments of the present disclosure;
[0039] Fig. 8A and Figure 8B is a diagram describing an update operation of a second matrix according to various embodiments of the present disclosure;
[0040] Fig. 9 is a diagram describing a method for merging a plurality of second matrices of a processor according to an embodiment of the present disclosure;
[0041] Fig.10 is a flowchart describing a control method of a decompression device according to an embodiment of the present disclosure;
[0042] Fig.11A , Fig. 11B , Fig. 11C and Fig.11D is a diagram describing the learning process of an artificial intelligence model according to various embodiments of the present disclosure;
[0043] Fig.12is a diagram describing a method for performing pruning in a learning process according to an embodiment of the present disclosure;
[0044] Fig.13 is a graph depicting the effect of the value of m according to an embodiment of the present disclosure; and
[0045] Fig.14A and Fig. 14B is a graph depicting improvement in learning speed according to various embodiments of the present disclosure.
[0046] The same reference numerals are used throughout the drawings to denote the same elements. DETAILED DESCRIPTION
[0047] The following description with reference to the accompanying drawings is provided to assist in a comprehensive understanding of the various embodiments of the present disclosure as defined by the claims and their equivalents. It includes various specific details to assist in understanding, but these details should be considered as merely exemplary. Therefore, it will be appreciated by those of ordinary skill in the art that various changes and modifications may be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, for the sake of clarity and conciseness, descriptions of well-known functions and configurations may be omitted.
[0048] The terms and words used in the following description and claims are not limited to the bibliographic meanings, but are merely used by the inventor to enable a clear and consistent understanding of the present disclosure. Therefore, it should be apparent to those skilled in the art that the following description of various embodiments of the present disclosure is provided for illustrative purposes only, and not for the purpose of limiting the present disclosure as defined by the attached claims and their equivalents.
[0049] It will be understood that singular forms include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a component surface" includes reference to one or more of such surfaces.
[0050] The present disclosure provides a decompression device and a control method thereof, which reduces memory usage in speech recognition or object recognition, uses an artificial intelligence model with reduced data capacity for high-speed processing, and decompresses the artificial intelligence model with reduced data capacity.
[0051] Considering the functions in the present disclosure, currently widely used general terms are selected as the terms used in the embodiments of the present disclosure, but may be changed according to the intention of those skilled in the art or the emergence of judicial precedents, new technologies, etc. In addition, in certain cases, there may be terms arbitrarily selected by the applicant. In this case, the meaning of such terms will be mentioned in detail in the corresponding description section of the present disclosure. Therefore, the terms used in the present disclosure should be defined based on the meaning of the terms and the content throughout the present disclosure rather than the simple names of the terms.
[0052] In the present disclosure, expressions “having”, “may have”, “including”, “may include”, etc. indicate the existence of corresponding features (e.g., values, functions, operations, components such as parts, etc.), and do not exclude the existence of additional features.
[0053] The expression “at least one of A and / or B” should be understood to mean any one of “A” or “B” or “A and B”.
[0054] The expressions “first,” “second,” etc. used in the present disclosure may indicate various components regardless of the order and / or importance of the components, and will only be used to distinguish one component from other components and will not limit the corresponding components.
[0055] Singular expressions include plural expressions unless the context clearly indicates otherwise. It should also be understood that the terms "comprising" or "consisting of" used in this application specify the presence of features, numbers, operations, components, parts or combinations thereof mentioned in the specification, but do not exclude the presence or addition of one or more other features, numbers, operations, components, parts or combinations thereof.
[0056] In the specification, the term "user" may be a person using an electronic device or a device using an electronic device (eg, an artificial intelligence electronic device).
[0057] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings.
[0058] Figure 1 is a block diagram showing a configuration of a decompression device according to an embodiment of the present disclosure.
[0059] refer to Figure 1 , the decompression device 100 may be a device for decompressing a compressed artificial intelligence model. Here, the compressed artificial intelligence model is data that can be used for neural network processing after decompression, and may be data that cannot be used for neural network processing before decompression.
[0060] For example, the decompression device 100 can be a device that decompresses the compressed data included in the compressed artificial intelligence model, finally obtains the neural network processable data (hereinafter referred to as the recovery data), and uses the recovery data to perform neural network processing. For example, the decompression device 100 can be implemented in the form of separate hardware (HW) between the memory and the chip present in the server, desktop personal computer (PC), notebook, smart phone, tablet PC, television (TV), wearable device, etc., and can also be implemented as a system on chip (SOC). Alternatively, the decompression device 100 can be implemented in the form of chips such as a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processing (DSP), a network processing unit (NPU), etc., can be implemented in the form of a circuit, and can be implemented in the form of a configuration inside the chip.
[0061] However, the type of decompression device 100 described above is merely an example, and any device may be used as long as it is a device that decompresses compressed data included in a compressed artificial intelligence model, finally obtains restored data, and performs neural network processing using the restored data.
[0062] like Figure 1 As shown, the decompression device 100 includes a memory 110 , a decoder 120 , and a processor 130 .
[0063] The memory 110 may store compressed data that is decompressed and used in the neural network processing of the artificial intelligence model. For example, the memory 110 may store compressed data received from a compression device. The compressed data and the data constituting the artificial intelligence model before compression may be expressed in at least one matrix form. Hereinafter, for ease of explanation, the compressed data will be described as a first matrix, and the data from which the first matrix is decompressed will be described as a second matrix. The second matrix may be converted into recovery data together with a representative value matrix to be described later.
[0064] Here, the first matrix may be a matrix that compresses the second matrix based on a compression method. For example, the first matrix may be a matrix compressed by a coding matrix formed based on a logic circuit such as an XOR gate, in which the second matrix constitutes the decoder 120. The second matrix may be a matrix in which the original matrix (weight matrix) included in the artificial intelligence model is compressed by pruning, etc., or a matrix in which the original matrix is quantized. Alternatively, the second matrix may be a matrix of binary data obtained in the quantization process of the pruned original matrix. The original matrix as a matrix included in the artificial intelligence model is a matrix obtained after the learning process of the artificial intelligence model is completed, and may be in a state where no compression is performed. On the other hand, the coding matrix is a matrix for compressing the second matrix, and may be a matrix for implementing the logic circuit constituting the decoder 120.
[0065] A detailed method of obtaining the first matrix by compressing the second matrix will be described later with reference to the accompanying drawings.
[0066] Decoder 120 may include multiple logic circuits related to the method for compressing compressed data. As an example of using a coding matrix to perform compression, decoder 120 may include multiple logic circuits formed based on the coding matrix. The logic circuit may be an XOR gate, but is not limited to this. For example, decoder 120 may include multiple XOR gates, wherein corresponding input / output terminals are connected to multiple XOR gates based on the coding matrix. That is, if binary data is input to multiple XOR gates, the corresponding input / output terminals may be connected to output the matrix multiplication result of the coding matrix and the binary data. Binary data refers to data that is binary quantized and displayed in 1 bit, but is not limited to this, and can be obtained by other quantization methods.
[0067] In the case of designing a compression method, the decoder 120 can be first implemented by randomly connecting a plurality of XOR gates, and the coding matrix can be obtained by comparing the input of each of the plurality of XOR gates and the output of each of the plurality of XOR gates. In addition, a coding matrix including 0 and 1 can be first generated, and the second matrix can be compressed based on the coding matrix to obtain the first matrix. In this case, when the decompression device 100 is implemented, the decoder 120 can also be implemented by connecting the input / output terminals of the plurality of XOR gates so as to correspond to the coding matrix.
[0068] When the first matrix is input, the decoder 120 can decompress the first matrix and output the decompressed data through a plurality of logic circuits. For example, when the binary data of the first number of cells included in the first matrix is input, the decoder 120 can output the binary data of the second number of cells larger than the first number through a plurality of XOR gates. For example, the decoder 120 can sequentially receive the binary data of the five cells included in the first matrix, and can sequentially output the binary data of nine cells through a plurality of XOR gates. For example, if the encoding matrix is in the form of 9×5, the decoder 120 may include five input terminals and nine output terminals. In addition, if the first matrix has 25 elements (parameters), the decoder 120 can receive the binary data of five cells five times and output the binary data of nine cells five times, thereby outputting a total of 45 elements.
[0069] The processor 130 may obtain the second matrix from the binary data output from the decoder 120. In the above example, the processor 130 may obtain the second matrix including 45 elements from the binary data of nine units output five times.
[0070] On the other hand, the memory 110 may also store a representative value matrix corresponding to the first matrix. For example, the memory 110 may store a representative value matrix corresponding to the first matrix received from the compression device. The representative value matrix may be obtained during the quantization process of the pruned original matrix, and may be a set of representative values representing a plurality of elements included in the pruned original matrix.
[0071] The processor 130 may obtain a recovery matrix based on the second matrix and the representative value matrix, and perform neural network processing using the recovery matrix. Here, the second matrix and the representative value matrix may be matrices obtained by quantizing the original matrix included in the artificial intelligence model. Quantization means changing data into one of a plurality of sample values, and in quantization, since the data is represented by a plurality of sample values, the overall data capacity may be reduced, but errors may occur due to differences between the changed sample values and the original data.
[0072] For example, the compression device can quantize the original matrix included in the artificial intelligence model to obtain a representative value matrix and a second matrix including binary quantized data. Here, binary quantized data refers to data represented by 1 bit. As a more specific example, the compression device can perform quantization by setting a representative value for a predetermined number of elements included in the original matrix and representing each element as binary quantized data. In this way, the compression device can obtain a representative value matrix including representative values for the entire original matrix, and can obtain a second matrix obtained by converting the elements included in the original matrix into binary quantized data. However, the compression method is not limited to this, and the compression device can use any other quantization method.
[0073] The processor 130 may obtain a restored matrix to be used for neural network processing from the second matrix and the representative value matrix. Here, the restored matrix may be different from the original matrix. That is, due to quantization error, the restored matrix may be different from the original matrix. However, the restored matrix may be obtained so that the result of the neural network processing using the original matrix and the result of the neural network processing using the restored matrix are not significantly different in the compression process to be described later, which will be described later.
[0074] The processor 130 may include a plurality of processing elements arranged in a matrix form, and may perform neural network processing of an artificial intelligence model using the plurality of processing elements.
[0075] On the other hand, the memory 110 may also store a pruning index matrix corresponding to the first matrix. For example, the memory 110 may also store a pruning index matrix corresponding to the first matrix received from a compression device.
[0076] The pruning index matrix is a matrix obtained in the pruning process of the original matrix and can be used in the process of compressing the second matrix into the first matrix.
[0077] First, pruning is a method of removing redundant weights, specifically changing the numbers of certain elements (specific deep learning parameters) in the original matrix included in the artificial intelligence model to zero. For example, the compression device can prune the m×n original matrix by changing a predetermined value or smaller element in the multiple elements included in the m×n original matrix to 0, and an m×n pruning index matrix representing 0 or 1 can be obtained, and the pruning index matrix indicates whether each of the multiple elements included in the m×n original matrix is pruned.
[0078] In addition, the compression device may use a pruning index matrix in the process of compressing the second matrix into the first matrix based on the encoding matrix. Specifically, the compression device may determine the value of the binary data of the first number of units to be included in the first matrix, so that the result of the matrix multiplication of the binary data of the first number of units to be included in the first matrix and the encoding matrix is equal to the corresponding binary data of the second number of units included in the second matrix, and in this case, because the first number is less than the second number, the value of the binary data of the first number of units to be included in the first matrix that satisfies the above conditions may not be derived. In this case, the compression device may determine the value of the binary data of the first number of units to be included in the first matrix by determining some corresponding binary data of the second number of units included in the second matrix as unnecessary data based on the pruning index matrix. It will be described in detail later.
[0079] The processor 130 may update the second matrix based on the pruning index matrix. For example, the processor 130 may change some of the plurality of elements included in the second matrix to 0 based on the pruning index matrix.
[0080] However, the processor 130 is not limited thereto, and the memory 110 may not store the pruning index matrix. That is, if the compression device does not provide the pruning index matrix to the decompression device 100, the decompression device 100 may recognize that the pruning index matrix is not used in the process of compressing the second matrix into the first matrix, and may omit the operation of changing some of the plurality of elements included in the second matrix to 0.
[0081] Alternatively, although the compression device may provide the pruning index matrix to the decompression device 100, the pruning index matrix may not be used in the process of compressing the second matrix into the first matrix. In this case, the compression device may provide the decompression device 100 with information indicating that the pruning index matrix is not used in the process of compressing the second matrix into the first matrix, and the decompression device 100 may omit the operation of changing some of the multiple elements included in the second matrix to 0.
[0082] On the other hand, the memory 110 may also store patch information corresponding to the first matrix. For example, the memory 110 may also store patch information corresponding to the first matrix received from a compression device.
[0083] Here, the patch information may include error information generated in the process of compressing the second matrix into the first matrix. Specifically, the compression device may determine the value of the binary data of the first number of units to be included in the first matrix by using the pruning index matrix in the process of compressing the second matrix into the first matrix. However, even if the pruning index matrix is used, the value of the binary data of the first number of units to be included in the first matrix may not be determined. In this case, the compression device may determine the value of the binary data of the first number of units to be included in the first matrix to minimize the number of errors. For example, the compression device may determine the value of the binary data of the first number of units to be included in the first matrix so as to minimize the number of bits of the difference between the binary data of the first number of units to be included in the first matrix and the matrix multiplication result of the encoding matrix and the corresponding binary data of the second number of units included in the second matrix. In addition, the compression device may generate information about the number of bits, positions, etc. of the difference as patch information. In addition, the compression device may provide patch information to the decompression device 100 to resolve errors that may occur during the decompression process.
[0084] The processor 130 may change the binary data values of some of the multiple elements included in the second matrix based on the patch information. For example, if the value of the element is 0, the processor 130 may change the value of the element indicated by the patch information among the multiple elements included in the second matrix to 1, and if the value of the element is 1, change it to 0.
[0085] On the other hand, the memory 110 may also store a first pruning index matrix corresponding to the compressed data and a second pruning index matrix corresponding to the first matrix. For example, the memory 110 may store a first pruning index matrix and a second pruning index matrix corresponding to the first matrix received from the compression device. In this case, the compression device may provide the first pruning index matrix and the second pruning index matrix to the decompression device 100 instead of the pruning index matrix.
[0086] First, as described above, the pruning index matrix may be a matrix obtained in the pruning process of the original matrix.
[0087] The first pruning index matrix and the second pruning index matrix can be obtained based on each of the first sub-matrix and the second sub-matrix obtained by factoring the original matrix. Factorization is a factorization that means dividing a matrix into two smaller matrices, and for example, a method such as non-negative matrix factorization (NMF) can be used. However, the method of obtaining the first pruning index matrix and the second pruning index matrix is not limited thereto, and various methods can be used.
[0088] The compression device may obtain the first pruning index matrix and the second pruning index matrix by factorizing the original matrix to obtain the first sub-matrix and the second sub-matrix and pruning the first sub-matrix and the second sub-matrix respectively. Thereafter, the compression device may update the first pruning index matrix and the second pruning index matrix by comparing the result of the neural network processing using the pruning index matrix with the result of the neural network processing using the first pruning index matrix and the second pruning index matrix. The updating method may be a method of changing the pruning rate of each of the first sub-matrix and the second sub-matrix. Finally, the compression device may obtain the first pruning index matrix and the second pruning index matrix, wherein the difference between the operation results of the two cases falls within a threshold.
[0089] The compression device can reduce the data capacity by converting the pruning index matrix into a first pruning index matrix and a second pruning index matrix. For example, the compression device can convert a 100×50 pruning index matrix into a first pruning index matrix of 100×10 and a second pruning index matrix of 10×50. In this case, the compression device can reduce 5000 data to 1000+500=1500 data.
[0090] The processor 130 may obtain a pruning index matrix based on the first pruning index matrix and the second pruning index matrix, and update the second matrix based on the pruning index matrix. Specifically, each of the first pruning index matrix and the second pruning index matrix may include binary data, and the processor 130 may obtain the pruning index matrix by performing a matrix multiplication operation on the first pruning index matrix and the second pruning index matrix.
[0091] The processor 130 may change some elements of the plurality of elements included in the second matrix to 0 based on the pruning index matrix.
[0092] On the other hand, the second matrix may be a matrix obtained by interleaving the original matrix and then quantizing the interleaved matrix. Interleaving means rearranging the order of data included in the matrix by a predetermined unit. That is, the compression device may also perform interleaving before quantizing the original matrix.
[0093] The processor 130 may deinterleave the restoration matrix according to a method corresponding to the interleaving, and perform a neural network process using the deinterleaved restoration matrix.
[0094] With the addition of interleaving and deinterleaving operations, the compression rate can be improved during the quantization process. For example, if the elements of the original matrix are not uniformly distributed, the compression rate or accuracy may be significantly reduced because the pruning index matrix is 1 or 0 continuous. In this case, when the matrix to be compressed is interleaved, the randomness of the pruning index matrix can be improved, thereby improving the compression rate and accuracy. There is no particular restriction on the type of interleaving method and deinterleaving method, and various methods can be used according to the decompression speed and randomness. For example, the method used in turbo code can be used, and there is no particular restriction as long as the interleaving method and the deinterleaving method correspond to each other.
[0095] On the other hand, the second matrix may be a matrix obtained by dividing the original matrix into a plurality of matrices having the same number of columns and the same number of rows and quantizing one of the plurality of divided matrices.
[0096] The advantage of dividing the original matrix into multiple matrices can be that, for example, in the original matrix of m×n, one of m and n is significantly larger than the other. For example, in the case of compressing the original matrix of 100×25 into a matrix of 100×r and a matrix of r×25, r is usually selected to be less than 25 and the compression ratio can be reduced. In this case, when the original matrix of 100×25 is divided into four matrices of 25×25 and each of the four matrices is compressed, the compression ratio can be improved. In addition, when the original matrix is divided into multiple matrices, the amount of calculation during the compression process can also be reduced. That is, it may be effective to perform compression after dividing the skew matrix into square matrices.
[0097] On the other hand, the memory 110 can also store a third matrix that is decompressed and used in the neural network processing of the artificial intelligence model. For example, the memory 110 can also store a third matrix received from a compression device.
[0098] The decompression device 100 includes a plurality of other logic circuits related to the compression method of the third matrix based on the coding matrix, and may also include another decoder, when the third matrix is input, the other decoder decompresses other compressed data through a plurality of other logic circuits and outputs other decompressed data. Here, the third matrix can be a matrix based on the coding matrix to compress the fourth matrix. The plurality of other logic circuits can be a plurality of other XOR gates.
[0099] However, a plurality of other XOR gates are not limited thereto, and a plurality of other XOR gates can be connected to each input / output terminal based on another coding matrix different from the coding matrix. In this case, the number of binary data input to a plurality of other XOR gates and the number of output binary data may not be the first number and the second number, respectively. Here, the third matrix may be a matrix based on another coding matrix compressing the fourth matrix.
[0100] The processor 130 may obtain a fourth matrix in a neural network operable form from data output by another decoder, and may combine the second matrix and the fourth matrix to obtain a matrix in which each element includes a plurality of binary data.
[0101] That is, the decompression device 100 may also include multiple decoders. This is because each element included in the original matrix may include multiple binary data. For example, the compression device may divide the original matrix into multiple matrices, in which each element is 1, and quantize and compress each of the multiple matrices. For example, if each element of the original matrix includes two binary data, the compression device may divide the original matrix into two matrices, in which each element includes one binary data, and quantize and compress each of the two matrices to obtain the above-mentioned first matrix and the third matrix. In addition, the compression device may provide the first matrix and the third matrix to the decompression device 100. The decompression device 100 may process the first matrix and the third matrix in parallel using a decoder 120 and another decoder, respectively, and the processor 130 may merge the second matrix and the fourth matrix to obtain a matrix in which each element includes multiple binary data.
[0102] Meanwhile, in the above, the decoder 120 has been described as being disposed between the memory 110 and the processor 130. In this case, because the internal memory provided in the processor 130 stores the decompressed data, a memory having a large capacity is required and the power consumption may be considerable. However, decompression may be performed while calculations are performed in the processing element unit inside the processor 130, and the impact on the processing execution time of the processing element unit may be made smaller without incurring overhead on the existing hardware. In addition, because the decoder 120 may be disposed between the memory 110 and the processor 130, it is possible to design it in the form of a memory package without modifying the content of the existing accelerator design. This configuration may be more suitable for a convolutional neural network (CNN) that reuses the entire decompressed data.
[0103] However, the decompression device 100 is not limited to this and can also be implemented as a chip. In this case, the memory 110 can only receive some compressed data from an external memory outside the chip and only store some compressed data. In addition, because the memory 110 performs decompression in operation whenever the processing element unit requests data, the memory 110 can use a memory with a small capacity and can also reduce power consumption. However, because the memory 110 only stores some compressed data, decompression and deinterleaving may be performed whenever the processing element unit requests data, thereby increasing waiting time and increasing long-term power consumption. In addition, the decoder 120 is added to the inside of the existing accelerator, and it may be necessary to modify many existing designs. This configuration may be more suitable for a recursive neural network (RNN) that uses some compressed data at a time.
[0104] The decompression device 100 may perform decompression by the above method to obtain a restoration matrix, and perform neural network processing using the obtained restoration matrix.
[0105] Figure 2 is a diagram for describing an electronic system according to an embodiment of the present disclosure.
[0106] refer to Figure 2 , the electronic system 1000 includes a compression device 50 and a decompression device 100 .
[0107] First, the compression device 50 may be a device for compressing an artificial intelligence model. For example, the compression device 50 is a device for compressing at least one original matrix included in the artificial intelligence model, and may be a device such as a server, a desktop PC, a notebook, a smart phone, a tablet PC, a TV, a wearable device, etc.
[0108] However, this is merely an example, and any device may be used as long as the compression device 50 may be a device that can reduce the data size of the artificial intelligence model by compressing the artificial intelligence model. Here, the original matrix may be a weight matrix.
[0109] The compression device 50 may quantize the original matrix included in the artificial intelligence model to obtain a representative value matrix and a second matrix. As described above, the quantization method is not particularly limited.
[0110] The compression device 50 may compress the second matrix into the first matrix based on the encoding matrix. Alternatively, the compression device 50 may compress the second matrix into the first matrix based on whether the plurality of elements included in the encoding matrix and the original matrix are pruned. In particular, as the compression device 50 further considers whether to perform pruning, the compression rate of the second matrix may be improved.
[0111] The compression device 50 may obtain only the first matrix from the second matrix. Alternatively, the compression device 50 may also obtain the first matrix and the patch information from the second matrix. In particular, as the compression device 50 further uses the patch information, the compression rate of the second matrix may be improved. However, as the size of the patch information increases, the compression rate may decrease.
[0112] At the same time, the compression device 50 may obtain a pruning index matrix indicating whether each element included in the original matrix is pruned by pruning the original matrix included in the artificial intelligence model. The compression device 50 may provide the pruning index matrix to the decompression device 100.
[0113] Alternatively, the compression device 50 may obtain a pruning index matrix indicating whether each element included in the original matrix is pruned by pruning the original matrix included in the artificial intelligence model, and may compress the pruning index matrix into a first pruning index matrix and a second pruning index matrix using the above method. The compression device 50 may provide the first pruning index matrix and the second pruning index matrix to the decompression device 100.
[0114] The decompression device 100 may receive the compressed artificial intelligence model from the compression device 50, perform decompression, and perform neural network processing.
[0115] However, the electronic system 1000 is not limited thereto and may be implemented as an electronic device. For example, when the electronic device compresses the artificial intelligence model in the same manner as the compression device 50 and performs neural network processing, the electronic device may perform decompression in the same manner as the decompression device 100.
[0116] Hereinafter, for convenience of explanation, the compression device 50 and the decompression device 100 will be described as separate. In addition, the compression operation of the compression device 50 will be described first, and the operation of the decompression device 100 will be described in more detail with reference to the drawings.
[0117] FIG. 3A to FIG. 3D is a diagram for describing a method for obtaining a first matrix, a first pruning index matrix, and a second pruning index matrix according to various embodiments of the present disclosure.
[0118] refer to FIG. 3A to FIG. 3D , Figure 3Ais a diagram showing an example of an original matrix included in an artificial intelligence model, and the original matrix may be in the form of m×n. For example, the original matrix may be in the form of 10000×8000. In addition, each of the multiple elements in the original matrix may be 32 bits. That is, the original matrix may include 10000×8000 elements with 32 bits. However, the original matrix is not limited thereto, and the size and number of bits of each element of the original matrix may vary.
[0119] Figure 3B It is shown that the trimming and quantization Figure 3A A plot of the result for the original matrix is shown.
[0120] The compression device 50 may prune each of the plurality of elements included in the original matrix based on the first threshold value, and obtain a pruning index matrix 310 indicating whether each of the plurality of elements is pruned as binary data.
[0121] For example, the compression device 50 may prune the original matrix by converting less than 30 elements among the plurality of elements included in the original matrix to 0 and keeping the remaining elements unchanged. In addition, the compression device 50 may obtain the pruning index matrix 310 by converting the elements converted to 0 among the plurality of elements to 0 and converting the remaining elements to 1. That is, the pruning index matrix 310 has the same size as the original matrix and may include 0 or 1.
[0122] In addition, the compression device 50 may quantize unpruned elements among a plurality of elements included in the original matrix to obtain a representative value matrix 330 and a second matrix 320 including binary quantized data.
[0123] The compression device 50 may use a representative value quantization Figure 3A The n elements in the original matrix of . Accordingly, Figure 3B A representative value matrix 330 including m elements is shown. Here, the number n of elements quantized into one representative value is only an example, and other values may be used, and if other values are used, the number of elements included in the representative value matrix 330 may also vary. In addition, the compression device 50 may obtain a second matrix 320 including binary quantized data, and the second matrix 320 has the same size as the original matrix and may include 0 or 1.
[0124] like Figure 3C As shown, the compression device 50 can compress the second matrix 320 into the first matrix 10 based on the encoding matrix. The number of elements included in the first matrix 10 is less than the number of elements included in the second matrix 320. Figure 4 A method of obtaining the first matrix 10 is described.
[0125] like Figure 3D As shown, the compression device 50 can compress the pruning index matrix 310 into the first pruning index matrix 20-1 and the second pruning index matrix 20-2. Specifically, the compression device 50 can obtain the first pruning index matrix 20-1 and the second pruning index matrix 20-2 by factoring the original matrix into the first sub-matrix and the second sub-matrix and pruning the first sub-matrix and the second sub-matrix respectively. In addition, the compression device 50 can update the first pruning index matrix 20-1 and the second pruning index matrix 20-2 so that the neural network processing result using the matrix multiplication result of the first pruning index matrix 20-1 and the second pruning index matrix 20-2 is close to the neural network processing result using the pruning index matrix 310. For example, if the accuracy of the neural network processing result using the matrix multiplication result of the first pruning index matrix 20-1 and the second pruning index matrix 20-2 becomes low, the accuracy can be improved by reducing the pruning rate used when pruning each of the first sub-matrix and the second sub-matrix. In this case, the pruning rates used to obtain the first pruning index matrix 20-1 and the second pruning index matrix 20-2 can be different from each other. The compression device 50 can improve the accuracy of the neural network processing result using the matrix multiplication result of the first pruning index matrix 20-1 and the second pruning index matrix 20-2 by repeating such a process.
[0126] exist FIG. 3A to FIG. 3D In the embodiment, the compression device 50 is described as performing both pruning and quantization, but is not limited thereto. For example, the compression device 50 may perform quantization without performing pruning.
[0127] Figure 4 is a diagram for describing a method for obtaining a first matrix according to an embodiment of the present disclosure.
[0128] refer to Figure 4 , Figure 4 A method of compressing a predetermined number of elements B included in a second matrix into x using a coding matrix A is shown. Here, x may be included in the first matrix. For example, nine binary data included in the second matrix may be compressed into five binary data. At this time, the trimmed data among the nine binary data is represented as "-" and does not need to be restored after compression. That is, according to Figure 4 The matrix multiplication of can generate nine equations, but there is no need to consider the equation including "-". Therefore, the nine binary data included in the second matrix can be compressed into five binary data.
[0129] That is, the matrix multiplication result of the binary data x of the first number of units included in the first matrix and the encoding matrix A can be the same as the corresponding binary data B of the second number of units included in the second matrix. Here, the second number is greater than the first number. During the compression process, the binary data of the second number of units included in the second matrix are converted into binary data of the first number of units smaller than the second number, and the converted binary data can form the first matrix. During the matrix multiplication process, the multiplication process between the corresponding binary data can be performed in the same manner as the AND gate, the addition process between the multiplication results can be performed in the same manner as the XOR gate, and the AND gate has a higher calculation priority than the XOR gate.
[0130] Here, the coding matrix may include a first type element and a second type element, and the number of the first type elements included in the coding matrix and the number of the second type elements included in the coding matrix may be the same as each other. For example, the coding matrix may include zeros and ones, and the number of zeros and the number of ones may be the same as each other. However, the coding matrix is not limited thereto, and when the number of elements included in the coding matrix is an odd number, the difference between the number of zeros and the number of ones may be within a predetermined number (e.g., one).
[0131] On the other hand, Figure 4 As shown, when the number of pruned data is three ("-" is three), the five binary data may not satisfy the remaining six equations. Therefore, the compression device 50 can obtain information about the number of equations that are not established by the five binary data and the position information of the last equation 410 as patch information. Figure 4 , 10110 represents compressed five binary data, 01 represents information about the number of unestablished equations, and 0110 may be position information of the last equation 410. Here, the position information of the last equation 410 has been described as indicating that it is the sixth based on the unpruned data, but is not limited thereto. For example, the position information of the last equation 410 may be obtained to indicate that it is the ninth based on nine data without considering pruning.
[0132] On the other hand, during matrix multiplication, the multiplication process between corresponding binary data can be performed in the same manner as an AND gate, the addition process between multiplication results can be performed in the same manner as an XOR gate, and the AND gate has a higher calculation priority than the XOR gate.
[0133] For ease of explanation, we will use 10110 to describe matrix multiplication, which is from Figure 4The value of x derived. The matrix multiplication of the first row of A and 10110 as the value of x is omitted because it is unnecessary data according to pruning. In the matrix multiplication of 10010 (the second row of A) and 10110 (the value of x), first, for each number, the multiplication operation between binary data is performed in the same manner as the AND gate. That is, 1, 0, 0, 1, 0 are obtained by the operation of 1×1=1, 0×0=0, 0×1=0, 1×1=1, and 0×0=0. Thereafter, an addition operation is performed on 1, 0, 0, 1, 0 in the same manner as the XOR gate, and 0 is finally obtained. Specifically, 1 can be obtained by the addition operation of the first binary data 1 and the second binary data 0, 1 can be obtained by the addition operation of the operation result 1 and the third binary data 0, 0 can be obtained by the addition operation of the accumulation operation 1 and the fourth binary data 1, and 0 can finally be obtained by the addition operation of the accumulation operation result 0 and the fifth binary data 0. Here, the order of operations can be changed as much as possible, and even if the order of operations is changed, the value finally obtained is the same. In this way, the matrix multiplication between the remaining rows of A and 10010 can be performed.
[0134] However, as described above, there may be an unestablished equation (for example, the last row of A (ie, the last equation 410)), and its operation results are as follows. In the matrix multiplication of 00011 (the last row of A) and 10110 (the value of x), first, for each number, the multiplication operation between binary data is performed in the same manner as the AND gate. That is, 0, 0, 0, 1, 0 are obtained by the operations of 0×1=0, 0×0=0, 0×1=0, 1×1=1, and 0×0=0. Thereafter, an addition operation is performed on 0, 0, 0, 1, 0 in the same manner as the XOR gate, and 1 is finally obtained. Specifically, 0 can be obtained by the addition operation of the first binary data 0 and the second binary data 0, 0 can be obtained by the addition operation of the operation result 0 and the third binary data 0, 1 can be obtained by the addition operation of the cumulative operation result 0 and the fourth binary data 1, and 1 can finally be obtained by the addition operation of the cumulative operation result 1 and the fifth binary data 0. This does not match the value 0 of the last row of B, and the compression device 50 provides it to the decompression device 100 as patch information, and the decompression device 100 can compensate for this based on the patch information. That is, the decompression device 100 can obtain the position information of the row where the equation is not established based on the patch information, and can convert the binary data of the row corresponding to the position information in the matrix multiplication result of the encoding matrix A and x into other binary data. Figure 4 In the example of , the decompression device 100 can convert the value of the last row of the matrix multiplication result of the encoding matrix A and x from 1 to 0 based on the patch information.
[0135] In this way, the compression device 50 can obtain the first matrix, the first pruning index matrix, the second pruning index matrix and the patch information from the original matrix.
[0136] However, the compression device 50 is not limited thereto, and the second matrix may be compressed into the first matrix by a method that does not use patch information. Figure 4 For example, the compression device 50 may determine x to be 6 bits to prevent generation of patch information.
[0137] Alternatively, the compression device 50 may also compress the second matrix into the first matrix without performing pruning. Figure 4 In the example of , the compression device 50 may determine the number of digits of x based on the dependencies between the nine equations. That is, if there is no dependency between the nine equations, the compression device 50 obtains x including nine digits, and in this case, compression may not be performed. Alternatively, if there is a dependency between the nine equations, the compression device 50 obtains x including less than nine digits, and in this case, compression may be performed.
[0138] Figure 5A and Figure 5B is a diagram for describing a decompression device according to various embodiments of the present disclosure.
[0139] Figure 6 is a diagram depicting an original matrix according to an embodiment of the present disclosure.
[0140] First, refer to Figure 5A , the decompression device 100 may include a plurality of decoders (D units). The decompression device 100 may also include an external memory 510, a plurality of deinterleavers 520, and a processor 530-1. Here, the external memory 510 and the processor 530-1 may be respectively Figure 1 The memory 110 and the processor 130 are the same.
[0141] refer to Figure 5A and Figure 6 , the external memory 510 can store a plurality of first matrices provided from the compression device 50. Here, the plurality of first matrices can be matrices in which the original matrix is quantized and compressed. This is actually because the data in the original matrix is very large, and for example, Figure 6 As shown, the compression device 50 may divide the original matrix into a plurality of matrices, quantize and compress each of the plurality of matrices to obtain a plurality of first matrices. The compression device 50 may provide the plurality of first matrices to the external memory 510.
[0142] Each of the plurality of decoders (D units) may receive one of the plurality of first matrices from the external memory 510 and output a decompressed second matrix. That is, the external memory 510 may decompress the plurality of first matrices in parallel by providing the plurality of first matrices to the plurality of decoders, respectively, and may improve parallelism.
[0143] However, the decompression device 100 is not limited thereto, and may also decompress the compressed data in sequence, such as Figure 6 The upper left corner of the matrix is compressed data decompressed and Figure 6 The upper left corner of the matrix adjacent to the right is the compressed data to be decompressed.
[0144] Each of the plurality of decoders may transmit the decompressed second matrix to the internal memory (on-chip memory) of the processor 530 - 1 . In this case, each of the plurality of decoders may transmit the second matrix via the plurality of deinterleavers 520 .
[0145] To describe the operation of the plurality of deinterleavers 520, the interleaving operation of the compression device 50 will first be described. The compression device 50 may interleave each of the plurality of partition matrices, such as Figure 6 The compression device 50 may quantize and compress each of the plurality of interleaved matrices.
[0146] The plurality of deinterleavers 520 may correspond to the interleaving operation of the compression apparatus 50. That is, the plurality of deinterleavers 520 may deinterleave the interleaved matrix to restore the matrix before interleaving.
[0147] exist Figure 5A In the embodiment, the decompression device 100 is shown as including a plurality of deinterleavers 520, but is not limited thereto. For example, the compression device 50 may interleave the original matrix itself without interleaving each of the plurality of matrices into which the original matrix is divided. In this case, the decompression device 100 may also include one deinterleaver.
[0148] On the other hand, the memory 110 may also store the first pruning index matrix, the second pruning index matrix and the patch information. In this case, the processor 530-1 may obtain the pruning index matrix from the first pruning index matrix and the second pruning index matrix, and update the second matrix based on the pruning index matrix and the patch information.
[0149] On the other hand, reference Figure 5B , the decompression device 100 can also be implemented as a chip. Figure 5B , the decompression device 100 is shown as a processor 530 - 2 , and the processor 530 - 2 may receive a plurality of first matrices from the external memory 510 .
[0150] Each of the plurality of decoders in the processor 530 - 2 may decompress the plurality of first matrices and transmit the plurality of second matrices to a processing element (PE) unit (ie, a PE array) included in the processor 530 - 2 .
[0151] As described above, since the original matrix is divided and interleaved during the compression process, the compression rate and accuracy can be improved, and decompression can be performed in parallel by a plurality of decoders, thereby efficiently performing decompression.
[0152] FIG. 7A to FIG. 7C is a diagram for describing a method for decompressing a first matrix according to various embodiments of the present disclosure.
[0153] refer to Fig. 7A , Fig. 7A An example of a 9×5 encoding matrix is shown, and it is shown that nine units of binary data are converted into five units of binary data during compression.
[0154] refer to Figure 7B , the decoder 120 can be implemented with multiple XOR gates, wherein the multiple XOR gates are connected with corresponding input / output terminals based on the coding matrix. However, the decoder 120 is not limited thereto, and can be implemented using a logic circuit different from a plurality of XOR gates, and other methods can be used as long as the operation corresponding to the coding matrix can be performed.
[0155] The decoder 120 may receive a plurality of binary data included in the first matrix in a first number of units, and output the plurality of binary data in a second number of units to be included in the second matrix.
[0156] refer to Figure 7B , the decoder 120 can convert the input x 0 to x 4 Output is O 1 To 9 For example, the decoder 120 may convert 10110 of the first matrix 710 into 001111001, as shown in FIG. Figure 7C shown.
[0157] The processor 130 may change the value of some data of 001111001 based on the patch information. Figure 7C , it is shown that the fourth and seventh data of 001111001 are changed, and 0 may be changed to 1 and 1 may be changed to 0.
[0158] Fig. 8A and Figure 8B is a diagram for describing an update operation of a second matrix according to various embodiments of the present disclosure.
[0159] refer to Fig. 8Aand Figure 8B , the processor 130 may receive a second matrix from each of the plurality of decoders, in which each element is 1 bit, and receive a pruning index matrix from the memory 110. Alternatively, the processor 130 may receive a first pruning index matrix and a second pruning index matrix from the memory 110, and obtain the pruning index matrix from the first pruning index matrix and the second pruning index matrix.
[0160] The processor 130 may identify the pruning elements in the second matrix based on the pruning index matrix, such as Fig. 8A In addition, the processor 130 may convert the identified elements to 0, such as Figure 8B shown.
[0161] Fig. 9 is a diagram for describing a method for merging a plurality of second matrices of a processor according to an embodiment of the present disclosure.
[0162] refer to Fig. 9 , for ease of explanation, the processor 130 will be described as combining two second matrices.
[0163] The decompression apparatus 100 may include a plurality of decoders, and each of the plurality of decoders may transmit the second matrix to the processor 130 .
[0164] like Fig. 9 As shown, the processor 130 may combine the second matrix 1 in which each element is 1 bit and the second matrix 2 in which each element is 1 bit to obtain a combined second matrix in which each element is 2 bits.
[0165] The processor 130 may convert the pruning elements in the combined second matrix, each element of which is 2 bits, into 0 based on the pruning index matrix.
[0166] However, the processor 130 is not limited thereto, and may combine three or more second matrices to obtain a combined second matrix, and convert the trimmed elements to 0.
[0167] Fig.10 is a flowchart describing a control method of a decompression device according to an embodiment of the present disclosure.
[0168] refer to Fig.10 First, at operation S1010, a plurality of logic circuits related to a compression method of compressed data receive compressed data that is decompressed and used in neural network processing of an artificial intelligence model. Furthermore, at operation S1020, a plurality of logic circuits decompress the compressed data and output the decompressed data. Furthermore, at operation S1030, data in a form processable by a neural network is obtained from the data output by the plurality of logic circuits.
[0169] Here, the control method may also include obtaining data in a form processable by a neural network based on the decompressed data and a representative value matrix corresponding to the compressed data, and performing neural network processing using the data in the form processable by the neural network, and the decompressed data and the representative value matrix may be matrices obtained by quantizing the original matrix included in the artificial intelligence model.
[0170] In addition, the control method may further include updating the decompressed data based on a pruning index matrix corresponding to the compressed data, and the pruning index matrix may be a matrix obtained in a pruning process of the original matrix and may be used in a process of obtaining the compressed data.
[0171] Here, the control method may further include changing binary data values of some of the plurality of elements included in the decompressed data based on patch information corresponding to the compressed data, and the patch information may include error information generated in a process of obtaining the compressed data.
[0172] At the same time, the control method may also include obtaining a pruning index matrix based on a first pruning index matrix corresponding to the compressed data and a second pruning index matrix corresponding to the compressed data, and updating the decompressed data based on the pruning index matrix, the pruning index matrix may be a matrix obtained in the pruning process of the original matrix and may be used in the process of obtaining the compressed data, and the first pruning index matrix and the second pruning index matrix may be obtained based on each of the first submatrix and the second submatrix obtained by factorizing the original matrix.
[0173] The decompressed data is a matrix obtained by interleaving an original matrix and then quantizing the interleaved matrix. The control method may further include deinterleaving the neural network processable data in a manner corresponding to the interleaving, and in performing the neural network processing, the neural network processing may be performed using the deinterleaved data.
[0174] On the other hand, in performing the neural network process at operation S1010 , the neural network process may be performed using a plurality of processing elements arranged in a matrix form.
[0175] Furthermore, the decompressed data may be a matrix obtained by dividing an original matrix into a plurality of matrices having the same number of columns and the same number of rows and quantizing one of the plurality of divided matrices.
[0176] On the other hand, the decompression device includes multiple other logic circuits related to the compression method of other compressed data, and the control method may also include receiving other compressed data that is decompressed and used in the neural network processing of the artificial intelligence model, decompressing the other compressed data through multiple other logic circuits and outputting the decompressed other data, obtaining other data in a neural network processable form output from the multiple other logic circuits, and obtaining a matrix in which each element includes multiple binary data by combining the neural network processable data and the other neural network processable data.
[0177] Furthermore, the decompression device may be implemented as one chip.
[0178] FIG. 11A to FIG. 11D is a diagram describing the learning process of an artificial intelligence model according to various embodiments of the present disclosure.
[0179] refer to Fig.11A , Fig.11A is a diagram showing an example of an artificial intelligence model before learning is completed. The artificial intelligence model includes two original matrices W12 and W23, and the compression device 50 can input the value of Li-1 into W12 to obtain the intermediate value of Li, and can input the intermediate value of Li into W23 to obtain the final value of Li+1. However, Fig.11A The AI model is shown very briefly, and can actually include Fig.11A Many matrices.
[0180] refer to Fig. 11B , Fig. 11B is a diagram showing an example of an original matrix included in an artificial intelligence model, which is Figure 3A However, Figure 3A The original matrix of can be in a state where learning is completed, and Fig. 11B The original matrix can be before the learning is completed.
[0181] refer to Fig. 11C , Fig. 11C It is shown that the quantization Fig. 11B A plot of the result for the original matrix is shown.
[0182] The compression device 50 may quantize each of the plurality of elements included in the original matrix to obtain a representative matrix 1120 and a second matrix 1110 including binary quantized data. In this case, different from Figure 3B , the compression device 50 may not perform pruning. However, the compression device 50 is not limited thereto, and may also perform pruning, and a method of performing pruning will be described later.
[0183] refer to Fig.11D , the compression device 50 can compress Fig.11DThe second matrix 1110 on the left side is obtained Fig.11D The number of elements included in the first matrix 1110-1 is smaller than the number of elements included in the second matrix 1110.
[0184] The compression device 50 may be based on Fig. 7A The encoding matrix shown obtains the first matrix 1110-1 from the second matrix 1110. All these operations are included in the learning process of the artificial intelligence model. FIG. 11B to FIG. 11D Only one original matrix has been described in , but in the learning process of the artificial intelligence model, all the multiple original matrices included in the artificial intelligence model can be compressed, such as FIG. 11B to FIG. 11D shown.
[0185] More specifically, for example, Fig.11A W12 can be quantized and compressed based on the first coding matrix and stored as Q 12 ',like Fig.11D In addition, Fig.11A W23 can be quantized and compressed based on the second coding matrix and stored as Q 23 ',like Fig.11D The compression device 50 can encode Q 12 'Decompress and dequantize to obtain W12, and Q can be encoded by the second encoding matrix 23 'Decompress and dequantize to obtain W23. In addition, the compression device 50 can perform a feed-forward operation by using W12 and W23. In this case, the first encoding matrix and the second encoding matrix can be implemented as an XOR gate, such as Figure 7B That is, when Q is decompressed by the first encoding matrix 12 ', the compression device 50 can 12 'Digitized to 0 or 1. In addition, when Q is decompressed by the second encoding matrix 23 ', the compression device 50 can 23 'Digitalized as 0 or 1.
[0186] Thereafter, the compression device 50 may perform a backward operation to update the elements included in the artificial intelligence model. However, since the operation using the XOR gate is an operation of a digital circuit, distinction is impossible, but distinction is required during the update process. Therefore, the compression device 50 may learn the artificial intelligence model by converting the operation using the XOR gate into a distinguishable form, as shown in the following mathematical expression 1. The input value 0 may be converted to -1 and input to the following mathematical expression 1.
[0187] XOR(a, b) = (-1) xtanh(a) xtanh(b) ... Mathematical expression 1
[0188] Mathematical expression 1 shows the case where the input values are a and b, but the input values may not actually be two. The input values may vary depending on the size of the encoding matrix, the number of 1s included in a row, etc. Therefore, the compression device 50 can use a more general mathematical expression (such as the following mathematical expression 2) to learn the artificial intelligence model.
[0189] ...Mathematical expression 2
[0190] Here, X is the input of the XOR gate and m is a variable for adjusting the learning speed, each of which can be expressed as follows. Fig.13 is a graph showing the output according to the value m, and as the gradient changes, the learning speed can vary. The value m can be set by the user.
[0191] X=[x 0 x 1 ...x n-1 ]
[0192] M=[m 0 m 1 ...m n-1 ],m i ∈{0, 1}
[0193] As described above, the compression device 50 can use the operation of the XOR gate analogically for learning an artificial intelligence model. That is, the input value of the XOR gate is stored as a real number, but the compression device 50 converts the negative number among the input values of the XOR gate to zero and converts the positive number to one during the reasoning process. That is, the compression device 50 digitizes the input value of the XOR gate to handle errors caused by using digital circuits such as XOR gates.
[0194] In addition, the compression device 50 can maintain full-precision values in the backward process and update internal variables in a distinguishable form of mathematical expressions. That is, even if an XOR gate is used in the decompression process, the compression device 50 can include an operation according to the XOR gate in the learning process of the artificial intelligence model to perform learning because the compression device 50 uses a distinguishable form of mathematical expressions.
[0195] Meanwhile, the loss value used in the learning process of the artificial intelligence model is expressed by the following Mathematical Expression 3. The compression device 50 can learn the artificial intelligence model through the processing shown in the following Mathematical Expression 3.
[0196]
[0197] ...Mathematical expression 3
[0198] Here, because tanh -1 (x j m j ) is a value between -1 and 1, so is getting closer and closer to zero. That is, as the number of inputs and outputs of the XOR gate increases, learning becomes more difficult. Therefore, the compression device 50 can also learn the artificial intelligence model by using the form of tanh for the distinction of itself (e.g., i) in mathematical expression 3 and converting tanh into a symbol for the distinction of the remaining terms (e.g., j≠i). In this case, regardless of the number of inputs and outputs of the XOR gate, the backward path can be simplified, thereby increasing the learning speed. Fig.14A The case where Mathematical Expression 3 is used is shown, and Fig. 14B The following figure shows how some tanh values of Mathematical Expression 3 are converted into symbols. As learning proceeds, Fig. 14B As shown, the values of 0 and 1 are clearly distinguished and the learning speed can be improved.
[0199] When the learning is completed as described above, the compression device 50 can obtain a plurality of first matrices corresponding to each of a plurality of original matrices included in the artificial intelligence model.
[0200] The compression device 50 can perform learning while including the operation of the XOR gate in the artificial intelligence model as described above, thereby ensuring a high level of compression rate while maintaining the accuracy of the artificial intelligence model. In addition, since the pruning process is omitted and patch information does not need to be used, the processing speed can be improved.
[0201] exist Figure 1 , Figure 2 , Figures 3A to 3D , Figure 4 , Figure 5A and Figure 5B , Figure 6 , FIG. 7A to FIG. 7C , Fig. 8A and Figure 8B , Fig. 9 as well as Fig.10 In the case of , after checking the pruning result, the size of the encoding matrix can be determined. On the other hand, when learning is performed while including the operation of the XOR gate in the artificial intelligence model, the size of the encoding matrix can be arbitrarily specified to set the fractional quantization bit value without pruning. For example, 0.7-bit quantization can also be used instead of the integer number.
[0202] The compression device 50 may use a plurality of encoding matrices corresponding to each of a plurality of original matrices included in the artificial intelligence model. For example, the compression device 50 may use the encoding matrix to perform compression with a relatively low compression rate on the first original matrix and the last original matrix of the artificial intelligence model, and may use the encoding matrix to perform compression with a relatively high compression rate on the remaining original matrices of the artificial intelligence model.
[0203] Even when compression is performed in the method as described above, the decompression device 100 can also Figure 1 The operation shown is decompressed, and a redundant description thereof will be omitted.
[0204] Fig.12 is a diagram describing a method for performing pruning in a learning process according to an embodiment of the present disclosure.
[0205] exist FIG. 11A to FIG. 11D In the present invention, a learning method has been described in which pruning is omitted in the artificial intelligence model and quantization using an XOR gate is included, but the compression device 50 can learn the artificial intelligence model by including pruning and quantization using an XOR gate in the artificial intelligence model.
[0206] Fig.13 is a graph for describing the influence of a value m according to an embodiment of the present disclosure.
[0207] Fig.14A and Fig. 14B is a diagram for describing improvement of learning speed according to various embodiments of the present disclosure.
[0208] refer to Fig.12 , Fig.13 , Fig.14A and Fig. 14B , the compression device 50 can use two XOR gates per weight to reflect the pruning. For example, the compression device 50 can obtain the final output from the two outputs (w, m) of the XOR gate. Here, w is the reference FIG. 11A to FIG. 11D The output of the XOR gate described above, and m can be a value used to reflect the trimming. For example, if (0,0), the compression device 50 can output -1, if (0,1) or (1,0), it outputs 0, and if (1,1), it outputs +1.
[0209] The compression device 50 can output three values from w and m through an equation such as (w+m) / 2, and to do this, the compression device 50 converts the input value into -1 when the input value is 0, and converts the input value into +1 when the input value is 1, and inputs the converted value into the equation.
[0210] The compression device 50 is referred to FIG. 11A to FIG. 11D The described method performs learning of w, and its redundant description will be omitted.
[0211] When the value w is a threshold value or less, the compression device 50 may set m to a value opposite to the value w, and finally convert w to 0 and output the result. Alternatively, when the value w exceeds the threshold value, the compression device 50 may set m to a value having the same sign as the value w, and finally convert w to +1 or -1 and output the result. In this way, a pruning effect achieved by converting a value w of a threshold value or less to 0 can be obtained.
[0212] When learning is completed as described above, the compression device 50 can obtain a plurality of first matrices corresponding to each of a plurality of original matrices included in the artificial intelligence model. Figure 1 , Figure 2 , Figures 3A to 3D , Figure 4 , Figure 5A and Figure 5B , Figure 6 , FIG. 7A to FIG. 7C , Fig. 8A and Figure 8B , Fig. 9 as well as Fig.10 , no separate pruning index matrix is generated.
[0213] As described above, the compression device 50 may include quantization using pruning and XOR gates in the artificial intelligence model to perform learning of the artificial intelligence model, thereby facilitating learning and improving accuracy.
[0214] Even when compression is performed in the method as described above, the decompression device 100 can also Figure 1 The operation shown is decompressed, and a redundant description thereof will be omitted.
[0215] FIG. 11A to FIG. 11D , Fig.12 , Fig.13 , Fig.14A and Fig. 14B The operation of the compression device 50 or the operation of the decompression device 100 can be performed by an electronic device such as a mobile device or a desktop PC. For example, the memory of the first electronic device can store the artificial intelligence model before the learning is completed and the sample data necessary for the learning process, and the processor of the first electronic device can be used with FIG. 11A to FIG. 11D , Fig.12 , Fig.13 , Fig.14A and Fig. 14B The data stored in the memory is learned in the same way as in the previous example, and the data is compressed at the same time.
[0216] In addition, the memory of the second electronic device can store the artificial intelligence model that has completed learning and compression, and the processor of the second electronic device can be like Figure 2The quantization acquisition matrix unit processes the data stored in the memory to decompress the data.
[0217] As described above, according to various embodiments of the present disclosure, the decompression device may perform a neural network process by decompressing a matrix using a decoder implemented using a plurality of logic circuits and obtaining a restored matrix from the decompressed matrix.
[0218] Meanwhile, according to the embodiments of the present disclosure, the various embodiments described above can be implemented by software including instructions stored in a machine (e.g., computer) readable storage medium. A machine is a device that calls stored instructions from a storage medium and can operate according to the called instructions, and may include an electronic device (e.g., electronic device A) according to the disclosed embodiments. When the instructions are executed by a processor, the processor may directly or use other components under the control of the processor to perform the function corresponding to the instructions. The instructions may include code generated or executed by a compiler or interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory" means that the storage medium does not include a signal and is tangible, but does not distinguish whether the data is stored in the storage medium semi-permanently or temporarily.
[0219] In addition, according to an embodiment of the present disclosure, the method according to the above various embodiments may be included and provided in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or may be downloaded through an application store (e.g., PlayStore). TM ) Online distribution. In the case of online distribution, at least a part of the computer program product may be at least temporarily stored in a storage medium (such as a memory of a manufacturer's server, an application store's server, or a relay server), or temporarily generated.
[0220] In addition, according to the embodiments of the present disclosure, the above various embodiments can be implemented in a computer or similar device readable recording medium using software, hardware or a combination thereof. In some cases, the embodiments described in the present disclosure can be implemented by the processor itself. According to the software implementation, the embodiments such as processes and functions described in the present disclosure can be implemented as separate software modules. Each software module can perform one or more functions and operations described in the present disclosure.
[0221] At the same time, computer instructions for performing the processing operations of the device according to the various embodiments described above may be stored in a non-transitory computer-readable medium. When executed by a processor of a specific device, the computer instructions stored in the non-transitory computer-readable medium allow the specific device to perform the processing operations of the device according to the various embodiments described above. Non-transitory computer-readable media are not media for short-term storage of data such as registers, cache memories, memories, etc., but are machine-readable media that semi-permanently store data. Specific examples of non-transitory computer-readable media may include compact discs (CDs), digital versatile discs (DVDs), hard disks, Blu-ray discs, universal serial buses (USBs), memory cards, read-only memories (ROMs), etc.
[0222] In addition, each component (e.g., module or program) according to the various embodiments described above may include a single entity or multiple entities, and some of the subcomponents described above may be omitted, or other subcomponents may be further included in various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into one entity to perform the same or similar functions performed by the various components prior to integration. The operations performed by modules, programs or other components according to the various embodiments may be performed in a sequential, parallel, iterative or heuristic manner, or at least some operations may be performed or omitted in a different order, or other operations may be added.
[0223] While the present disclosure has been shown and described with reference to various embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present disclosure as defined by the appended claims and their equivalents.
Claims
1. A decompression device, comprising: a memory configured to store compressed data and patch information corresponding to the compressed data, wherein the compressed data is to be decompressed and used in a neural network processing of an artificial intelligence model, and the patch information includes error information generated in a process of obtaining the compressed data; A decoder corresponding to the encoding matrix and comprising a plurality of logic circuits related to the compression method of the compressed data, wherein the decoder is configured to: decompressing the compressed data through the plurality of logic circuits based on the input of the compressed data, and outputting the decompressed data; and Processor, configured as: based on the patch information, changing a binary data value of at least one element of a plurality of elements included in the decompressed data output from the decoder so that the binary data value is determined to minimize the number of errors, and Based on the decompressed data and the representative value matrix corresponding to the compressed data, data in a form processable by a neural network is obtained; wherein, using the encoding matrix, a first number of elements included in the decompressed data are compressed into a second number of elements included in the compressed data, The second number is smaller than the first number.
2. The decompression device according to claim 1, wherein: The memory is further configured to store the representative value matrix, Wherein, the processor is further configured to: obtaining data in a form processable by the neural network based on the decompressed data, and The neural network processing is performed using the data in a form processable by the neural network.
3. The decompression device according to claim 2, in, The memory is further configured to store a pruning index matrix corresponding to the compressed data, The processor is further configured to update the decompressed data based on the pruning index matrix, The pruning index matrix includes a matrix obtained during the pruning process of the original matrix, and The pruning index matrix is used in the process of obtaining the compressed data.
4. The decompression device according to claim 2, wherein: The memory is further configured as: storing a first pruning index matrix corresponding to the compressed data, and storing a second pruning index matrix corresponding to the compressed data, Wherein, the processor is further configured to: obtaining a pruning index matrix based on the first pruning index matrix and the second pruning index matrix, and updating the decompressed data based on the pruning index matrix, Wherein, the pruning index matrix includes a matrix obtained in the pruning process of the original matrix, wherein the pruning index matrix is used in the process of obtaining the compressed data, and The first pruning index matrix and the second pruning index matrix are respectively obtained based on each of a first sub-matrix and a second sub-matrix obtained by factorizing the original matrix.
5. The decompression device according to claim 2, wherein: The decompressed data includes a matrix obtained by interleaving an original matrix and then quantizing the interleaved matrix, and Wherein, the processor is further configured to: deinterleaving the data in a form processable by the neural network in a manner corresponding to the interleaving, and The neural network processing is performed using the deinterleaved data.
6. The decompression device according to claim 2, in, The processor includes a plurality of processing elements arranged in a matrix, and Wherein, the processor is further configured to use the plurality of processing elements to perform the neural network processing.
7. The decompression device according to claim 2, wherein the decompressed data comprises: A matrix obtained by dividing an original matrix into multiple matrices having the same number of columns and rows and quantizing one of the divided multiple matrices.
8. The decompression device according to claim 1, wherein: The memory is further configured to store other compressed data to be decompressed and used in neural network processing of the artificial intelligence model, Wherein, the decompression device further includes another decoder, and the another decoder is configured as follows: including a plurality of other logic circuits related to the compression method of the other compressed data, decompressing the other compressed data through the plurality of other logic circuits based on input of the other compressed data, and Output other decompressed data, and Wherein, the processor is further configured to: obtaining from said other decompressed data output by said other decoder other data in a form processable by a neural network, and A matrix including a plurality of binary data in each element is obtained by coupling the data processable by the neural network and other data in a form processable by the neural network.
9. The decompression device of claim 1, wherein the decompression device is implemented as a chip.
10. A control method for a decompression device, the decompression device comprising a plurality of logic circuits related to a compression method of compressed data, the control method comprising: receiving, by the plurality of logic circuits, the compressed data and patch information corresponding to the compressed data, wherein the compressed data is to be decompressed and used in a neural network processing of an artificial intelligence model, and the patch information includes error information generated in a process of obtaining the compressed data; decompressing the compressed data by the plurality of logic circuits corresponding to the encoding matrix; Output decompressed data; based on the patch information, changing a binary data value of at least one element of a plurality of elements included in the decompressed data output from the plurality of logic circuits so that the binary data value is determined to minimize the number of errors, and Based on the decompressed data and the representative value matrix corresponding to the compressed data, data in a form processable by a neural network is obtained; wherein, using the encoding matrix, a first number of elements included in the decompressed data are compressed into a second number of elements included in the compressed data, The second number is smaller than the first number.
11. The control method according to claim 10, further comprising: Based on the decompressed data, obtaining data in a form processable by the neural network; as well as The neural network processing is performed using the data in a form processable by the neural network.
12. The control method according to claim 11, further comprising: updating the decompressed data based on a pruning index matrix corresponding to the compressed data, The pruning index matrix includes a matrix obtained during the pruning process of the original matrix, and The pruning index matrix is used in the process of obtaining the compressed data.
13. The control method according to claim 11, further comprising: Obtaining a pruning index matrix based on a first pruning index matrix corresponding to the compressed data and a second pruning index matrix corresponding to the compressed data; as well as updating the decompressed data based on the pruning index matrix, Wherein, the pruning index matrix includes a matrix obtained in the pruning process of the original matrix, wherein the pruning index matrix is used in the process of obtaining the compressed data, and The first pruning index matrix and the second pruning index matrix are respectively obtained based on each of a first sub-matrix and a second sub-matrix obtained by factorizing the original matrix.
Citation Information
Patent Citations
Light source unit, and display and lighting devices using it
KR1020190060991A
Odorous susbstances removal system
KR1020190117081A
Method and device for modelling and producing implant for orbital wall
KR1020190140720A