Model lossless compression method, model decompression method and device
By analyzing the data distribution of the large model parameter matrix, determining the appropriate number of coded bits and encoding it, the problem of deploying large models on resource-limited hardware is solved, and lossless compression and resource saving of model parameters are achieved.
Patent Information
- Application Number
- CN202410946976.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-15
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-07-15
AI Technical Summary
Large models occupy a large amount of computing and storage resources during training and inference, making it difficult to deploy on resource-limited hardware and limited performance optimization.
By analyzing the data distribution of the model parameter matrix, determining the number of coded bits smaller than the original number of digits, encoding the parameter matrix, generating encoding results and encoding tables, and realizing lossless compression of model parameters.
Without losing model accuracy, the storage space and memory access overhead of the parameter matrix are saved, and the resource requirements for model deployment and inference are reduced.
Smart Images

Figure CN118860285B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a model lossless compression method, a model decompression method and a device. Background Art
[0002] In recent years, with the rapid development of computer technology and big data, large models based on deep learning have achieved remarkable results in various fields. A large model refers to a deep learning model with tens of millions or even hundreds of millions of parameters. It uses a large amount of data and computing resources to train a neural network model with a large number of parameters. By continuously adjusting the model parameters, the model can achieve the best performance in various tasks. The "big" characteristics of large models are reflected in: a large number of parameters, a large amount of training data, and high computing resource requirements. The increasing number of model parameters also makes its generalization performance and accuracy better and better. On the other hand, the "big" characteristics also bring some problems to the implementation and application of technology. For example, a large amount of computing and storage resources will be occupied during model training and inference, making it difficult to deploy on hardware with limited resources, and its performance optimization is also greatly challenged.
[0003] In order to solve the storage and computing problems caused by too many model parameters while ensuring model performance, model compression technology is a commonly used and important method that can reduce the size of deep learning models, improve operating efficiency, and reduce deployment costs. Model compression technologies mainly include: pruning, quantization, knowledge distillation, low-rank decomposition, parameter sharing, neural network architecture search, etc. Although these methods change the complexity or structure of the existing model to a certain extent, they will also cause varying degrees of model accuracy loss during the model compression process. Summary of the invention
[0004] The present application provides a model lossless compression method, which can achieve lossless compression of model parameters.
[0005] This application provides the following solutions:
[0006] According to a first aspect, a model lossless compression method is provided, the method comprising: obtaining parameter matrices corresponding to one or more network layers respectively contained in the model; determining data distribution of the parameter matrices respectively corresponding to the one or more network layers; determining a number of coding bits using the data distribution, the number of coding bits being less than the number of original bits used for the parameters in the parameter matrix; encoding the parameter matrices respectively corresponding to the one or more network layers according to the number of coding bits, to obtain encoding results of the parameter matrices corresponding to each network layer; storing compressed model data, the compressed model data comprising encoding results of the parameter matrices corresponding to each network layer and an encoding table, the encoding table comprising a mapping relationship between the parameter matrix and the encoding result.
[0007] According to an achievable method in an embodiment of the present application, the data distribution includes a unique numerical quantity in the parameter matrix, and the number of coding bits is determined using the unique numerical quantity; or, the data distribution includes the probability of occurrence of the parameters in the parameter matrix, and the number of coding bits is negatively correlated with the probability of occurrence of the parameters.
[0008] According to an implementable method in an embodiment of the present application, the use of the unique numerical quantity to determine the number of coding bits includes: determining multiple block sizes; respectively determining the unique numerical quantity in each block obtained by dividing the parameter matrix into blocks using each block size and the number of coding bits required; using the unique numerical quantity in each block corresponding to each block size and the number of coding bits required, select a block size from the multiple block sizes and determine the number of coding bits required.
[0009] According to an implementable method in an embodiment of the present application, the parameter matrices corresponding to the one or more network layers are respectively encoded according to the number of coding bits, and the encoding results of the parameter matrices corresponding to each network layer are obtained, including: the parameter matrices corresponding to the one or more network layers are respectively blocked using a selected block size to obtain multiple matrix blocks; each matrix block is respectively encoded using a determined number of coding bits to obtain the encoding results corresponding to each matrix block and the encoding table corresponding to each matrix block as the compressed model data.
[0010] According to an implementable method in the embodiment of the present application, the encoding result includes identification information indicating the network layer corresponding to the matrix block; or, the encoding result is stored in a storage space according to the corresponding network layer.
[0011] According to an implementable method in an embodiment of the present application, the encoding result and the encoding table are stored in a global memory.
[0012] According to the second aspect, a model decompression method is provided, the method comprising: obtaining compressed model data, the compressed model data comprising encoding results of parameter matrices corresponding to each network layer contained in the model and a coding table, the coding table comprising a mapping relationship between the parameter matrix and the encoding results; using the coding table, respectively decoding the encoding results of the parameter matrices corresponding to each network layer, and obtaining the parameter matrices corresponding to each network layer for reading by an inference execution unit.
[0013] According to an implementable method in an embodiment of the present application, the encoding results of the parameter matrices corresponding to each network layer and the encoding table include: the encoding results of each matrix block and the corresponding encoding table, the matrix block is obtained by dividing the parameter matrix corresponding to each network layer into blocks; using the encoding table, the encoding results of the parameter matrices corresponding to each network layer are decoded respectively to obtain the parameter matrices corresponding to each network layer, including: using the encoding table corresponding to each matrix block, respectively decoding the encoding results of each matrix block to obtain data of each matrix block; according to the network layer corresponding to each matrix block, the data of the parameter matrix corresponding to each network layer is obtained.
[0014] According to an implementable method in the embodiment of the present application, the network layer corresponding to each matrix block is determined based on the identification information contained in each encoding result or the storage position of the encoding result in the storage space.
[0015] According to an achievable method in an embodiment of the present application, obtaining the compressed model data includes: obtaining the encoding results of the parameter matrices corresponding to each network layer contained in the model from the global memory and loading them into the first register; obtaining the encoding table from the global memory and loading it into the cache; using the encoding table, respectively decoding the encoding results of the parameter matrices corresponding to each network layer, including: querying the encoding table in the cache, using the encoding table to respectively decode the encoding results of the parameter matrices corresponding to each network layer, and storing the decoding results in the second register, and the second register is read by the inference execution unit.
[0016] According to a third aspect, a model lossless compression device is provided, the device comprising: a parameter acquisition unit, acquiring parameter matrices corresponding to one or more network layers respectively contained in the model; a data analysis unit, determining data distribution of the parameter matrices respectively corresponding to the one or more network layers; a coding bit determination unit, determining the number of coding bits using the data distribution, the number of coding bits being less than the original number of bits used for the parameters in the parameter matrix; an encoding unit, encoding the parameter matrices respectively corresponding to the one or more network layers according to the number of coding bits, to obtain encoding results of the parameter matrices corresponding to each network layer; a data storage unit, storing compressed model data, the compressed model data comprising encoding results of the parameter matrices corresponding to each network layer and an encoding table, the encoding table comprising a mapping relationship between the parameter matrix and the encoding result.
[0017] According to a fourth aspect, a model decompression device is provided, the device comprising: a data acquisition unit, acquiring compressed model data, the compressed model data comprising encoding results of parameter matrices corresponding to each network layer contained in the model and a coding table, the coding table comprising a mapping relationship between the parameter matrix and the encoding results; a decoding unit, utilizing the coding table to respectively decode the encoding results of the parameter matrices corresponding to each network layer, and obtain the parameter matrices corresponding to each network layer for reading by an inference execution unit.
[0018] According to a fifth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any one of the first and second aspects are implemented.
[0019] According to a sixth aspect, there is provided an electronic device, comprising:
[0020] one or more processors; and
[0021] A memory associated with the one or more processors, the memory being used to store program instructions, wherein when the program instructions are read and executed by the one or more processors, the steps of the method described in any one of the first and second aspects are executed.
[0022] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0023] 1) The present application uses the data distribution of the parameter matrix to determine the number of coding bits, which is less than the original number of bits used by the parameters in the parameter matrix, and uses the number of coding bits to encode the parameter matrix to obtain the encoding result of the parameter matrix and the encoding table. The present application compresses the model parameters without losing the model accuracy, saving the storage space and memory access overhead of the parameter matrix.
[0024] 2) The present application determines the number of coding bits according to the unique numerical value in the parameter matrix, or determines the number of coding bits according to the probability of occurrence of the parameters in the parameter matrix, and performs fixed-length or variable-length encoding on the parameter matrix, thereby reducing the storage space of the parameter matrix and achieving lossless compression of the model parameters.
[0025] 3) The present application analyzes the unique numerical quantities in each block under different block sizes and the required number of coding bits, and uses the analysis results to select a suitable block size from multiple block sizes and determine the number of coding bits required. When determining the number of coding bits, the number can be weighed against the block size to make the determined number of coding bits more reasonable.
[0026] 4) This application further reduces the required number of coding bits and improves the compression effect of the model by dividing the model parameter matrix into blocks and encoding the divided matrix blocks.
[0027] 5) By including identification information indicating the network layer corresponding to the matrix block in the encoding result, or storing the encoding result in a storage space according to the corresponding network layer, the present application can quickly and accurately locate the corresponding coding table during decoding, thereby improving the efficiency and accuracy of decoding.
[0028] 6) The present application stores the encoding result and the encoding table in the global memory, loads the encoding table into the cache during decoding, loads the encoding result of the parameter matrix from the global memory to the first register, decodes the encoding result respectively using the encoding table in the cache, and stores the decoding result in the second register for reading by the inference execution unit. The decoding process of the present application can be completed in the register, without writing the result back to the video memory, thereby improving the decoding efficiency and saving the memory access overhead.
[0029] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0031] Figure 1 A system architecture diagram applicable to the embodiments of the present application.
[0032] Figure 2 A flowchart of a model lossless compression method provided in an embodiment of the present application.
[0033] Figure 3 Schematic diagram of the method for representing the number of bits of half-precision floating-point numbers.
[0034] Figure 4 This is a distribution diagram showing the relationship between the uniform distribution of half-precision floating-point numbers and the actual numerical mapping distribution.
[0035] Figure 5a Data distribution diagram of the parameter matrix of layer0.
[0036] Figure 5b This is the data distribution diagram of the parameter matrix of layer1.
[0037] Figure 6 A flowchart of a model decompression method provided in an embodiment of the present application.
[0038] Figure 7 A schematic diagram of the decoding process provided in an embodiment of the present application.
[0039] Figure 8 A schematic block diagram of a model lossless compression device provided in an embodiment of the present application.
[0040] Fig. 9 A schematic block diagram of a model decompression device provided in an embodiment of the present application.
[0041] Fig.10 A schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.
[0043] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.
[0044] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0045] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0046] In the process of model reasoning, memory access capability refers to the ability of the processor in a computer system to access memory (such as RAM, hard disk, etc.). Memory access performance is one of the important performance limitations. Generally speaking, there is an inherent imbalance between the theoretical peak performance and memory access bandwidth of a chip. Ideally, the computational access-to-memory ratio should reach tens or hundreds to achieve a balance. However, the reality is that many high-performance operations involve loop-based processing that requires a large amount of data transfer, and the actual computational access-to-memory ratio is far lower than the ideal situation, which directly results in inefficient use of chip resources and reduced program performance.
[0047] The main source of memory consumption is the access to the parameter matrices of each network layer of the model, which are the main content of the model file. Since data access is in a streaming manner during the inference decoding process, with almost no temporal and spatial locality, the memory pressure cannot be solved by increasing the utilization of cache or shared memory. Therefore, reducing the memory space occupied by the above parameter matrix data becomes the primary optimization point to consider.
[0048] At present, there are some model compression technologies, such as pruning, low-rank decomposition, quantization and other methods. These methods change the complexity or structure of the existing model to a certain extent, which will affect the model accuracy to varying degrees. In the existing approximate invention technologies, such as the method of introducing index matrix and reference matrix, compressing the differential data of model parameters, etc., there are also different degrees of model accuracy loss in the compression process.
[0049] In view of this, the present application provides a new idea. In order to facilitate the understanding of the present application, the system architecture on which the present application is based is first described. Figure 1 An exemplary system architecture to which the embodiments of the present application can be applied is shown. Figure 1 As shown in , the system architecture may include a user device and a model lossless compression device and a model decompression device located on the server side.
[0050] User devices may include, but are not limited to, smart mobile terminals, smart home devices, wearable devices, PCs (Personal Computers), etc. Smart mobile devices may include mobile phones, tablet computers, laptops, PDAs (Personal Digital Assistants), Internet cars, etc. Smart home devices may include smart TVs, smart refrigerators, etc. Wearable devices may include smart watches, smart glasses, virtual reality devices, augmented reality devices, mixed reality devices (i.e., devices that can support virtual reality and augmented reality), etc.
[0051] The model lossless compression device can adopt the method provided in the embodiment of the present application to encode the parameter matrix to obtain the encoding result and the encoding table.
[0052] The model decompression device can use the method provided in the embodiment of the present application to decode the compressed model parameters for use in large model reasoning.
[0053] The model lossless compression device, model decompression device and large model can be set up as an independent server, or in a server group, or in a cloud server. A cloud server, also known as a cloud computing server or cloud host, is a host product in the cloud computing service system to solve the defects of difficult management and weak service scalability in traditional physical hosts and virtual private servers (VPS). Figure 1 In addition to the shown architecture, the model parameter encoding device, the model parameter decoding device and the large model can also be set in a computer terminal with strong computing power.
[0054] As one of the feasible ways, the model lossless compression device can pre-encode the parameter matrix of the model and store the compressed model data, and the compressed model data includes the encoding results and encoding tables of the parameter matrices corresponding to each network layer. When the user needs to use the large model for reasoning, the user can input the requirement information through the user device, and the user device sends the requirement information to the server. The large model reasoning execution unit uses the model data of the large model obtained after decompression to perform reasoning and returns the reasoning result to the user device.
[0055] It should be understood that Figure 1 The user equipment, model lossless compression device, and model decompression device in the figure are only illustrative. According to the implementation requirements, there may be any number of user equipment, model lossless compression device, and model decompression device.
[0056] Figure 2 A flow chart of a model lossless compression method provided in an embodiment of the present application. The method can be performed by Figure 1 The model lossless compression device in the system shown is implemented. Figure 2 As shown in , the method may include the following steps:
[0057] Step 201, obtaining parameter matrices corresponding to one or more network layers included in the model.
[0058] Step 202: determine the data distribution of the parameter matrices corresponding to more than one network layer.
[0059] Step 203, determining the number of coding bits using data distribution, where the number of coding bits is smaller than the original number of bits used by the parameters in the parameter matrix.
[0060] Step 204 , encoding the parameter matrices corresponding to more than one network layer respectively according to the number of coding bits, to obtain the coding results of the parameter matrices corresponding to each network layer.
[0061] Step 205, storing the compressed model data, the compressed model data including the encoding results of the parameter matrices corresponding to each network layer and the encoding table, the encoding table including the mapping relationship between the parameter matrices and the encoding results.
[0062] As can be seen from the above process, the present application uses the data distribution of the parameter matrix to determine the number of coding bits, which is less than the original number of bits used by the parameters in the parameter matrix, and uses the coding bit number to encode the parameter matrix to obtain the encoding result of the parameter matrix and the encoding table. The present application compresses the model parameters without losing the model accuracy, saving the storage space and memory access overhead of the parameter matrix.
[0063] The following describes in detail each step in the above process and the effects that can be further produced in conjunction with the embodiments.
[0064] First, the above step 201, namely "obtaining parameter matrices corresponding to more than one network layers included in the model", is described in detail in conjunction with the embodiment.
[0065] In deep learning, the model usually consists of multiple network layers, each of which contains a set of specific parameter matrices that define the connection strength between layers. The parameter matrix is learned during the model training process and is used to perform forward calculations to generate prediction results when given input data. As the model size increases, the number of parameters will also increase significantly, requiring a large amount of data and computing resources. Use its parallel processing capabilities to speed up the model's calculation process.
[0066] The model lossless compression method of the present application is to compress the parameter matrix of the model. Before compressing the parameter matrix, it is necessary to obtain the parameter matrix corresponding to each network layer of the model. The parameter matrix may be stored in the video memory or other dedicated hardware, and the parameter matrix can be read according to the storage location of the parameter matrix. In the process of lossless compression of the model, the parameter matrices of all network layers of the model can be compressed, or only the parameter matrices of certain specific layers can be compressed. Those skilled in the art can obtain the parameter matrix required for compression according to actual needs.
[0067] The model lossless compression method of the present application is performed on the parameter matrix obtained from model training. It does not require intervention in the model training process, is easy to handle, and does not generate training computing power consumption.
[0068] The above step 202, namely "determining the data distribution of the parameter matrices respectively corresponding to more than one network layer", is described in detail below in conjunction with an embodiment.
[0069] The model parameter matrix is composed of a large number of parameters, and the parameter data distribution has certain rules. This application formulates a suitable encoding scheme by analyzing the data distribution rules in the parameter matrix. The data distribution of the parameter matrix can be statistically analyzed from multiple angles. For example, the probability of occurrence of parameters in the parameter matrix, the range of values of the parameters, the symmetry or concentration of the parameter distribution, the unique number value in the parameter matrix, etc. can be analyzed. Among them, the unique number value is obtained by deduplicating the parameters in the parameter matrix and counting the number of parameters.
[0070] The data distribution of the model parameter matrix is also related to the encoding method used. The parameter matrix values may be encoded in a variety of ways, such as floating-point encoding, half-precision floating-point encoding, fixed-point encoding, etc. LLaMA (Large Language Model Meta AI) is a series of large language models developed by Meta. The LLaMA model family includes models of different sizes, designed to perform various natural language processing tasks. Taking the LLaMA2-7b model in the LLaMA model family as an example, the data format of its parameter matrix is FP16 (16-bit floating-point, half-precision floating point). Figure 3 This is a schematic diagram of the method for representing the number of bits of a half-precision floating point number, such as Figure 3 As shown, the floating point representation method of FP16 is: a 16-bit value is used to represent a real number, including 1 sign bit, 5 exponent bits, and 10 mantissa bits.
[0071] According to the encoding principle of FP16, the effective value range of FP16 is about 10 -4 to 10 4 The floating point number expression precision is not uniformly distributed. Its uniform distribution and actual numerical mapping distribution are as follows: Figure 4 As shown in the figure, the horizontal axis is the average value size, and the vertical axis is the value size actually expressed by FP16. It can be seen that the closer the number represented by the floating point number is to the extreme value range, the lower the fidelity.
[0072] Taking the LLaMA 2-7b model as an example, the parameter matrix Wq (matrix size: 4096*4096, data format: FP16) of its layer0 and layer1 is statistically analyzed, and the parameter matrix data distribution is obtained as follows: Figure 5a and Figure 5b As shown, Figure 5aThis is the statistical result of the parameter matrix data distribution of layer0. The parameter distribution range of the parameter matrix is [-0.7734357, 0.72265625], and the unique number value is 5753; Figure 5b This is the statistical result of the parameter matrix data distribution of layer1. The parameter distribution range of the parameter matrix is [-0.40820312, 0.47265625], and the unique number value is 5270.
[0073] The above step 203, namely "determining the number of coding bits by using data distribution, where the number of coding bits is smaller than the original number of bits used by the parameters in the parameter matrix", is described in detail below in conjunction with an embodiment.
[0074] The present application reduces the storage space required for the parameter matrix by encoding the parameter matrix, so the number of encoding bits must be less than the original number of bits used for the parameters in the parameter matrix. Among them, the number of encoding bits is determined using data distribution. For a model, the data distribution of all parameter matrices of the model can be statistically analyzed, and a number of encoding bits can be jointly determined based on the statistical results for encoding all parameter matrices in the model. Different encoding bits can also be used for different parameter matrices based on the data distribution of different parameter matrices.
[0075] As an implementable method, the number of coding bits can be determined according to the probability of occurrence of the parameters in the parameter matrix, and the number of coding bits is negatively correlated with the probability of occurrence of the parameters. Specifically, a variable-length coding method is used, and a shorter code is used for parameters with a higher probability of occurrence, and a longer code is used for parameters with a lower probability of occurrence, which reduces the average length and expected value of the encoded string, thereby achieving the purpose of lossless data compression.
[0076] As another feasible method, the number of coding bits can be determined by using unique numerical quantities. The statistical results of Figure 5 show that in a 16-bit coded half-precision parameter matrix, there are approximately 6000 unique numerical quantities, and the coding expression utilization rate is less than 10% (6000 / 2^16), and there is obviously redundant coding space. After obtaining the unique numerical quantity of the parameter matrix, the present application can select the minimum number of bits that can represent the unique numerical quantity as the number of coding bits. Taking the LLaMA 2-7b model as an example, after statistics of each parameter matrix in the model, it can be obtained that: the maximum unique numerical quantity is 6452, then it can be considered to use 13 bits for encoding for all parameter matrices of the model. The use of unique numerical quantities to determine the number of coding bits in this embodiment is a fixed-length coding method, which can give play to the advantages of fixed-length coding such as compact storage, fast access speed, and simple data processing.
[0077] Furthermore, the number of coding bits can be further reduced by dividing the parameter matrix into blocks. Specifically, multiple block sizes are determined; the unique numerical quantities in each block obtained by dividing the parameter matrix into blocks using each block size and the number of coding bits required are determined respectively; using the unique numerical quantities in each block corresponding to each block size and the number of coding bits required, a block size is selected from multiple block sizes and the number of coding bits required is determined.
[0078] The multiple block sizes determined may be several pre-set block sizes, or may be block sizes calculated according to a preset rule or algorithm.
[0079] In this embodiment, the number of coding bits is determined by weighing the block size and the required number of coding bits. Specifically, several feasible block sizes can be selected in advance, and the unique numerical quantities under each block size can be counted to determine the required number of coding bits. Under the acceptable number of coding bits, a larger block size can be selected for block division to obtain a lower number of coding bits while ensuring a larger block size. Taking the LLaMA 2-7b model as an example, the comparison of block size, unique numerical quantity, and required number of coding bits is shown in Table 1. It can be seen from Table 1 that when the block size is 512K, 256K and 128K, the unique numerical quantity of each block is about 3000. In this case, the block size of 512K is preferably selected for 12-bit encoding. Compared with the original 16-bit encoding, this embodiment can achieve a compression rate of about 25%.
[0080] Table 1: Example block size trade-off comparison table
[0081] Block size Unique numeric value Number of bits required 512K About 3000 12bit 256K About 3000 12bit 128K About 3000 12bit
[0082] The above step 204, namely "encoding the parameter matrices corresponding to more than one network layer respectively according to the number of coding bits to obtain the encoding results of the parameter matrices corresponding to each network layer" is described in detail below in conjunction with the embodiments.
[0083] After determining the number of coding bits, the parameters in the parameter matrix are encoded according to the number of coding bits to obtain the encoding result of the parameter matrix, so that the number of bits of the encoding result is consistent with the number of coding bits. Each parameter matrix can be directly encoded, and the parameter matrix obtained after encoding is used as the encoding result.
[0084] The parameter matrix can also be divided into blocks and each matrix block can be encoded. Specifically, the parameter matrices corresponding to more than one network layer are divided into blocks using the selected block size to obtain multiple matrix blocks; each matrix block is encoded using the determined number of coding bits to obtain the encoding results corresponding to each matrix block and the encoding table corresponding to each matrix block as compressed model data.
[0085] In the embodiment of the present application, there is no limitation on the specific encoding method, and methods such as Huffman encoding, Shannon-Fernaud encoding, arithmetic encoding, etc. may be used, the purpose of which is to map the floating point numbers used by the model matrix to a coding space with a smaller number of bits. For example, compressing the original 16-bit floating point number to a 12-bit coding space.
[0086] When encoding, single-byte encoding, double-byte encoding, multi-byte encoding, etc. can be used.
[0087] The above step 205, i.e., "storing the compressed model data, wherein the compressed model data includes the encoding results of the parameter matrices corresponding to each network layer and the encoding table, wherein the encoding table includes the mapping relationship between the parameter matrices and the encoding results" is described in detail below in conjunction with the embodiments.
[0088] When encoding the parameter matrix, a coding table is created according to the principle of coding. The coding table includes the mapping relationship between the parameter matrix and the coding result. The coding table is used to decode the coding result in the model inference stage. A coding table can be built for all parameter matrices of the entire model, or a corresponding coding table can be created for each parameter matrix. When the parameter matrix is block-coded, a coding table can be built for each matrix block.
[0089] Since the model contains a large number of network layers, there are multiple encoding results after the parameter matrix of the network layer is encoded. Therefore, in order to identify the corresponding relationship between the network layer and the encoding result of the parameter matrix, the identification information of the corresponding network layer can be set in the encoding result. When the model is decompressed in the subsequent embodiments, the network layer corresponding to the encoding result can be determined according to the identification information in the encoding result. Alternatively, the encoding result is stored in the storage space according to the corresponding network layer, for example, the storage address is distinguished according to different network layers. When the model is decompressed in the subsequent embodiments, the encoding result corresponding to the network layer can be obtained from the storage address corresponding to the network layer.
[0090] Similarly, when the parameter matrix is encoded in blocks, the encoding result includes identification information indicating the network layer corresponding to the matrix block; or, the encoding result is stored in the storage space according to the corresponding network layer. The encoding results of each block corresponding to the same network layer can be stored in sequence, so that when the model is decompressed during inference in subsequent embodiments, the encoding results of each block corresponding to the same network layer can be obtained in sequence to decompress and obtain the parameter matrix of the network layer.
[0091] The compressed model data of the present application can be stored in a variety of storage units, can be stored in a local hard disk, or can be stored in a cloud storage server. The encoding result and the encoding table can be stored in the same location or in different locations.
[0092] As a preferred embodiment, the encoding result and the encoding table can be stored in the global memory, such as the video memory, which can effectively save the space of the global memory and the bandwidth consumption of accessing the global memory, and can effectively improve the computing memory access ratio.
[0093] Taking the Wq parameter matrix of layer0 in the LLaMA 2-7b model as an example, the 16-bit coded data is compressed into 12-bit coded data. The encoded 12-bit weight matrix data occupies 24M (4096*4096*12bit) Global Memory (32M before encoding). The memory space occupied by the coding table is 8K (2^12*16bit), which is much smaller than the block size of 512k.
[0094] The compressed model data obtained by the model lossless compression method in this application can be decompressed to complete the model reasoning in the model reasoning stage. The model decompression method is as follows: Figure 6 As shown in , the following steps are included:
[0095] Step 601: Obtain compressed model data, the compressed model data including encoding results of parameter matrices corresponding to each network layer included in the model and an encoding table, the encoding table including a mapping relationship between the parameter matrix and the encoding results.
[0096] Step 602: Using the coding table, the coding results of the parameter matrices corresponding to each network layer are decoded respectively to obtain the parameter matrices corresponding to each network layer for reading by the inference execution unit.
[0097] The above step 601, i.e., "obtaining compressed model data, wherein the compressed model data includes encoding results of parameter matrices corresponding to each network layer included in the model and a coding table, wherein the coding table includes a mapping relationship between the parameter matrix and the encoding results" is described in detail below in conjunction with an embodiment.
[0098] The compressed model data includes the encoding results and encoding tables of the parameter matrices corresponding to each network layer contained in the model. The encoding results of the parameter matrices corresponding to each network layer can be the encoding results for the entire matrix or the encoding results for each matrix block, where the matrix block is obtained by dividing the parameter matrix corresponding to each network layer into blocks.
[0099] The coding table is created according to the coding principle when encoding the parameter matrix. The coding table includes the mapping relationship between the parameter matrix and the coding result. A coding table can be constructed for all parameter matrices of the entire model, or a corresponding coding table can be created for each parameter matrix. When the parameter matrix is block-coded, a coding table can be constructed for each matrix block.
[0100] The present application obtains the compressed model data according to the storage location of the compressed model data. The compressed model data can be stored in a variety of storage units, can be stored in a local hard disk, or can be stored in a cloud storage server. The encoding result and the encoding table can be stored in the same location or in different locations.
[0101] Figure 7 The decoding process diagram provided for the embodiment of the present application is a preferred implementation mode, in which the encoding result is stored in the global memory, and the encoding result of the parameter matrix corresponding to each network layer contained in the model is obtained from the global memory and loaded into the first register.
[0102] The coding table can be pre-stored in the Global Memory, and loaded into a high-speed cache such as a shared memory or a constant memory when decoding the parameter matrix. In the above LLaMA 2 decoding process, to access 12-bit encoded data, take 8 12-bit (96-bit in total) compressed data as a group as an example, load them from the Global Memory into 3 32-bit registers, and obtain 8 16-bit decoded data by querying the coding table in the high-speed cache, and put them into 4 32-bit registers to participate in the calculation.
[0103] The above step 602, i.e., "using the coding table to respectively decode the coding results of the parameter matrices corresponding to each network layer to obtain the parameter matrices corresponding to each network layer for reading by the inference execution unit", is described in detail below in conjunction with the embodiments.
[0104] After obtaining the compressed model data, it is necessary to decode the compressed model data to obtain the original parameter matrix for use in model inference.
[0105] The decoding process uses the coding table, and the coding table corresponding to the parameter matrix of different network layers needs to be matched during decoding. In order to identify the correspondence between the coding table and the encoding result of the parameter matrix, the identification information of the corresponding network layer can be set in the coding result and the coding table, or matching can be performed based on the coding result and the storage address of the coding table.
[0106] When the encoding result of the parameter matrix is the encoding for the matrix block, for the decoding of the matrix block, it is necessary to use the encoding table corresponding to each matrix block to decode the encoding result of each matrix block respectively to obtain the data of each matrix block; according to the network layer corresponding to each matrix block, the data of the parameter matrix corresponding to each network layer is obtained. Among them, the network layer corresponding to each matrix block is determined according to the identification information contained in each encoding result or the storage position of the encoding result in the storage space.
[0107] As an implementable method, the coding table in the cache is queried, and the coding results of the parameter matrices corresponding to each network layer are decoded respectively using the coding table, and the decoding results are stored in the second register, and the second register is read by the inference execution unit.
[0108] The above method provided by the embodiment of the present application can be applied to a variety of application scenarios, including but not limited to: the deep learning model deployed on resource-constrained devices such as smart phones, tablet computers, embedded systems, etc., and the method provided by the embodiment of the present application can utilize the model lossless compression to reduce the model size and reduce memory and storage requirements. It can also be applied to cloud services and data centers. Although cloud services and data centers usually have higher computing resources, model lossless compression can reduce the amount of data transmission, reduce bandwidth requirements and costs. On some small or low-cost servers, the model lossless compression method provided by the present application can also be used, which can enable these servers to run complex models that were originally unable to run due to resource limitations.
[0109] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0110] According to an embodiment of another aspect, a model lossless compression apparatus is provided. Figure 8 A schematic block diagram of a model lossless compression device according to an embodiment is shown, wherein the device is arranged at Figure 1 Servers in the architecture shown. Figure 8 As shown, the device 800 includes:
[0111] The parameter acquisition unit 801 is configured to acquire parameter matrices corresponding to one or more network layers included in the model.
[0112] The data analysis unit 802 is configured to determine data distribution of parameter matrices corresponding to more than one network layer.
[0113] The coding bit number determining unit 803 is configured to determine the coding bit number by using the data distribution, where the coding bit number is smaller than the original bit number used by the parameters in the parameter matrix.
[0114] The encoding unit 804 is configured to encode the parameter matrices corresponding to more than one network layer respectively according to the number of encoding bits, and obtain the encoding results of the parameter matrices corresponding to each network layer.
[0115] The data storage unit 805 is configured to store compressed model data, wherein the compressed model data includes encoding results of parameter matrices corresponding to each network layer and an encoding table, wherein the encoding table includes a mapping relationship between the parameter matrix and the encoding results.
[0116] As one possible implementation method, the data distribution includes the probability of occurrence of parameters in the parameter matrix, and the number of coding bits is negatively correlated with the probability of occurrence of parameters; or, the data distribution includes unique numerical quantities in the parameter matrix, and the number of coding bits is determined using the unique numerical quantities.
[0117] As one of the feasible ways, when determining the number of coding bits using a unique numerical value, the coding bit determination unit 803 can be configured as follows: determining a plurality of block sizes; respectively determining the unique numerical values in each block obtained by dividing the parameter matrix into blocks using each block size and the number of coding bits required; and selecting a block size from a plurality of block sizes and determining the number of coding bits required using the unique numerical values in each block corresponding to each block size and the number of coding bits required.
[0118] As one of the feasible ways, when the encoding unit 804 uses the number of coding bits to encode the parameter matrices corresponding to more than one network layers respectively and obtains the encoding results of the parameter matrices corresponding to each network layer, it can be configured as follows: the parameter matrices corresponding to the more than one network layers are divided into blocks using the selected block size to obtain multiple matrix blocks; each matrix block is encoded using the determined number of coding bits to obtain the encoding results corresponding to each matrix block and the encoding table corresponding to each matrix block as compressed model data.
[0119] As one of the implementable ways, the encoding result includes identification information indicating the network layer corresponding to the matrix block; or, the encoding result is stored in a storage space according to the corresponding network layer.
[0120] As one of the possible implementations, the encoding results and the encoding table are stored in the global memory.
[0121] According to an embodiment of another aspect, a model decompression device is provided. Fig. 9 A schematic block diagram of a model decompression device according to an embodiment is shown, wherein the device is arranged at Figure 1Servers in the architecture shown. Fig. 9 As shown, the device 900 includes:
[0122] The data acquisition unit 901 is configured to acquire compressed model data, wherein the compressed model data includes encoding results of parameter matrices corresponding to each network layer included in the model and an encoding table, wherein the encoding table includes a mapping relationship between the parameter matrix and the encoding result;
[0123] The decoding unit 902 is configured to use the coding table to decode the encoding results of the parameter matrix corresponding to each network layer respectively, and obtain the parameter matrix corresponding to each network layer for reading by the inference execution unit.
[0124] As one of the implementable ways, the encoding result and encoding table of the parameter matrix corresponding to each network layer include: the encoding result of each matrix block and the corresponding encoding table, and the matrix block is obtained by dividing the parameter matrix corresponding to each network layer into blocks. The decoding unit 902 uses the encoding table to decode the encoding results of the parameter matrix corresponding to each network layer respectively, and obtains the parameter matrix corresponding to each network layer. It can be configured as follows: using the encoding table corresponding to each matrix block, respectively decode the encoding results of each matrix block to obtain the data of each matrix block; according to the network layer corresponding to each matrix block, obtain the data of the parameter matrix corresponding to each network layer.
[0125] As one of the implementable ways, the network layer corresponding to each matrix block is determined according to the identification information contained in each encoding result or the storage position of the encoding result in the storage space.
[0126] As one of the achievable methods, the data acquisition unit 901 can be configured to: obtain the encoding results of the parameter matrices corresponding to each network layer contained in the model from the global memory and load them into the first register; obtain the encoding table from the global memory and load it into the cache. The decoding unit 902 can be configured to: query the encoding table in the cache, use the encoding table to decode the encoding results of the parameter matrices corresponding to each network layer, and store the decoding results in the second register, which is read by the inference execution unit.
[0127] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative work.
[0128] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0129] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.
[0130] And an electronic device, comprising:
[0131] one or more processors; and
[0132] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the aforementioned method embodiments.
[0133] The present application also provides a computer program product, including a computer program, which implements the steps of any one of the methods in the aforementioned method embodiments when executed by a processor.
[0134] in, Fig.10The electronic device architecture is shown as an example, which may include a processor 1010, a video display adapter 1011, a disk drive 1012, an input / output interface 1013, a network interface 1014, and a memory 1020. The processor 1010, the video display adapter 1011, the disk drive 1012, the input / output interface 1013, the network interface 1014, and the memory 1020 may be communicatively connected via a communication bus 1030.
[0135] Among them, the processor 1010 can be implemented by a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., to execute relevant programs to implement the technical solutions provided in this application.
[0136] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store an operating system 1021 for controlling the operation of the electronic device 1000, and a basic input and output system (BIOS) 1022 for controlling the low-level operation of the electronic device 1000. In addition, a web browser 1023, a data storage management system 1024, and a model lossless compression device / model decompression device 1025, etc. can also be stored. The above-mentioned model lossless compression device / model decompression device 1025 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided by the present application is implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0137] The input / output interface 1013 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0138] The network interface 1014 is used to connect to a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).
[0139] The bus 1030 comprises a pathway for transmitting information between the various components of the device (eg, the processor 1010, the video display adapter 1011, the disk drive 1012, the input / output interface 1013, the network interface 1014, and the memory 1020).
[0140] It should be noted that, although the above device only shows a processor 1010, a video display adapter 1011, a disk drive 1012, an input / output interface 1013, a network interface 1014, a memory 1020, a bus 1030, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include components necessary for implementing the solution of the present application, and does not necessarily include all the components shown in the figure.
[0141] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a computer program product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.
[0142] The technical solution provided by the present application is described in detail above. The principle and implementation method of the present application are described in detail using specific examples. The description of the above embodiments is only used to help understand the method and core idea of the present application. At the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as limiting the present application.
Claims
1. A model lossless compression method, characterized in that: The method comprises: Get the parameter matrices corresponding to one or more network layers contained in the model; Determine data distribution of parameter matrices corresponding to the one or more network layers respectively; Determining a number of coding bits using the data distribution, the number of coding bits being less than an original number of bits used by parameters in the parameter matrix; According to the number of coding bits, the parameter matrices corresponding to the more than one network layers are respectively encoded to obtain encoding results of the parameter matrices corresponding to the respective network layers; The compressed model data is stored, wherein the compressed model data includes the encoding results of the parameter matrices corresponding to each network layer and a coding table, wherein the coding table includes a mapping relationship between the parameter matrices and the encoding results.
2. The method according to claim 1, characterized in that The data distribution includes a unique numerical value in the parameter matrix, and the number of encoding bits is determined using the unique numerical value; or, The data distribution includes the probability of occurrence of parameters in the parameter matrix, and the number of coding bits is negatively correlated with the probability of occurrence of the parameters.
3. The method according to claim 2, characterized in that Determining the number of encoding bits by using the unique numerical value comprises: Determine multiple block sizes; Determining respectively the number of unique values in each block obtained by dividing the parameter matrix into blocks using each block size and the number of encoding bits required; By using the unique numerical quantities in the blocks corresponding to the block sizes and the required number of coding bits, a block size is selected from the multiple block sizes and the required number of coding bits is determined.
4. The method according to claim 3, characterized in that According to the number of coding bits, the parameter matrices corresponding to the more than one network layers are respectively encoded, and the encoding results of the parameter matrices corresponding to the network layers are obtained, including: Using the selected block size, the parameter matrices corresponding to the one or more network layers are respectively divided into blocks to obtain a plurality of matrix blocks; Each matrix block is encoded respectively using the determined number of encoding bits to obtain the encoding result corresponding to each matrix block and the encoding table corresponding to each matrix block as the compressed model data.
5. The method according to claim 4, characterized in that The encoding result includes identification information indicating the network layer corresponding to the matrix block; or, The encoding results are stored in a storage space according to the corresponding network layer.
6. The method according to any one of claims 1 to 5, further comprising: The encoding result and the encoding table are stored in a global memory.
7. A model decompression method, characterized in that: The method comprises: Obtaining compressed model data, the compressed model data including encoding results of parameter matrices corresponding to each network layer contained in the model and a coding table, the coding table including a mapping relationship between the parameter matrix and the encoding result, the encoding result being obtained by encoding the parameter matrices corresponding to each network layer according to the number of encoding bits, and the number of encoding bits being determined according to data distribution of the parameter matrix; The encoding table is used to decode the encoding results of the parameter matrices corresponding to the network layers respectively, and the parameter matrices corresponding to the network layers are obtained for reading by the inference execution unit.
8. The method according to claim 7, characterized in that The encoding results and encoding tables of the parameter matrices corresponding to the network layers include: encoding results of each matrix block and a corresponding encoding table, wherein the matrix block is obtained by dividing the parameter matrix corresponding to each network layer into blocks; Using the coding table, the coding results of the parameter matrices corresponding to the network layers are decoded respectively, and the parameter matrices corresponding to the network layers are obtained, including: Using the coding table corresponding to each matrix block, the coding results of each matrix block are decoded respectively to obtain the data of each matrix block; According to the network layers corresponding to each matrix block, the data of the parameter matrix corresponding to each network layer is obtained.
9. The method according to claim 8, characterized in that The network layer corresponding to each matrix block is determined according to the identification information contained in each encoding result or the storage position of the encoding result in the storage space.
10. The method according to any one of claims 7 to 9, wherein obtaining the compressed model data comprises: Obtain the encoding result of the parameter matrix corresponding to each network layer contained in the model from the global memory and load it into the first register; Obtaining the encoding table from global memory and loading it into cache; Using the coding table, decoding the coding results of the parameter matrices corresponding to the network layers respectively includes: The coding table in the cache is queried, and the coding results of the parameter matrices corresponding to the network layers are respectively decoded using the coding table, and the decoding results are stored in a second register, and the second register is read by the inference execution unit.
11. A model lossless compression device, characterized in that: The device comprises: A parameter acquisition unit, which acquires parameter matrices corresponding to one or more network layers included in the model; A data analysis unit, determining data distribution of parameter matrices corresponding to the one or more network layers respectively; A coding bit determination unit, which determines the number of coding bits using the data distribution, wherein the number of coding bits is smaller than the original number of bits used by the parameters in the parameter matrix; An encoding unit, encoding the parameter matrices corresponding to the one or more network layers respectively according to the number of encoding bits, to obtain encoding results of the parameter matrices corresponding to the network layers; A data storage unit stores compressed model data, wherein the compressed model data includes encoding results of parameter matrices corresponding to each network layer and an encoding table, wherein the encoding table includes a mapping relationship between the parameter matrix and the encoding result.
12. A model decompression device, characterized in that: The device comprises: A data acquisition unit, acquiring compressed model data, wherein the compressed model data includes encoding results of parameter matrices corresponding to each network layer contained in the model and a coding table, wherein the coding table includes a mapping relationship between the parameter matrix and the encoding result, wherein the encoding result is obtained by encoding the parameter matrix corresponding to each network layer according to the number of coding bits, and the number of coding bits is determined according to the data distribution of the parameter matrix; The decoding unit uses the coding table to decode the coding results of the parameter matrices corresponding to the network layers respectively, and obtains the parameter matrices corresponding to the network layers for reading by the inference execution unit.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
14. An electronic device, characterized in that: include: one or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of claims 1 to 10.
Citation Information
Patent Citations
Named entity identification method based on neural network and vehicle machine
CN111274816A
Data processing method and device, electronic equipment and storage medium
CN116306879A
Model compression method and device, electronic equipment and storage medium
CN117371508A