Neural network differential compression method and system

Through the neural network differential compression method, N:M pruning and differential computing technology are used, combined with FPGA hardware design decompression module, the scale limitation problem of neural network models when deploying on embedded platforms is solved, and effective lightweighting of model weight parameters and optimization of storage resources are achieved.

CN120218138APending Publication Date: 2025-06-27SHANDONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510305594.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The deployment of existing neural network models on embedded platforms is limited by scale. Traditional model lightweighting techniques such as pruning and quantization can reduce model size, but are not effective in reducing actual storage resources.

Method used

A neural network differential compression method is proposed. The index data obtained by N:M pruning is established to establish the correlation between weight data in blocks, perform differential calculations, and provide three compression methods for the weight data after differential calculations. Finally, the decompression module is designed using FPGA hardware.

Benefits of technology

It realizes lightweighting of the weight parameters of the neural network model, reduces the consumption of actual storage resources, and designs an efficient decompression module on FPGA hardware, suitable for the deployment of embedded platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218138A_ABST
    Figure CN120218138A_ABST
Patent Text Reader

Abstract

The invention discloses a neural network differential compression method and system, and the method comprises the steps: obtaining a sparse weight matrix and an index matrix, and carrying out the sorting and differential calculation of the weight of each line in the sparse weight matrix through the index matrix; each row in the difference matrix is used as a block, a bit width is additionally arranged behind the head data of the block, and the bit width represents the minimum bit width required by the binary system of the maximum difference value except the head data in the block; according to the minimum bit width and the distribution frequency required by binary representation of the maximum differential value except the header data in the block, determining the value of the bit width in the block and the storage bit width of other differential values except the header data; according to the storage bit width in the blocks, other differential values are converted into binary character strings meeting the storage bit width requirement, compression is completed after the binary character strings are connected with head data and binary representation of the bit width, lightweight of a network model is achieved, and finally corresponding decompression is designed according to actual operation of a compressed network by means of the programmable characteristic of FPGA hardware.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of neural network model compression, and particularly to a neural network differential compression method and system. Background Art

[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] With the improvement of the performance of neural networks, the scale of network models has been continuously increasing, which has led to certain limitations in their deployment on embedded platforms. Researchers have proposed various model lightweight technologies to address this problem, including model pruning, model quantization, network architecture search, and low-rank decomposition. Neural network models can achieve lightweighting of network models by using a combination of the above single or multiple technologies. However, lightweighting at the model level does not necessarily mean a reduction in the actual storage resources during deployment.

[0004] Taking pruning as an example, most network model pruning is achieved through masking technology, and the deleted 0 values are still retained. Although the network model is theoretically lightweighted, the actual parameters of the network model have not decreased. In subsequent research, some researchers have achieved a reduction in actual parameters by deleting 0 values and retaining non-0 values and their corresponding indices. Such index records include the CSR (Compressed Sparse row) method, the CSC (Compressed Sparse Column) method, and the COO (Coordinate Format) method. However, these storage methods may lead to the introduction of index data being much larger than the amount of weight data being deleted. Summary of the Invention

[0005] To solve the problem of compressing network models, the present invention proposes a neural network differential compression method and system. By using the index data within each block obtained by N:M pruning, the correlation between the weight data within the block is established. Then, differential calculation is performed on the sparse weight matrix, and the weight data after differential calculation is compressed, thereby achieving lightweighting of the weight parameters of the network model. Finally, using the characteristics of FPGA hardware programmability, a corresponding decompression module is designed for the actual operation of the compressed network.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a neural network differential compression method, including:

[0008] Obtain the sparse weight matrix and index matrix of the pruned neural network. Sort the weights in each row of the sparse weight matrix through the index matrix, and perform differential calculation on each row of the sorted sparse weight matrix to obtain a differential matrix;

[0009] Take each row in the differential matrix as a block, and add a bit width after the head data of each block. The bit width represents the minimum bit width required for the binary representation of the maximum difference value in the block except the head data;

[0010] Determine the value of the bit width in each block and the storage bit width of other difference values except the head data according to the minimum bit width required for the binary representation of the maximum difference value in the block except the head data and its distribution frequency;

[0011] According to the storage bit width in each block, convert other difference values in the block into binary strings that meet the storage bit width requirements, and connect them with the head data and the binary representation of the bit width to complete compression.

[0012] As an alternative implementation, use the value of the bit width in each block as the number of bit widths required for other difference values in the current block during compression, that is, the storage bit width of other difference values. During compression, if the actual binary representation bit width of other difference values is less than the storage bit width, expand it to the required storage bit width by padding with zeros.

[0013] As an alternative implementation, during decompression, set the bit width size required for the binary representation of the bit width in each block to a fixed value, that is b is the base bit width, and the base bit width is determined according to the distribution frequency.

[0014] As an alternative implementation, determine the base bit width according to the distribution frequency. For blocks where the value of the bit width in the block is less than the base bit width, use the base bit width as the storage bit width. After converting other difference values in the block into binary representation, expand them to the number of bit widths required by the base bit width by padding with zeros;

[0015] For blocks where the value of the bit width in the block is greater than or equal to the base bit width, use the value of the bit width of each block as the storage bit width, and convert other difference values into binary strings.

[0016] As an alternative implementation, the base bit width is:

[0017] b * =argmax b (f1(b)-f2(b));

[0018]

[0019] Among them, f1(b) represents the reduction in the binary representation of the bit width when the basic bit width is b; f2(b) represents the additional storage amount when expanding the binary representation of other difference values inside a block with a minimum bit width less than the basic bit width required for the binary representation of the maximum difference value except the head data in the block; w is the bit width value after quantization of each difference value in the difference matrix; b is the basic bit width, 0 ≤ b < w; x i,j is the j-th value in the i-th block; m i is the minimum bit width required for the binary representation of the maximum difference value except the head data in the i-th block; BN is the number of blocks; N is the number of weight values in each block.

[0020] As an alternative implementation, the head data and the bit width are each converted into binary representations, and after being concatenated with the binary representations of other difference values converted to meet the storage bit width requirements, compression is completed.

[0021] As an alternative implementation, starting from the low bit of the head data, the part with the same width as the binary representation of the bit width is replaced with the binary representation of the bit width, thereby determining the binary representation of the head data. After being concatenated with the binary representations of other difference values converted to meet the storage bit width requirements, compression is completed.

[0022] As an alternative implementation, after introducing the basic bit width BASE_WIDTH, the bit width required for the binary representation of the bit width DATA_WIDTH in each block is updated to w is the bit width value after quantization of each difference value in the difference matrix; b is the basic bit width;

[0023] The corresponding relationship between the updated bit width value and the minimum bit width required for the binary representation of the maximum difference value in the corresponding block is DATA_WIDTH new = DATA_WIDTH old - BASE_WIDTH; where DATA_WIDTH new is the updated DATA_WIDTH, and DATA_WIDTH old is the DATA_WIDTH before update.

[0024] As an alternative implementation, during decompression, the bit width size required for the binary representation of the bit width in each block is set to

[0025] In a second aspect, the present invention provides a neural network differential compression system, including:

[0026] A difference module, configured to obtain a sparse weight matrix and an index matrix of a pruned neural network, sort the weights of each row in the sparse weight matrix through the index matrix, and perform difference calculation on each row of the sorted sparse weight matrix to obtain a difference matrix;

[0027] A first bit-width analysis module, configured to use each row in the difference matrix as a block, and add a bit-width after the head data of each block, where the bit-width represents the minimum bit-width required for the binary representation of the maximum difference value in the block except the head data;

[0028] A second bit-width analysis module, configured to determine the value of the bit-width in each block and the storage bit-width of other difference values except the head data according to the minimum bit-width required for the binary representation of the maximum difference value in the block except the head data and its distribution frequency;

[0029] A compression module, configured to convert other difference values in the block into a binary string that meets the storage bit-width requirements according to the storage bit-width in each block, and complete the compression after connecting with the head data and the binary representation of the bit-width.

[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0031] The present invention provides a neural network difference compression method and system, which uses the index data in each block obtained by N:M pruning to establish the correlation between the weight data in the block. By increasing the correlation between these adjacent data, the weight matrix obtained by N:M pruning is suitable for difference compression; then, difference calculation is performed on the sparse weight matrix, and three compression methods are provided for the weight data after difference calculation for compression, so as to realize the lightweight of the network model; finally, using the characteristics of FPGA hardware programmability, a corresponding decompression module is designed for the actual operation of the compressed network.

[0032] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0034] Figure 1 It is a flowchart of the neural network difference compression method provided in Embodiment 1 of the present invention;

[0035] Figure 2Schematic diagram of sorting and differencing for the sparse weight matrix provided in Embodiment 1 of the present invention;

[0036] Figure 3 Schematic diagram of key data to be stored for compression in the difference matrix provided in Embodiment 1 of the present invention;

[0037] Figure 4 Frequency diagram of bit widths of the maximum difference values provided in Embodiment 1 of the present invention;

[0038] Figure 5 Schematic diagram of the lossy difference compression method provided in Embodiment 1 of the present invention;

[0039] Figure 6 Schematic diagram of the overall decompression process provided in Embodiment 1 of the present invention;

[0040] Figure 7 Design concept of the decompression module in FPGA taking the lossless difference compression method as an example provided in Embodiment 1 of the present invention. Detailed implementation manners

[0041] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0042] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0043] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0044] In the case of no conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0045] Embodiment 1

[0046] The block storage method divides the weight matrix into blocks and separately records the positions of non-zero values within each block inside the block. By introducing position indices, a connection is established with the weight data. Inside a small block, the sparse weight data can be associated using the index data, and differential compression of the sparse weights can be achieved by reasonably utilizing this correlation.

[0047] In the field of neural network model optimization, N:M pruning, as an efficient sparsification technique, achieves structured sparse processing by forcing no more than (M - N) non-zero values to be retained among consecutive M weight elements. To achieve substantial compression of model parameters, this technique adopts a dual matrix representation strategy: converting the original sparse weight matrix into a compact non-zero value weight matrix and supplementing it with a corresponding relative position index matrix. Thanks to the characteristic of pruning by block (with M elements as the processing unit), the index matrix only needs to record the relative offsets (in the range of 0 to M - 1) of the retained elements within their respective blocks, and the storage overhead introduced is significantly lower than the amount of redundant zero-value data eliminated during the pruning process. The subsequent research of this invention will be deeply explored based on this technical framework.

[0048] As Figure 1 shown, this embodiment provides a neural network differential compression method, including:

[0049] Obtain the sparse weight matrix and index matrix of the pruned neural network, sort the weights in each row of the sparse weight matrix through the index matrix, perform differential calculation on each row of the sorted sparse weight matrix to obtain a differential matrix;

[0050] Take each row in the differential matrix as a block, and add a bit width after the header data of each block. The bit width represents the minimum bit width required for the binary representation of the maximum differential value in the block except the header data;

[0051] Determine the value of the bit width in each block and the storage bit width of other differential values except the header data according to the minimum bit width required for the binary representation of the maximum differential value in the block except the header data and its distribution frequency;

[0052] According to the storage bit width within each block, convert other differential values in the block into binary strings that meet the storage bit width requirements, and connect them with the header data and the binary representation of the bit width to complete the compression.

[0053] In this embodiment, taking N:M pruning as an example, after pruning, a sparse weight matrix and an index matrix of the neural network are obtained. The effectiveness of differential compression needs to be based on the similarity between adjacent data, while there is no correlation between adjacent weight data inside the obtained sparse weight matrix after pruning. Therefore, first, by using the index matrix to sort the weights in each row of the sparse weight matrix, it can be in ascending or descending order. The similarity between adjacent weight data after sorting increases, thus adapting to differential compression.

[0054] As Figure 2 shown in (a) of [reference], it is the sparse weight matrix and the index matrix of the corresponding weight data obtained after N:M pruning. The left part represents the weight data, and the right part represents the index data. By sorting the weight data inside each row, the sparse weight matrix and the index matrix shown in (b) of [reference] are obtained. Figure 2 shown in (b) of [reference].

[0055] Finally, perform differential calculation on each row of the sparse weight matrix sorted in ascending order to obtain the differential matrix shown in (c) of [reference]. Each row represents a block. The first data inside the block is the head data HEAD_DATA, and the other differential values are DIFF_DATA. Figure 2 shown in (c) of [reference].

[0056] Weight sorting and differencing only theoretically reduce the data volume. However, without compression and when the bit width of the actually stored differenced data remains unchanged, the storage consumption does not decrease. Therefore, in order to truly reduce the data volume, the sorted and differenced data needs to be compressed.

[0057] Before compression, it is first necessary to analyze the key data that needs to be stored during compression for the differential matrix. Since the maximum values of the other differential values DIFF_DATA in each row inside the differential matrix after sorting and differencing are different, this leads to inconsistent storage bit widths required for the other differential values. If the maximum value among the other differential values except the head data HEAD_DATA (the first data) in the differential matrix of the entire sparse weight matrix after differencing is used as the unified bit width representation standard, in the face of extreme cases, such as the maximum value DIFF_DATA after 8-bit quantization differencing being 255 (binary representation is 11111111), this compression is completely ineffective. Therefore, to avoid ineffective compression, it is necessary to determine the corresponding storage bit width according to the actual situation of the data inside each block.

[0058] Therefore, based on the above analysis, in addition to the head data HEAD_DATA and other difference values DIFF_DATA, each block also requires a value representing the number of bit widths required for storing DIFF_DATA within each block during compressed storage. Thus, a bit width DATA_WIDTH is added after the head data of each block. The value of the bit width DATA_WIDTH is determined as the minimum bit width required for the binary representation of the maximum difference value among the other difference values except the head data in each block.

[0059] Therefore, the key data to be stored, such as Figure 3 shown, includes the binary representation of the head data HEAD_DATA, the bit width DATA_WIDTH, and the binary representation of other difference values DIFF_DATA. Figure 3 In the first column, it is the binary representation of the head data HEAD_DATA, and in the second column, it is the binary representation of the bit widths of other difference values DIFF_DATA in each block. For example, if the maximum value among the other difference values in the current block is 25, then the required bit width is 5. Therefore, the bit width required for storing other difference values in the current block during compressed storage is 5. Then, the binary representation of the bit width DATA_WIDTH is 0101, and the following part is the binary representation of other difference values DIFF_DATA. If the actual bit width of the binary representation of DIFF_DATA is less than 5, it is padded with zeros.

[0060] Thus, through the analysis of the above key data, the data necessary for compressing the sorted and differenced sparse weight matrix is obtained. Subsequently, three compression methods are proposed in this embodiment based on the sorted and differenced weight data.

[0061] The first compression method is to analyze the key data that needs to be stored for the difference matrix and perform simple stacking. As Figure 3 shown in the upper part of, the key data is concatenated in the order shown in the figure to obtain the first compression method.

[0062] The second compression method is a lossless differential compression method. By analyzing the distribution frequency of the maximum bit widths within the data of each block in the difference matrix using a function, the storage consumption of DATA_WIDTH is reduced using the analysis result. The first and second compression methods are both lossless compressions and can recover the original data.

[0063] The third compression method is a lossy differential compression method. By integrating DATA_WIDTH into the HEAD_DATA internally, the bit width requirement of DATA_WIDTH is directly deleted, but it will cause certain data loss.

[0064] Specifically:

[0065] (1) Lossless differential compression method.

[0066] In the first compression scheme, although the DIFF_DATA elements within each block are variable and their specific bit widths are determined by the bit width DATA_WIDTH within the block, the meta-parameter DATA_WIDTH itself is stored using a fixed bit width.

[0067] For example, when the bit width of the binary representation of the maximum difference value in the data block except for the header data is 8 bits, the DATA_WIDTH field needs to be fixedly allocated 4 bits (binary representation: 1000) of storage space to represent 9 cases where the width of DIFF_DATA is 0, 1, 2, 3, 4, 5, 6, 7, and 8. This design adopts a fixed bit width for DATA_WIDTH and a variable bit width strategy for DIFF_DATA.

[0068] By analyzing the distribution frequency of the bit widths of the maximum difference values in the DIFF_DATA of each block in the difference matrix, as Figure 4 shown, it can be found that the distribution of the bit widths of the maximum difference values in this figure mainly concentrates on several values (5, 6, 7 bits). Based on this observation, in this embodiment, it is considered to utilize this discovery. By analyzing the distribution frequency of the minimum bit width required for the binary representation of the maximum difference values except for the header data in all blocks, the base bit width BASE_WIDTH is introduced to also compress DATA_WIDTH to a certain extent and improve the compression effect.

[0069] The calculation formula for the base bit width BASE_WIDTH is:

[0070] b * = argmax b (f1(b) - f2(b));

[0071]

[0072] where f1(b) represents the reduction in the binary representation of the bit width DATA_WIDTH when the size of the base bit width is b; f2(b) represents the additional storage amount when expanding the binary representations of other difference values within the blocks where the minimum bit width required for the binary representation of the maximum difference values except for the header data in the block is less than the base bit width when the size of the base bit width is b; w is the quantized bit width value of each difference value in the difference matrix; b is the base bit width, 0 ≤ b < w; x i,j represents the j-th value within the i-th block, starting from 0, and 0 represents the first element of the 5th block, that is, x 5,1 , as Figure 2 shown in (c) of 5,1 is 32; m irepresents the minimum bit width required for the binary representation of the maximum difference value in the i-th block except for the head data; BN is the number of blocks; N is the number of weight values within each block.

[0073] After determining the base bit width BASE_WIDTH, special processing is performed on the blocks whose minimum bit width required for the binary representation of the maximum DIFF_DATA within the block is less than BASE_WIDTH. That is, after converting the DIFF_DATA within these blocks into a binary string, it needs to be padded with zeros to expand it to the size represented by BASE_WIDTH.

[0074] For the blocks whose minimum bit width required for the binary representation of the maximum DIFF_DATA within the block is greater than or equal to the base bit width, they are converted into binary strings according to the minimum bit width required for the binary representation of the maximum DIFF_DATA within their respective blocks.

[0075] At the same time, since the introduction of the base bit width BASE_WIDTH is to reduce the storage occupied by DATA_WIDTH, the bit width occupied by DATA_WIDTH is no longer but becomes

[0076]

[0077] The corresponding relationship between the updated value of DATA_WIDTH and the minimum bit width required for the binary representation of the maximum DIFF_DATA within the corresponding block is DATA_WIDTH new = DATA_WIDTH old - BASE_WIDTH; where DATA_WIDTH new represents the updated DATA_WIDTH, and DATA_WIDTH old is the DATA_WIDTH before update.

[0078] (2) Lossy differential compression method.

[0079] In practical applications, sometimes there is no high requirement for the accuracy of the network model. In this case, lossy compression can be adopted, such as Figure 5As shown, DATA_WIDTH is incorporated into HEAD_DATA, and at the same time, combined with the base width BASE_WIDTH of the lossless differential compression method. In the lossless differential compression method, the analysis and calculation of the base width BASE_WIDTH are mainly to reduce the amount of stored data. In the lossy differential compression method, since DATA_WIDTH has been incorporated into HEAD_DATA, therefore, the binary representation width of DATA_WIDTH has no impact on the compression efficiency. However, a reasonable DATA_WIDTH can reduce the reduction of the network model accuracy.

[0080] Among them, the process of incorporating DATA_WIDTH into HEAD_DATA is to replace the part of the head data starting from the low bit and having the same width as the binary representation of the width with the binary representation of the width. As Figure 5 shown, replace the 1101 starting from the low bit in the head data with the 0101 of the width.

[0081] In binary, the definitions of the low bit and the high bit are as follows: The low bit refers to the rightmost bit in the binary number, with the smallest weight and the smallest represented value. The high bit refers to the leftmost bit in the binary number, with the largest weight and the largest represented value. For example, in the binary number 1011, the rightmost 1 is the low bit, and the leftmost 1 is the high bit.

[0082] Among them, compared with the lossless compression method, the difference of this method is that DATA_WIDTH is incorporated into HEAD_DATA, and the rest of the process is the same. The compression process is to concatenate the binary representations of HEAD_DATA, DATA_WIDTH, and DIFF_DATA.

[0083] The purpose of compression is to reduce the actual storage requirement. The three compression methods mentioned in this embodiment are, in order, each compression method is an improvement on the basis of the previous compression method. In the first and second compression methods, DATA_WIDTH independently occupies a certain storage space. In order to further improve the compression effect, consider making DATA_WIDTH not occupy storage space, and in actual requirements, in some cases, not a very high accuracy is required, then lossy compression can be adopted, such as incorporating DATA_WIDTH into HEAD_DATA, thereby improving the compression effect.

[0084] This embodiment uses the index data inside each block obtained by N:M pruning to establish the correlation between the weight data inside the block. By increasing the correlation between these adjacent data, the weight matrix obtained by N:M pruning is made suitable for differential compression, and three compression methods are provided for the weight data after differential calculation, thereby realizing the lightweight of the network model.

[0085] In this embodiment, the differential compression of the neural network is completed through the above method. By using the index data obtained by N:M pruning, the sparse weight data is sorted and differentiated to achieve the recompression of the sparse weight data. However, the compressed weights form a set of binary strings, which cannot be used to calculate with the activation values in actual network calculations. In order to extract the required weight data from the compressed strings, this embodiment designs a decompression method with an FPGA as the practical platform to achieve the decompression of the binary string data.

[0086] According to the differences in the selected compression methods, the decompression process will have slight differences. Figure 6 The decompression process corresponding to the first compression method is shown in the dashed box. The process of extracting DATA_WIDTH corresponding to the first compression method and the second compression method is similar, and both can be obtained according to the Figure 6 way of obtaining DATA_WIDTH within the dashed box in. However, since the base width BASE_WIDTH in the second compression method is no longer a fixed value of 0, but will be dynamically adjusted according to the characteristics of the network weights. Therefore, there will be slight differences in the bit width DIFF_DATA_REPRESENT_WIDTH required for the binary representation of DATA_WIDTH within each block corresponding to the first compression method and the second compression method.

[0087] For the first compression method, DIFF_DATA_REPRESENT_WIDTH is set to a fixed value, and its size is

[0088] For the second compression algorithm, the size of DIFF_DATA_REPRESENT_WIDTH is

[0089]

[0090] The third compression method is obtained by integrating DATA_WIDTH into the header data HEAD_DATA on the basis of the second compression method. Therefore, compared with the second compression method, when parsing the binary string obtained by the third compression method, the calculation of DATA_WIDTH needs to be changed to a certain extent, specifically:

[0091] DATA_WIDTH = HEAD_DATA[0:DIFF_DATA_REPRESENT_WIDTH - 1];

[0092] Among them, the way to obtain HEAD_DATA is as Figure 6As shown in the dashed box, the size of DIFF_DATA_REPRESENT_WIDTH is

[0093] According to the above decompression process, the compressed binary string data is restored to the pruned sparse weight matrix.

[0094] As Figure 6 As shown in the dashed box, in order to reduce the impact of the decompression process on the network operation speed, when designing the decompression module in the FPGA, the Figure 6 decompression in the right dashed box is not serially calculated, but the decompression is designed by utilizing the high parallelism characteristics of the FPGA.

[0095] The parallel design idea is as Figure 7 shown. The Buffer internally stores the compressed binary string read from the storage area. As long as there is free space in the Buffer, the compressed binary string can be read. When the data in the Buffer reaches the amount that can be decompressed, decompression is performed. These two processes are carried out simultaneously.

[0096] In order to reduce the time for decompressing a group of weights, according to the characteristics of the compression algorithm, in the first clock cycle, HEAD_DATA and DATA_WIDTH are read, and in the second clock cycle, all DIFF_DATA are read according to DATA_WIDTH. At the same time, all the read DIFF_DATA are added according to the Figure 7 process shown in, and all the weight data WEIGHT_DATA of the current block are obtained at the same time.

[0097] Specifically, for the solution of the first weight data WEIGHT_DATA0, the input is HEAD_DATA; for the solution of the second weight data WEIGHT_DATA1, the inputs are HEAD_DATA(WEIGHT_DATA0) and DIFF_DATA0, and their sum is WEIGHT_DATA1; for the solution of the third weight data WEIGHT_DATA2, the inputs are HEAD_DATA, DIFF_DATA0 (i.e., WEIGHT_DATA1) and DIFF_DATA1, and the sum of WEIGHT_DATA1 and DIFF_DATA1 is WEIGHT_DATA2; and so on to obtain the decompressed sparse weight matrix.

[0098] Embodiment 2

[0099] This embodiment provides a neural network differential compression system, including:

[0100] A difference module, configured to obtain a sparse weight matrix and an index matrix of a pruned neural network, sort the weights of each row in the sparse weight matrix through the index matrix, and perform difference calculation on each row of the sorted sparse weight matrix to obtain a difference matrix;

[0101] A first bit-width analysis module, configured to take each row in the difference matrix as a block, and add a bit-width after the header data of each block, where the bit-width represents the minimum bit-width required for the binary representation of the maximum difference value in the block except the header data;

[0102] A second bit-width analysis module, configured to determine the value of the bit-width in each block and the storage bit-width of other difference values except the header data according to the minimum bit-width required for the binary representation of the maximum difference value in the block except the header data and its distribution frequency;

[0103] A compression module, configured to convert other difference values in the block into a binary string that meets the storage bit-width requirements according to the storage bit-width in each block, and complete the compression after connecting with the header data and the binary representation of the bit-width.

[0104] It should be noted here that the above modules correspond to the steps described in Embodiment 1. The above modules and the corresponding steps have the same examples and application scenarios, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0105] In more embodiments, there is also provided:

[0106] An electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.

[0107] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0108] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random memory. For example, the memory may also store information about the device type.

[0109] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the method described in Embodiment 1.

[0110] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0111] A computer program product includes a computer program, which, when executed by a processor, implements the method described in Embodiment 1.

[0112] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to perform the process / method as described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided as needed. The machine-executable instructions for program modules can be executed locally or within a distributed device. In a distributed device, program modules can be located in local and remote storage media.

[0113] The computer program code for implementing the method of the present invention can be written in one or more programming languages. These computer program codes can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program code is executed by the computer or other programmable data processing devices, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as an independent software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.

[0114] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, etc. Examples of signals can include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, etc.

[0115] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0116] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, they do not limit the protection scope of the present invention. Those skilled in the art should understand that based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.

Claims

1. A neural network differential compression method, characterized in that: include: Obtain the sparse weight matrix and index matrix of the pruned neural network, sort the weights of each row in the sparse weight matrix by the index matrix, perform differential calculation on each row of the sorted sparse weight matrix, and obtain a differential matrix; Each row in the difference matrix is ​​a block, and a bit width is added after the header data of each block, wherein the bit width represents the minimum bit width required for the binary representation of the maximum difference value in the block except the header data; Determine the value of the bit width in each block and the storage bit width of other differential values ​​excluding the header data according to the minimum bit width required for the binary representation of the maximum differential value excluding the header data in the block and its distribution frequency; According to the storage bit width in each block, other differential values ​​in the block are converted into binary strings that meet the storage bit width requirements, and are connected with the header data and the binary representation of the bit width to complete the compression.

2. A neural network differential compression method as claimed in claim 1, characterized in that: The bit width value in each block is used as the bit width required for other differential values ​​in the current block during compression, that is, the storage bit width of other differential values. During compression, if the actual binary representation bit width of other differential values ​​is smaller than the storage bit width, it is expanded to the required storage bit width by padding with zeros.

3. A neural network differential compression method as claimed in claim 2, characterized in that: During decompression, the bit width required for the binary representation of the bit width within each block is set to a fixed value, which is b is the basic bit width, which is determined according to the distribution frequency.

4. A neural network differential compression method as claimed in claim 1, characterized in that: Determine the basic bit width according to the distribution frequency, take the basic bit width as the storage bit width for the blocks with bit width values ​​smaller than the basic bit width, convert other differential values ​​in the blocks into binary representation, and expand them to the bit width required by the basic bit width by zero padding; For blocks whose bit width is greater than or equal to the base bit width, use the bit width of each block as the storage bit width, and convert other differential values ​​into binary strings.

5. A neural network differential compression method as claimed in claim 4, characterized in that: The basic bit width is: b * =argmax b (f1(b)-f2(b)); Where f1(b) represents the reduction in the binary representation of the bit width when the base bit width is b; f2(b) represents the additional storage capacity when the binary representation of other differential values ​​in the block whose minimum bit width required for the binary representation of the maximum differential value excluding the header data is less than the base bit width when the base bit width is b; w is the quantized bit width of each differential value in the differential matrix; b is the base bit width, 0≤b <w;x i,j is the jth value in the i-th block; m i is the minimum bit width required for the binary representation of the maximum difference value excluding the header data in the i-th block; BN is the number of blocks; N is the number of weight values ​​in each block.

6. A neural network differential compression method as claimed in claim 4, characterized in that: The header data and the bit width are each converted into a binary representation, and then concatenated with the binary representation of the other differential values ​​converted to the storage bit width requirement to complete the compression.

7. A neural network differential compression method as claimed in claim 4, characterized in that: The portion of the header data starting from the low bit and having the same width as the binary representation of the bit width is replaced with the binary representation of the bit width, thereby determining the binary representation of the header data, and then connecting it with the binary representation of other differential values ​​converted into the storage bit width requirement to complete the compression.

8. A neural network differential compression method as claimed in claim 6 or 7, characterized in that: After the introduction of the base bit width BASE_WIDTH, the bit width required for the binary representation of the bit width DATA_WIDTH in each block is updated to w is the bit width value of each differential value in the differential matrix after quantization; b is the basic bit width; The correspondence between the updated bit width value and the minimum bit width required for the binary representation of the maximum difference value in the corresponding block is DATA_WIDTH new =DATA_WIDTH old -BASE_WIDTH; where DATA_WIDTH new is the updated DATA_WIDTH, DATA_WIDTH old This is the DATA_WIDTH before the update.

9. A neural network differential compression method as claimed in claim 8, characterized in that: During decompression, the bit width required for the binary representation of the bit width within each block is set to 10. A neural network differential compression system, characterized in that: include: A differential module is configured to obtain a sparse weight matrix and an index matrix of the pruned neural network, sort the weights of each row in the sparse weight matrix by the index matrix, and perform differential calculation on each row of the sorted sparse weight matrix to obtain a differential matrix; A first bit width analysis module is configured to use each row in the difference matrix as a block, and to add a bit width after the header data of each block, wherein the bit width represents the minimum bit width required for the binary representation of the maximum difference value in the block excluding the header data; A second bit width analysis module is configured to determine the value of the bit width in each block and the storage bit width of other differential values ​​excluding the header data according to the minimum bit width required for the binary representation of the maximum differential value excluding the header data in the block and its distribution frequency; The compression module is configured to convert other differential values ​​in the block into a binary string that meets the storage bit width requirements according to the storage bit width in each block, and complete the compression after connecting with the header data and the binary representation of the bit width.

Citation Information

Cited By

  • Sparse identification scheduling method, device, equipment and medium

    CN121979581A

  • A sparse identification scheduling method, device, apparatus and medium

    CN121979581B