Hardware-oriented matrix data lossless compression and decompression method and system

By adopting a multi-stage dynamic approximation frequency table model and a data block tail encoding processing method suitable for rangeCode dynamic expansion method in the data compression method, the problems of inefficient hardware efficiency and incomplete encoding information output caused by division operations in the prior art are solved, and more efficient hardware resource utilization and complete encoding information output are achieved.

CN120074536APending Publication Date: 2025-05-30HEFEI JUNZHENG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311632686.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, there is division operation in the data compression method, resulting in low hardware efficiency and high resource consumption; in addition, the traditional rangeCode method may cause incomplete output of the encoded information and lead to decoding errors during the encoding process at the end of the data block.

Method used

The multi-stage dynamic approximation frequency table model is adopted to eliminate the division operation in the traditional rangeCode method; at the same time, a data block tail encoding processing method is proposed suitable for the rangeCode dynamic expansion method. Through DOHB and DEHB operations, dynamic expansion and normalization of the data interval are realized to ensure the complete output of the encoded information.

Benefits of technology

Improves hardware efficiency and reduces resource consumption; avoids wasting data intervals, ensures the complete output of encoded information, and avoids decoding errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120074536A_ABST
    Figure CN120074536A_ABST
Patent Text Reader

Abstract

The invention provides a hardware-oriented matrix data lossless compression and decompression method and a hardware-oriented matrix data lossless compression and decompression system. The method comprises the following steps: S1, judging a compression process or a decompression process; if compression is carried out, S2 is carried out; if decompression is carried out, S3 is carried out; s2, executing a matrix data lossless compression method; and S3, executing the matrix data lossless decompression method. The system implements the method in NNA hardware, and comprises a plurality of parallel hardware modules corresponding to the lossless compression method; the lossless decompression method corresponds to a plurality of parallel hardware modules; after the NNA generates FMs data, the FMs data are firstly placed in the DRAM, the lossless data compression hardware modules compress the data of different layers / channels in parallel, and then the compressed data are transmitted to the large-capacity SRAM; when the FMs data need to be used, the data of different layers / channels in the SRAM are decompressed in parallel to the DRAM through a plurality of lossless data decompression hardware modules; in the process, the data volume of the FMs is reduced through the data compression module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of lossless data compression and decompression, and particularly relates to a hardware-oriented method and system for lossless compression and decompression of matrix data. Background Art

[0002] In the prior art, data compression, as a method for reducing the capacity of original data, is widely used in modern information storage / transmission systems. For example, in modern computer systems, a large amount of original information data detected / produced is processed by a compression system and then placed in a storage system or transmitted to the next information processing node. This process aims to save storage space or data transmission bandwidth by reducing the capacity of the original data. For a data compression system, there must be a corresponding data decompression system for restoring / partially restoring the corresponding original information data from the compressed data. Data compression / decompression depends on two important parts: compression / decompression methods and systems.

[0003] The data compression / decompression method encodes data through specific strategies (such as reducing data redundancy) to reduce the data capacity. In terms of type, data compression / decompression methods can be divided into lossy data compression / decompression methods and lossless data compression / decompression methods. The compression result obtained from the original data to be compressed through a lossy data compression method cannot completely restore the original data to be compressed after passing through the corresponding decompression method. In contrast, the compression result obtained from the original data to be compressed through a lossless data compression method can completely restore the original data to be compressed after passing through the corresponding decompression method. What this application focuses on is the lossless data compression / decompression method.

[0004] The implementation of the compression / decompression method depends on the corresponding system. For each compression / decompression method, different implementation systems may result in specific efficiencies. For the current mainstream compression / decompression methods, general von Neumann-type computing systems can achieve the corresponding functions. However, through the design of application specific integrated circuits (ASICs), each compression / decompression method can theoretically obtain a specific hardware circuit, thereby achieving the optimal state of execution efficiency and resource consumption. In this regard, the execution efficiency and resource consumption of the ASICs corresponding to different compression / decompression methods may vary greatly. Generally speaking, the lower the compression ratio of a method, the more complex it is, so the corresponding ASIC has lower execution efficiency and higher resource consumption. However, in any case, a lower compression ratio, higher execution efficiency, and less resource consumption are always the goals and pursuits of compression / decompression methods and corresponding ASIC designs.

[0005] In the prior art, for example, in the overall framework of the traditional rangeCode lossless compression method, essentially speaking, the rangeCode method is similar to arithmetic coding (Salomon D. Data compression: the complete reference. Springer, London, 2007, 24(10): 635 - 42). The basic principle of the rangeCode method is to map each data element to be encoded to a sub - interval of an overall interval according to the cumulative counting distribution (CCD) of the elements. By this method, high - frequency elements will be mapped to larger sub - intervals, thus delaying the contraction speed of the overall interval. This also means that at this time, a fixed - length overall interval can accommodate more elements, so the overall compression ratio is lower. On the contrary, low - frequency elements will be mapped to smaller sub - intervals, thus accelerating the contraction speed of the overall interval. This also means that at this time, a fixed - length overall interval can accommodate fewer elements, so the overall compression ratio is higher.

[0006] Assume that the bit - width of each element is n, then the range of unsigned integers that each element can represent is k ∈ [0, 2 n ). For the i - th element to be encoded, the corresponding CCD can be expressed as the following formula (1), where V(i) represents the specific unsigned value of the i - th element to be encoded. When k ≤ V(i), δ(k, V(i)) = 1, otherwise, δ(k, V(i)) = 0:

[0007]

[0008] In formula (1), it can be noted that the right boundary of the values of k and V(i) (2 n ) has exceeded the unsigned value limit of an n - bit element. Here, 2 n is used as the end of file (EOF) and applied at the end of each data block.

[0009] According to the CCD, when encoding the i - th element to be encoded, the cumulative probability distribution function (CDF) of the unsigned value k of each element is represented by the following formula (2):

[0010]

[0011] In the rangeCode method, there is a global interval. As the encoding progresses, the range of this interval also changes continuously. The range of this interval is dynamically represented by its left and right endpoints rangeLeft (RL) and rangeRight (RR). RL and RR are represented by (Y + 1)-bit data, denoted here as RL[0:Y] and RR[0:Y]. For the sake of convenience of expression, here we use RL i [0:Y] and RR i [0:Y] to represent the values of the corresponding left and right endpoints when encoding the i-th element respectively.

[0012] Initially, when the first (i = 0) generation of encoded elements is input, RL 0 [0:Y] and RR 0 [0:Y] are respectively set to b’000...00 and b’111...11. In other words, at this time RL 0 [0:Y] = 0, RR 0 [0:Y] = 2 Y+1 -1. For the subsequent i-th input element, assuming its corresponding value is V(i), then RL i [0:Y] and RR i [0:Y] will be updated to RL i +1 [0:Y] and RR i+1 [0:Y] according to formula (3):

[0013]

[0014] Here, RW i [0:Y] represents the interval width before encoding the i-th element. Obviously, when the interval transformation is completed through formula (3), this width will be updated to RW i+1 [0:Y] = RR i+1 [0:Y] - RL i+1 [0:Y], and the new width will be used for encoding the (i + 1)-th element.

[0015] Figure 1 Shows the change situation of the left and right endpoints RL i [0:Y], RR i [0:Y] and the interval width RW i [0:Y] during the encoding process. When RL N [0:Y] ≥ RR N [0:Y], the rangeCode method will output a (Y + 1)-bit encoding result (in fact, it is the left interval RL NThe value of [0:Y]), and at the same time, the current coding unit (from the 0th element to the (N - 1)th element) will also be terminated. Here, the Nth element will serve as the dividing point between the previous coding unit (from the 0th element to the (N - 1)th element) and the next coding unit (from the Nth element to subsequent elements). For the next coding unit (from the Nth element to subsequent elements), the rangeCode method will reset RL N [0:Y] = b'000...00, RR N [0:Y] = b'111...11 to complete the initialization.

[0016] For the data decompression process, for each coding unit, there is a corresponding decoding unit. For each decoding unit, a storage space with a capacity of Y + 1 bits needs to be pre-allocated, denoted here as currentCode[0:Y]. During the initialization process, Y + 1 bits need to be read from the encoded data into currentCode[0:Y]. For the sake of convenient description, currentCode i [0:Y] is used to represent the relevant encoded data read when decoding to the ith element.

[0017] Data decompression actually utilizes the following feature shown in formula (4) formed during the data compression process:

[0018]

[0019] As can be seen from formula (4), RL N [0:Y] ≤ currentCode 0 [0:Y] < RR N [0:Y]. Therefore, using the following steps, the first element (i = 0) in the first coding unit can be decompressed according to currentCode 0 [0:Y]:

[0020] Step0: During the initialization process of each decoding unit, read Y + 1 bits from the encoded data to fill currentCode 0 [0:Y]. At the same time, set RL 0 [0:Y] = b'000...00, RR 0 [0:Y] = b'111...11. Set CCD(k; i, V(i)) to CCD(k; 0). Set P(k; i, V(i)) to P(k; 0);

[0021] Step1: For each k ∈ [0, 2 n ), calculate according to the following formula (5) and If then the first decoded element is k0.

[0022]

[0023] Step2: Update

[0024] Step3: According to formulas (1) and (2), use the decoded element k0 to update P(k; 0) to P(k; 1, k0).

[0025] Step4: Update currentCode 1 [0:Y] = currentCode 0 [0:Y].

[0026] Repeat Step1 to Step4, then all the data included in currentCode 0 [0:Y] can be decoded.

[0027] It should be noted here that for the first decoding unit, currentCode 0 [0:Y] == currentCode 1 [0:Y] ==.... And for the next decoding unit, the rangeCode method will execute Step0 to read the next Y + 1 bits of data from the encoded data and fill it into currentCode N [0:Y].

[0028] From the compression process of each encoding unit, it can be seen that as the compression progresses, the corresponding width RW i [0:Y] = RR i [0:Y] - RL i [0:Y] is gradually shrinking. And at the end of each encoding unit, inevitably, there is RW N [0:Y] ≥ 0, which actually leads to waste of the interval. Here is an example to illustrate: RL i [0:9] = b’1111 1100 00 and RR i [0:9] = b’1111 1111 11, then RW i [0:9] = b’0000 0011 11. At this time, assume that P(V(i) - 1; i - 1, V(i - 1)) = 0.1 and P(V(i); i - 1, V(i - 1)) = 0.10001. According to formula (3), RL i [0:Y] == RR i[0: Y], which means that the encoding of the current unit must be interrupted and 10-bit encoding is output, that is, output RL i [0: Y]. It can be seen that 4-bit data intervals will be wasted at this time.

[0029] To avoid the problem of wasted data intervals mentioned above, there is an improvement scheme here. This improvement scheme is an existing technical solution, and this improvement scheme provides a partial theoretical basis for the technical solution of this application. For the convenience of description, it is called the rangeCode dynamic expansion strategy.

[0030] The above rangeCode method interrupts the encoding of the current encoding unit and outputs the complete encoding RL of the unit when the critical condition RL i [0: Y] ≥ RR i [0: Y] is triggered. Different from this, the rangeCode dynamic expansion strategy will try to output partial encoding and perform re-normalization operations on the interval after completing the encoding process of each element. As i [0: Y]. Figure 2 shows the simplest re-normalization operation in the rangeCode dynamic expansion strategy.

[0031] In Figure 2 , Y is set to 9. The re-normalization process is very intuitive: compare the bit data at each same address (from 0 to Y) of RL[0: Y] and RR[0: Y]. If the same bit data is encountered, output the corresponding bit data at that address and shift RL[0: Y] and RR[0: Y] to the left. This process continues until different bit data of RL[0: Y] and RR[0: Y] is encountered at a certain address, and the process is interrupted. During the above process of shifting RL[0: Y] and RR[0: Y] to the left, for each bit shifted, the tail of RL[0: Y] will be filled with bit "0" and the tail of RR[0: Y] will be filled with bit "1". For the Figure 2 example shown, after the above normalization operation, RW[0: Y] = RR[0: Y] - RL[0: Y] will be expanded from 62 to 251, which means that only at the cost of 2 bits (outputting "10") can the interval range be expanded to 255 / 63 ≈ 4.05 times. Therefore, in the example described currently, the re-normalization operation will bring the following two potential benefits:

[0032] (1) Expand the interval range from 62 to 251, which can, to a certain extent, avoid the problem that the critical condition is triggered due to the too small interval, resulting in the forced interruption of the current encoding unit and the output of 10-bit encoding, thus causing the waste of 8-bit data intervals (the interval of size 62 is not utilized);

[0033] ​(2) Since 2-bit data usually can only ensure that the interval size is expanded by 4 times, while the interval here is expanded by about 4.05 times. Therefore, it shows that this re-normalization operation obtains a larger data interval at a smaller cost.

[0034] For the sake of convenient expression, the re-normalization method that detects and outputs each bit data in the above-mentioned sequential detections RL[0:Y] and RR[0:Y] is called DOHB here. Denote the bit length detected and output during the DOHB process as L DOHB (i).

[0035] In fact, RL[0:Y] and RR[0:Y] can be further re-normalized. Here at the beginning, define an internal variable L DEHB (i) and set it to 0. As Figure 3 shown, detect the highest 2-bit numbers of RL[0:Y] and RR[0:Y], f2bL = RL[0:1], f2bR = RR[0:1]. If f2bL = b'01 and f2bR = b'10, then the second bit data in RL[0:Y] and RR[0:Y] will be removed respectively. At the same time, similar to the DOHB operation, the tail of RL[0:Y] will be filled with the bit "0" while the tail of RR[0:Y] will be filled with the bit "1". The above detection and removal operations will continue until the following pattern is encountered: when f2bL = b'01 and f2bR = b'10, the operation terminates. The difference from the DOHB operation here is that when each bit data is removed, no corresponding output will be generated, but the internal variable L DEHB (i) is incremented by 1. For the sake of convenient narration, this strategy is called DEHB.

[0036] Combining DOHB and DEHB, the following normalization strategy can be obtained (for the sake of convenient expression, it is simply called the combined DOHB and DEHB strategy here):

[0037] (1) For the i-th coding element, if the updated RL i+1 [0:Y] and RR i+1 [0:Y] are hit by the DOHB operation, that is, L DOHB (i)>0, which means RL i+1 [0:L DOHB (i)-1]==RR i+1 [0:L DOHB (i)-1]. In the current situation, L DOHB (i)+L DEHB (i-1) bit data will be output: And after the output, set L DEHB (i-1)=0;

[0038] (2) For the i-th encoded element, regardless of whether the DOHB operation hits or not, if L DEHB (i) ≥ 0, then the DEHB operation must be executed, and let L DEHB (i) += L DEHB (i - 1).

[0039] Under the current strategy of combining DOHB and DEHB operations, RW[0:Y] = RR[0:Y] - RL[0:Y] will be further extended from 62 to 251 and improved to 62 to 959. This means that the data interval range is further expanded, resulting in a further reduction in the probability of triggering the interruption critical condition during the encoding process of the next element, thereby potentially reducing the probability of waste of the data interval. It is particularly noteworthy that after combining DOHB and DEHB operations, it can be inferred that the data interval width RW i [0:Y] = RR i [o:Y] - RL i [0:Y] has a minimum value in the following two cases:

[0040] 1) RL i [0∶Y] = b′00111..111 = 2 Y - 1 -1 and RR i [0:Y] = b′10000...000 = 2 Y :

[0041] 2) RL i [0:Y] = b′01111..111 = 2 Y -1 and RR i [0:Y] = b′11000...000 = 2 Y +2 Y-1 ;

[0042] Therefore, a very useful conclusion can be inferred:

[0043] For any i, RW i [0:Y] ≥ 2 Y-1 +1. This means that by combining DOHB and DEHB operations, the left and right endpoints RL[0:Y] and RR[0:Y] of the interval represented by Y + 1 bits can ensure that the lower limit of the data interval RW[0:Y] for each element is 2 Y-2 +1. This conclusion provides some theoretical basis for the subsequent method innovation based on the rangeCode method in this application.

[0044] In contrast, the upper limit of RW[0:Y] must exist in the following case: RL i[0:Y] = b'00000..000 = 0 and RR i [0:Y] = b'11111...111 = 2 Y+1 One - 1. Therefore, the upper limit of RW[0:Y] is 2 Y+1 -1. In summary, the range of RW[0:Y] can be obtained as:

[0045] 2 Y-2 +1 ≤ RW[0:Y] ≤ 2 Y+1 -1 (6).

[0046] After all the original data elements are encoded, at the end, an additional value EOF = 2 needs to be encoded n , since the count of EOF remains 1, it is expected that the encoding of this low - frequency value EOF can make the corresponding L DOHB >0, thus triggering the operation (1) in the combined DOHB and DEHB strategy obtained by combining DOHB and DEHB above, and outputting all the previously accumulated L DEHB bit data, and further expecting to ensure that all the encoded information is output.

[0047] However, the deficiencies of the existing technology are as follows:

[0048] (1), when encoding / decoding each element, the formula (2) needs to be used to calculate the CDF, but the formula (2) itself contains a division operation. In hardware - related methods, the existence of division will greatly reduce the hardware efficiency and increase resource consumption. For example, in the ASIC design process, the existence of a divider will greatly increase the circuit delay and area, resulting in an increase in data flow blockage and chip cost. In fact, for the current mainstream neural network accelerator (NNA) with a frequency of GHz, if the above rangeCode method is used to design a data compression hardware module to compress neural network feature data, the existence of the divider in it will increase the overall delay of the method by dozens of clock cycles, thus causing the entire rangeCode method hardware module to block the entire hardware system.

[0049] (2), the existing RangeCode method encodes an additional element EOF = 2 n to achieve the purpose of processing the end of the data, but this approach has the following deficiencies:

[0050] A Because the additional element EOF = 2 n has a very low counting frequency, this will cause a long string of bit data to be output at the end, thus reducing the compression effect of the method (increasing the compression ratio);

[0051] B Due to the existence of the extra element EOF = 2 n , the types of all elements in the hardware system will change from 2 n to misaligned 2 n +1, which will waste hardware space. For example, n + 1 bits are required to record the numerical size of each element. In addition, this misaligned state will also reduce the efficiency of some data processing methods, such as the binary search method during the decoding process;

[0052] C In hardware, to maximize the use of storage space, the encoded results of different data blocks may be stored continuously together. If the data blocks corresponding to these adjacent encoded results use independent CDF or CCD during the encoding process, to ensure the correctness of decoding, some additional information (such as the starting address of the encoded result corresponding to each data block) must be recorded, and some complex masking operations must be performed based on this information, which increases the complexity of the overall hardware implementation;

[0053] D It can be shown that the above rangeCode method will cause the encoded information at the end of the data block not to be fully output in some cases, resulting in potential decoding errors:

[0054] Assume that EOF is the (i + 1)-th element. According to formulas (1) and (2), it can be found that P(EOF; i - V(i)) ≡ 1. Then, according to formula (3), there must be:

[0055] RR i+1 [0:Y] = RL i [[0:Y] + RW i [0:Y] × P(EOF; i, V(i)) - 1 = RR i [0:Y] (7)

[0056] Furthermore, according to formulas (2) and (3):

[0057]

[0058] Considering that the count of EOF remains 1, that is, CCD(EOF; i, V(i)) - CCD(EOF - 1; i, V(i)) ≡ 1, combining inequality (6) with formula (7), then formula (8) becomes:

[0059]

[0060] According to inequality (6), combined with formula (9), it can be inferred that:

[0061]

[0062] In fact, it can be observed that RL cannot be guaranteed solely based on formula (10). i+1 [0:0] == RR i [0:0] == RR i+1 [0:0], thus L cannot be guaranteed DOHB (i + 1)>0, which in turn triggers operation (1) in the combined DOHB and DEHB strategies. For example, according to the combined DOHB and DEHB strategies, RR i [0:1] == b′10 or RR i [0:1] == b′11. When RR i [0:1] == b′10 ≥ 2 Y At this time, there must be RL i [0:1] == b′00 ≥ 0, thus RW i [0:Y] == RR i [0:1] - RL i [0:1] ≥ 2 Y-1 . Assume that at this time RR i+1 [0:Y] == RR i [0:Y] == 2 Y , RW i [[0:Y] == 2 Y-1 , i = 1, n = 1, then according to formula (9), RL can be obtained as i+1 [0:Y] == 2 Y -2 Y-3 == 7×2 Y-3 . Obviously, RR i+1 [0:0] ≠ RL i+1 [0:0], that is, L DOHB (i + 1) == 0, so operation (1) in the combined DOHB and DEHB strategies cannot be triggered, and as a result, the previously accumulated L DEHB (i - 1) bits of data cannot be output, which may lead to decoding errors at the end.

[0063] In addition, common terms in the prior art include:

[0064] Data compression: Encoding the original data through specific method steps to obtain encoded data. The purpose is to make the total capacity of the encoded data different from that of the original data.

[0065] Data decompression: Decoding the encoded data through a set of inverse methods corresponding to data compression to obtain decompressed data. During lossless compression / decompression, the decompressed data should be exactly the same as the original data.

[0066] Compression ratio: Total capacity of encoded data / Total capacity of original data. It can be seen that the lower the compression ratio, the better the data compression effect.

[0067] Positive compression: Compression ratio is less than 1.

[0068] Negative compression: Compression ratio is greater than 1.

[0069] Arithmetic coding: Arithmetic coding is one of the main methods of image compression. It is a lossless data compression method and also a method of entropy coding.

[0070] Feature maps (FMs): Two-dimensional or three-dimensional matrix data containing the features of the current layer obtained by each layer of the neural network through convolution or other feature operations.

[0071] SRAM: static random access memory, that is, static random access memory.

[0072] DRAM: dynamic random access memory, that is, dynamic random access memory.

[0073] ASIC: application specific integrated circuit, that is, application specific integrated circuit.

[0074] CCD: cumulative counting distribution, that is, cumulative frequency distribution.

[0075] CDF: cumulative probability distribution function, cumulative probability distribution function.

[0076] CPU: central processing unit, that is, central processing unit.

[0077] MCU: micro-control unit, that is, micro-control unit. Summary of the Invention

[0078] In order to solve the disadvantages and problems existing in the above technologies, the purpose of this application is to disclose a hardware-oriented data lossless compression and decompression method and system:

[0079] (1) Disclosed is a multi-stage dynamic approximate frequency table model and its application. Using this model, while realizing the dynamic corresponding data change law, the division operation included in the frequency calculation in the traditional rangeCode method can be eliminated.

[0080] (2) discloses a method for processing the tail coding of data blocks applicable to the rangeCode dynamic expansion method. Compared with the above-mentioned traditional data block tail coding processing technology, the implementation complexity of this method is quite low (which is beneficial to hardware implementation), and this method can avoid the potential problem of incomplete output of coding information caused by the above-mentioned traditional tail coding.

[0081] (3) discloses a dynamic frequency table partial retention strategy applicable to the rangeCode dynamic expansion method. In combination with the above point (2), it can make better use of the spatial correlation between data.

[0082] Specifically, the present invention provides a hardware-oriented lossless compression and decompression method for matrix data, characterized in that the method includes the following steps:

[0083] S1. Determine whether to perform a compression or decompression process according to the instruction code issued by the CPU / MCU? If it is lossless compression of matrix data, further proceed to step S2; if it is lossless decompression of matrix data, further proceed to step S3;

[0084] S2. The lossless compression method for matrix data includes:

[0085] E1. Obtain the starting address of the data block;

[0086] E2. Initialize the relevant variables required for encoding the current data block;

[0087] E3. Sequentially read the value of the i-th element from the data block; execute step E4 (which includes two sub-steps E4-1 and E4-2 that need to be executed serially) and step E5 (which includes two sub-steps E5-1 and E5-2 that need to be executed serially) in parallel;

[0088] E4:

[0089] E4-1. Calculate RL i+1 [0:Y] and RR i+1 [0:Y];

[0090] E4-2. Re-normalize and output;

[0091] E5:

[0092] E5-1. Update CCD;

[0093] E5-2. Update the relevant information of the approximate frequency table;

[0094] E6. After both E4 (including sub-steps E4-1 and E4-2) and step E5 (including sub-steps E5-1 and E5-2) are completed, determine whether the end of the data block is reached. If so, proceed to step E7; otherwise, proceed to step E8;

[0095] E7. Process the tail encoding of the data block; return to step E1;

[0096] E8. The index i of the element to be compressed is incremented by 1; return to step E3;

[0097] S3. The matrix data lossless decompression method includes:

[0098] D1. Initialize the relevant variables required for decoding the current data block;

[0099] D2. Read Y + 1 bits of data from the encoding result to fill the variable or register array currentCode[0:Y];

[0100] D3. Search and find the decoded value, and output the obtained current decoded value; execute step D4 (which includes two sub-steps D4-1 and D4-2 that need to be executed serially) and step D5 (which includes two sub-steps D5-1 and D5-2 that need to be executed serially) in parallel;

[0101] D4:

[0102] D4-1. Calculate RL i+1 [0:Y] and RR i+1 [0:Y];

[0103] D4-2. Re-normalize;

[0104] D5:

[0105] D5-1. Update CCD;

[0106] D5-2. Update the relevant information of the approximate frequency table;

[0107] D6. After both D4 (including sub-steps D4-1 and D4-2) and step D5 (including sub-steps D5-1 and D5-2) are completed, determine whether the end of the data block is reached. If so, proceed to step D7; if not, proceed to step D8;

[0108] D7. Process the tail encoding of the data block; return to step D1;

[0109] D8. Increment the decompression element index i by 1; return to step D3.

[0110] The said step S2 further includes the following detailed steps:

[0111] E1. Obtain the starting address of the data block;

[0112] All data is divided into multiple independent data blocks. For example, in the NNA, each channel on the feature maps (FMs) will be regarded as an independent data block for compression due to different data distribution characteristics; multiple data compression / decompression modules are used on the hardware to compress the data in parallel, so that each data block can be mapped to an independent compression / decompression module during the parallel computing process; inside each independent data block, the data is continuously stored. Therefore, before compression, the corresponding starting data address needs to be obtained, denoted as addr_begin;

[0113] E2. Initialize the relevant variables required for encoding the current data block;

[0114] Before compressing each data block, relevant initialization work is required, including the following sub-steps:

[0115] E2-1. Use the following formula (11) to initialize the cumulative frequency distribution CCD:

[0116] CCD(k; 0) = k + 1, k ∈ [0, 2 n ) (11)

[0117] E2-2. Initialize the values of the left endpoint RL and the right endpoint RR: RL 0 [0:Y] = b’000...00, RR 0 [0:Y] = b’111...11;

[0118] E2-3. Initialize the values of L DOHB and L DEHB : L DOHB (0) = 0 and L DEHB (0) = 0;

[0119] E2-4. Initialize the relevant data of the multi-stage dynamic approximate frequency table model;

[0120] The relevant construction values of this multi-stage dynamic approximate frequency table model are shown in Table 1 below:

[0121] Table 1:

[0122]

[0123] To apply this table, the user needs to manually select a T 0 ≥5 to obtain table entries with different precisions; initially, set u = 0, h = 0, T = T 0 ; in this way, according to Table 1, the initial mapping relationship between and can be obtained;

[0124] E3. Sequentially read the value of the i-th element from the data block;

[0125] Assume that each element is represented by n bits. Then the read data is V(i)[0:n-1] = memory_ori[addr_begin:addr_begin+n-1], where memory_ori[] represents the physical address space / virtual address space range where the original generation code data is stored;

[0126] E4. Calculate RL i+1 [0:Y] and RR i+1 [0:Y], and renormalize, including the following two sub-steps E4-1 and E4-2 (which can be executed in parallel with the following E5, and the E4-1 and E4-2 steps inside E4 are executed serially);

[0127] E4-1: Calculate RL[0:Y] and RR[0:Y] according to the following formula (12): i+1 [0:Y] and RR i+1 [0:Y]:

[0128]

[0129] In fact, the calculation of RL[0:Y] and RR[0:Y] shown in formula (12) can be further transformed by the following formula (13) to reduce the bit width of the data in the multiplication operations involved, which will be beneficial to selecting a lower-bit-width multiplier in the ASIC implementation process to improve the hardware efficiency, that is, reduce the latency of the hardware multiplier and save hardware resources: i+1 [0:Y] and RR i+1 [0:Y]:

[0130]

[0131] In formula (13), S 0 、S 1 、S 2 、S 3 are user-adjustable parameters, determined offline by the user;

[0132] E4-2. Renormalize and output the data; this process needs to perform two basic operations: DOHB and DEHB;

[0133] E5. Update CCD and update the relevant information of the approximate frequency table, including the following two sub-steps E5-1 and E5-2 (which can be executed in parallel with the above E4, and the E5-1 and E5-2 steps inside E5 are executed serially):

[0134] E5-1: Update the CCD, which includes the following two sub-steps E5-1-1 and E5-1-2 executed serially:

[0135] E5-1-1: For each k ∈ [0, 2 n ), update CCD(k; i, V(i)) = CCD(k; i - 1, V(i - 1)) + δ(k, V(i));

[0136] E5-1-2: If CCD(2 n - 1; i, V(i)) == 2 T , then for each k ∈ [0, 2 n ), perform the following two calculation operations: CCD(k; i, V(i)) <<= 1, CCD(k; i, V(i)) += 1;

[0137] E5-2. Update the relevant information of the approximate frequency table, including the following two sub-steps E5-2-1 and E5-2-2 executed sequentially:

[0138] E5-2-1: If then, for each h, update u = u + 1;

[0139] E5-2-2: If then update h = h + 1;

[0140] E6. When both E4 (including sub-steps E4-1 and E4-2) and step E5 (including sub-steps E5-1 and E5-2) are completed, determine whether the end of the data block is reached? If so, proceed to step E7, otherwise proceed to step E8;

[0141] E7. Process the tail coding of the data block, including:

[0142] E7-1: Output a single-bit FB = (RL i [1:1] | RR i [1:1]);

[0143] E7-2: Output L DEHB (i)

[0144] E7-3: Output a single-bit

[0145] Through this data block tail coding process, the coding results of different data blocks using different frequency tables can be continuously stored together, and the correctness of the decoding process can be guaranteed without storing additional auxiliary information;

[0146] E8. Update the index i of the element to be compressed: i = i + 1.

[0147] In the said step E2,

[0148] In Table 1, if T 0 = 5 is selected, then initially corresponding to If T 0 = 8 is selected, then initially then corresponding to

[0149] In the said step E4-1,

[0150] The parameters S 0 、S 1 、S 2 、S 3 must satisfy the constraint conditions shown in the following formula (14):

[0151]

[0152] It is not difficult to find from formula (13) that, under the condition of satisfying the constraint condition (14), by adjusting the parameters S 0 、S 1 、S 2 、S 3 , the bit width of the multiplication operation involved in formula (13) can be made the shortest, that is, in the process of ASIC implementation of the method, a multiplier with the shortest bit width is selected.

[0153] In the said step E4-2,

[0154] The description of the DOHB operation is as follows: Compare the bit data at each same address of RL i+1 [0:Y] and RR i+1 [0:Y]. The same address ranges from 0 to Y. If the same bit data is encountered, output the corresponding bit data at this address and shift RL[0:Y] and RR[0:Y] one bit to the left. This process continues until different bit data of RL[0:Y] and RR[0:Y] is encountered at a certain address, and then the process is interrupted; in the above process of shifting RL[0:Y] and RR[0:Y] to the left, each time one bit is shifted, a bit '0' will be filled at the tail of RL[0:Y] and a bit '1' will be filled at the tail of RR[0:Y];

[0155] The description of the DEHB operation is as follows: Detect RL i+1 [0:Y] and RR i+1The top 2 bits of [0:Y], f2bL = RL[0:1], f2bR = RR[0:1]; if f2bL = b'01 and f2bR = b'10, then the second bit data in RL[0:Y] and RR[0:Y] will be removed respectively; at the same time, the tail of RL[0:Y] will be filled with bit "0" and the tail of RR[0:Y] will be filled with bit "1"; the above detection and removal operations will continue until the following pattern is encountered: when f2bL = b'01 and f2bR = b'10, the operation terminates; different from the DOHB operation here, when removing each bit data, no corresponding output will be generated, but the internal variable L DEHB (i) Increment by 1;

[0156] According to the DOHB and DEHB operations, the steps of renormalization and output are as follows:

[0157] (1). For the i-th element to be compressed, if the updated RL i+1 [0:Y] and RR i+1 [0:Y] are hit by the DOHB operation, that is, L DOHB (i) > 0, which means RL i+1 [0:L DOHB (i)-1] == RR i+1 [0:L DOHB (i)-1]; in the current case, output L DOHB (i)+L DEHB (i-1) bit data: And set L DEHB (i-1) = 0 after output;

[0158] (2). For the i-th element to be compressed, regardless of whether the DOHB operation is hit, if L DEHB (i) > 0, then the DEHB operation must be performed, and let L DEHB (i)+ = L DEHB (i-1).

[0159] Step S3, the matrix data lossless decompression method, is the reverse decompression method supporting the matrix data lossless compression method in step S2; the detailed steps of the method are described as follows:

[0160] D1. Set the initial values of the variables involved in the current data block decoding process, including:

[0161] D1-1. In the step S2, the complete data is segmented into multiple independent data blocks and the encoding results of each data block may be stored continuously; therefore, during the decoding process, it is necessary to read the initial address of the encoded data corresponding to the corresponding data block, denoted here as addr_begin_dec;

[0162] D1-2. Initialize the CCD according to the following formula (15):

[0163] CCD(k; 0) = k + 1, k ∈ [0, 2 n ) (15)

[0164] D1-3. Initialize the values of RL and RR: RL 0 [0:Y] = b’000...00, RR 0 [0:Y] = b’111...11;

[0165] D1-4. Initialize the values of L DOHB and L DEHB : L DOHB (0) = 0 and L DEHB (0) = 0;

[0166] D1-5. Initialize the value of the Y+1-bit variable currentCode: currentCode[0:Y] = b’000...00;

[0167] D1-6. Initialize the data related to the multi-stage dynamic approximate frequency table model;

[0168] The related construction values of the multi-stage dynamic approximate frequency table model are shown in Table 1;

[0169] Applying this table requires the user to manually select a T 0 ≥5, so as to obtain table entries with different precisions; initially, set u = 0, h = 0, T = T 0 ;

[0170] D2. According to addr_begin_dec, read Y+1-bit data from the encoded data corresponding to the current data block and fill it into the variable currentCode[0:Y]:

[0171] currentCode[0:Y] = memory-enc[addrbegin_dec:addrbegin_dec+Y];

[0172] Here, memory_enc[] represents the physical address space / virtual address space range where the original encoded data is stored; after the reading is completed, addrbegin_dec += shiftBits;

[0173] D3. Search for and find the decoded value, and output the obtained current decoded value:

[0174] For each k ∈ [0, 2 n ), calculate the corresponding and If there exists a unique k such that the condition: is satisfied, then the value obtained from the current decoding is k, i.e., V(i) = k;

[0175] According to formula (16), since and increase monotonically with the increase of k, this process can be carried out using binary search:

[0176]

[0177] D4. Calculate RL i+1 [0:Y] and RR i+1 [0∶Y], and renormalize, including the following two sub-steps D4-1 and D4-2 (which can be executed in parallel with the following D5, and the steps D4-1 and D4-2 inside D4 are executed serially);

[0178] D4-1: According to the following formula (17), use the decoded value V(i) to update the RL and RR values:

[0179]

[0180]

[0181] D4-2. Renormalize; this process requires two basic operations to be executed successively: DOHB and DEHB;

[0182] D5. Update the CCD and related information of the approximate frequency table, including the following two sub-steps D5-1 and D5-2 (which can be executed in parallel with the above

[0183] D4, and the sub-steps D5-1 and D5-2 inside D5 are executed serially):

[0184] D5-1: Update the CCD, including the following two sub-steps D5-1-1 and D5-1-2 that are executed serially:

[0185] D5-1-1: For each k ∈ [0, 2 n ), update CCD(k; i, V(i)) = CCD(k; i - 1, V(i - 1)) + δ(k, V(i));

[0186] D5-1-2: If CCD(2n -1; i, V(i)) == 2 T , then for each k ∈ [0, 2 n ), the following two calculation operations are performed:

[0187] CCD(k; i, V(i)) <<= 1, CCD(k; i, V(i)) += 1;

[0188] D5-2. Update the relevant information of the approximate frequency table, including the following two sub-steps D5-2-1 and D5-2-2 executed in sequence:

[0189] D5-2-1: If then, for each h, update u = u + 1;

[0190] D5-2-2: If then update h = h + 1;

[0191] D6. When both D4 (including sub-steps D4-1 and D4-2) and step D5 (including sub-steps D5-1 and D5-2) are completed, determine whether the end of the data block is reached? If so, perform step D7, if not, perform step D8;

[0192] D7. Process the tail coding of the data block: Different from the coding process, here only need to directly shift currentCode[0:Y] two bits to the left;

[0193] D8. Update the value of the decompression element index i: i = i + 1.

[0194] In the said step D4-2, it further includes:

[0195] The description of the said DOHB operation is as follows: Compare RL i+1 [0:Y] with each bit data at the same address of RR i+1 [0:Y], the same address is from 0 to Y. If the same bit data is encountered, shift RL i+1 [0:Y], RR i+1 [0:Y] and currentCode[0:Y] one bit to the left each. This process continues until at a certain address, RL i+1 [0:Y] and RR i+1 [0:Y] have different bit data, then this process interrupts; During the above process of shifting RL i+1 [0:Y], RR i+1 [0:Y] and currentCode[0:Y], each time one bit is shifted to the left, the tail of RL i+1 [0:Y] will be filled with the bit "0" while RR i+1The tail will be filled with bit "1"; meanwhile, read a bit of data: memory_enc[addrbegin_dec: addr_begin - dec] is filled to the tail of currentCode[0:Y] and addr_begin_dec += 1:

[0196] The description of the DEHB operation is as follows: Detect RL i+1 [0:Y] and RR i+1 [0:Y] of the highest 2 bits f2bL = RL i +1 [0:1], f2bR = RR i+1 [0:1]; if f2bL = b'01 and f2bR = b'10, then RL i+1 [0:Y], RR i+1 [0:Y] and the second bit data in currentCode[0:Y] will be removed respectively; the above detection and removal operations will continue until the following pattern is encountered: when f2bL = b'01 and f2bR = b'10, the operation terminates; during the above left shift of RL i+1 [o:Y], RR i+1 [0:Y] and currentCode[0:Y], for each left shift of one bit, the tail of RL i+1 [0:Y] will be filled with bit "0" while the tail of RR i+1 will be filled with bit "1"; meanwhile, read a bit of data: memory_enc[addrbegin_dec: addr_begin_dec] is filled to the tail of currentCode[0:Y] and addr_begin_dec += 1.

[0197] This application also relates to an application system for a hardware - oriented matrix data lossless compression and decompression method. The application system implements any of the above - mentioned methods in a neural network acceleration dedicated chip NNA. The system includes:

[0198] Multiple parallel hardware modules ED1...EDN corresponding to the lossless compression method;

[0199] Multiple parallel hardware modules DD1...DDN corresponding to the lossless decompression method;

[0200] After the NNA generates FMs data, it is first placed in a dynamic random access memory DRAM. Multiple lossless data compression modules DE1...EDN perform parallel compression on data of different layers / channels therein. Then, the compressed data will be transferred to a large - capacity static random access memory SRAM;

[0201] When these FMs data need to be used, the data of different layers / channels in the SRAM are decompressed in parallel to the DRAM by multiple lossless data decompression modules DD1...DDN; in this process, the data volume of the FMs is reduced by the data compression module.

[0202] Therefore, the advantages of this application are as follows:

[0203] (1) This application discloses a multi-stage dynamic approximate frequency table model and its application. Using this model, while realizing the dynamic response to the data change law, the division operation included in the frequency calculation in the traditional rangeCode method can be eliminated, which will greatly reduce the hardware time delay and resource consumption.

[0204] (2) This application discloses a method for processing the tail coding of data blocks applicable to the rangeCode dynamic expansion method. Compared with the above traditional data block tail coding processing technology, the implementation complexity of this method is quite low (which is beneficial to the implementation of hardware), the compression effect is better, and more importantly, this method can avoid the potential problem of incomplete output of coding information caused by the above traditional tail coding.

[0205] (3) This application discloses a dynamic frequency table partial retention strategy applicable to the rangeCode dynamic expansion method. Combined with the above point (2), it can better utilize the spatial correlation between data to further improve the compression effect. Description of the Drawings

[0206] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not constitute a limitation to the present invention.

[0207] Figure 1 It is a schematic diagram of the interval change situation in the implementation process of the existing rangeCode method.

[0208] Figure 2 It is an example diagram of the process of re-normalizing and outputting RL[0:Y] and RR[0:Y] by the existing DOHB operation.

[0209] Figure 3 It is an example diagram of the process of re-normalizing and outputting RL[0:Y] and RR[0:Y] by the existing DEHB operation through detecting and eliminating the head "10" "01" patterns.

[0210] Figure 4 It is a flowchart of the matrix data lossless compression method in step S2 of the method of this application.

[0211] Figure 5 It is a flowchart of the matrix data lossless decompression method in step S3 of the method of this application.

[0212] Figure 6 It is a schematic diagram of the original hardware architecture for implementing a method of lossless compression and decompression of matrix data in the NNA of the system proposed in this application. Detailed implementation manners

[0213] In order to more clearly understand the technical content and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.

[0214] The embodiment of this application relates to a method and system for lossless compression and decompression of matrix data for hardware, and the method includes the following steps:

[0215] S1. Determine whether to perform a compression or decompression process according to the instruction code issued by the CPU / MCU? If it is lossless compression of matrix data, step S2 is further performed; if it is lossless decompression of matrix data, step S3 is further performed;

[0216] S2. Method for lossless compression of matrix data;

[0217] S3. Method for lossless decompression of matrix data.

[0218] As Figure 4 shown, it is a flowchart corresponding to the method of lossless compression of matrix data included in the present invention. According to Figure 4 the flowchart shown, the detailed steps of the method are described as follows:

[0219] E1. Obtain the starting address of the data block. In the method of lossless data compression provided by the present invention, all the data generated by the system may be divided into multiple data blocks. For example, in the overall data of the feature maps (FMs) generated in the middle of a Convolution Neural Network (CNN), the frequency distributions of data in different layers / channels may vary greatly. Therefore, the data in different layers / channels are divided here to obtain multiple data blocks, and independent CCD / CDF are used for these data blocks respectively, which can maximize the use of the unique data distribution characteristics of each data block, so as to expect to obtain better compression effects. In fact, in order to improve the data throughput, multiple data compression / decompression modules can often be used on the hardware to perform parallel compression / decompression on the data, which also requires the data to be divided into multiple independent data blocks, so that each data block can be mapped to an independent compression / decompression module during the parallel computing process, thereby improving the parallelism of the system. Inside each independent data block, the data is continuously stored. Therefore, before compression, the method needs to obtain the corresponding starting data address, denoted as addr_begin;

[0220] Before compressing each data block, relevant initialization work is required, including the following sub-steps:

[0221] E2-1. Initialize CCD using the following formula (11):

[0222] CCD(k; 0) = k + 1, k ∈ [0, 2 n ) (11)

[0223] E2-2. Initialize the values of RL and RR: RL 0 [0:Y] = b’000...00, RR 0 [0:Y] = b’111...11;

[0224] E2-3. Initialize the values of L DOHB and L DEHB : L DOHB (0) = 0 and L DEHB (0) = 0;

[0225] E2-4. Initialize the data related to the multi-stage dynamic approximate frequency table model. This multi-stage dynamic approximate frequency table model and its application are the most important parts of the present invention. The relevant construction values of this multi-stage dynamic approximate frequency table are shown in Table 1 below:

[0226] Table 1: Construction table of a multi-stage dynamic approximate frequency table:

[0227]

[0228]

[0229] To apply this table, the user needs to manually select a T 0 ≥5 to obtain table entries with different precisions. Initially, set u = 0, h = 0, T = T 0 . In this way, according to Table 1, the initial mapping relationship between and can be obtained. For example, if T 0 = 5, then initially corresponds to For another example, if T 0 = 8, then initially corresponds to

[0230] E3. Sequentially read the value of the \(i\)-th element from the data block. Assuming each element is represented by \(n\) bits, the read data is \(V(i)[0:n - 1]=\text{memory_ori}[addrbegin:addr_begin + n - 1]\), where \(\text{memory_ori}[]\) represents the physical address space / virtual address space range where the original generation code data is stored;

[0231] E4. Calculate \(RL\) i+1 [0:Y] and \(RR\) i+1 [0:Y], and re-normalize, including the following sub-steps E4-1 and E4-2 (which can be executed in parallel with the following E5, and the steps E4-1 and E4-2 inside E4 are executed serially);

[0232] E4-1. Calculate \(RL\) according to the following formula (12) i+1 [0:Y] and \(RR\) i+1 [0:Y]:

[0233]

[0234] In fact, the calculation of \(RL\) shown in formula (12) i+1 [0:Y] and \(RR\) i+1 [0:Y] can be further transformed by the following formula (13) to reduce the bit width of the data in the multiplication operations involved, which will be beneficial to selecting a lower-bit-width multiplier in the ASIC implementation process to improve the hardware efficiency (reduce the latency of the hardware multiplier) and save hardware resources:

[0235]

[0236] In formula (13), \(S\) 0 、\(S\) 1 、\(S\) 2 、\(S\) 3 are user-adjustable parameters, determined offline by the user. However, the parameters \(S\) 0 、\(S\) 1 、\(S\) 2 、\(S\) 3 must satisfy the constraint conditions shown in the following formula (14):

[0237]

[0238] It is not difficult to find from formula (13) that, under the condition of satisfying the constraint condition (14), by adjusting the parameters \(S\) 0 、\(S\) 1 、\(S\) 2 、\(S\) 3, it can achieve the shortest bit width for the multiplication operation involved in formula (13), that is, during the ASIC implementation process, a multiplier with the shortest bit width is selected, thereby making the hardware execution efficiency the highest and the resource consumption the smallest.

[0239] E4-2. Re-normalize and output the data. This process requires two basic operations to be performed successively: DOHB and DEHB.

[0240] The DOHB operation is described as follows: Compare RL i+1 [0:Y] with RR i+1 for each bit data at the same address (from 0 to Y) of [0:Y]. If the same bit data is encountered, output the corresponding bit data at that address and shift RL[0:Y] and RR[0:Y] one bit to the left. This process continues until different bit data is encountered at a certain address between RL[0:Y] and RR[0:Y], at which point the process is interrupted. During the above left shift of RL[0:Y] and RR[0:Y], each time a bit is shifted left, the tail of RL[0:Y] will be filled with bit '0' and the tail of RR[0:Y] will be filled with bit '1'.

[0241] The description of the DEHB operation is as follows: Detect RL i+1 [0:Y] and RR i+1 The highest 2-bit numbers of [0:Y] are f2bL = RL[0:1], f2bR = RR[0:1]. If f2bL = b'01 and f2bR = b'10, then the second bit data in RL[0:Y] and RR[0:Y] will be removed respectively. At the same time, the tail of RL[0:Y] will be filled with bit '0' and the tail of RR[0:Y] will be filled with bit '1'. The above detection and removal operations will continue until the following pattern is encountered: when f2bL = b'01 and f2bR = b'10, the operation terminates. Here, different from the DOHB operation, when each bit data is removed, no corresponding output will be generated, but the internal variable L DEHB (i) is incremented by 1.

[0242] According to the DOHB and DEHB operations, the steps for re-normalization and output are as follows:

[0243] (1) For the i-th coded element, if the updated RL i+1 [0:Y] and RR i+1 [0:Y] are hit by the DOHB operation, that is, L DOHB (i) > 0, which means RL i+1 [0:L DOHB (i)-1] == RR i+1 [0:L DOHB(i)-1]. In the current situation, the output will be L DOHB (i)+L DEHB (i - 1) - bit data: And set L after the output DEHB (i - 1) = 0;

[0244] (2) For the i-th coding element, regardless of whether the DOHB operation hits, if L DEHB (i) > 0, then the DEHB operation must be executed, and let L DEHB (i) += L DEHB (i - 1);

[0245] E5. Update CCD and update the relevant information of the approximate frequency table, including the following sub-steps E5-1 and E5-2 (which can be executed in parallel with the above E4, and the steps E5-1 and E5-2 inside E5 are executed serially):

[0246] E5-1: Update CCD, including the following two sub-steps E5-1-1 and E5-1-2 that are executed serially:

[0247] E5-1-1: For each k ∈ [0, 2 n ), update CCD(k; i, V(i)) = CCD(k; i - 1, V(i - 1)) + δ(k, V(i));

[0248] E5-1-2: If CCD(2 n -1; i, V(i)) == 2 T , then for each k ∈ [0, 2 n ), perform the following two calculation operations: CCD(k; i, V(i)) <<= 1, CCD(k; i, V(i)) += 1;

[0249] E5-2. Update the relevant information of the approximate frequency table, including the following two sub-steps E5-2-1 and E5-2-2 that are executed sequentially:

[0250] E5-2-1: If Then, for each h, update u = u + 1

[0251] E5-2-2: If Then update h = h + 1;

[0252] E6. When both E4 (including sub-steps E4-1 and E4-2) and step E5 (including sub-steps E5-1 and E5-2) are completed, determine whether the end of the data block is reached? If so, proceed to step E7, otherwise proceed to step E8;

[0253] E7. Process the tail coding of the data block:

[0254] E7-1: Output single-bit FB = (RL i [1:1]|RR i [1:1]);

[0255] E7-2: Output L DEHB (i)

[0256] E7-3: Output single-bit

[0257] Through the tail encoding process of the data block, the encoding results of different data blocks using different frequency tables can be continuously stored together, and the correctness of the decoding process can be ensured without storing additional auxiliary information, that is, the encoding result of the adjacent next data block will not affect the decoding process of the previous data block without any special processing at the tail of the data block encoding result. More importantly, this strategy ensures that all encoding information is output, and there will be no problem of incomplete output of potential encoding information (resulting in decoding errors) existing in the tail processing method of the traditional rangeCode method.

[0258] E8. Update the index i of the element to be compressed: i = i + 1.

[0259] Generally speaking, the matrix data lossless decompression method provided in this application is actually the reverse decompression method supporting the matrix data lossless compression method in step S2 included in this application. A matrix data lossless decompression method provided in this application is consistent with a matrix data lossless compression method provided in many implementation steps, but there are also some differences in some steps. Figure 5 Shows the flowchart of a matrix data lossless decompression method in step S3 provided in this application. According to Figure 5 the flowchart shown, the detailed steps of the method are described as follows:

[0260] D1. Set the initial values of some variables involved in the current data block decoding process:

[0261] D1-1. In a data compression method provided in this application, the complete data may be divided into multiple independent data blocks and the encoding results of each data block may be continuously stored. Therefore, during the decoding process, it is necessary to read the initial address of the encoded data corresponding to the corresponding data block, denoted here as addr_begin_dec;

[0262] D1-2. Initialize CCD according to the following formula (15):

[0263] CCD(k; 0) = k + 1, k ∈ [0, 2 n) (15)

[0264] D1-3. Initialize the values of RL and RR: RL 0 [0:Y] = b’000...00, RR 0 [0:Y] = b’111...11;

[0265] D1-4. Initialize L DOHB and L DEHB values: L DOHB (0) = 0 and L DEHB (0) = 0;

[0266] D1-5. Initialize the value of the Y+1-bit variable currentCode: currentCode[0:Y] = b’000...00;

[0267] D1-6. Initialize the data related to the multi-stage dynamic approximate frequency table model. This multi-stage dynamic approximate frequency table model and its application are the most important parts of the present invention. The related construction values of this multi-stage dynamic approximate frequency table model are shown in Table 1 below. To apply this table, the user needs to manually select a T 0 ≥5, so as to obtain table entries with different precisions. Initially, set u = 0, h = 0, T = T 0 .

[0268] D2. According to addr_begin_dec, read the Y+1-bit data from the encoded data corresponding to the current data block and fill it into the variable currentCode[0:Y]:

[0269] currentCode[0:Y] = memory_enc[addr_begin_dec:addr_begin_dec+Y]. Here, memory_enc[] represents the physical address space / virtual address space range where the original encoded data is stored; after the reading is completed, addr_begin_dec += shiftBits;

[0270] D3. Search and find the decoded value, and output the obtained current decoded value:

[0271] For each k ∈ [0, 2 n ), calculate the corresponding and according to formula (16). If there exists a unique k such that the condition: is satisfied, then the obtained current decoded value is k, that is, V(i) = k. According to formula (16), since and It monotonically increases as k increases, so this process can be carried out using binary search.

[0272]

[0273] D4. Calculate RL i+1 [0:Y] and RR i+1 [0:Y], and re-normalize, including two sub-steps D4-1 and D4-2 (which can be executed in parallel with the following D5, and the D4-1 and D4-2 steps inside D4 are executed serially);

[0274] D4-1. According to the following formula (17), use the decoded value V(i) to update the RL and RR values:

[0275] RW i [0:Y] = RR i [0:Y] - RL i [0:Y] (17)

[0276]

[0277]

[0278] D4-2. Re-normalize. Successively perform the following DOHB operation and DEHB operation:

[0279] The description of the DOHB operation is as follows: Compare RL i+1 [0:Y] with RR i+1 The bit data at each same address (from 0 to Y) of [0:Y]. If the same bit data is encountered, shift RL i+1 [0:Y], RR i+1 [0:Y] and currentCode[0:Y] one bit to the left. This process continues until different bit data is encountered at a certain address between RL i+1 [0:Y] and RR i+1 [0:Y]. At this time, the process interrupts. During the above process of shifting RL i+1 [0:Y], RR i+1 [0:Y] and currentCode[0:Y] one bit to the left, every time one bit is shifted to the left, the tail of RL i+1 [0:Y] will be filled with the bit "0" while the tail of RR i+1 will be filled with the bit "1". At the same time, read one bit of data: memory_enc[addrbegin_dec:addr_begin_dec] is filled into the tail of currentCode[0:Y] and addr_begin_dec += 1:

[0280] The description of the DEHB operation is as follows: Detect RL i+1 [0: Y] and RR i+1 The highest 2 bits f2bL of [0: Y] = RL i+1 [0: 1], f2bR = RR i+1 [0: 1]. If f2bL = b'01 and f2bR = b'10, then RL i+1 [0: Y], RR i+1 The second bit data in [0: Y], RR i+1 [0: Y] and currentCode[0: Y] will be removed respectively. The above detection and removal operations will continue until the following pattern is encountered: when f2bL = b'01 and f2bR = b'10, the operation terminates. In the above left shift of RL i+1 [0: Y], RR i+1 [0: Y] and currentCode[0: Y], each time a bit is left shifted, the tail of RL i+1 The tail of [0: Y] will be filled with the bit "0" while the tail of RR

[0281] D5. Update the relevant information of CCD and the approximate frequency table, including sub-steps D5-1 and D5-2 (which can be executed in parallel with the above D4, and the internal sub-steps D5-1 and D5-2 of D5 are executed serially):

[0282] D5-1: Update CCD, including the following two sub-steps D5-1-1 and D5-1-2 executed serially:

[0283] D5-1-1: For each k ∈ [0, 2 n ), update CCD(k; i, V(i)) = CCD(k; i - 1, V(i - 1)) + δ(k, V(i));

[0284] D5-1-2: If CCD(2 n -1; i, V(i)) == 2 T , then for each k ∈ [0, 2 n ), perform the following two calculation operations: CCD(k; i, V(i)) <<= 1, CCD(k; i, V(i)) += 1;

[0285] D5-2. Update the relevant information of the approximate frequency table, including the following two sub-steps D5-2-1 and D5-2-2 executed sequentially:

[0286] D5-2-1: If Then, for each h, update u = u + 1;

[0287] D5-2-2: If Then update h = h + 1;

[0288] D6. When both D4 (including sub-steps D4-1 and D4-2) and step D5 (including sub-steps D5-1 and D5-2) are completed, determine whether the end of the data block is reached? If so, proceed to step D7, if not, proceed to step D8;

[0289] D7. Process the tail coding of the data block: Different from the coding process, here only need to directly shift currentCode[0:Y] two bits to the left. This also reflects the simplicity of the tail processing method included in this application. This will be very suitable for hardware implementation scenarios.

[0290] D8. Update the decoding element index i: i = i + 1.

[0291] This application also provides an application system for a hardware-oriented matrix data lossless compression and decompression method. The application system implements any of the above methods in a neural network acceleration dedicated chip NNA. The system includes:

[0292] Multiple parallel hardware modules ED1...EDN corresponding to the lossless compression method;

[0293] Multiple parallel hardware modules DD1...DDN corresponding to the lossless decompression method;

[0294] Such as Figure 6 shows the original hardware architecture for implementing a matrix data lossless compression and decompression method proposed in this application in NNA. Figure 6 ED1...EDN in represents multiple parallel hardware modules corresponding to the lossless compression method proposed in the present invention. Figure 6 DD1...DDN in represents multiple parallel hardware modules corresponding to the lossless decompression method proposed in the present invention.

[0295] From Figure 6As can be seen, after the NNA generates the FMs data, it is first placed in a dynamic random access memory (DRAM). Multiple lossless data compression modules proposed by this application perform parallel compression on the data of different layers / channels therein. Then, the compressed data will be transferred to a large-capacity static random access memory (SRAM). When these FMs data are needed, multiple lossless data decompression modules proposed by this application perform parallel decompression on the data of different layers / channels in the SRAM to the DRAM. In this process, it is expected to reduce the data volume of the FMs through the data compression module, thereby reducing the overall data bandwidth, especially the bandwidth of the DRAM.

[0296] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A hardware-oriented method for lossless compression and decompression of matrix data, characterized in that, the method comprises the following steps: S1. Determine whether to perform a compression or decompression process according to the instruction code issued by the CPU / MCU. If it is lossless compression of matrix data, further proceed to step S2; If it is lossless decompression of matrix data, further proceed to step S3; S2. The method for lossless compression of matrix data includes: E1. Obtain the starting address of the data block; E2. Initialize the relevant variables required for encoding the current data block; E3. Sequentially read the value of the i-th element from the data block; execute steps E4 and E5 in parallel; E4. Internally includes two sub-steps E4-1 and E4-2 that need to be executed serially: E4-1, calculate the left and right endpoint values RL i+1 [0:Y] and RR i+1 [0:Y]; E4-2. Re-normalize and output; E5. Internally includes two sub-steps E5-1 and E5-2 that need to be executed serially: E5-1. Update the CCD; E5-2. Update the relevant information of the approximate frequency table; E6. When both steps E4 and E5 are completed, determine whether the end of the data block is reached. If so, proceed to step E7, otherwise proceed to step E8; E7. Process the tail encoding of the data block; return to step E1; E8. Update the index i of the element to be compressed, i = i + 1; return to step E3; S3. The method for lossless decompression of matrix data includes: D1. Initialize the relevant variables required for decoding the current data block; D2. Read Y + 1-bit data from the encoding result to fill the variable or register array currentCode[0:Y]; D3. Search and find the decoded value, and output the obtained current decoded value; execute steps D4 and D5 in parallel; D4. Internally includes two sub-steps D4-1 and D4-2 that need to be executed serially: D4-1, calculate the left and right endpoint values RL i+1 [0:Y] and RR i+1 [0:Y]; D4-2. Re-normalize; D5. Internally includes two sub-steps D5-1 and D5-2 that need to be executed serially: D5-1. Update the CCD; D5-2. Update the relevant information of the approximate frequency table; D6. When both steps D4 and D5 are completed, determine whether the end of the data block is reached. If so, proceed to step D7. If not, proceed to step D8; D7. Process the tail encoding of the data block; return to step D1; D8. Update the index i of the decompressed element, i = i + 1; return to step D3.

2. A hardware-oriented method for lossless compression and decompression of matrix data according to claim 1, characterized in that, the step S2 further includes the following detailed steps: E1. Obtain the starting address of the data block; All data is divided into multiple independent data blocks. Since the data distribution characteristics of each channel on the feature map in the NNA are different, each channel will be regarded as an independent data block for compression; multiple data compression / decompression modules are used on the hardware to perform parallel compression on the data, so that each data block can be mapped to an independent compression / decompression module during the parallel computing process; inside each independent data block, the data is continuously stored. Therefore, before compression, the corresponding starting data address needs to be obtained, denoted as addr_begin; E2. Initialize the relevant variables required for encoding the current data block; Before compressing each data block, relevant initialization work is required, including the following sub-steps: E2-1. Initialize the cumulative frequency distribution CCD using the following formula (11): CCD(k, 0) = k + 1, k ∈ [0, 2 n ) (11) E2-2. Initialize the values of the left endpoint RL and the right endpoint RR: RL 0 [0:Y] = b'000...00, RR 0 [0:Y] = b'111...11; E2-3. Initialize L DOHB and L DEHB value: L DOHB (0) = 0, L DEHB (0) = 0; E2-4. Initialize the data related to the multi-stage dynamic approximate frequency table model; The relevant construction values of the multi-stage dynamic approximate frequency table model are shown in Table 1 below: Table 1: To apply this table, the user needs to manually select a T 0 ≥5, so as to obtain table entries with different precisions; initially, set u = 0, h = 0, and T = T 0 ; in this way, according to Table 1, the and initial mapping relationship between them can be obtained; E3. Sequentially read the value of the i-th element from the data block; Assume that each element is represented by n bits, then the read data is V(i)[0:n-1] = memory_ori[addr_begin:addr_begin+n-1], where memory_ori[] represents the physical address space / virtual address space range where the original generation-coded data is stored; E4: Calculate the left and right endpoint values RL i+1 [0: Y] and RR i+1 [0: Y], and renormalize and output, including the following two sub-steps E4-1 and E4-2, which can be executed in parallel with the following step E5. The internal sub-steps E4-1 and E4-2 of step E4 are executed serially; Calculate the left and right endpoint values RL according to the following formula (12) i+1 [0: Y] and RR i+1 [0: Y]: In fact, for RL i+1 [0:Y] and RR i+1 The calculation of [0:Y] can be further transformed by the following formula (13) to reduce the bit width of the data in the multiplication operations involved therein: In formula (13), S 0 , S 1 , S 2 , S 3 are user-adjustable parameters and are determined offline by the user; E4-2. Renormalize and output the data; this process requires two basic operations to be performed successively: DOHB and DEHB; E5. Update the information related to CCD and the approximate frequency table, including the following two sub-steps E5-1 and E5-2, which can be executed in parallel with the above step E4, and the internal sub-steps E5-1 and E5-2 of E5 are executed serially: E5-1: Update CCD, including the following two sub-steps E5-1-1 and E5-1-2 that are executed serially: E5-1-1: For each k ∈ [0, 2 n ), update CCD(k; i, V(i)) = CCD(k; i - 1, V(i - 1)) + δ(k, V(i)); E5-1-2: If CCD(2 n -1; i, V(i)) == 2 T , then for each k ∈ [0, 2 n ), perform the following two calculation operations: CCD(k; i, V(i)) <<= 1, CCD(k; i, V(i) += 1; E5-2. Update the information related to the approximate frequency table, including the following two sub-steps E5-2-1 and E5-2-2 that are executed serially: E5-2-1: If then, for each h, update E5-2-2: If then update h = h + 1; E6. When both E4 (including sub-steps E4-1 and E4-2) and step E5 (including sub-steps E5-1 and E5-2) are completed, determine whether the end of the data block is reached? If so, perform step E7, otherwise perform step E8; E7. Process the tail coding of the data block, including: E7-1: Output single-bit FB = (RL i [1:1] | RR i [1:1]); E7-2: Output L DEHB (i) E7-3: Output single bit Through this data block tail coding process, the coding results of different data blocks using different frequency tables can be continuously stored together, and the correctness of the decoding process can be ensured without storing additional auxiliary information; E8. Update the value of the compressed element index i: i = i + 1.

3. According to the hardware-oriented matrix data lossless compression and decompression method described in claim 2, characterized in that, in the step E2, In Table 1, if T 0 is selected as 5, then initially corresponds to If T 0 is selected as 8, then initially it corresponds to 4. According to the hardware-oriented matrix data lossless compression and decompression method described in claim 2, characterized in that, in the step E4-1, Parameter S 0 , S 1 , S 2 , S 3 must satisfy the constraint conditions shown in the following formula (14): It is not difficult to find from formula (13) that, when the constraint condition (14) is satisfied, by adjusting the parameters S 0 、S 1 、S 2 、S 3 , the bit width of the multiplication operation involved in formula (13) can be made the shortest, that is, in the process of implementing the method on ASIC, a multiplier with the shortest bit width is selected.

5. According to the hardware-oriented matrix data lossless compression and decompression method described in claim 2, characterized in that, in the step E4-2, The DOHB operation is described as follows: Compare RL i+1 [0:Y] with RR i+1 for the bit data at each same address from 0 to Y. The same address ranges from 0 to Y. If the same bit data is encountered, output the corresponding bit data at that address and shift RL[0:Y] and RR[0:Y] left by one bit. This process continues until different bit data is encountered at a certain address between RL[0:Y] and RR[0:Y], at which point the process is interrupted. During the above process of shifting RL[0:Y] and RR[0:Y] left, each time a bit is shifted, a bit '0' will be filled at the tail of RL[0:Y] and a bit '1' will be filled at the tail of RR[0:Y]. The description of the DEHB operation is as follows: Detect RL i+1 [0:Y] and RR i+1 The highest 2-bit numbers of [0:Y] are f2bL = RL[0:1], f2bR = RR[0:1]; if f2bL = b'01 and f2bR = b'10, then the second bit data in RL[0:Y] and RR[0:Y] will be removed respectively; meanwhile, the tail of RL[0:Y] will be filled with bit "0" while the tail of RR[0:Y] will be filled with bit "1"; the above detection and removal operations will continue until the following pattern is encountered: when f2bL = b'01 and f2bR = b'10, the operation terminates; here, different from the DOHB operation, when each bit data is removed, no corresponding output will be generated, but the internal variable L DEHB (i) is incremented by 1; According to the DOHB and DEHB operations, the renormalization and output operation steps are as follows: (1). For the i-th element to be compressed, if the updated RL i+1 [0:Y] and RR i+1 [0:Y] are hit by the DOHB operation, that is, L DOHB (i) > 0, which means RL i+1 [0:L DOHB (i)-1] == RR i+1 [0:L DOHB (i)-1]; in the current case, output L DOHB (i)+L DEHB (i - 1) bits of data: {RL i+1 [0:0], (repeated L DEHB (i - 1) times), RL i+1 [1:L DOHB (i)-1]}, and set L DEHB (i - 1) = 0 after the output; (2). For the i-th element to be compressed, regardless of whether the DOHB operation hits or not, if L DEHB (i) > 0, then the DEHB operation must be executed, and let L DEHB (i) += L DEHB (i - 1).

6. According to the hardware-oriented matrix data lossless compression and decompression method described in claim 2, characterized in that, in the step S3, the matrix data lossless decompression method is the reverse decompression method supporting the matrix data lossless compression method in the step S2; the detailed steps of the method are described as follows: D1. Set the initial values of the variables involved in the current data block decoding process, including: D1-1. In the step S2, the complete data is split into multiple independent data blocks and the encoding results of each data block may be stored continuously; therefore, during the decoding process, the initial address of the encoded data corresponding to the corresponding data block needs to be read, denoted here as addr_begin_dec; D1-2. Initialize the CCD according to the following formula (15): CCD(k; 0) = k + 1, k ∈ [0, 2 n ) (15) D1-3. Initialize the values of RL and RR: RL 0 [0:Y] = b’000...00, RR 0 [0:Y] = b’111...11; D1 - 4. Initialize L DOHB and L DEHB value of: L DOHB (0) = 0 and L DEHB (0) = 0; D1-5. Initialize the value of the Y+1-bit variable currentCode: currentCode[0:Y] = b’000...00; D1-6. Initialize the relevant data of the multi-stage dynamic approximate frequency table model; The relevant construction values of the multi-stage dynamic approximate frequency table model are shown in Table 1; Applying this table requires the user to manually select a T 0 ≥5, so as to obtain table entries with different precisions; initially, set u = 0, h = 0, and T = T 0 ; D2. According to addr_begin_dec, read Y+1-bit data from the encoded data corresponding to the current data block and fill it into the variable currentCode[0:Y]: currentCode[0:Y] = memory_enc[addr_begin_dec:addr_begin_dec+Y]; Here, memory_enc[] represents the physical address space / virtual address space range where the original encoded data is stored; after the reading is completed, addr_begin_dec += shiftBits; D3. Search for and find the decoded value, and output the obtained current decoded value: For each k ∈ [0, 2 n ), the corresponding and are calculated according to formula (16). If there exists a unique k such that the condition: is satisfied, then the value obtained by the current decoding is k, that is, V(i) = k; According to formula (16), since and monotonically increase as k increases, this process can be carried out using binary search: D4. Calculate the left and right endpoint values PL i-1 [0: Y] and PR i+1 [0: Y], and re-normalize, including sub-steps D4-1 and D4-2, which can be executed in parallel with D5 below. Steps D4-1 and D4-2 inside D4 are executed serially: D4-1: According to the following formula (17), use the decoded value to update the RL and RR values of V(i): RW i [0:Y] = RR i [0:Y] - RL i [0:Y] (17) D4-2: Re-normalize; this process requires two basic operations to be performed successively: DOHB and DEHB; D5. Update the relevant information of the CCD and the approximate frequency table, including sub-steps D5-1 and D5-2, which can be executed in parallel with the above D4, and the internal sub-steps D5-1 and D5-2 of D5 are executed serially: D5-1: Update the CCD, including the following two serially executed sub-steps D5-1-1 and D5-1-2: D5-1-1: For each k ∈ [0, 2 n ), update CCD(k; i, V(i)) = CCD(k; i - 1, V(i - 1)) + δ(k, V(i)); D5-1-2: If CCD(2 n -1; i, V(i)) == 2 T , then for each k ∈ [0, 2 n ), perform the following two calculation operations: CCD(k; i, V(i)) <<= 1 CCD(k; i, V(i)) += 1. D5-2. Update the relevant information of the approximate frequency table, including the following two serially executed sub-steps E5-2-1 and E5-2-2: D5-2-1: If Then, for each h, update D5-2-2: If Then update h = h + 1; D6. When both step D4 and step E5 are completed, determine whether the end of the data block is reached? If so, proceed to step D7, otherwise proceed to step D8; D7. Process the tail encoding of the data block: Different from the encoding process, here only need to directly shift currentCode[0:Y] two bits to the left; D8. Update the value of the decompressed data index i: i = i + 1.

7. According to a hardware-oriented matrix data lossless compression and decompression method as claimed in claim 6, wherein, in the step D4-2, it further includes: The description of the DOHB operation is as follows: Compare RL i+1 [0: The gate and RR i+1 [0: Y] for the bit data at each same address, where the same address ranges from 0 to Y. If the same bit data is encountered, shift RL to the left i+1 [0: Y], RR i+1 [0: Y] and one bit each of currentCode[0: Y]. This process continues until different bit data is encountered at a certain address and RR i+1 [0: Y], at which point the process interrupts; during the above-mentioned left shift of RL i+1 [0: Y], RR i+1 [0: Y] and currentCode[0: Y], for each left shift of one bit, the tail of RL i+1 [0: Y] will be filled with the bit "0" while the tail of RR i+1 will be filled with the bit "1"; meanwhile, read one bit of data: memory_enc[addrbegin_dec: addr_begin_dec] is filled to the tail of currentCode[0: Y] and addr_begin_dec += 1; The description of the DEHB operation is as follows: Detect RL i+1 [0: Y] and RR i+1 The highest 2-bit number f2bL of [0: Y] = RL i+1 [0: 1], f2bR = RR i+1 [0: 1]; If f2bL = b'01 and f2bR = b'10, then RL i+1 [0: Y], RR i+1 The second bit data in [0: Y], RR i+1 [0: Y] and currentCode[0: Y] will be removed respectively; The above detection and removal operations will continue until the following pattern is encountered: When f2bL = b'01 and f2bR = b'10, the operation terminates; In the above left shift of RL i+1 [0: Y], RR i+1 [0: Y] and currentCode[0: Y], for each left shift of one bit, the tail of RL i+1 will be filled with the bit "0" while the tail of RR will be filled with the bit "1"; At the same time, read one bit of data: memory_enc[addr_begin_dec: addr_begin_dec] is filled to the tail of CurrentCode[0: Y] and addr_begin_dec += 1.

8. An application system of a hardware-oriented matrix data lossless compression and decompression method, wherein, the application system implements the method as claimed in any one of claims 1 to 7 in the neural network acceleration dedicated chip NNA, and the system includes: A plurality of parallel hardware modules ED1...EDN corresponding to the lossless compression method; A plurality of parallel hardware modules DD1...DDN corresponding to the lossless decompression method; After the NNA generates the FMs data, it is first placed in the dynamic random access memory DRAM. A plurality of lossless data compression modules DE1...EDN perform parallel compression on the data of different layers / channels therein. Then, the compressed data will be transferred to the large-capacity static random access memory SRAM; When these FMs data are needed, a plurality of lossless data decompression modules DD1...DDN perform parallel decompression on the data of different layers / channels in the SRAM to the DRAM; in this process, the data volume of the FMs is reduced by the data compression module.