Reduce the latch count to save hardware area for dynamic Huffman table generation

By mapping 24-bit symbol counts to 10-bit floating point representations, generating 5-bit shift fields and 5-bit mantissa, the problem of excessive latch counts required for symbol sorting during dynamic Hoffman table generation is solved, and the area, power consumption and wiring costs are reduced.

CN113366765BActive Publication Date: 2025-07-01INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080012147.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-02-14
Filing Date
2020-02-11
Publication Date
2025-07-01
Estimated Expiration
2040-02-11

AI Technical Summary

Technical Problem

In digital computer systems, the symbol count needs to be sorted when generating dynamic Hoffman tables, resulting in an increase in latch counts and a higher cost of occupancy, power, and timing/wiring.

Method used

By mapping 24-bit symbol counts to 10-bit class floating point representations, a 5-bit shift field and a 5-bit mantissa are generated, reducing the latch count required for symbol sorting.

Benefits of technology

Reduces latch counts required for symbol sorting, frees up wafer area, reduces power consumption, and simplifies timing/wiring of accelerator hardware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113366765B_ABST
    Figure CN113366765B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention relate to a DEFLATE compression accelerator and a method for reducing the latch count required for symbol sorting when generating a dynamic Huffman table. The accelerator includes an input buffer and a Lempel-Ziv 77 (LZ77) compressor communicatively coupled to the output of the input buffer. The accelerator further includes a Huffman encoder communicatively coupled to the LZ77 compressor. The Huffman encoder includes a one-bit converter. The accelerator further includes an output buffer communicatively coupled to the output of the Huffman encoder.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] The present invention relates to digital computer systems, and more particularly, to digital data compression and decompression schemes employed in digital computer systems.

[0002] Digital computer systems perform data compression to achieve more efficient use of limited storage space. Computer systems typically include a hardware component called a compression accelerator that receives work requests or data requests from a host system to compress or decompress one or more blocks of the requested data. When designing an accelerator to perform compression, there is a trade-off between the size of the input data to be compressed and the latency resulting from the compressed data compared to the possible compression ratio.

[0003] Compression accelerators typically utilize the "DEFLATE" algorithm, which is a lossless compression scheme that combines the Lempel-Ziv (e.g., LZ77) compression algorithm with the Huffman coding algorithm to perform compression. The computational output from the Huffman algorithm can be regarded as a variable-length code table for encoding source symbols (e.g., characters in a file). The Huffman algorithm derives this table from the estimated occurrence probabilities or frequencies (weights) of each possible value of the source symbols.

[0004] To maximize the compression ratio obtained using the DEFLATE algorithm, symbols are encoded into a variable-length code table based on their occurrence frequencies. In other words, the most frequent symbols are encoded with the fewest bits, while relatively less common symbols are encoded with relatively more bits. This results in a direct reduction in the storage space required for the compressed data stream. Since symbols are encoded based on their relative frequencies, the occurrence counts of each symbol must be sorted. Sorting the symbol counts (frequencies) during this process is expensive in terms of area (the number of latches and width comparators required), power, and timing / wiring considerations. SUMMARY OF THE INVENTION

[0005] Embodiments of the present invention are directed to an accelerator, such as a DEFLATE compression accelerator, that is configured to reduce the required latch count during dynamic Huffman table generation. Non-limiting examples of accelerators include an input buffer and a Lempel-Ziv 77 (LZ77) compressor communicatively coupled to the output of the input buffer. The accelerator further includes a Huffman encoder communicatively coupled to the LZ77 compressor. The Huffman encoder includes a bit converter. The accelerator also includes an output buffer communicatively coupled to the output of the Huffman encoder.

[0006] In some embodiments of the present invention, the bit converter is a 24-bit to 10-bit converter.

[0007] In some embodiments of the present invention, the bit converter is configured to generate a 5-bit shift field and a 5-bit mantissa based on a first symbol count.

[0008] In some embodiments of the present invention, the bit converter is further configured to concatenate the 5-bit shift field and the 5-bit mantissa to generate a second symbol count.

[0009] Embodiments of the present invention relate to a method for reducing the latch count required for symbol sorting when generating a dynamic Huffman table. Non-limiting examples of the method include determining a plurality of first symbol counts. Each of the first symbol counts includes a first bit width. The method further includes generating a plurality of second symbol counts. The second symbol counts are based on a reduced bit mapping of the first symbol counts. The plurality of second symbol counts are sorted by frequency and used to generate a dynamic Huffman tree.

[0010] In some embodiments of the present invention, a 5-bit shift field and a 5-bit mantissa are generated based on a first symbol having a plurality of first symbol counts.

[0011] In some embodiments of the present invention, the 5-bit shift field encodes the position of the most significant non-zero bit of the first symbol.

[0012] In some embodiments of the present invention, the 5-bit mantissa encodes the most significant non-zero bit of the first symbol and the next four bits.

[0013] In some embodiments of the present invention, the 5-bit mantissa encodes the next five bits after the most significant non-zero bit of the first symbol.

[0014] Embodiments of the present invention are directed to a computer program product for reducing the latch count required for symbol sorting when generating a dynamic Huffman table. Non-limiting examples of the computer program product include program instructions executable by an electronic computer processor to control a computer system to perform operations. The operations may include determining a plurality of first symbol counts. Each of the first symbol counts includes a first bit width. The operations may further include generating a plurality of second symbol counts. The second symbol counts are based on a reduced bit mapping of the first symbol counts. The plurality of second symbol counts are sorted by frequency and used to generate a dynamic Huffman tree.

[0015] Embodiments of the present invention are directed to a system for reducing the latch count required for symbol sorting when generating a dynamic Huffman table. Non-limiting examples of the system include an accelerator, a memory having computer-readable instructions, and a processor configured to execute the computer-readable instructions. The computer-readable instructions, when executed by the processor, cause the accelerator to perform a method. The method may include determining a plurality of first symbol counts, each of the first symbol counts including a first bit width. A plurality of second symbol counts may be generated. Each of the second symbol counts may be a mapping based on the symbol counts among the plurality of first symbol counts. The second symbol counts may include a second bit width that is less than the first bit width. The method may further include sorting the plurality of second symbol counts by frequency and generating a dynamic Huffman tree based on the sorted plurality of second symbol counts.

[0016] Embodiments of the present invention relate to a method. Non-limiting examples of the method include receiving a data stream including a first symbol from an input buffer. A first symbol count having a first bit width may be determined based on the first symbol. The method may include generating a 5-bit shift field and a 5-bit mantissa based on the first symbol count. By concatenating the 5-bit shift field and the 5-bit mantissa, a second symbol count having a second bit width may be generated. The method may include sorting the frequency of the second symbol count.

[0017] Additional technical features and benefits are achieved through the techniques of the present invention. Embodiments and aspects of the present invention are described in detail herein and are considered to be part of the claimed subject matter. For a better understanding, reference is made to the detailed description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The details of the exclusive rights described herein are particularly pointed out and clearly claimed in the claims at the end of the specification. From the following detailed description taken in conjunction with the drawings, the foregoing and other features and advantages of embodiments of the present invention are apparent, in which:

[0019] Figure 1A and 1B Huffman trees generated according to various embodiments of the present invention are depicted;

[0020] Figure 2 A block diagram of a computer system capable of compressing and decompressing data according to various embodiments of the present invention is shown;

[0021] Figure 3 A block diagram of an accelerator according to one or more embodiments is shown;

[0022] Figure 4 Shown Figure 3 A portion of the Huffman encoder of the accelerator shown;

[0023] Figure 5shows Figure 4 a portion of the sorting module of the DHT generator of the Huffman encoder shown;

[0024] Figure 6 is a flowchart showing a method according to a non - restrictive embodiment; and

[0025] Figure 7 is a flowchart showing a method according to another non - restrictive embodiment.

[0026] The figures described herein are illustrative. Without departing from the spirit of the present invention, there can be multiple variations to the figures or operations described therein. For example, actions can be performed in a different order, or actions can be added, deleted, or modified. Additionally, the term "coupled" and its variations describe having a communication path between two elements, but do not imply a direct connection between the elements, without an intermediate element / connection between them. All such variations are considered to be part of the specification.

[0027] In the drawings and the following detailed description of the disclosed embodiments, the various elements shown in the drawings have two or three - digit reference numerals. With minor exceptions, the left - most digit of each reference number corresponds to the figure in which its element is first shown. Detailed Description

[0028] Various embodiments of the present invention are described herein with reference to the related drawings. Alternative embodiments of the present invention can be designed without departing from the scope of the present invention. In the following description and drawings, various connection and positional relationships (e.g., above, below, adjacent, etc.) are set forth between elements. Unless otherwise stated, these connections and / or positional relationships can be direct or indirect, and the present invention is not intended to be limited in this regard. Thus, the coupling of entities can refer to direct or indirect coupling, and the positional relationship between entities can be direct or indirect positional relationship. Additionally, the various tasks and process steps described herein can be incorporated into a more comprehensive program or process having additional steps or functionality not detailed herein.

[0029] The following definitions and abbreviations are used to interpret the claims and the specification. As used herein, the term "comprising", "including", "having", "containing" or any other variation thereof is intended to cover a non - exclusive inclusion. For example, a composition, mixture, process, method, article, or apparatus that includes a series of elements is not necessarily limited to those elements, but can include other elements not expressly listed or inherent to such composition, mixture, process, method, article, or apparatus.

[0030] Additionally, the term "exemplary" is used herein to mean "serving as an example, instance, or illustration", and any embodiment or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or designs. The terms "at least one" and "one or more" can be understood to include any integer greater than or equal to one, i.e., one, two, three, four, etc. The term "plurality" can be understood to include any integer greater than or equal to two, i.e., two, three, four, five, etc. The term "connected" can include both indirect "connection" and direct "connection".

[0031] The terms "about", "substantially", "approximately" and their variants are intended to include the degree of error associated with a particular quantity measurement based on the equipment available at the time of filing the present application. For example, "about" can include a range of ±8% or 5% or 2% of a given value.

[0032] For the sake of brevity, conventional techniques related to aspects of manufacturing and using the present invention may or may not be described in detail herein. In particular, aspects of the computing systems and specific computer programs for implementing the various technical features described herein are well known. Therefore, for the sake of brevity, many conventional implementation details are only briefly mentioned or completely omitted herein without providing well-known system and / or process details.

[0033] Turning now to an overview of the techniques more specifically related to aspects of the present invention, the reduction in the size of data representations produced by an applied data compression algorithm is generally referred to as the compression ratio (C / R). The compression ratio can be defined as the ratio between the uncompressed size and the compressed size. Thus, as the compression ratio increases, a more efficient use of the storage space of the computer system is achieved, thereby improving the overall performance of the computer system.

[0034] The DEFLATE data compression algorithm is a commonly used method for compressing data. When compressing data, the DEFLATE algorithm has two main parts: (1) LZ77 compression for identifying repeated strings and (2) Huffman coding of this information.

[0035] The LZ77 compression stage attempts to find repeated strings in the previously encoded source operand. When a match is found, instead of outputting the literal characters of the repeated string, the LZ77 compression stage outputs the "distance" from the repeated string to the original (matching) string in the previous dataset history, as well as the matching "length" of the data. For example, assume the input operand contains the following symbols: ABBACBABBAABBABBA. This operand can be encoded as:

[0036] Literal byte A; literal byte B; literal byte B; literal byte A; literal byte C; literal byte B; distance 6, length 4 (which encodes "ABBA"); distance 4, length 8 (which encodes "ABBAABBA")

[0037] It can be seen that the more duplicate strings that can be found in the input operand data, the more the output can be compressed. There are two ways to examine the input operand history to find matching strings: inline history, and via a circular history buffer. For inline history, the LZ77 compressor only looks at previous inputs from the source operand. For the circular history buffer, the input data is copied (either actually or conceptually) into the circular history buffer, and then the data in the buffer is searched for matches. In either case, the DEFLATE standard allows backtracking up to 32KB to match strings.

[0038] The Huffman coding stage is based on the probabilities and distributions of the symbols produced by the LZ77 compressor. The idea behind Huffman coding is that symbols can be encoded with variable bit lengths such that frequent symbols are encoded with few bits, while rare symbols are encoded with many bits. In this way, further compression of the data obtained from the LZ77 compressor is possible.

[0039] For this encoding process, the DEFLATE standard supports three types of compressed data blocks: literal copy blocks, fixed Huffman tables (FHTs), and dynamic Huffman tables (DHTs). FHT blocks are static, while DHT blocks include: a highly compressed version of the Huffman tree, followed by the symbols representing the compressed data encoded using that tree.

[0040] An example Huffman tree is shown in Figure 1A and 1B As Figure 1A shown, the Huffman tree can be highly asymmetric, where most of the nodes (also called leaves) occur along a single branch of the tree. Alternatively, the Huffman tree can be compressed as Figure 1B shown, where the leaf distribution traverses all available branches. In either case, the Huffman tree is constructed such that the depth of the leaves (nodes) is determined by the frequency of the symbols corresponding to each leaf. In other words, the depth of a leaf is determined by its symbol frequency.

[0041] Table 1 illustrates an exemplary DHT corresponding to the Huffman tree described in Figure 1A The DHT shown in Table 1 is constructed such that symbols with relatively high counts / frequencies are encoded using relatively short code lengths.

[0042] Table 1: Dynamic Huffman Table

[0043]

[0044]

[0045] As shown in Table 1, the "A" and "E" symbols have the lowest frequencies, each occurring only 100 times. The "D" symbol has the second highest frequency and appears 200 times in the dataset. The "C" symbol appears 400 times in the dataset, while the "B" symbol appears most frequently, 800 times. As further shown in Table 1, the "A" symbol is encoded as the binary number "1110", the "B" symbol is encoded as "0", the "C" symbol is encoded as "10", the "D" symbol is encoded as "110", and the "E" symbol is encoded as "1111".

[0046] Encoding the most frequent symbols (e.g., "B" in the above example) with the fewest bits results in a direct reduction in the storage space required to compress the data stream. For example, the "B" symbol, which appears 800 times, can be represented by a single "0" bit each time it appears. Thus, only 800 bits (100 bytes) are needed to store each occurrence of the "B" symbol. The less frequent "E" symbol can be represented by a longer binary code, such as "1111". As a result, 100 occurrences of the "E" symbol require 400 bits (50 bytes) of storage. Continuing with this example, the symbols depicted in Table 1 can be encoded using a total of 375 bytes. Without using a DHT, this same data would require 1600 bytes of storage.

[0047] To increase the speed of DEFLATE compression, this Huffman tree generation process can be implemented in hardware. The LZ77 algorithm in DEFLATE uses 256 literals (ASCII values 0x00 - 0xFF), 29 length symbols, and 30 distance symbols for compression. The length symbols and distance symbols represent the distance and length of matching strings in the data stream (data history). Since a length is always followed by a distance, a DHT can be built to encode the literals, end-of-block symbol, and length symbols. This requires a total of 286 symbol letters. A second DHT can be built for the distance symbols. This requires a total of 30 symbol letters.

[0048] One challenge associated with the Huffman tree generation process is that it is actually difficult to fill each DHT leaf with the correct symbol. For each leaf, the symbol with the second highest frequency is required. In other words, the frequency of each symbol must be determined, stored, and sorted. This sorting process can be expensive in terms of area (the number of latches and width comparators required), power, and timing / wiring considerations.

[0049] To illustrate this, consider 2 NLZ77 compression of data in bytes. To fully (uniquely) encode all 286 alphabetic symbols into the first DHT of the Huffman encoder (i.e., the DHT encoding literal, end-of-block, and length symbols) would require an N-bit counter. For example, LZ77 compressing 16 MB of data using all 286 symbols would require a 24-bit counter. In another example, LZ77 compressing 32 MB of data using all 286 symbols would require a 25-bit counter.

[0050] To store the counts associated with each of these 286 symbols, sorted blocks can be used to store 286 "symbol, count" pairs. In a hardware implementation, these pairs are stored in latches. Continuing with the previous example, to store 286 symbols with a 24-bit counter would require 6,864 latches (sometimes called flip-flops). While this latch requirement is already area-intensive, for each additional bit required for the counter, the number of latches required increases by N. For example, storing 286 symbols with a 25-bit counter (for a 32 MB data stream) requires 7,150 latches. Similarly, storing 286 symbols with a 26-bit counter (for a 64 MB data stream) requires 7,436 latches.

[0051] Now turning to an overview of aspects of the teachings of the present invention, one or more embodiments address the above disadvantages of the prior art by providing new accelerator hardware and software implementations for reducing the latch count required for symbol sorting when generating dynamic Huffman tables. The latch count is reduced by mapping the X-bit symbol frequencies (sometimes called "LZ counts") received from the LZ77 compressor to a Y-bit floating-point-like representation that requires fewer than X bits (i.e., X is greater than Y) before sorting. The following process is illustrated specifically for a 24-bit counter, however, it should be understood that the low-count mapping can be adapted to work with any N-bit counter. The 24-bit counter is chosen only for ease of discussion.

[0052] In some embodiments of the present invention, a 24-bit counter (for 16 MB of data) can be mapped to a 10-bit value. To achieve this, the 24-bit value is mapped to a 5-bit exponent (also called a shift field) and a 5-bit mantissa (also called the most significant bits).

[0053] The 5-bit exponent represents the position of the first "1" in the 24-bit counter (this bit is called the shift bit). Mathematically, the 5-bit exponent is the amount of shift required to obtain the original value. For example, the first (most significant) "1" in the 24-bit value "0000000 1 0110111100010101" appears in the 17th digit (read from the right). The 17th digit can be encoded as the 5-bit binary number "10001".

[0054] Once this shift is known, the "0" bits to the left of the shifted bit can be discarded without losing any information. Note that 5 bits of exponent are needed to store each possible position of the shifted bit in a 24-bit counter (5 binary digits are needed to uniquely encode 24 shift possibilities). Although shown as 5 bits of exponent, the number of bits can be more or fewer, depending on the underlying counter being mapped. For example, a 32-bit counter would require 6 bits of exponent for an exhaustive mapping of the shift.

[0055] The 5-bit mantissa contains the five most significant bits of the non-zero data present in the 24-bit counter. In some embodiments of the present invention, the 5-bit mantissa includes the shifted bit, while in other embodiments the shifted bit is skipped. For example, from the previous example "0000000 1 0110111100010101" the resulting 5-bit mantissa is " 1 0110" (when including the shifted bit and the next four digits) and "01101" (when skipping the shifted bit and including the next five digits).

[0056] In either case, these 5-bit values are then combined to provide a 10-bit many-to-one mapping of the 24-bit counter. A "many-to-one" mapping means any mapping in which two or more input values will map to the same output value. Continuing with the previous example, multiple 24-bit counters will map to the same 10-bit value.

[0057] Although both methods are possible and within the scope of the present invention, the second method utilizes an additional data bit (the shifted bit is not reused). Thus, the second method can reduce the number of many-to-one mappings generated using the first method. Continuing with the previous 24-bit example, the first method (where the shifted bit is the first digit of the mantissa) results in a 32-1 mapping, while the second method (ignoring the shifted bit) results in a 16-1 mapping. For illustration, for an LZ count with the value "1____XXXXX" (where "_" represents the same bit value in all LZ counts and "X" represents different bit values), all 32 of these numbers are mapped to 1 number (i.e., a 32:1 mapping). Alternatively, for an LZ count with the value "1____XXXX", only 16 of these numbers are mapped to 1 number (i.e., a 16:1 mapping).

[0058] To further illustrate this, consider the 10-bit mappings of the 24-bit representations of the numbers 929 and 959 respectively " 1 110100000" and " 1"110111111" (leading zeros discarded). Reusing the shifted bit (here, the 10th digit from the right, having binary value "01010") results in the same 10 - digit number: "01010,11101" and "01010,11101". However, ignoring the shifted bit in the mantissa results in unique 10 - digit numbers "01010,11010" and "01010,11011".

[0059] Constructing the many - to - one mapping in this way (shift, mantissa) results in the loss of the exact count (or frequency) of each symbol, but preserves the relative frequency distribution of the symbols. For example, consider symbols "A", "B", "C", and "D" having frequency counts 11, 104, 418, and 1117 respectively in a 16 - MB data stream. A 24 - bit counter can fully encode the exact "symbol, count" pairs for all 286 symbols in a sorted block. A 10 - bit mapping (5 - bit shift, 5 - bit mantissa) will lose the exact count values of these symbols, but will preserve the relative frequencies (i.e., D count >= C count >= B count >= A count).

[0060] Since the relative symbol frequencies are preserved, the latch count can be reduced without affecting the DHT tree quality. In other words, the present disclosure allows the Huffman tree to be populated without knowing the exact frequencies of the symbols. Also, since the deflate algorithm does not allow the DHT tree to exceed a depth of 15 levels (i.e., the encoding length should be 15 bits or less), allowing a many - to - one mapping for high - frequency symbols does not introduce errors into the DHT tree.

[0061] Reducing the number of latches used for a given sorted block frees up valuable die area, reduces power consumption, and simplifies the timing / wiring of the accelerator hardware. Continuing the previous example, mapping a 24 - bit counter to a 10 - bit value before sorting the block reduces the number of required latches from 6,864 latches (24×286) to 2,860 latches (10×286). Additionally, the use of 10 - bit values simplifies the later sorting step because 10 - bit comparators can replace conventional 24 - bit comparators. This results in further area savings.

[0062] In some embodiments of the present invention, the widths of the exponent (shift) and mantissa are fixed (e.g., as described above, each 5 bits). In some embodiments of the present invention, the widths of the exponent (shift) and mantissa can be adjusted dynamically. For example, the widths can be adjusted according to the LZ count range.

[0063] For illustration, consider “K” bits that are implemented to represent the LZ count in a “shift, mantissa” format (i.e., in the previous example using a 5-bit exponent and a 5-bit mantissa, “K” is 10). Depending on the upper limit of the LZ count, “i” bits can be allocated to the shift bits and “K - i” bits can be allocated to the mantissa. This results in a limited improvement in sorting accuracy for the same fixed hardware cost.

[0064] Table 2 shows exemplary dynamic widths based on various LZ count ranges. As shown in Table 2, by dynamically allocating additional bits to the mantissa, the many-to-one mapping can be reduced as the LZ count range increases. Although Table 1 shows shifting a single bit from the shift field to the mantissa, other dynamic adjustments are possible.

[0065] Table 2: Dynamic Shift and Mantissa Width

[0066]

[0067] In some embodiments of the present invention, based on the LZ count range, the width of the shift field is made as small as possible, thereby freeing up additional bits for the mantissa. The width of the shift field can be reduced until the loss of bits would cause some shift bit positions to no longer be uniquely allocable.

[0068] Now refer Figure 2 , a computer system 10 is shown according to a non-limiting embodiment of the present disclosure. The computer system 10 can be based on, for example, the z / architecture provided by International Business Machines Corporation (IBM). However, this architecture is only one example of the computer system 10 and is not intended to impose any limitation on the scope of use or functionality of the embodiments described herein. Other system configurations are possible. In any case, the computer system 10 is capable of implementing and / or performing any of the functions set forth above.

[0069] The computer system 10 can operate with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for operating with the computer system / server 12 include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, cellular telephones, hand-held or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems or devices.

[0070] The computer system 10 may be described in the general context of computer system-executable instructions, such as program modules, executed by a computer system 10. Generally, program modules may include routines, programs, object programs, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. The computer system 10 may be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer system storage media including memory storage devices.

[0071] As Figure 2 shown, the computer system 10 is depicted in the form of a general-purpose computing device, which is also referred to as a processing device. The components of the computer system 10 may include, but are not limited to, one or more processors or processing units 16, a compression accelerator 17, a system memory 28, and a bus 18 that couples various system components including the system memory 28 to the processing unit 16.

[0072] The compression accelerator 17 may be implemented as hardware or both hardware and software, and may include functions and modules for compressing data using the DEFLATE data compression algorithm according to one or more embodiments. In some embodiments of the present invention, the compression accelerator 17 may receive data on an input buffer, process the data using an LZ77 compressor, encode the data using a Huffman encoder, and output the data to an output buffer. Figure 3 An embodiment of the compression accelerator 17 is depicted in

[0073] In some embodiments of the present invention, the compression accelerator 17 may be directly connected to the bus 18 (as depicted). In some embodiments of the present invention, the compression accelerator 17 is connected to the bus 18 between the RAM 30 / cache 32 and the processing unit 16. In some embodiments of the present invention, the compression accelerator 17 is directly connected to the cache 32 (e.g., L3 cache), rather than the bus 18. In some embodiments of the present invention, the compression accelerator 17 is directly connected to the processing unit 16.

[0074] The bus 18 represents any one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a local bus using any of a variety of bus architectures. By way of example and not limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0075] The computer system 10 may include a variety of computer system-readable media. These media can be any available media that can be accessed by the computer system / server 10, including volatile and non-volatile media, removable and non-removable media.

[0076] The system memory 28 may include an operating system (OS) 50 and computer system-readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache 32. The computer system 10 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used for reading and writing on non-removable, non-volatile magnetic media (not shown and commonly referred to as a "hard disk drive"). Although not shown, a disk drive for reading and writing on removable non-volatile disks (e.g., "floppy disks") and an optical disk drive for reading and writing on removable non-volatile optical disks (such as CD-ROM, DVD-ROM, or other optical media) may be provided. In such cases, each drive may be connected to the bus 18 through one or more data media interfaces. As will be further depicted and described below, the memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present disclosure.

[0077] The OS 50 controls the execution of other computer programs and provides scheduling, input-output control, file and data management, memory management, and communication control and related services. The OS 50 may also include a library API (not shown in FIG. 1). The library API is a software library that includes APIs for performing data manipulation functions provided by a dedicated hardware device such as an accelerator (not shown in FIG. 1).

[0078] The storage system 34 may store a basic input / output system (BIOS). The BIOS is a set of basic routines that initialize and test the hardware at startup, start the execution of the OS 50, and support data transfer between hardware devices. When the computer system 10 is in operation, one or more processing units 16 are configured to execute instructions stored in the storage system 34 to transfer data to and from the memory 28 and generally control the operation of the computer system 10 according to the instructions.

[0079] One or more processing units 16 may also access an internal microcode (not shown) and the data stored therein. The internal microcode (sometimes referred to as firmware) may be regarded as a data storage area separate and distinct from the main memory 28 and may be accessed or controlled independently of the OS. The internal microcode may contain a part of the complexly constructed instructions of the computer system 10. Complex instructions may be defined as a single instruction to the programmer; however, it may also include internally licensed code that breaks down a complex instruction into many less complex instructions. The microcode contains algorithms that have been specifically designed and tested for the computer system 10 and may provide full control over the hardware. In at least one embodiment, the microcode may also be used to store one or more compression dictionaries, which may be delivered to the hardware to facilitate data decompression as described in more detail below.

[0080] By way of example and not limitation, a program / utilities 40 having a set (at least one) of program modules 42, as well as the OS 50, one or more application programs, other program modules, and program data may be stored in the memory 28. Each or some combination of the operating system, one or more application programs, other program modules, and program data may include an implementation of the network environment. The program modules 42 generally execute the functions and / or methods of embodiments of the present invention as described herein.

[0081] The computer system 10 may also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.); may also communicate with one or more devices that enable a user to interact with the computer system / server 10; and / or may communicate with any device that enables the computer system / server 10 to communicate with one or more other computing devices (such as, for example, a network card, a modem, etc.). Such communication may be through an input / output (I / O) interface 22. In addition, the computer system 10 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network (such as, for example, the Internet)) through a network adapter 20. As shown, the network adapter 20 communicates with other components of the computer system 10 through a bus 18. It should be understood that although not shown, other hardware and / or software components may be used in conjunction with the computer system 10. Examples include but are not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, data backup storage systems, etc.

[0082] In computer system 10, various types of compression algorithms can be used, such as, for example, the Adaptive Lossless Data Compression (ALDC) family of products that use derivative methods of Lempel-Ziv coding to compress data. As a general compression technique, the Lempel-Ziv 77 (LZ77) algorithm is well integrated into systems required to process many different data types. The algorithm processes a sequence of bytes by maintaining the most recent history of the bytes processed and pointing to matching sequences within the history. Compression is achieved by replacing the matching byte sequence with a copy pointer and a length code, which together are smaller in size than the replaced byte sequence.

[0083] The compression algorithm can also include the "DEFLATE" compression format, which uses a combination of the LZ77 algorithm (which removes duplicates from the data) and Huffman coding. Huffman coding is an entropy coding based on a "Huffman tree". To perform Huffman coding and decoding on data, the system must know in advance that a Huffman tree is being used. To accommodate decompression (e.g., the "Inflate" operation), the Huffman tree is written at the head of each compressed block. In one embodiment, two options are provided for the Huffman tree in the Deflate standard. One option is the "static" tree, which is a single hard-coded Huffman tree that is known to all compressors and decompressors. The advantage of using this static tree is that its description does not have to be written to the head of the compressed block and is ready for immediate decompression. On the other hand, the "dynamic" tree is customized for the current data block, and therefore, the exact description of the dynamic tree must be written to the output.

[0084] Huffman coding can also use an entropy-based variable-length code table to encode source symbols and, as mentioned above, is defined as static or dynamic. In static Huffman coding, a fixed table (FHT) defined in the RFC is used to encode each literal or distance. However, in dynamic Huffman coding, a special coding table (DHT) is constructed to better fit the statistics of the data being compressed. In most cases, compared to the FHT, using the DHT results in a better compression ratio (e.g., quality), at the cost of a lower compression rate (e.g., performance) and increased design complexity. The fixed Huffman coding method and the dynamic Huffman coding method best reflect the inherent trade-off between the compression rate and the compression ratio. The static Huffman method can achieve a lower compression ratio than the compression ratio possible with dynamic Huffman coding. This is due to the use of a fixed coding table regardless of the content of the input data block. For example, random data and four-letter DNA sequences will be encoded using the same Huffman table.

[0085] In some embodiments of the present invention, computer system 10 includes a compression library, which can be implemented as a software library for deflation / inflation and can be an extraction of a compression algorithm. In at least one embodiment, the compression library allows computer system 10 and / or compression accelerator 17 to decompose input data to be deflated / inflated across multiple requests in any manner and provides an output buffer of any size to store the results of the deflation / inflation operation.

[0086] Figure 3 depicts a block diagram of compression accelerator 17 in accordance with one or more embodiments. Figure 2 Compression accelerator 17 can include, for example, input buffer 302, LZ77 compressor 304, Huffman encoder 306 (sometimes referred to as a DEFLATE Huffman encoder), and output buffer 308. As Figure 3 shown, input buffer 302 can be communicatively coupled to LZ77 compressor 304, and the output of LZ77 compressor 304 can be directly connected to the input of Huffman encoder 306. In this way, DEFLATE accelerator 200 is configured to facilitate data compression using the DEFLATE algorithm.

[0087] In some embodiments of the present invention, uncompressed data is obtained by compression accelerator 17 on input buffer 302 (sometimes referred to as the input data buffer). In some embodiments of the present invention, compression accelerator 17 performs LZ77 compression on the data provided to input buffer 302. In some embodiments of the present invention, the compressed data is received and encoded by Huffman encoder 306. In some embodiments of the present invention, the compressed and encoded data can be stored in output buffer 308 (sometimes referred to as the output data buffer).

[0088] To initiate data compression, compression accelerator 17 can receive one or more requests to compress target data or a target data stream in input buffer 302. In some embodiments of the present invention, a request block (not shown) can be used to facilitate the request. In some embodiments of the present invention, the request block is transmitted to the compression interface of OS 50. For each request, computer system 10 can provide the data to be processed to an input buffer (e.g., input buffer 302) and provide an output buffer (e.g., output buffer 308) in which to store the results of the processed data.

[0089] In some embodiments of the present invention, to start processing a compression request, the compression accelerator 17 reads the request block and processes the data in the input buffer 302 to generate compressed and / or decompressed data. As described herein, various compression algorithms can be employed, including but not limited to the DEFLATE compression algorithm and the ALDC algorithm. The resulting compressed data can be saved in the output buffer 308.

[0090] Figure 4 Describes a Figure 3 Block diagram of the DHT generator 400 of the Huffman encoder 306 shown. As Figure 4 shown, the DHT generator 400 can include a sorting module 402, a Huffman tree module 404, a tree static random access memory (SRAM) 406, a tree traversal module 408, a code length SRAM 410, and a coding length module 412. In some embodiments of the present invention, the DHT generator 400 is the first stage of a Huffman encoder (e.g., Figure 3 the Huffman encoder 306 shown in

[0091] The sorting module 402 receives the symbol frequency counter ("LZ count", an X-bit counter) of each symbol compressed by the LZ77 compressor 304. According to one or more embodiments, the sorting module 402 then maps the X-bit counter to a compressed many-to-one Y-bit value. In some embodiments of the present invention, the Y-bit values are sorted (as previously discussed herein to generate the relative frequency distribution of the symbols).

[0092] In some embodiments of the present invention, the Y-bit mapping can be decompressed back to the X-bit value after sorting but before the Huffman tree module 404. In this way, the Huffman tree module 404 can receive the complete X-bit value and does not need to be modified. Similarly, any remaining downstream modules, including the Huffman tree module 404, the tree SRAM 406, the tree traversal module 408, the code length SRAM 410, and the coding length module 412, do not need to be modified. In other words, the Huffman tree module 404, the tree SRAM 406, the tree traversal module 408, the code length SRAM 410, and the coding length module 412 can be implemented using known DEFLATE compression implementations and are not meant to be limiting. Although described as having separate modules for ease of discussion, it should be understood that the DHT generator 400 can include more or fewer modules. For example, the output of the sorting module 402 can be received by a single Huffman tree module and encoded into the DHT, and may or may not include a separate tree SRAM and / or code length SRAM.

[0093] Figure 5 Depicts a Figure 4Block diagram of the classification module 402 as shown. As Figure 5 shown, the sorting module 402 (also referred to as a sorting block) may include a bit converter. For ease of discussion, a 24-bit to 10-bit converter 502 is described; as previously discussed herein, other X-bit to Y-bit conversions are possible.

[0094] In some embodiments of the present invention, the 24-bit to 10-bit converter 502 receives a 24-bit counter from an LZ77 compressor (e.g., Figure 3 the LZ77 compressor 304 as shown). In some embodiments of the present invention, the 24-bit to 10-bit converter 502 generates a 5-bit exponent and a 5-bit mantissa based on the 24-bit counter according to the following algorithm:

[0095] Step 1: Determine the leading zero bit (LZB) index of the 24-bit counter, where the index is from 1 to 24 from the least significant bit to the most significant bit (1 to 25 for a 25-bit counter, etc.).

[0096] Step 2: Generate a 29-bit vector by concatenating the 24-bit counter with "00000". For example, the 24-bit value "000000010110111100010101" can be concatenated with "00000" to form "000000010110111100010101.00000".

[0097] Step 3: Shift the 29-bit vector by the LZB index.

[0098] Step 4: Store the shift amount (i.e., the shift bit position) as a 5-bit exponent. For example, the 17th bit (read from the right, underlined for emphasis) of the 24-bit value "0000000 1 0110111100010101" can be stored as the 5-bit binary number "10001".

[0099] Step 5: Store the five most significant digits as a 5-bit mantissa. In some embodiments of the present invention, the five most significant digits include the shift bit and the next four digits. For example, the 5-bit mantissa generated from the 24-bit value "0000000 1 0110111100010101" can be "10110". In some embodiments of the present invention, the five most significant bits include the five digits immediately following the shift bit. For example, the 5-bit mantissa generated from the 24-bit value "0000000 1 0110111100010101" can be "01101".

[0100] Step 6: Concatenate the 5-bit exponent and the 5-bit mantissa to generate a 10-bit value. Continuing the previous example, where the shift bit is ignored in the mantissa, the 10-bit value is "10001,01101" (shift, mantissa).

[0101] In some embodiments of the present invention, the 24-bit to 10-bit converter 502 receives 24-bit counters for each symbol in the data stream from the LZ77 compressor (e.g., 286 24-bit counters for each of the 286 symbols in the DHT). In some embodiments of the present invention, a 10-bit value is generated for each of the 24-bit counters. These 10-bit values can be passed to the sorting module 504.

[0102] In some embodiments of the present invention, the sorting module 504 sorts the values of the 286 10-bit values. Any known suitable method for a DEFLATE accelerator can be used to sort the 10-bit values. In some embodiments of the present invention, the sorting module 504 stores 286 "symbol, count" pairs in 2,860 latches and uses 2-D shearsort for fast execution. For 2-D shearsort, the 286 "symbol, count" pairs can be arranged in an 18×16 matrix filled with 143 comparators. The comparators are spaced such that no two comparators are horizontally or vertically adjacent (directly left, right, up, or down). Instead, each comparator is diagonally adjacent to one or more other comparators. Advantageously, 10-bit comparators can be used instead of 24-bit comparators, further increasing the area savings provided by the 10-bit mapping. In some embodiments of the present invention, the sorted 10-bit values can then be used to generate a dynamic Huffman tree.

[0103] In some embodiments of the present invention, downstream processing (after classification) requires conversion back to 24-bit values. This allows, for example, easy addition of the LZ count based on 2 ascending symbols and comparison of the LZ count of the next symbol. In some embodiments of the present invention, the 10-bit to 24-bit decompressor 506 receives each 10-bit number from the sorting module 504 and converts each number back to a 24-bit number. For ease of discussion, a 10-bit to 24-bit decompressor is described; as previously discussed, other Y-bit to X-bit decompressors are also possible.

[0104] The 24-bit number can be constructed from the 10-bit number according to the following algorithm: Step 1. Generate a 29-bit field in which all digits are set to "0". Step 2. Copy the mantissa from the 10-bit number to the least significant digits of the 24-bit number. Step 3. Shift the value of the shift bit (or the shift bit minus one if the shift bit is ignored in the mantissa) and insert the shift bit if it is not included in the mantissa. Step 4. Discard the five leading bits (structurally always "0") to convert the 29-bit field to a 24-bit field.

[0105] For illustration, consider, for example, the 10 - bit number "01010,11010" (mantissa ignored shift) generated from the compression of the number 928 as previously discussed herein. In step 2, the 29 - bit field is set to "00……0011010" (leading zeros truncated). In step 3, the 29 - bit field is shifted by 10 digits (10 is the decimal value of the shift bit "01010") and the shift bit is inserted, resulting in "00……001110100000.00000". In step 4, five of the leading "0"s (the left - most bits) are discarded, resulting in the 24 - bit number "000000000 1 11010000000000". Although the previous example was provided in the context of a 10 - bit to 24 - bit decompressor, the same scheme can be used to decompress LZ counts having any initial bit - width (e.g., 11 bits, 12 bits, 20 bits, etc.).

[0106] Figure 6 FIG. 600 is a flow chart depicting a method for reducing the latch count required for symbol sorting when generating a dynamic Huffman table according to a non - limiting embodiment. As shown in block 602, a plurality of first symbol counts are determined. Each of the first symbol counts may include a first bit - width. In some embodiments of the present invention, each of the first symbol counts is encoded as a 24 - bit number.

[0107] At block 604, a plurality of second symbol counts are generated based on the mapping of the plurality of first symbol counts. The second symbol counts may include a second bit - width that is less than the first bit - width. In some embodiments of the present invention, each of the second symbol counts is encoded as a 10 - bit number.

[0108] In some embodiments of the present invention, generating each of the second symbol counts includes generating a 5 - bit shift field and a 5 - bit mantissa according to one or more embodiments. In some embodiments of the present invention, the 5 - bit shift field encodes the position of the most significant non - zero bit of the first symbol (i.e., the shift bit as previously discussed herein). In some embodiments of the present invention, the 5 - bit mantissa encodes the most significant non - zero bit of the first symbol and the next four bits (i.e., the shift bit is reused as the first bit in the mantissa). In some embodiments of the present invention, the 5 - bit mantissa encodes the next five bits of the first symbol after the most significant non - zero bit (i.e., the shift bit is not reused in the mantissa). In some embodiments of the present invention, the 5 - bit shift field and the 5 - bit mantissa are concatenated to form a 10 - digit number.

[0109] At block 606, multiple second symbol counts are sorted by frequency. At block 608, according to one or more embodiments, a dynamic Huffman tree is generated based on the sorted multiple second symbol counts. In some embodiments of the present invention, as previously discussed herein, before generating the dynamic Huffman tree, a 10-bit mapping is decompressed back to a 24-bit number.

[0110] Figure 7 A flowchart 700 depicting a method according to a non-limiting embodiment is shown. As shown at block 702, a data stream including first symbols can be received from an input buffer.

[0111] At block 704, a first symbol count having a first bit width can be determined based on the first symbols. In some embodiments of the present invention, the first bit width is 24 bits.

[0112] At block 706, a 5-bit shift field is generated based on the first symbol count. In some embodiments of the present invention, the 5-bit shift field encodes the position of the most significant non-zero bit of the first symbol.

[0113] At block 708, a 5-bit mantissa is generated based on the first symbol count. In some embodiments of the present invention, the 5-bit mantissa encodes the next five bits after the most significant non-zero bit of the first symbol.

[0114] At block 710, a second symbol count having a second bit width is generated by concatenating the 5-bit shift field and the 5-bit mantissa. At block 712, the frequencies of the second symbol counts are sorted.

[0115] The present invention can be a system, method, and / or computer program product at any possible technical detail integration level. The computer program product can include a computer-readable storage medium (or multiple computer-readable storage media) having computer-readable program instructions thereon for causing a processor to perform aspects of the present invention.

[0116] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium can be, by way of example and not limitation, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punched card or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as being a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse passing through an optical fiber cable) or an electrical signal transmitted through a wire.

[0117] The computer-readable program instructions described herein can be downloaded to respective computing / processing devices from a computer-readable storage medium or can be downloaded to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.

[0118] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by using the state information of the computer-readable program instructions, and the electronic circuit executes the computer-readable program instructions to perform aspects of the present invention.

[0119] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0120] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus create a means for implementing the functions / actions specified in the flowchart and / or one or more blocks of the block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium, which may cause a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium having the instructions stored therein includes an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in the flowchart and / or one or more blocks of the block diagram.

[0121] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other devices implement the functions / acts specified in the flowchart and / or one or more block(s) of the block diagram.

[0122] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which comprises one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially in parallel, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0123] The description of the various embodiments of the present invention has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein were chosen to best explain the principles of the embodiments, the practical application, or the technical improvement over technologies found in the marketplace, or to enable other ordinary skilled in the art to understand the embodiments described herein.

Claims

1. An accelerator, comprising: An input buffer; A Lempel-Ziv 77 (i.e., LZ77) compressor communicatively coupled to an output of the input buffer; A Huffman encoder communicatively coupled to the LZ77 compressor, the Huffman encoder including a bit converter; and An output buffer communicatively coupled to the Huffman encoder; Wherein the bit converter is configured to map a first symbol count including a first bit width to a second symbol count including a second bit width and generate a 5-bit shift field and a 5-bit mantissa based on the first symbol count; And wherein the second bit width is less than the first bit width to reduce a latch count required for symbol sorting when generating a dynamic Huffman table.

2. The accelerator according to claim 1, wherein The bit converter includes a 24-bit to 10-bit converter, the first bit width includes 24 bits, and the second bit width includes 10 bits.

3. The accelerator according to claim 1, wherein, The bit converter is further configured to concatenate the 5-bit shift field and the 5-bit mantissa to generate the second symbol count.

4. The accelerator according to claim 1, wherein, The accelerator includes a DEFLATE hardware accelerator.

5. A method for reducing a latch count required for symbol sorting when generating a dynamic Huffman table, the method comprising: Determining a plurality of first symbol counts, each of the first symbol counts including a first bit width; Generating a plurality of second symbol counts, each of the second symbol counts being a mapping based on a symbol count among the plurality of first symbol counts, the second symbol counts including a second bit width less than the first bit width to reduce a latch count required for symbol sorting when generating a dynamic Huffman table, wherein generating each second symbol count among the plurality of second symbol counts includes generating a 5-bit shift field and a 5-bit mantissa based on a first symbol among the plurality of first symbol counts; Sorting the plurality of second symbol counts by frequency; and Generating a dynamic Huffman tree based on the sorted plurality of second symbol counts.

6. The method according to claim 5, wherein, The first bit width includes 24 bits, and the second bit width includes 10 bits.

7. The method according to claim 5 further comprises: Concatenating the 5-bit shift field and the 5-bit mantissa.

8. The method according to claim 5, wherein The 5-bit shift field encodes a position of a most significant non-zero bit of the first symbol.

9. The method according to claim 8, wherein The 5-bit mantissa encodes the most significant non-zero bit of the first symbol and the next four bits.

10. The method according to claim 8, wherein, The 5-bit mantissa encodes the next five bits after the most significant non-zero bit of the first symbol.

11. The method according to claim 6 further comprises: Generating a 29-bit field by concatenating a first count among the plurality of first symbol counts with a 5-bit field.

12. A computer program product for reducing a latch count required for symbol sorting when generating a dynamic Huffman table, the computer program product including a computer-readable storage medium having program instructions embodied therein, the program instructions executable by an electronic computer processor to control a computer system to perform operations, the operations including: Determining a plurality of first symbol counts, each of the first symbol counts including a first bit width; Generate multiple second symbol counts, each of the second symbol counts being a mapping based on the symbol counts among the multiple first symbol counts, the second symbol counts including a second bit width less than the first bit width to reduce the latch count required for symbol sorting when generating a dynamic Huffman table, wherein generating each second symbol count among the multiple second symbol counts includes generating a 5-bit shift field and a 5-bit mantissa based on a first symbol among the multiple first symbol counts; Sort the multiple second symbol counts by frequency; and Generate a dynamic Huffman tree based on the sorted multiple second symbol counts.

13. The computer program product according to claim 12, further comprising: Concatenate the 5-bit shift field and the 5-bit mantissa.

14. The computer program product according to claim 12, wherein, The 5-bit shift field encodes the position of the most significant non-zero bit of the first symbol.

15. The computer program product according to claim 14, wherein, The 5-bit mantissa encodes the next five bits after the most significant non-zero bit of the first symbol.

16. A system for reducing the latch count required for symbol sorting when generating a dynamic Huffman table, the system comprising: An accelerator; A memory having computer-readable instructions; And A processor configured to execute the computer-readable instructions, wherein the computer-readable instructions, when executed by the processor, cause the accelerator to execute a method that includes: Determine multiple first symbol counts, each of the first symbol counts including a first bit width; Generate multiple second symbol counts, each of the second symbol counts being a mapping based on the symbol counts among the multiple first symbol counts, the second symbol counts including a second bit width less than the first bit width to reduce the latch count required for symbol sorting when generating a dynamic Huffman table, wherein generating each second symbol count among the multiple second symbol counts includes generating a 5-bit shift field and a 5-bit mantissa based on a first symbol among the multiple first symbol counts; Sort the multiple second symbol counts by frequency; and Generate a dynamic Huffman tree based on the sorted multiple second symbol counts.

17. The system according to claim 16, wherein The 5-bit shift field encodes the position of the most significant non-zero bit of the first symbol.

18. The system according to claim 17, wherein, The 5-bit mantissa encodes the next five bits after the most significant non-zero bit of the first symbol.

19. A method for reducing the latch count required for symbol sorting when generating a dynamic Huffman table, comprising: Receive a data stream including first symbols from an input buffer; Based on the first symbols, determine first symbol counts having a first bit width; Based on the first symbol counts, generate a 5-bit shift field; Based on the first symbol counts, generate a 5-bit mantissa; Determine second symbol counts having a second bit width by concatenating the 5-bit shift field and the 5-bit mantissa, wherein the second bit width is less than the first bit width to reduce the latch count required for symbol sorting when generating a dynamic Huffman table; Sort the frequencies of the second symbol counts; and Generate a dynamic Huffman tree based on the sorted second symbol counts.