Data compression using a static dictionary

By implementing a lazy matching algorithm that efficiently compares data buffers and uses a static dictionary, the method addresses the challenges of computational complexity and power consumption in data compression, achieving improved efficiency and compression ratios.

WO2025128605A1PCT designated stage expired Publication Date: 2025-06-19MAXLINEAR INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/059427
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-11
Filing Date
2024-12-10
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing data compression methods, particularly those using lazy matching algorithms, face challenges in reducing computational complexity and power consumption while maintaining effective compression ratios.

Method used

The implementation of a lazy matching algorithm in hardware or software that improves computational efficiency by comparing strings from a look-ahead buffer to a history buffer, extending substring matches without additional scans, and using a static dictionary for efficient matching.

Benefits of technology

This approach reduces CPU cycles and power consumption, improves throughput, and decreases latency in data compression operations, while maintaining or improving compression ratios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024059427_19062025_PF_FP_ABST
    Figure US2024059427_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A method includes storing input data in a first buffer where the input data includes a first set of symbols. The method also includes obtaining one or more keys from a second buffer where the one or more keys individually include buffer substrings. The method further includes comparing a portion of the first set of symbols to the buffer substrings to identify one or more substring matches. The method also includes identifying a longest substring match of the one or more substring matches. The longest substring match includes substring match symbols. The method further includes retaining a key of the one or more keys associated with the longest substring match. The method also includes performing a compression operation to the substring match symbols. The method further includes removing the portion of the first set of symbols from the first buffer.
Need to check novelty before this filing date? Find Prior Art

Description

DATA COMPRESSION USING A STATIC DICTIONARYCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This U.S. Patent Application claims priority to U.S. Provisional Patent Application No. 63 / 608,813, titled “LAZY MATCHING ALGORITHM FOR DATA COMPRESSION,” and filed on December 11, 2023, the disclosure of which is hereby incorporated by reference in its entirety.TECHNICAL FIELD

[0002] This disclosure relates to lossless data compression, and more specifically, to a static dictionary-based lossless data compression.BACKGROUND

[0003] Unless otherwise indicated herein, the materials described herein are not prior art to the claims in the present application and are not admitted to be prior art by inclusion in this section.

[0004] Data compression may be lossy or lossless. In lossless compression, data compression algorithms (or “compression algorithms”) may reduce the size of data by identifying and removing redundancies in the data. Further, information content included in the data may not be removed, such that when a decompression operation may be applied to the compressed data, the decompressed data may be restored to be the same as the data before compression. Many data transform accelerators (DTAs), computational storage devices (CSDs), data processing units (DPUs), network interface controllers (NICs), central processing units (CPUs), and field-programmable gate arrays (FPGAs) used in storage or cryptographic appliances, use lossless compression methods.

[0005] Many lossless data compression techniques may use dictionary-based compression methods that may extract substrings from the original data string. The substrings may be variable length or fixed length, and the substrings may be used to index a dictionary that maps the substrings into tokens. When the tokens can be represented using reduced number of bits relative to the substrings the tokens are mapped from, data reduction may be achieved. Many texts are not a random sequence of symbols and a substring of symbols occurring in the input data may be more likely to appear again in the input data. In instances in which the tokens in the dictionary are references to previous occurrences of the substrings, the tokens may be represented using a reduced number of bits that would otherwise be required to represent original substring, such that data reduction may be achieved.

[0006] Many widely used dictionary-based compression methods may build the dictionary in an adaptive fashion using a sliding window on already processed symbols and on symbols to be compressed. For example, the LZ77 compression algorithm that may be used in LZ4, LZS, Deflate, GZIP, ZLIB, and / or XP10 may utilize an adaptive dictionary as described. The compression algorithm may maintain a history buffer that contains input data already processed by the compression algorithm. The history buffer may operate as a dictionary that may be built adaptively. For example, to process new input data contained in a look-ahead buffer, a pointer may be moved back through the history buffer until a match (e.g., a matched symbol) is found in the history buffer with the first symbol of the new input data included in the look-ahead buffer. Once the match is found, a next symbol in the look-ahead buffer may be compared with a next symbol in history buffer, where the next symbol in the history buffer may be located adjacent to the matched symbol, to determine if additional matches may be obtained. The matching process continues comparing subsequent symbols from the look-ahead buffer with the consecutive symbolsin the history buffer until the match ends. In this fashion, the compression algorithm searches the entire history buffer to determine the longest substring match for a substring in the look-ahead buffer.

[0007] Once the longest match is found, the compression algorithm encodes the longest match with a <distance, length> pair. Distance may be the distance of the beginning of the longest matched substring in the history buffer from the beginning of the look-ahead buffer. Length may be the length of the longest substring match. Different definitions of distance may be considered, such as if distance indicates the location of the substring in the history buffer. Once a substring in the look-ahead buffer is matched, the substring slides into the front of the history buffer from the look-ahead buffer. In instances in which the history buffer is full, the oldest data in the tail of the history buffer may be discarded. In instances in which there is no match, the symbol in the look-ahead buffer may be emitted as a literal token. Some dictionary-based compression algorithms (e.g., XP10) may create a cache of recently used <distance, length> pairs in a look up table and instead of using <distance, length> pair as a token directly, the compression algorithm may use an index in the lookup table associated with the <distance, length> pair to increase compression ratio (commonly known as Move to Front (MTF) coding).

[0008] Some dictionary-based methods may improve a compression ratio by using a lazy evaluation technique, where the compression ratio may be a ratio of a size of the input data relative to a size of the compressed data. After finding the longest substring match using substrings from the beginning of the look-ahead buffer, the compression algorithm may consider a longest substring match, skipping the first symbol of the look-ahead buffer and starting a matching process from the second symbol of the look-ahead buffer with the substrings in the history buffer.

[0009] If a longer match is found, the compression algorithm emits the first symbol in look-ahead buffer as a literal token and a subsequent substring match may be emitted as <distance, length> pair token. Otherwise, the compression algorithm may emit the first longest substring match as <distance, length> pair token. As such, the compression algorithm may consider multiple candidate substrings for longest substring match, while skipping a first few symbols (instead of just one symbol) one at a time from the beginning of the look-ahead buffer and longest substring match may be selected from the candidate substrings. Such compression algorithm may be commonly called lazy evaluation or lazy matching. Alternatively, or additionally, the number of symbols the substring match process can defer from the beginning of look-ahead buffer (to select a substring for match) may be called a lazy match window or a delayed match window. For example, a delayed match window value of two may indicate a longest substring match may be considered among candidates of substring matches i) from the beginning of the look-ahead buffer (including the first symbol therein), ii) skipping the first symbol, and ii) also skipping the first two symbols. Larger value lazy match windows may improve compression ratio at the expense of encode latency and / or computational complexity.

[0010] When compressing a small data block, there may not be enough history or a sufficiently populated dictionary to use some of the compression algorithms described. In such instances, some compression algorithms may use static pre-shared dictionaries. Alternatively, or additionally, some compression algorithms may use a static pre-initialized dictionary and / or a sliding window, where the pre-initialized dictionary may be referenced at any time while scanning the input data for a substring match.

[0011] Once the input data is converted to a set of tokens (e.g., <distance, length> pairs, index in dictionary, literals, etc.), a variable length coding algorithm, such as Huffman code, may be used to encode the tokens. The Huffman code may use fixed codebooks, and / orconstructed codebooks based on frequency of occurrence of the different tokens encountered (also known as retrospective or dynamic coding), that may be generated during the compression process. Multiple codebooks may be used for the same alphabets in the same compressed data block (as in context modeling). Alternatively, or additionally, asymmetric numeral system (ANS) may be used instead of Huffman code, such as may be used in standard.

[0012] The subject matter claimed in the present disclosure is not limited to implementations that solve any disadvantages or that operate only in environments such as those described above. Rather, this background is only provided to illustrate one example technology area where some implementations described in the present disclosure may be practiced.SUMMARY

[0013] In an example embodiment, a method may include storing input data in a first buffer. The input data may include a first set of symbols. The method may also include obtaining one or more keys from a second buffer. The one or more keys may individually include buffer substrings. The method may further include comparing a portion of the first set of symbols to the buffer substrings to identify one or more substring matches. The method may also include identifying a longest substring match of the one or more substring matches. The longest substring match may include substring match symbols. The method may further include retaining a key of the one or more keys associated with the longest substring match. The method may also include performing a compression operation to the substring match symbols. The method may further include removing the portion of the first set of symbols from the first buffer.

[0014] In another embodiment, a compression device may include a first buffer, a second buffer, and a data compression module. The data compression module may beoperable to store input data in the first buffer. The input data may include a first set of symbols. The data compression module may also be operable to obtain one or more keys from a second buffer. The one or more keys may individually include buffer substrings. The data compression module may further be operable to compare a portion of the first set of symbols to the buffer substrings to identify one or more substring matches. The data compression module may also be operable to identify a longest substring match of the one or more substring matches. The longest substring match may include substring match symbols. The data compression module may further be operable to retain a key of the one or more keys associated with the longest substring match. The data compression module may also be operable to perform a compression operation to the substring match symbols. The data compression module may further be operable to removing the portion of the first set of symbols from the first buffer.

[0015] The objects and advantages of the embodiments will be realized and achieved at least by the elements, features, and combinations particularly pointed out in the claims.

[0016] Both the foregoing general description and the following detailed description are given as examples and are explanatory and not restrictive of the invention, as claimed.DESCRIPTION OF DRAWINGS

[0017] Example implementations will be described and explained with additional specificity and detail using the accompanying drawings in which:

[0018] FIG. 1 illustrates a block diagram of an example system for data compression using a lazy matching algorithm;

[0019] FIG. 2 illustrates a flowchart of an example method of data compression using a lazy matching algorithm;

[0020] FIG. 3 illustrates a flowchart of an example aspect of the method of FIG. 2;

[0021] FIG. 4 illustrates a flowchart of an example aspect of the method of FIG. 2;

[0022] FIG. 5 illustrates a flowchart of an example aspect of the method of FIG. 2;

[0023] FIG. 6 illustrates a block diagram of an example look-ahead buffer and an example history buffer that may be used in data compression using a lazy matching algorithm;

[0024] FIG. 7 illustrates a flowchart of an example method of data compression using a lazy matching algorithm;

[0025] FIG. 8 illustrates an example computing device; and

[0026] FIG. 9 illustrates an example suffix tree.DETAILED DESCRIPTION

[0027] Various implementations exist for compressing data, and each of them may vary in a compression ratio of the input data, where the compression ratio may be a ratio of a size of the input data relative to a size of the compressed data. In many instances, improvements to the compression ratio in a particular compression operation may be at the expense of increased computational complexity and / or increased latency in the compression operation. Some prior approaches may use a lazy matching data compression operation, which may be computationally expensive and / or which may cause increases in the latency of the compression operation.

[0028] Lazy matching may use multiple searches in a history buffer or the dictionary (or just “history buffer”) which may make lazy matching computationally expensive and / or increase the latency of the compression operation. As such, more CPU cycles may be used when implemented in software on a CPU. When lazy matching is implemented on hardware (such as reconfigurable hardware such as FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuits), the lazy matching operation may use more clock cycles and / or more hardware resources (e.g., circuits) resulting in higher power consumption.

[0029] Aspects of the present disclosure address these and other limitations by implementing a lazy matching algorithm in hardware or in software that may improve on computational complexity, reduce consumption of computational resources, such as CPU cycles on a processor or clock cycles, and / or reduce power consumption to process data in a hardware-based implementation. Aspects of the present disclosure may also improve throughput and / or latency of the data compression operation. In such implementations, the lazy matching algorithm may compare strings from a look-ahead buffer to a history buffer to obtain substring matches. The lazy matching algorithm may also extend the substring matches to include additional symbols without performing additional scans of the history buffer, such that the computational expense and / or the latency associated with the lazy matching algorithm may be reduced, compared to other lazy matching algorithms and / or other data compression operations.

[0030] FIG. 1 illustrates a block diagram of an example system 100 for data compressions using a lazy matching algorithm, in accordance with at least one embodiment of the present disclosure. The system 100 may include a computing device 110, a look- ahead buffer 120, and a history buffer 125. The computing device 110 may include a data compressions module 115 and the computing device 110 may be operable to output compressed data 130.

[0031] The computing device 110 may be any computing device operable to perform at least data compression operations and obtain input data 102 and generate and / or transmit the compressed data 130. The computing device 110 may be communicatively coupled to the look-ahead buffer 120 and / or the history buffer 125, such that the computing device 110 may direct data to be stored in the look-ahead buffer 120 and / or the history buffer 125 (e.g., the input data 102), direct operations associated with the data stored in the look-ahead buffer 120 and / or the history buffer 125, and / or use the look-ahead buffer 120 and / or thehistory buffer 125 to generate the compressed data 130, such as via the lazy matching algorithm described herein.

[0032] The look-ahead buffer 120 and the history buffer 125 may be storage devices operable to store at least the input data 102. For example, the input data 102 may be obtained by the computing device 110 and may be directed to be stored in the look-ahead buffer 120. The history buffer 125 may include a dictionary that may be used in a comparison with the input data 102 in the look-ahead buffer 120.

[0033] The history buffer 125 may be a static dictionary that may include one or more keys. In some instances, the keys in the history buffer may be substrings and / or values that may be used to match the symbols in the look-ahead buffer 120. In some instances, the history buffer 125 may assembled based on previously processed data (e.g., based on a frequency of substrings that may have been present in the input data 102), and / or based on a machine learning output from a dataset. In these and other embodiments, the symbols and / or the keys in the history buffer 125 may be predetermined.

[0034] In some instances, multiple history buffers may be constructed, where the history buffers may vary based on a class of data that the history buffer may be used with. For example, a first history buffer may be constructed and used with a first class of data, a second history buffer may be constructed and used with a second class of data, and so forth. In these and other embodiments, the history buffer 125 may be arranged for efficient implementation. For example, the history buffer 125 may be initialized prior to performance of any compression operations.

[0035] In some instances, one or more suffixes may be constructed for each key in the history buffer 125. In some instances, the suffixes may be arranged in a data structure, such as a suffix tree and / or a suffix array. In these and other embodiments, the suffixes may be constructed based on a length of the key and / or a length of a delayed match window (DMW)window, as described herein. For example, in instances in which a particular key is X1X2X3X4. . . Xnand the DMW has a length of three, the suffixes may include at least X1X2X3X4. . . Xn, X2X3X4. . . Xn, X3X4. . . Xn, and X4. . . Xn.

[0036] In instances in which the suffixes are arranged in a suffix tree, each edge of the suffix tree may be associated with a substring of the key in the history buffer 125 when the history buffer 125 is a static dictionary. Alternatively, or additionally, each leaf in the suffix tree may be marked with the starting position of the suffix in the key and / or a particular key identifier. An example suffix tree 900 is illustrated in FIG. 9. In the suffix tree 900, two keys are considered where the first key is X1X2X3X4X5X6X7 and the second key is X1X2X3X4Y1Y2Y3. The DMW in the suffix tree 900 may have a length of 3.

[0037] The substrings in the look-ahead buffer 120 at the coding position aligned with the symbol after the DMW length number of symbols may be matched with the suffixes. After matching, the matched substrings (Si) may be retained in the look-ahead buffer 120, represented by (lmax - DWM length) < |Si| < lmax, where lmax may be the maximum length of the matched substring, represented by lmax = arg max { |Si|} and |Si| may be the length of the zthmatched substring Si.

[0038] For the suffixes associated with the matched substrings Si described above and for the suffixes that include a starting position greater than 1 (e.g., all of the suffixes in the suffix tree 900 except for X1X2X3X4X5X6X7 and X1X2X3X4Y1Y2Y3), two operations may be performed. The first operation may include comparing the prefix of each matched suffix of the keys to symbols in front of the coding position in the look-ahead buffer 120. The second operation includes retaining a particular key when a match in the first operation is obtained, or dropping the particular key when no match is obtained.

[0039] In an example, in instances in which the look-ahead buffer 120 includes symbols Si, S2, ... S10, with Si at the start of the look-ahead buffer 120 and a DMW length of 3, thecoding position may be at S4. In instances in which a matched suffix is the third position of the second key described above (e.g., X3X4Y1Y2Y3), the prefix of the second key (e.g., X1X2) may be compared to the two symbols previous in the look-ahead buffer 120 (e.g., S2 and S3). In instances in which the prefixes match the symbols, the key may be retained. In instances in which the prefixes do not match the symbols, the key may be dropped. In instances in which the second key is retained, Si may be emitted as a literal and symbols S2 through S?may be represented by the second key, which may be emitted for a next coding stage.

[0040] In another example and using the same look-ahead buffer 120 as the previous example and the same DMW length, in instances in which the matched suffix is the first position of the second key (e.g., X1X2X3X4Y1Y2Y3), there may not be a prefix to compare to additional symbols in the look-ahead buffer 120. In such instances, and if the second key is retained, Si, S2, and S3 may be emitted as literals symbols S4 through S10 may be represented by the second key, which may be emitted for a next coding stage.

[0041] In these and other embodiments, for all key matches that may be retained, a key match having the longest length may be used to replace the substring of symbols (the matched symbols) in the look-ahead buffer 120. Symbols in the look-ahead buffer 120 that are in front of the matched symbols but are not matched may be emitted as literals. In instances in which there is a tie for the longest length of the matched keys, a selection criteria may be used to determine which of the more than one matched keys may be selected. For example, a first key that may facilitate better compression may be selected over a second key (e.g., the first key may have been used more recently in a compression operation). Once the matched key with the longest length and / or literals are emitted, the look-ahead buffer 120 may be adjusted such that the emitted matched symbols and / orliterals may be removed from the look-ahead buffer 120 and the process may continue until all symbols in the look-ahead buffer 120 have been emitted.

[0042] In instances in which no matched keys are determined (e.g., no match was found, or the length of the match failed to satisfy a threshold), the first symbol in the look- ahead buffer 120 may be emitted as a literal. Subsequently, the look-ahead buffer 120 may slide by one symbol (e.g., such that the first symbol may be removed therefrom) and the process of matching keys and / or suffixes from the history buffer 125 to the symbols in the look-ahead buffer 120 may be repeated until all symbols in the look-ahead buffer 120 may be processed.

[0043] The computing device 110 may be operable to implement dictionary -based compression methods, which may include generating fixed length substrings and / or variable length substrings from the input data 102 and the substrings may be used to index a dictionary that maps the substrings into tokens. In instances in which the mapping results in reduction of number of bits to represents the tokens relative to the input data, data reduction may be achieved.

[0044] The data compression module 115 may be operable to perform and / or direct operations to be performed to generate the compressed data 130 based on the input data 102. The data compression module 115 may be operable to perform the lazy matching algorithm described herein. The data compression module 115 may be operable to search the history buffer 125 (e.g., such as by adjusting a location of a pointer relative to the symbols in the history buffer 125) to attempt to obtain a match relative to a symbol in the look-ahead buffer 120. In instances in which a match is found (e.g., between the symbol in the look-ahead buffer 120 and the history buffer 125), the data compression module 115 may attempt to extend the match by comparing an adjacent symbol (e.g., adjacent to the matched symbol) in the look-ahead buffer 120 with an adjacent symbol in the history buffer125. The data compression module 115 may continue in such manner until an adjacent symbol in the look-ahead buffer 120 fails to match with an adjacent symbol in the history buffer 125. In this manner, the data compression module 115 may be operable to continue matching symbols in the look-ahead buffer 120 with symbols in the history buffer 125 until the match ends. The matched symbols between the look-ahead buffer 120 and the history buffer 125, as described, may be referred to as a substring match, unless indicated otherwise.

[0045] In such manner, the data compression module 115 may be operable to search the entire history buffer 125 to determine the longest substring match for a substring in the look-ahead buffer 120. In some instances, the data compression module 115 may consider a length of the substring match in view of a minimum substring match length, where the minimum substring match length may indicate a minimum number of symbols to be included in the substring for data compression. In some instances, the minimum substring match length may be input by a user of the system 100, may be based on the input data 102 (e.g., a data type associated with the input data 102, an amount of symbols included in the input data 102, etc.), may be dynamically adjusted by the system 100 and / or the computing device 110 based on the input data 102 (e.g., a number of substring matches satisfying the minimum substring match length is less than a desired / expected threshold, such that the system 100 may adjust the minimum substring match length such that more substring matches may satisfy the minimum substring match length), etc.

[0046] Once the data compression module 115 determines a longest substring match (where the longest substring match may include a length that is greater than or equal to the minimum match length), the data compression module 115 may encode the longest substring match with a token, where the token may include at least a distance reference and / or a length reference. In such instances, the distance reference may be the number ofsymbols from the first symbol in the history buffer 125 to the first symbol in the longest substring match, and length reference may be the number of symbols included in the longest substring match. In instances in which no substring match is determined between the symbols in the look-ahead buffer 120 and the history buffer 125, the symbols in the look- ahead buffer 120 may be emitted as literals.

[0047] The data compression module 115 may be operable to perform one or more iterations of determining the longest substring match, by adjusting a pointer in the look- ahead buffer 120 to skip one or more symbols before performing the matching as described. In some instances, the data compression module 115 may skip a number of symbols based on a size of a delayed match window that may be used in the data compression operation. A detailed description of the algorithm and steps there may be further described and illustrated relative to the flowcharts in FIGS. 2-6.

[0048] As the data compression module 115 implements the above described algorithm, or in other words a lazy matching algorithm with a delayed match window, the data compression module 115 may compress the input data 102 into the compressed data 130. In some instances, as the size of the delayed match window increases, the compression ratio may improve and may correspondingly increase the encode latency in the system 100. In the present disclosure, the data compression module 115 may implement a lazy matching algorithm that may be lower complexity relative to other lazy matching algorithms, reduce CPU cycles when implemented in software, and / or reduce clock cycles and / or reduce power consumption of the system 100 when implemented in hardware.

[0049] The system 100 may be implemented in various devices, such as, but not limited to, computation storage devices, DPUs, NICs, general purpose CPUs, GPUs, microcontrollers, FPGAs, ASICs, and / or data transform accelerators, any of which may be used for data compression. In instances in which the system 100 is a storage andcryptographic data transform accelerator, the system 100 may contain one or more data transform engines as compute resources for data transform operations such as data compression, cryptographic operations, and / or other transform operations. The data transform engines included in the system 100 can operate on the data in a highly parallel fashion. Such data transform accelerators may be connected to a host computer or a server platform using PCI Express (PCIe), CXL, and / or USB. A host computer or server may submit commands to the data transform accelerator along with source data to transform. Alternatively, or additionally, the data transform accelerator may provide control information and / or metadata that describes the specific algorithmic transformation to be applied on the data. Based on the metadata, the data transform engines may perform operations including data compression on the data. The data transform engines for dictionary-based data compression in the system 100 that may be a data transform accelerator may implement the low complexity algorithm provided herein.

[0050] The following examples of two iterations of lazy math (delayed match window size of 1) provides at least some motivation behind the algorithm described herein.

[0051] Set A may include the substring matches where the match length is greater than or equal to the minimum match length when the search begins from the first symbol in the look-ahead buffer (e.g., the DMW = 0 symbol). The above described set (e.g., Set A) may be referred to as the DMW = 0 iteration match set. Each element of Set A may be a matched substring in the history buffer (e.g., that includes a length greater than or equal to minimum match length).

[0052] Set B may include the substring matches where match length is greater than or equal to the minimum match length when the search begins from the second symbol in the look-ahead buffer (e.g., the DMW = 1 symbol). The above described set (e.g., Set B) may be referred to as the DMW = 1 iteration match set. Each element of Set B may be a matchedsubstring in history buffer (e.g., that includes a length greater than or equal to minimum match length).

[0053] Considering the two iterations described above (e.g., Set A where DMW = 0 and Set B where DMW = 1), at least the following three cases are possible: Case 1, Case 2, and Case 3, all described below.

[0054] In Case 1, Set A may be NULL and regardless of what may be in Set B, the DMW = 0 symbol may be emitted as a literal and the symbol pointer may be advanced by one symbol (e.g., the next symbol). In instances in which Set B is constructed to include all matches greater than or equal to the minimum match length minus one, then Set B may be NULL. In such instances, Set B may be NULL as a match of the minimum match length in the DMW = 0 iteration may yield a match of the minimum match length minus one in the DMW = 1 iteration. Therefore, Set A may not be constructed as the NULL value of Set A may be inferred from the NULL-ity of Set B when Set B is constructed for a substring match of a minimum match length minus one.

[0055] In Case 2, Set B may be NULL and Set A may be not NULL (e.g., there is no substring match in DMW = 1 iteration of length greater than or equal to minimum match length, but there may be substring matches in the history buffer in the DMW = 0 iteration). Such a configuration may occur when at least one of the longest substring matches in the DMW = 0 iteration is the minimum match length. Alternatively, if Set B is constructed to include all matches that are greater than or equal to the minimum match length minus one, then Set B may contain elements having a length equal to the minimum match length minus one. because the foregoing may be due to a match having the minimum match length in the DMW = 0 iteration may yield a match having the minimum match length minus one in the DMW = 1 iteration. In such circumstances, constructing Set A may be unnecessary as the elements having the minimum match length in Set A may be inferred from the lengths ofthe elements in Set B. Additionally, the first symbol in the look-ahead buffer may be verified as to whether it matches the symbol preceding the matched substrings in the history buffer corresponding to the matches in Set B.

[0056] Alternatively, or additionally, in Case 2, a substring search in the DMW = 1 iteration may be performed for one less than the minimum match length (e.g., minimum match length - 1) and the results may be checked to determine if any of the matches can be extended, by comparing the first symbol in the look-ahead buffer with the symbol before the matched substrings in history buffer. All substring matches with the one less than the minimum match length obtained in the DMW = 1 iteration, and that have a symbol in the history buffer before the substring matching with the first symbol of the look-ahead buffer, may be candidate substring matches in DMW = 0 iteration. One of these candidate matches may be chosen. The criteria for selecting the candidate matches may include, but not limited to, the candidate with least distance from the beginning of the look-ahead buffer. Or a different criterion can be chosen to break the tie.

[0057] The search process may be terminated once a match is determined (e.g., so long as the match satisfies the minimum match length). In such instances, the match may be emitted using distance and length values, and the symbol pointer may be advanced to a next symbol based on the length value. In such instances, the length value may satisfy the minimum match length.

[0058] In Case 3, Set A may not be NULL and Set B may not be NULL. There may be one or more substring matches in the DMW = 0 iteration, where the substring matches may include a length greater than or equal to a minimum match length. Alternatively, or additionally, there may be one or more substring matches in the history buffer in the DMW = 1 iteration, where the substring matches may include a length greater than or equal to the minimum match length. In instances in which Set B is constructed to include all matchesgreater than or equal to the minimum match length minus one, then set B may not be NULL. A match of the length ‘L’ in the DMW = 0 iteration may result in a match of the match length ‘L-T in the DMW = 1 iteration. Therefore, Set A may not be constructed by scanning the history buffer and dictionary again as the elements of Set B may be considered with the longest length and the elements with the longest length minus one. Alternatively, or additionally, the selected elements included in Set B may be considered as to whether they may be extended further by comparing the first symbol in the look-ahead buffer with the symbol preceding the substring matches in Set B within the history buffer.

[0059] In this case, the substring matches in the DMW = 1 iteration that have the longest match length and / or the longest match length minus one may be selected. Alternatively, or additionally, the substring matches in the DMW = 1 iteration having a length one symbol less than longest match length may be selected.

[0060] In response to the selected substring matches included in the history buffer, the selected substring matches may be examined for a determination as to whether any of the selected substring matches may be further extended. For example, a first symbol in the look-ahead buffer may be compared with a symbol before the matched substrings in history buffer. Based on the content of the history buffer, some or none of the selected substring matches may be extended. After attempting to extend the substring matches, the substring matches having the longest length may be selected. In instances in which there are two or more matched substrings having an equal longest length, a tie-breaking criteria to select one matched substring over the other(s) may be implemented. For example, the tie-breaking criteria may include selecting the matched substring having a shortest distance from the beginning of the look-ahead buffer. Other tie-breaking may be implemented, any of which may be used to select a particular matched substring from multiple matched substrings having an equal longest length, as described.

[0061] Modifications, additions, or omissions may be made to the system 100 without departing from the scope of the present disclosure. For example, the designations of different elements in the manner described is meant to help explain concepts described herein and is not limiting. Further, the system 100 may include any number of other elements or may be implemented within other systems or contexts than those described. For example, any of the components of FIG. 1 may be divided into additional or combined into fewer components.

[0062] FIG. 2 illustrates a flowchart of an example method 200 of data compression using a lazy matching algorithm, in accordance with at least one embodiment of the present disclosure. The method 200, or subsequently described method 300, method 400, method 500, and / or method 700, may be performed by processing logic that may include hardware (circuitry, dedicated logic, etc.), software (such as is run on a general purpose computer system or a dedicated machine), or a combination of both, which processing logic may be included in any computer system or device such as computing device 110 or data compression module 115 of FIG. 1.

[0063] For simplicity of explanation, methods (e.g., the method 200, the method 300, the method 400, the method 500, and / or the method 700) described herein are depicted and described as a series of acts. However, acts in accordance with this disclosure may occur in various orders and / or concurrently, and with other acts not presented and described herein. Further, not all illustrated acts may be used to implement the methods in accordance with the disclosed subject matter. In addition, those skilled in the art will understand and appreciate that the methods may alternatively be represented as a series of interrelated states via a state diagram or events. Additionally, the methods disclosed in this specification may be capable of being stored on an article of manufacture, such as a non-transitory computer- readable medium, to facilitate transporting and transferring such methods to computingdevices. The term article of manufacture, as used herein, is intended to encompass a computer program accessible from any computer-readable device or storage media. Although illustrated as discrete blocks, various blocks may be divided into additional blocks, combined into fewer blocks, or eliminated, depending on the desired implementation.

[0064] At block 210, input data to be compressed may be obtained by a computing device (e.g., the computing device 110 of FIG. 1). The input data may be obtained from a remote device (e.g., transmitted to the computing device) and / or the input data may be obtained from a storage device, such as a server, a cloud-storage system, and the like.

[0065] At block 220, the obtained input data may be stored in a buffer (e.g., the look- ahead buffer 120 of FIG. 1). The input data in the buffer may be compared to data in a history buffer (e.g., the history buffer 125 of FIG. 1) to determine whether a substring of the input data in the buffer matches at least a portion of the data in the history buffer. For example, a symbol-wise comparison between the symbols in the buffer and the symbols in the history buffer may be performed to determine whether a substring match exists between the symbols in the buffer and the symbols in the history buffer. Additional details associated with storing the input data and obtaining a substring match may be further discussed relative to the method 300 of FIG. 3.

[0066] In some instances, the substrings in the buffer may be compared to the keys and / or suffixes in the history buffer to identify one or more matches. In instances in which the matched substrings satisfy a threshold length, the matched substrings may be retained. For example, in instances in which a length of the matched substring is greater than zero and less than or equal to a maximum substring length less the DMW window length, the matched substring may be retained.

[0067] At block 230, the substring match determined at block 220 may be extended by comparing additional symbols in the buffer to additional symbols in the history buffer. In some instances, the additional symbols may be adjacent to the symbols already matched in the buffer and / or the history buffer. For example, a substring in the buffer (e.g., having at least one symbol) having matched a substring in the history buffer, may be extended by comparing an adjacent symbol in the buffer to an adjacent symbol in the history buffer. Additional details associated with extending the substring match may be further discussed relative to the method 400 of FIG. 4.

[0068] At block 240, a compression operation may be performed on the substring match identified in the previous blocks. In some instances, some symbols from the buffer may be emitted as literals and compressed separately from the substring match. Additional details associated with performing the compression operation may be further discussed relative to the method 500 of FIG. 5.

[0069] At block 250, the buffer and / or the history buffer may be updated. The buffer may include removing the symbols associated with the substring match (which may include symbols that may be emitted as literals) from the buffer and adjusting the remaining symbols in the buffer to be used for matching at another time. Alternatively, or additionally, the history buffer may be updated to include the symbols removed from the buffer. For example, the symbols included in the substring match may be removed from the buffer and may be included in the history buffer to be used in subsequent matching operations. In instances in which the history buffer is a static buffer (e.g., a static dictionary), the symbols of the substring match removed from the buffer may not be added to the history buffer. Additional details associated with updating the buffer may be further discussed relative to the method 500 of FIG. 5.

[0070] At block 260, compressed data obtained from the compression operation on the substring match (and associated literals) may be generated and / or output. For example, the compressed data may be stored in a data storage, stored in a buffer, and / or transmitted to another system or device for subsequent handling.

[0071] Modifications, additions, or omissions may be made to the method 200 without departing from the scope of the present disclosure. For example, the designations of different elements in the manner described is meant to help explain concepts described herein and is not limiting. Further, the method 200 may include any number of other elements or may be implemented within other systems or contexts than those described.

[0072] FIG. 3 illustrates a flowchart of an example method 300 of an aspect of the method 200 of FIG. 2, in accordance with at least one embodiment of the present disclosure. In particular, the method 300 may detail the storing of the input data in a look-ahead buffer and obtaining a substring match between symbols in the look-ahead buffer and symbols in a history buffer.

[0073] At block 302, input data may be obtained and may be stored in the look-ahead buffer. The input data may include one or more symbols, that may be compared and / or matched with symbols in the history buffer, as described herein.

[0074] At block 304, a delayed match window (DMW) having a DMW length may be obtained for use in the data compression operation. In some instances, the DMW length may be input by a user of the lazy matching algorithm. Alternatively, or additionally, the DMW length may correspond to a data type of the input data, a number of symbols included in the input data, and / or other factors associated with the input data. Alternatively, or additionally, the DMW length may be determined based on a desired compression ratio, as adjusting the DMW length may cause variations to the compression ratio.

[0075] At block 306, a pointer in the look-ahead buffer may be set to be equal to the DMW length. The pointer may be used in the lazy matching algorithm to point to a symbol to be matched with symbols in the history buffer (e.g., a pointer symbol), and in response to being set equal to the DMW length, the pointer may skip a number of symbol in the look- ahead buffer based on the DMW length. For example, in instances in which the DMW length is 2, the pointer may skip the first symbol and the second symbol in the look-ahead buffer and may point to the third symbol in the look-ahead buffer (e.g., the pointer symbol). In some instances, the pointer may also be referred to as an iteration (e.g., a DMW iteration) of the lazy matching algorithm. For example, the pointer may first point to the third symbol on a first iteration (described in the previous example), and on a second iteration, the pointer may be updated to point to the second symbol, and so forth. As such, the DMW iteration (e.g., i) may be associated with a current pointer value.

[0076] At block 308, a match between the pointer symbol and symbols in the history buffer may be obtained. For example, a substring in the look-ahead buffer (e.g., the substring starting with the pointer symbol) may be searched in the history buffer. In some instances, to be considered a valid substring match, the substring match may satisfy a threshold length (e.g., length ‘ Z’). The threshold length of the substring match may be greater than a minimum substring match length (e.g., Lminimum) minus the iteration (e.g., the current pointer value). In some instances, the search for a match between the look-ahead buffer and the history buffer may be performed using a parallel look-up in the hardware and / or using a rolling hash algorithm in a software and / or hardware implementation.

[0077] In instances in which the search by a rolling hash is utilized, an incrementally increased length of substring matches may be located. In such cases (e.g., when increased length substring matches are found), a maximum length of a substring in the look-ahead buffer that matches with a substring in the history buffer in zthDMW iteration may beupdated. The maximum length for an zthDMW iteration may be represented as lmax,i and the substring matches whose lengths fall in the range lmax,iI > lmax,i ~ * may be retained. In instances in which identified substring matches have a length that falls outside the above range, the identified substring matches may be dropped.

[0078] At block 310, the identified substring matches may be included in a set of matches, based on the DMW iteration. For example, a set of substring matches of length 7’ in the zthDMW iteration may be represented as S£;£. Alternatively, or additionally, the substrings may be kept in other sets, where the position of the matched substring may be retained in the individual sets of substring matches (in the history buffer).

[0079] As such, an ordered set (S£) may be generated and in some instances, the substring matches in the ordered set may be arranged in decreasing order based on the length of the substring matches. For example, the ordered set may be represented by St= Si j, lmnxj > I > lmnxj — tj. In some embodiments, the ordered set may include, but not limited to, a list, a hash table, and / or a tree data structure, all of which may be stored in the memory of a CPU and / or a microcontroller. Alternatively, or additionally, in instances in which the algorithm is implemented in hardware, one or more hardware elements may be used as a container to store the ordered set(s).

[0080] In instances in which a static dictionary is used, the symbols at the beginning of the look-ahead buffer may be compared to the prefix of the keys in the history buffer. For example, the prefixes of the keys and / or suffixes associated with the keys that were previously retained may be used in the comparison with the symbols in the look-ahead buffer. In some instances, the keys with the longest match (e.g., the most matching symbols in the match) may be selected. In instances in which more than one key is identified as a longest match, an additional consideration may determine which of the longest matched keys are selected. For example, a compression ratio associated with the keys may determinewhich of the longest match keys may be selected. In another example, a particular key that may have been more recently selected as the longest match may be selected as the longest match relative to another key that may not have been selected as recently as a longest match.

[0081] In some instances, no substring match may be found between the look-ahead buffer and the history buffer or the substring match may fail to satisfy the minimum substring match length. In one instance in which no substring match is found, a first symbol may be transmitted from the beginning of the look-ahead buffer as a literal to a subsequent coding state, such as a Huffman code. Alternatively, or additionally, in a second instance in which no substring match is found, the ‘z’ symbols from the beginning of the look-ahead buffer may be transmitted as literals to a subsequent coding state, such as a Huffman code. Subsequent to the symbol(s) being emitted as literals to a coding state, the look-ahead buffer may be adjusted to slide past the first symbol (e.g., the first instance) or the ‘z’ symbols (e.g., the second instance) and the first symbol or the ‘z’ symbols may be included as part of the history buffer, as further described in method 500.

[0082] Modifications, additions, or omissions may be made to the method 300 without departing from the scope of the present disclosure. For example, the designations of different elements in the manner described is meant to help explain concepts described herein and is not limiting. Further, the method 300 may include any number of other elements or may be implemented within other systems or contexts than those described.

[0083] FIG. 4 illustrates a flowchart of an example method 400 of an aspect of the method 200 of FIG. 2, in accordance with at least one embodiment of the present disclosure. In particular, the method 400 may detail iterating through matches in the set of matches to determine if a substring match may be extended.

[0084] At block 402, the pointer in the look-ahead buffer may be adjusted for use in a subsequent iteration of obtaining matches between symbols in the look-ahead buffer andsymbols in the history buffer. For example, a first DMW iteration may have been performed with the pointer pointing to a third symbol in the look-ahead buffer (e.g., the pointer equal to the DMW length) and the pointer may be adjusted for a second iteration to be equal to one less that the DMW length, such that the second DMW iteration may be performed with the pointer pointing to a second symbol in the look-ahead buffer (e.g., the pointer equal to one less than the DMW length). In such instances, the pointer may cause symbols to be skipped in the look-ahead buffer, which may be one less symbol skipped than a first iteration of finding substring matches between the symbols in the look-ahead buffer and the symbols in the history buffer. As described herein, a full scan of the history buffer and / or a dictionary search may be performed once, irrespective of the number of DMW iterations. Alternatively, or additionally, subsequent iterations may look for an expansion of the identified match, as described herein.

[0085] At block 404, a substring match from a set of matches may be obtained. The set of matches may be obtained from another portion of the lazy matching algorithm, such as block 310 described in FIG. 3. In instances in which the set of matches is empty, the algorithm may proceed to check the previous iteration set of matches (e.g., the current DMW iteration plus one) to determine whether one or more symbols may be emitted as literals to a subsequent coding state and thereafter, to be removed from the look-ahead buffer and / or added to the history buffer. In instances in which a static dictionary is used, symbols may not be added to the history buffer (e.g., as the history buffer is a static dictionary). In instances in which the substring match is obtained, the method 400 may continue to block 406.

[0086] At block 406, an adjacent look-ahead buffer symbol disposed adjacent to the symbols in the substring match may be compared to an adjacent history buffer symbol, which may be utilized to extend the substring match. Extending the substring match mayrefer to adding at least one or more symbols to the symbols included in the substring match, which may increase the count of symbols included within the substring match.

[0087] At block 408, a determination as to whether the substring match was able to be extended may be performed. For example, if the adjacent look-ahead buffer symbol matches the adjacent history buffer symbol, the substring match may be extended to an extended substring match. In instances in which the substring match was extended, the method may continue at block 410. In instances in which the substring match was not extended, the method may continue at block 412.

[0088] At block 410, the extended substring match may be added to a new set of substring matches where the length of the matches in the new set of substring matches may be at least one symbol longer than the set of matches (e.g., that included the substring match, that had not yet been extended). In some instances, once the new set of substring matches is generated, the set of matches may be discarded as the new set of substring matches may include the extended substring match that includes more symbols than the substring match in the set of matches. After generating the new set of substring matches, the method 400 may continue at block 412.

[0089] At block 412, the process of extending the substring matches may be repeated. For example, in instances in which a first substring match was determined to not be extended in block 408, a second substring match in the set of matches may be obtained and determined if the second substring match may be extended. In another example, in instances in which a first substring match was determined to be extended in block 408, a second substring match in the set of matches may be obtained and determined if the second substring match may be extended. The process may be repeated for some or all of the substring matches included in the set of matches.

[0090] Modifications, additions, or omissions may be made to the method 400 without departing from the scope of the present disclosure. For example, the designations of different elements in the manner described is meant to help explain concepts described herein and is not limiting. Further, the method 400 may include any number of other elements or may be implemented within other systems or contexts than those described.

[0091] FIG. 5 illustrates a flowchart of an example method 500 of an aspect of the method 200 of FIG. 2, in accordance with at least one embodiment of the present disclosure. In particular, the method 500 may detail the compression operation on the substring match symbols and / or literal symbols in the look-ahead buffer and the subsequent updates to the look-ahead buffer and / or the history buffer.

[0092] At block 502, a substring match may be obtained from the set of matches. The substring match may be the substring match from the set of matches and / or may be the extended substring match from the new set of matches, both of which as described herein. In some instances, the obtained substring match may include a longest length relative to other substring matches in the set of matches and / or in other sets of matches.

[0093] At block 504, a length of the obtained substring match may be compared to a minimum length to determine the substring match length satisfies the minimum length. The minimum length may be based on a user input, based on characteristics of a processing device performing the lazy matching algorithm, and / or other factors. In instances in which the substring match length satisfies the minimum length, the method 500 may continue to block 506. In instances in which the substring match length fails to satisfy the minimum length, the method 500 may continue to block 508.

[0094] At block 506, the symbols in the substring match and / or preceding symbols may be emitted to a compression module, such as a coding state (e.g., Huffman code), to be compressed by a compression operation. In instances in which the substring match doesnot include one or more symbols at the beginning of the look-ahead buffer (e.g., the substring match begins at a later symbol, such as a second symbol or a third symbol in the look-ahead buffer), the preceding symbols (or extra literals) may be individually emitted as literals to the compression module along with the symbols in the substring match, all of which may be compressed by the compression module.

[0095] At block 508, the symbols in the substring match may be individually emitted as literals to the compression module and may be compressed by the compression operation.

[0096] At block 510, the pointer in the look-ahead buffer may be adjusted in response to the symbols being emitted therefrom (e.g., as the substring match, as literals and the substring match, and / or as literals). For example, the look-ahead buffer may be adjusted such that the symbols emitted for compression may be removed and a new symbol may be disposed at the start of the look-ahead buffer. Correspondingly, the pointer may be adjusted relative to the new symbol and based on the DMW length. In instances in which there are no more symbols in the look-ahead buffer following emission of symbols therefrom (or the number of symbols remaining in the look-ahead buffer may not satisfy the minimum length for a substring match), the lazy matching algorithm may be completed and be halted.

[0097] At block 512, the symbols emitted from the look-ahead buffer as part of the compression operation may be added to the history buffer. The emitted symbols may be used in subsequent matching operations between symbols in the look-ahead buffer and symbols in the history buffer to obtain a substring match as part of the lazy matching algorithm. In some instances, the emitted symbols may be added to the beginning of the history buffer (e.g., the first symbols compared for a match with symbols from the look- ahead buffer). Alternatively, or additionally, the history buffer may be a static buffer (e.g., a dictionary-like buffer) that may not update using the emitted symbols and the emittedsymbols may be discarded. In instances in which the history buffer is static, the symbols and / or keys in the history buffer may be predetermined, as described herein.

[0098] Modifications, additions, or omissions may be made to the method 500 without departing from the scope of the present disclosure. For example, the designations of different elements in the manner described is meant to help explain concepts described herein and is not limiting. Further, the method 500 may include any number of other elements or may be implemented within other systems or contexts than those described.

[0099] FIG. 6 illustrates a block diagram 600 of an example look-ahead buffer 605 and an example history buffer 610 that may be used in data compression using a lazy matching algorithm, in accordance with at least one embodiment of the present disclosure. As illustrated, the look-ahead buffer 605 may include multiple symbols and the history buffer 610 may include multiple symbols, and identical symbols (e.g., symbols that may match) may be illustrated as having an identical pattern. For example, the leftmost symbol in the look-ahead buffer 605 (e.g., the first symbol in the look-ahead buffer 605), illustrated with a first pattern, may be identical to the leftmost symbol in the history buffer 610 (e.g., the last symbol in the history buffer 610), illustrated with the first pattern.

[0100] As the lazy matching algorithm identifies one or more substring matches between the symbols in the look-ahead buffer 605 and the history buffer 610, an initial position 624 in the history buffer 610 of a first symbol in the substring match may be obtained. The initial position 624 may be reflective of a count of the number of symbols from the first symbol in the history buffer 610 to the first symbol in the substring match in the history buffer 610.

[0101] In instances in which the lazy matching algorithm attempts to extend the substring match between the symbols in the look-ahead buffer 605 and the symbols in the history buffer 610, an updated position 626 may be obtained. The updated position 626may be reflective of a count of an additional symbol and the number of symbols from the first symbol in the history buffer 610 to the first symbol in the substring match in the history buffer 610 (e.g., the additional symbol in addition to the initial position 624).

[0102] In instances in which one or more matches are obtained between a substring in the look-ahead buffer 605 and one or more substrings in the history buffer 610, a position and / or a length of the matched substring may be obtained. For example, a first substring match 628 in the history buffer 610 may be determined to match the substring in the look- ahead buffer 605, where the first substring match 628 may be a first distance from the substring in the look-ahead buffer 605 and the first substring match 628 may include a first distance (e.g., measured from a first symbol in the history buffer 610) and / or a length associated therewith. Alternatively, or additionally, a second substring match 630 in the history buffer 610 may be determined to match the substring in the look-ahead buffer 605, where the second substring match 630 may be a second distance from the substring in the look-ahead buffer 605 and the second substring match 630 may include a second distance and / or the same length as the first substring match 628.

[0103] In instances in which the lazy matching algorithm attempts to extend the substring match (e.g., by comparing an input adjacent symbol in the look-ahead buffer 605 to a history adjacent symbol in the history buffer 610), such as the second substring match 630, and determines the second substring match 630 may be extended (e.g., the input adjacent symbol matches the history adjacent symbol), the extended substring match 632 may be the second distance from the substring in the look-ahead buffer 605 and the extended substring match 632 may include the second distance and / or a second length, where the second length may be at least one symbol more than the second substring match 630 (and / or at least one symbol longer than the first substring match 628).

[0104] Modifications, additions, or omissions may be made to the block diagram 600 without departing from the scope of the present disclosure. For example, the designations of different elements in the manner described is meant to help explain concepts described herein and is not limiting. Further, the block diagram 600 may include any number of other elements or may be implemented within other systems or contexts than those described.

[0105] FIG. 7 illustrates a flowchart of an example method 700 of data compression using a static dictionary and a lazy matching algorithm, in accordance with at least one embodiment of the present disclosure. At block 702, input data may be stored in a first buffer. The input data may include a first set of symbols.

[0106] At block 704, one or more keys may be obtained from a second buffer. The one or more keys may individually include buffer substrings. In some instances, the second buffer may be a static dictionary. The buffer substrings may be predetermined symbols stored in the second buffer.

[0107] In some instances, the buffer substrings may be one or more suffixes associated with each of the one or more keys. In some instances, the one or more suffixes may be arranged in a suffix tree. Alternatively, or additionally, the one or more suffixes may be arranged in a suffix array. In some instances, a length of the one or more suffixes may be less than a length of an associated key of the one or more keys.

[0108] At block 706, a portion of the first set of symbols may be compared to the buffer substrings to identify one or more substring matches. A portion of the first set of symbols may be compared to the one or more suffixes to identify the one or more substring matches. In some instances, the one or more substring matches may each include a substring length that may be defined by an amount of the substring match symbols in the one or more substring matches. In some instances, the substring length may be greater than or equal toa minimum length. The minimum length may be a combination of the substring length and a delayed match window length.

[0109] At block 708, a longest substring match of the one or more substring matches may be identified. The longest substring match may include substring match symbols. In some instances, a first length of a first longest substring match and a second length of a second longest substring match may be equal. Further, in some instances, the first longest substring match may have a better compressibility ratio relative to the second longest substring match. In such instances, the first longest substring match may be selected as the longest substring match.

[0110] At block 710, a key of the one or more keys associated with the longest substring match may be retained.

[0111] At block 712, a compression operation may be performed to the substring match symbols. At block 714, the portion of the first set of symbols may be removed from the first buffer.

[0112] Modifications, additions, or omissions may be made to the method 700 without departing from the scope of the present disclosure. For example, the compressed data may be output, such as to an output buffer, where the compressed data may include at least a portion of the input data that may have been compressed by the compression operation.

[0113] In another example, the designations of different elements in the manner described is meant to help explain concepts described herein and is not limiting. Further, the method 700 may include any number of other elements or may be implemented within other systems or contexts than those described.

[0114] FIG. 8 illustrates an example computing device 800 within which a set of instructions, for causing the machine to perform any one or more of the methods discussed herein, may be executed. The computing device 800 may include a mobile phone, a smartphone, a netbook computer, a rackmount server, a router computer, a server computer, a personal computer, a mainframe computer, a laptop computer, a tablet computer, a desktop computer, or any computing device with at least one processor, etc., within which a set of instructions, for causing the machine to perform any one or more of the methods discussed herein, may be executed. In alternative implementations, the machine may be connected (e.g., networked) to other machines in a LAN, an intranet, an extranet, or the Internet. The machine may operate in the capacity of a server machine in client-server network environment. The machine may include a personal computer (PC), a set-top box (STB), a server, a network router, switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine. Further, while only a single machine is illustrated, the term “machine” may also include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods discussed herein.

[0115] The computing device 800 includes a processing device 802 (e.g., a processor), a main memory 804 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM)), a static memory 806 (e.g., flash memory, static random access memory (SRAM)) and a data storage device 816, which communicate with each other via a bus 808.

[0116] The processing device 802 represents one or more general-purpose processing devices such as a microprocessor, central processing unit, or the like. More particularly, the processing device 802 may include a complex instruction set computing (CISC) microprocessor, reduced instruction set computing (RISC) microprocessor, very long instruction word (VLIW) microprocessor, or a processor implementing other instruction sets or processors implementing a combination of instruction sets. The processing device 802 may also include one or more special-purpose processing devices such as anapplication specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), network processor, or the like. The processing device 802 is configured to execute instructions 826 for performing the operations and steps discussed herein.

[0117] The computing device 800 may further include a network interface device 822 which may communicate with a network 818. The computing device 800 also may include a display device 810 (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), an alphanumeric input device 812 (e.g., a keyboard), a cursor control device 814 (e.g., a mouse) and a signal generation device 820 (e.g., a speaker). In at least one implementation, the display device 810, the alphanumeric input device 812, and the cursor control device 814 may be combined into a single component or device (e.g., an LCD touch screen).

[0118] The data storage device 816 may include a computer-readable storage medium 824 on which is stored one or more sets of instructions 826 embodying any one or more of the methods or functions described herein. The instructions 826 may also reside, completely or at least partially, within the main memory 804 and / or within the processing device 802 during execution thereof by the computing device 800, the main memory 804 and the processing device 802 also constituting computer-readable media. The instructions may further be transmitted or received over a network 818 via the network interface device 822.

[0119] While the computer-readable storage medium 824 is shown in an example implementation to be a single medium, the term “computer-readable storage medium” may include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) that store the one or more sets of instructions. The term “computer-readable storage medium” may also include any medium that is capable of storing, encoding or carrying a set of instructions for execution by the machine and thatcause the machine to perform any one or more of the methods of the present disclosure. The term “computer-readable storage medium” may accordingly be taken to include, but not be limited to, solid-state memories, optical media and magnetic media.

[0120] Terms used in the present disclosure and especially in the appended claims (e.g., bodies of the appended claims) are generally intended as “open terms” (e.g., the term “including” should be interpreted as “including, but not limited to ”).

[0121] Additionally, if a specific number of an introduced claim recitation is intended, such an intent will be explicitly recited in the claim, and in the absence of such recitation no such intent is present. For example, as an aid to understanding, the following appended claims may contain usage of the introductory phrases “at least one” and “one or more” to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles “a” or “an” limits any particular claim containing such introduced claim recitation to implementations containing only one such recitation, even when the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “an” (e.g., “a” and / or “an” should be interpreted to mean “at least one” or “one or more”); the same holds true for the use of definite articles used to introduce claim recitations.

[0122] In addition, even if a specific number of an introduced claim recitation is expressly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number (e.g., the bare recitation of “two recitations,” without other modifiers, means at least two recitations, or two or more recitations). Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, etc.” or “one or more of A, B, and C, etc.” is used, in general such a construction is intended to include A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together, etc.

[0123] Further, any disjunctive word or phrase preceding two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both of the terms. For example, the phrase “A or B” should be understood to include the possibilities of “A” or “B” or “A and B.”

[0124] All examples and conditional language recited in the present disclosure are intended for pedagogical objects to aid the reader in understanding the present disclosure and the concepts contributed by the inventor to furthering the art, and are to be construed as being without limitation to such specifically recited examples and conditions. Although implementations of the present disclosure have been described in detail, various changes, substitutions, and alterations could be made hereto without departing from the spirit and scope of the present disclosure.

Claims

CLAIMSWHAT IS CLAIMED IS:

1. A method, comprising: storing input data in a first buffer, the input data comprising a first set of symbols; obtaining one or more keys from a second buffer, the one or more keys individually comprising buffer substrings; comparing a portion of the first set of symbols to the buffer substrings to identify one or more substring matches; identifying a longest substring match of the one or more substring matches, the longest substring match comprising substring match symbols; retaining a key of the one or more keys associated with the longest substring match; performing a compression operation to the substring match symbols; and removing the portion of the first set of symbols from the first buffer.

2. The method of claim 1, wherein the buffer substrings are one or more suffixes associated with each of the one or more keys, and a portion of the first set of symbols are compared to the one or more suffixes to identify the one or more substring matches.

3. The method of claim 2, wherein the one or more suffixes are arranged in a suffix tree.

4. The method of claim 2, wherein the one or more suffixes are arranged in a suffix array.

5. The method of claim 2, wherein a length of the one or more suffixes is less than a length of an associated key of the one or more keys.

6. The method of claim 1, further comprising outputting compressed data, the compressed data comprising at least one symbol of the input data compressed by the compression operation.

7. The method of claim 4, wherein the buffer substrings are predetermined symbols stored in the second buffer.

8. The method of claim 1, wherein the second buffer is a static dictionary.

9. The method of claim 1, wherein the one or more substring matches each comprise a substring length defined by an amount of the substring match symbols in the one or more substring matches.

10. The method of claim 8, wherein the substring length is greater than or equal to a minimum length.

11. The method of claim 9, wherein the minimum length is a combination of the substring length and a delayed match window length.

12. The method of claim 1, wherein a first length of a first longest substring match and a second length of a second longest substring match is equal and the first longest substringmatch has a better compressibility ratio relative to the second longest substring match, the first longest substring match is selected as the longest substring match.

13. A compression device, comprising: a first buffer; a second buffer; and a data compression module operable to: store input data in the first buffer, the input data comprising a first set of symbols; obtain one or more keys from a second buffer, the one or more keys individually comprising buffer substrings; compare a portion of the first set of symbols to the buffer substrings to identify one or more substring matches; identify a longest substring match of the one or more substring matches, the longest substring match comprising substring match symbols; retain a key of the one or more keys associated with the longest substring match; perform a compression operation to the substring match symbols; and remove the portion of the first set of symbols from the first buffer.

14. The compression device of claim 13, wherein the buffer substrings are one or more suffixes associated with each of the one or more keys, and a portion of the first set of symbols are compared to the one or more suffixes to identify one or more substring matches.

15. The compression device of claim 14, wherein the one or more suffixes are arranged in a suffix tree.

16. The compression device of claim 14, wherein the one or more suffixes are arranged in a suffix array.

17. The compression device of claim 14, wherein a length of the one or more suffixes is less than a length of an associated key of the one or more keys.

18. The compression device of claim 13, wherein the data compression module is further operable to output compressed data, the compressed data comprising at least a portion of the input data compressed by the compression operation.

19. The compression device of claim 13, wherein the buffer substrings are predetermined symbols stored in the second buffer.

20. The compression device of claim 13, wherein a first length of a first longest substring match and a second length of a second longest substring match is equal and the first longest substring match has a better compressibility ratio relative to the second longest substring match, the first longest substring match is selected as the longest substring match.

Citation Information

Patent Citations

  • Static dictionary-based compression hardware pipeline for data compression accelerator of a data processing unit

    US20200169268A1

  • Data structures and operations for searching, computing, and indexing in DNA-based data storage

    US20200357483A1