Data compression using static dictionary

CN122680686APending Publication Date: 2026-09-01MAXLINEAR INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480087272.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-11
Filing Date
2024-12-10
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

更大的惰性匹配窗口值可以提高压缩比,但代价是编码延迟和/或计算复杂度的增加

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122680686A_ABST
    Figure CN122680686A_ABST
Patent Text Reader

Abstract

A method includes storing input data in a first buffer, wherein the input data includes a first set of symbols. The method further includes retrieving one or more keys from a second buffer, wherein each of the one or more keys includes a substring of the buffer. The method further includes comparing a portion of the first set of symbols with the buffer substrings to identify one or more substring matches. The method further includes identifying the longest substring match among the one or more substring matches. The longest substring match includes the substring matching symbols. The method further includes retaining the keys from the one or more keys associated with the longest substring match. The method further includes performing a compression operation on the substring matching symbols. The method further includes removing the portion of the first set of symbols from the first buffer.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing related applications This U.S. patent application claims priority to U.S. Provisional Patent Application No. 63 / 608,813, filed December 11, 2023, entitled “LAZY MATCHINGALGORITHM FOR DATA COMPRESSION,” the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0002] This disclosure relates to lossless data compression, and more specifically, to lossless data compression based on static dictionaries. Background Technology

[0003] Unless otherwise stated herein, the materials described herein do not constitute prior art for the claims of this application, and are not acknowledged as prior art by virtue of their inclusion in this section.

[0004] Data compression can be lossy or lossless. In lossless compression, a data compression algorithm (or "compression algorithm") reduces the size of the data by identifying and removing redundancy. Furthermore, the information content contained within the data is not removed, so when decompression is applied to compressed data, the decompressed data can be restored to its original state. Many data transform accelerators (DTAs), computational storage devices (CSDs), data processing units (DPUs), network interface controllers (NICs), central processing units (CPUs), and field-programmable gate arrays (FPGAs) used in storage or encryption devices employ lossless compression methods.

[0005] Many lossless data compression techniques use dictionary-based compression methods that extract substrings from the original data string. Substrings can be of variable or fixed length, and they can be used to index a dictionary that maps substrings to tokens. Data reduction is achieved when a token can be represented with fewer bits relative to its mapped substring. Much text is not a random sequence of symbols, and substrings of symbols appearing in the input data are more likely to reappear. In cases where tokens in the dictionary are references to previous occurrences of a substring, tokens can be represented with fewer bits than would be needed to represent the original substring, thus achieving data reduction.

[0006] Many widely used dictionary-based compression methods adaptively construct a dictionary using a sliding window on both processed and unprocessed symbols. For example, the LZ77 compression algorithm, which may be used in LZ4, LZS, Deflate, GZIP, ZLIB, and / or XP10, can utilize an adaptive dictionary as described above. The compression algorithm can maintain a history buffer containing the input data processed by the algorithm. This history buffer can operate as an adaptively constructed dictionary. For example, to process new input data contained in the lookahead buffer, a pointer can be moved back along the history buffer until a match (e.g., a matching symbol) is found that matches the first symbol of the new input data contained in the lookahead buffer. Once a match is found, the next symbol in the lookahead buffer is compared to the next symbol in the history buffer, which may be adjacent to the matching symbol, to determine if an additional match can be obtained. The matching process continues, comparing subsequent symbols in the lookahead buffer with consecutive symbols in the history buffer until a match is found. In this way, the compression algorithm searches the entire history buffer to determine the longest substring match in the lookahead buffer.

[0007] Once the longest match is found, the compression algorithm encodes it using <distance, length> pairs. Distance can be the distance between the start of the longest matching substring in the history buffer and the start of the lookahead buffer. Length can be the length of the longest substring match. Different definitions of distance can be considered, such as representing the position of the substring in the history buffer. Once a substring in the lookahead buffer is matched, it slides to the front of the history buffer. If the history buffer is full, the oldest data at the tail of the history buffer can be discarded. In the absence of a match, the symbol in the lookahead buffer can be emitted as a literal character token. Some dictionary-based compression algorithms (e.g., XP10) may create a cache of recently used <distance, length> pairs in a lookup table, and the compression algorithm may use an index in the lookup table associated with the <distance, length> pair instead of directly using the <distance, length> pair as a token to improve the compression ratio (often called Move to Front (MTF) encoding).

[0008] Some dictionary-based methods can improve the compression ratio by using lazy evaluation techniques, where the compression ratio can be the ratio of the input data size to the compressed data size. After finding the longest substring match using the substring at the beginning of the lookahead buffer, the compression algorithm can consider the longest substring match, skip the first symbol of the lookahead buffer, and start the matching process with the substring in the history buffer from the second symbol of the lookahead buffer.

[0009] If a longer match is found, the compression algorithm emits the first symbol in the lookahead buffer as a literal character marker, and subsequent substring matches can be emitted as <distance, length> markers. Otherwise, the compression algorithm emits the first longest substring match as a <distance, length> marker. Therefore, the compression algorithm can consider multiple candidate substrings for the longest substring match, skipping the first few symbols sequentially from the beginning of the lookahead buffer (not just one symbol), and can select the longest substring match from the candidate substrings. This type of compression algorithm is often called lazy evaluation or lazy matching. Alternatively, the number of symbols that the substring matching process can be delayed from the beginning of the lookahead buffer (to select the substring for matching) can be called the lazy matching window or delayed matching window. For example, a delayed matching window value of two can indicate that the longest substring match can be considered among the following candidate substring matches: i) from the beginning of the lookahead buffer (including the first symbol), ii) skipping the first symbol, and iii) also skipping the first two symbols. A larger lazy matching window value can improve the compression ratio, but at the cost of increased encoding latency and / or computational complexity.

[0010] When compressing small data blocks, there may not be enough historical data or a sufficiently padded dictionary to use some of the compression algorithms described. In such cases, some compression algorithms can use a static pre-shared dictionary. Alternatively, some compression algorithms use a static pre-initialized dictionary and / or a sliding window, where the pre-initialized dictionary can be referenced at any time while scanning the input data to find substring matches.

[0011] Once the input data is converted into token sets (e.g., <distance, length> pairs, dictionary indices, literal characters, etc.), these tokens can be encoded using variable-length encoding algorithms, such as Huffman coding. Huffman coding can use a fixed codebook and / or a codebook built based on the frequency of occurrence of different tokens encountered (also known as backtracking or dynamic coding), which can be generated during compression. Multiple codebooks can be used for the same alphabet within the same compressed data block (as in context modeling). Alternatively, or alternatively, an asymmetric numeral system (ANS) can be used instead of Huffman coding, as might be used in Zstandard.

[0012] The subject matter claimed in this disclosure is not limited to embodiments that address any drawbacks, nor is it limited to embodiments that operate only in the environment described above. Rather, this background is provided merely to illustrate an example technical field in which some of the embodiments described in this disclosure can be practiced. Summary of the Invention

[0013] In an example embodiment, a method may include storing input data in a first buffer. The input data may include a first set of symbols. The method may further include retrieving one or more keys from a second buffer. Each of the one or more keys may comprise a substring of the buffer. The method may further include comparing a portion of the first set of symbols with the buffer substring to identify one or more substring matches. The method may further include identifying the longest substring match among the one or more substring matches. The longest substring match may include the substring matching symbols. The method may further include retaining the keys from the one or more keys associated with the longest substring match. The method may further include performing a compression operation on the substring matching symbols. The method may further include removing the portion of the first set of symbols from the first buffer.

[0014] In another embodiment, the compression apparatus may include a first buffer, a second buffer, and a data compression module. The data compression module is operable to store input data in the first buffer. The input data may include a first set of symbols. The data compression module is also operable to retrieve one or more keys from the second buffer. Each of the one or more keys may include a buffer substring. The data compression module is also operable to compare a portion of the first set of symbols with the buffer substring to identify one or more substring matches. The data compression module is also operable to identify the longest substring match among the one or more substring matches. The longest substring match may include a substring matching symbol. The data compression module is also operable to retain the key from the one or more keys associated with the longest substring match. The data compression module is also operable to perform a compression operation on the substring matching symbol. The data compression module is also operable to remove the portion of the first set of symbols from the first buffer.

[0015] The objectives and advantages of the embodiments will be realized and achieved, at least by means of the elements, features and combinations particularly pointed out in the claims.

[0016] The preceding general description and the following detailed description are given as examples and are illustrative, and do not constitute a limitation on the claimed invention. Attached Figure Description

[0017] Exemplary embodiments will be described and explained in more detail with reference to the accompanying drawings, wherein: Figure 1 A block diagram of an exemplary system for data compression using a lazy matching algorithm is shown; Figure 2 A flowchart illustrating an exemplary method for data compression using a lazy matching algorithm is shown; Figure 3 It shows Figure 2 A flowchart illustrating an exemplary aspect of the method; Figure 4 It shows Figure 2 A flowchart illustrating an exemplary aspect of the method; Figure 5 It shows Figure 2 A flowchart illustrating an exemplary aspect of the method; Figure 6 A block diagram of an exemplary look-ahead buffer and an exemplary history buffer that can be used in data compression using a lazy matching algorithm is shown; Figure 7 A flowchart illustrating an exemplary method for data compression using a lazy matching algorithm is shown; Figure 8 An exemplary computing device is shown; and Figure 9 An exemplary suffix tree is shown. Detailed Implementation

[0018] Various implementations exist for compressing data, each potentially differing in its compression ratio of the input data, where the compression ratio is the ratio of the input data size to the compressed data size. In many cases, improvements in compression ratio within a particular compression operation may come at the cost of increased computational complexity and / or increased latency. Some previous methods may have used lazy matching data compression operations, which can be computationally expensive and / or lead to increased latency in the compression operation.

[0019] Lazy matching may involve multiple searches in the history buffer or dictionary (or just the "history buffer"), which can make it computationally expensive and / or increase latency in the compression operation. Therefore, when implemented in software on the CPU, it may require more CPU cycles. When lazy matching is implemented in hardware (e.g., reconfigurable hardware such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits)), the lazy matching operation may use more clock cycles and / or more hardware resources (e.g., circuitry), resulting in higher power consumption.

[0020] The aspects of this disclosure address these and other limitations by implementing lazy matching algorithms in hardware or software. Lazy matching algorithms can optimize computational complexity, reduce the consumption of computational resources (e.g., CPU cycles on a processor or clock cycle), and / or reduce power consumption when processing data in hardware-based implementations. The aspects of this disclosure can also improve the throughput and / or latency of data compression operations. In such implementations, lazy matching algorithms can compare a string from a lookahead buffer with a history buffer to obtain a substring match. Lazy matching algorithms can also extend the substring match to include additional symbols without performing an additional scan of the history buffer, thereby reducing the computational overhead and / or latency associated with lazy matching algorithms compared to other lazy matching algorithms and / or other data compression operations.

[0021] Figure 1 A block diagram of an exemplary system 100 for data compression using a lazy matching algorithm according to at least one embodiment of the present disclosure is shown. System 100 may include a computing device 110, a look-ahead buffer 120, and a history buffer 125. The computing device 110 may include a data compression module 115 and is operable to output compressed data 130.

[0022] Computing device 110 may be any computing device operable to perform at least data compression operations and acquire input data 102 and generate and / or transmit compressed data 130. Computing device 110 may be communicatively coupled to lookahead buffer 120 and / or history buffer 125 such that computing device 110 can instruct data stored in lookahead buffer 120 and / or history buffer 125 (e.g., input data 102), instruct operations related to the data stored in lookahead buffer 120 and / or history buffer 125, and / or generate compressed data 130 using lookahead buffer 120 and / or history buffer 125, for example, via the lazy matching algorithm described herein.

[0023] The lookahead buffer 120 and the history buffer 125 may be storage devices operable to store at least the input data 102. For example, the input data 102 may be acquired by the computing device 110 and directed to be stored in the lookahead buffer 120. The history buffer 125 may contain a dictionary that can be used to compare with the input data 102 within the lookahead buffer 120.

[0024] The history buffer 125 can be a static dictionary that may include one or more keys. In some instances, the keys in the history buffer can be substrings and / or values ​​that can be used to match symbols in the lookahead buffer 120. In some instances, the history buffer 125 can be assembled based on previously processed data (e.g., based on the frequency of substrings that may exist in the input data 102) and / or based on machine learning output from the dataset. In these and other embodiments, the symbols and / or keys in the history buffer 125 can be predetermined.

[0025] In some instances, multiple history buffers can be constructed, where the history buffers can differ based on the types of data they might use. For example, a first history buffer can be constructed and used with a first type of data, a second history buffer can be constructed and used with a second type of data, and so on. In these and other embodiments, history buffer 125 can be arranged for efficient implementation. For example, history buffer 125 can be initialized before performing any compression operations.

[0026] In some instances, one or more suffixes can be constructed for each key in the history buffer 125. In some instances, the suffixes can be arranged in a data structure, such as a suffix tree and / or a suffix array. In these and other embodiments, the suffixes can be constructed based on the length of the key and / or the length of the delayed matching window (DMW) window, as described herein. For example, for a particular key X1X2X3X4…X n Furthermore, when the DMW length is three, the suffix can include at least X1X2X3X4…X n X2X3X4…X n X3X4…X n And X4…X n .

[0027] With suffixes arranged in a suffix tree, when history buffer 125 is a static dictionary, each edge of the suffix tree can be associated with a substring of a key in history buffer 125. Alternatively, each leaf node in the suffix tree can be labeled with the starting position of the suffix in the key and / or a specific key identifier. An example of suffix tree 900 is shown in... Figure 9 As shown in the diagram. In suffix tree 900, two keys are considered, where the first key is X1X2X3X4X5X6X7 and the second key is X1X2X3X4Y1Y2Y3. The DMW in suffix tree 900 can have a length of 3.

[0028] The substring aligned with the symbol after the DMW length in lookahead buffer 120 at the encoded position can be matched with the suffix. After matching, the matched substring (S) i ) can be retained in the lookahead buffer 120, by (l max – DWM length)≤ |S i | ≤ l max It means that l max It can be the maximum length of the matched substring, determined by l. max = arg max {|S i |} indicates that |S i | can be the i-th matching substring S i The length.

[0029] For the matching substring S above i The associated suffixes, including suffixes with a starting position greater than 1 (e.g., all suffixes in suffix tree 900 except X1X2X3X4X5X6X7 and X1X2X3X4Y1Y2Y3), can perform two operations. The first operation may include comparing the prefix of each matching suffix of the key with the symbol preceding the encoded position in lookahead buffer 120. The second operation may include retaining the specific key if a match is found in the first operation, or discarding the specific key if no match is found.

[0030] In the example, the lookahead buffer 120 includes symbols S1, S2, ... S 10 If S1 is located at the beginning of lookahead buffer 120 and the DMW length is 3, the encoding position can be at S4. If the matching suffix is ​​the third position of the second key (e.g., X3X4Y1Y2Y3), the prefix of the second key (e.g., X1X2) can be compared with the two preceding symbols (e.g., S2 and S3) in lookahead buffer 120. If the prefix matches the symbol, the key can be retained. If the prefix does not match the symbol, the key can be discarded. If the second key is retained, S1 can be emitted as a literal character, and symbols S2 to S7 can be represented by the second key, which can be emitted in the next encoding stage.

[0031] In another example, using the same lookahead buffer 120 and the same DMW length as the previous example, if the matching suffix is ​​the first position of the second key (e.g., X1X2X3X4Y1Y2Y3), there might be no prefix to compare with the extra symbols in lookahead buffer 120. In this case, and if the second key is preserved, S1, S2, and S3 can be emitted as literal characters, and symbols S4 through S... 10 It can be represented by a second key, which can be issued in the next encoding stage.

[0032] In these and other embodiments, for all possible key matches that may be retained, the symbol substring (matched symbol) in lookahead buffer 120 can be replaced with the key match of the longest length. Unmatched symbols preceding the matched symbols in lookahead buffer 120 can be emitted as literal characters. In the case of a tie in the longest length of matched keys, a selection criterion can be used to determine which of more than one matched key can be selected. For example, the first key that might contribute to better compression can be preferred over the second key (e.g., the first key might have been used most recently in the compression operation). Once the matched key and / or literal character of the longest length has been emitted, lookahead buffer 120 can be adjusted such that the emitted matched symbols and / or literal characters can be removed from lookahead buffer 120, and this process can continue until all symbols in lookahead buffer 120 have been emitted.

[0033] If no matching key is determined (e.g., no match is found, or the length of the match does not meet a threshold), the first symbol in lookahead buffer 120 can be emitted as a literal character. Subsequently, lookahead buffer 120 can slide a symbol (e.g., so that the first symbol can be removed from it), and the process of matching keys and / or suffixes in history buffer 125 with symbols in lookahead buffer 120 can be repeated until all symbols in lookahead buffer 120 have been processed.

[0034] The computing device 110 can operate to implement a dictionary-based compression method, which may include generating fixed-length and / or variable-length substrings from the input data 102, and these substrings can be used to index a dictionary that maps the substrings to tags. Data reduction can be achieved where the mapping results in a reduction in the number of bits required to represent the tags relative to the input data.

[0035] Data compression module 115 is operable to perform and / or instruct operations to be performed to generate compressed data 130 based on input data 102. Data compression module 115 is operable to perform the lazy matching algorithm described herein. Data compression module 115 is operable to search history buffer 125 (e.g., by adjusting the position of a pointer relative to symbols in history buffer 125) to attempt to obtain a match relative to symbols in lookup buffer 120. If a match is found (e.g., between symbols in lookup buffer 120 and symbols in history buffer 125), data compression module 115 may attempt to expand the match by comparing adjacent symbols in lookup buffer 120 (e.g., symbols adjacent to the matching symbol) with adjacent symbols in history buffer 125. Data compression module 115 may continue in this manner until adjacent symbols in lookup buffer 120 and adjacent symbols in history buffer 125 can no longer be matched. In this way, data compression module 115 is operable to continue matching symbols in lookup buffer 120 with symbols in history buffer 125 until the matching ends. As previously stated, unless otherwise noted, the matched symbols between lookahead buffer 120 and history buffer 125 may be referred to as substring matching.

[0036] In this way, the data compression module 115 can be operated to search the entire history buffer 125 to determine the longest substring match in the lookahead buffer 120. In some cases, the data compression module 115 can consider the length of the substring match in conjunction with the minimum substring match length, where the minimum substring match length can indicate the minimum number of symbols contained in the substring during data compression. In some cases, the minimum substring match length can be user input from the system 100, can be based on input data 102 (e.g., the data type associated with input data 102, the number of symbols contained in input data 102, etc.), or can be dynamically adjusted by the system 100 and / or computing device 110 based on input data 102 (e.g., if the number of substring matches satisfying the minimum substring match length is less than the expected / anticipated threshold, the system 100 can adjust the minimum substring match length so that more substring matches can satisfy the minimum substring match length), etc.

[0037] Once the data compression module 115 determines the longest substring match (where the longest substring match can include a length greater than or equal to the minimum match length), the data compression module 115 can encode the longest substring match using markers, where the markers can include at least a distance reference and / or a length reference. In this case, the distance reference can be the number of symbols between the first symbol in the history buffer 125 and the first symbol in the longest substring match, and the length reference can be the number of symbols included in the longest substring match. If no substring match is determined between a symbol in the lookahead buffer 120 and the history buffer 125, the symbol in the lookahead buffer 120 can be emitted as a literal character.

[0038] The data compression module 115 is operable to perform one or more iterations to determine the longest substring match by adjusting pointers in the lookahead buffer 120 to skip one or more symbols before performing the matching as described. In some cases, the data compression module 115 may skip a certain number of symbols based on the size of the delayed matching window that may be used in the data compression operation. A detailed description of the algorithm and its steps can be further provided in the section on Figure 2-6 The flowchart will describe and explain this process.

[0039] When the data compression module 115 implements the above algorithm, or the lazy matching algorithm with a delayed matching window, the data compression module 115 can compress the input data 102 into compressed data 130. In some cases, as the size of the delayed matching window increases, the compression ratio may increase, and the encoding latency in the system 100 may increase accordingly. In this disclosure, the data compression module 115 can implement a lazy matching algorithm with lower complexity than other lazy matching algorithms, reducing CPU cycles in software implementation and / or reducing the clock cycles and / or power consumption of the system 100 in hardware implementation.

[0040] System 100 can be implemented in various devices, such as, but not limited to, computing storage devices, DPUs, NICs, general-purpose CPUs, GPUs, microcontrollers, FPGAs, ASICs, and / or data conversion accelerators, any of which can be used for data compression. In the case where System 100 is a storage and encryption data conversion accelerator, System 100 may include one or more data conversion engines as computing resources for data conversion operations, such as data compression, encryption operations, and / or other conversion operations. The data conversion engines included in System 100 can operate on data in a highly parallel manner. Such data conversion accelerators can be connected to a host or server platform using Peripheral Component Interconnect High Speed ​​(PCIe), CXL, and / or USB. The host or server can submit commands to the data conversion accelerator along with the source data to be converted. Alternatively or additionally, the data conversion accelerator can provide control information and / or metadata describing the specific algorithmic conversions to be applied to the data. Based on the metadata, the data conversion engine can perform operations including data compression. The data conversion engine in System 100 for dictionary-based data compression (potentially a data conversion accelerator) can implement the low-complexity algorithms described herein.

[0041] The following two examples of lazy matching iterations (with a delayed matching window size of 1) provide at least some of the motivations behind the algorithm described in this paper.

[0042] Set A may include substring matches where the match length is greater than or equal to the minimum match length when the search begins from the first symbol in the lookahead buffer (e.g., the symbol DMW = 0). The set described above (e.g., set A) may be referred to as the DMW = 0 iterative match set. Each element of set A may be a matching substring in the history buffer (e.g., the matching substring includes a length greater than or equal to the minimum match length).

[0043] Set B may include substring matches where the match length is greater than or equal to the minimum match length when the search begins from the second symbol in the lookahead buffer (e.g., the symbol DMW = 1). The set described above (e.g., set B) may be referred to as the DMW = 1 iterative match set. Each element of set B may be a matching substring in the history buffer (e.g., the matching substring includes a length greater than or equal to the minimum match length).

[0044] Considering the two iterations above (e.g., set A with DMW = 0 and set B with DMW = 1), at least three cases are possible: Case 1, Case 2, and Case 3, all of which are described below.

[0045] In case 1, set A may be empty. Regardless of what set B might contain, the DMW = 0 symbol may be emitted as a literal character, and the symbol pointer may advance one symbol (e.g., the next symbol). Set B can be empty if it is constructed to contain all matches greater than or equal to the minimum match length minus one. In this case, set B can be empty because a match of the minimum match length in the DMW = 0 iteration can produce a match of the minimum match length minus one in the DMW = 1 iteration. Therefore, it may not be necessary to construct set A, because when constructing set B for substring matches of the minimum match length minus one, the emptiness of set A can be inferred from the emptiness of set B.

[0046] In case 2, set B might be empty, while set A might be non-empty (e.g., there might be no substring match with a length greater than or equal to the minimum match length in the DMW = 1 iteration, but a substring match might exist in the history buffer of the DMW = 0 iteration). This configuration might occur when at least one of the longest substring matches in the DMW = 0 iteration is the minimum match length. Alternatively, if set B is constructed to include all matches greater than or equal to the minimum match length minus one, then set B might contain elements with a length equal to the minimum match length minus one. This is because the above situation might arise from a match with the minimum match length in the DMW = 0 iteration, which could potentially produce a match with the minimum match length minus one in the DMW = 1 iteration. In this case, constructing set A might be unnecessary, as the elements with the minimum match length in set A can be inferred from the length of the elements in set B. Furthermore, it can be verified whether the first symbol in the lookahead buffer matches the symbol preceding the substring in the history buffer that corresponds to the match in set B.

[0047] Alternatively, in case 2, a substring search can be performed in the DMW = 1 iteration with a length one less than the minimum match length (e.g., minimum match length minus one), and the results can be checked by comparing the first symbol in the lookahead buffer with the symbols preceding the matching substring in the history buffer to determine if any matches can be expanded. All substring matches obtained in the DMW = 1 iteration that are one less than the minimum match length, and that have a symbol preceding the substring in the history buffer that matches the first symbol in the lookahead buffer, can be candidate substring matches in the DMW = 0 iteration. One of these candidate matches can be selected. Criteria for selecting a candidate match can include, but are not limited to, the candidate closest to the start of the lookahead buffer. Alternatively, different criteria can be selected to choose the optimal match from among the tie-breakers.

[0048] Once a match is found, the search process can terminate (e.g., if the match satisfies the minimum match length). In this case, a match can be issued using both distance and length values, and the symbol pointer can advance to the next symbol based on the length value. In this case, the length value satisfies the minimum match length.

[0049] In case 3, set A may not be empty, and set B may not be empty either. There may be one or more substring matches in the DMW = 0 iteration, where substring matches can include matches with a length greater than or equal to the minimum match length. Alternatively, there may be one or more substring matches in the history buffer of the DMW = 1 iteration, where substring matches can include matches with a length greater than or equal to the minimum match length. Set B can be non-empty if it is constructed to include all matches greater than or equal to the minimum match length minus one. A match of length 'L' in the DMW = 0 iteration may result in a match of length 'L-1' in the DMW = 1 iteration. Therefore, it is not necessary to construct set A by rescanning the history buffer and dictionary, since the elements of set B can consider the longest length and the longest length minus one. Alternatively, it can be considered whether the selected elements included in set B can be further expanded by comparing the first symbol in the lookahead buffer with the symbols in the history buffer preceding the substring match within set B.

[0050] In this case, you can choose the substring match with the longest matching length and / or the longest matching length minus one in the DMW = 1 iteration. Alternatively, you can choose the substring match with a length one sign less than the longest matching length in the DMW = 1 iteration.

[0051] In response to the presence of selected substring matches in the history buffer, these selected substring matches can be examined to determine if any of them can be further expanded. For example, the first symbol in the lookahead buffer can be compared to the symbol preceding the matching substring in the history buffer. Depending on the contents of the history buffer, a portion of the selected substring matches may be expanded, or no selected substring matches may be expanded. After attempting to expand the substring matches, the substring match with the longest length can be selected. In the presence of two or more matching substrings with equal longest lengths, criteria for selecting the best match among the parallel items can be implemented to choose one matching substring over others. For example, criteria for selecting the best match among the parallel items could include selecting the matching substring with the shortest distance to the beginning of the lookahead buffer. Other methods for selecting the best match among the parallel items can be implemented, any of which can be used to select a specific matching substring from multiple matching substrings with equal longest lengths, as described above.

[0052] System 100 may be modified, added to, or omitted without departing from the scope of this disclosure. For example, the designation of different elements in the manner described is intended to help explain the concepts presented herein and is not restrictive. Furthermore, system 100 may include any number of other elements, or may be implemented in a system or context other than those described herein. For example, Figure 1 Any component can be divided into more components or merged into fewer components.

[0053] Figure 2 A flowchart illustrating an exemplary method 200 for data compression using a lazy matching algorithm according to at least one embodiment of the present disclosure is shown. Method 200, or methods 300, 400, 500, and / or 700 described subsequently, may be executed by processing logic, which may include hardware (circuit, dedicated logic, etc.), software (e.g., running on a general-purpose computer system or a dedicated machine), or a combination of both. This processing logic may be contained in any computer system or device, such as… Figure 1 The computing device 110 or the data compression module 115.

[0054] For simplicity, the methods described herein (e.g., method 200, method 300, method 400, method 500, and / or method 700) are depicted and described as a series of actions. However, actions according to this disclosure can occur in various orders and / or concurrently, accompanied by other actions not presented and described herein. Furthermore, not all actions shown are applicable to implementing the methods according to the disclosed subject matter. Moreover, those skilled in the art will understand and appreciate that the method may alternatively be represented by a state diagram or events as a series of interrelated states. Furthermore, the methods disclosed in this specification can be stored on an article of art, such as a non-transitory computer-readable medium, to facilitate the transfer and assignment of these methods to a computing device. As used herein, the term "article of art" is intended to encompass a computer program accessible from any computer-readable device or storage medium. Although illustrated as discrete boxes, various boxes may be divided into additional boxes, merged into fewer boxes, or eliminated, depending on the desired implementation.

[0055] At box 210, computing devices (e.g.) Figure 1 The computing device 110 can obtain the input data to be compressed. The input data can be obtained from a remote device (e.g., transmitted to the computing device), and / or the input data can be obtained from a storage device, such as a server, cloud storage system, etc.

[0056] At box 220, the obtained input data can be stored in a buffer (e.g., Figure 1 The lookahead buffer 120). Input data in the buffer can be compared with historical buffers (e.g., ...). Figure 1The input data in the buffer (125) is compared to determine whether a substring of the input data in the buffer matches at least partially with the data in the historical buffer. For example, a symbol-level comparison can be performed between symbols in the buffer and symbols in the historical buffer to determine if a substring match exists between the symbols in the buffer and the symbols in the historical buffer. Additional details related to storing input data and obtaining substring matches can be combined with... Figure 3 Method 300 will be discussed further.

[0057] In some instances, substrings in the buffer can be compared with keys and / or suffixes in the history buffer to identify one or more matches. Matched substrings may be preserved if they meet a threshold length. For example, a matched substring may be preserved if its length is greater than zero and less than or equal to the maximum substring length minus the DMW window length.

[0058] At box 230, the substring match determined at box 220 can be expanded by comparing the additional symbols in the buffer with the additional symbols in the history buffer. In some cases, the additional symbols may be adjacent to symbols that have already been matched in the buffer and / or the history buffer. For example, a substring in the buffer that matches a substring in the history buffer (e.g., has at least one symbol) can be expanded by comparing adjacent symbols in the buffer with adjacent symbols in the history buffer. Additional details related to expanding the substring match can be combined... Figure 4 Method 400 will be discussed further.

[0059] At box 240, compression can be performed on the substring matches identified in the aforementioned boxes. In some cases, certain symbols in the buffer can be emitted as literal characters and compressed separately from the substring matches. Additional details related to performing the compression operation can be found in [the relevant documentation / details]. Figure 5 Method 500 will be discussed further.

[0060] At box 250, the buffer and / or history buffer can be updated. Buffer updates may include removing symbols associated with substring matching (which may include symbols that can be emitted as literal characters) from the buffer and adjusting the remaining symbols in the buffer for use in subsequent matches. Alternatively, the history buffer may be updated to include symbols removed from the buffer. For example, symbols contained in a substring match may be removed from the buffer and included in the history buffer for use in subsequent matching operations. In the case of a static history buffer (e.g., a static dictionary), substring matching symbols removed from the buffer may not be added to the history buffer. Additional details related to updating the buffer can be found in [the relevant documentation / details]. Figure 5 Method 500 will be discussed further.

[0061] At box 260, compressed data obtained from the compression operation of substring matching (and related literal characters) can be generated and / or output. For example, the compressed data can be stored in a data memory, stored in a buffer, and / or transferred to other systems or devices for subsequent processing.

[0062] Method 200 may be modified, added to, or omitted without departing from the scope of this disclosure. For example, the designation of different elements in the manner described is intended to help explain the concepts presented herein and is not restrictive. Furthermore, method 200 may include any number of other elements, or may be implemented in a system or context other than those described herein.

[0063] Figure 3 At least one embodiment according to the present disclosure is shown. Figure 2 The flowchart illustrates an exemplary method 300 of method 200. Specifically, method 300 may detail storing input data in a lookahead buffer and obtaining substring matches between symbols in the lookahead buffer and symbols in the history buffer.

[0064] At box 302, input data can be acquired and stored in a lookahead buffer. The input data may contain one or more symbols that can be compared and / or matched with symbols in the history buffer, as described herein.

[0065] At box 304, a delayed match window (DMW) with a DMW length can be obtained for data compression operations. In some cases, the DMW length can be user-inputted by the lazy matching algorithm. Alternatively, the DMW length can correspond to the data type of the input data, the number of symbols contained in the input data, and / or other factors related to the input data. Alternatively, since adjusting the DMW length may cause changes in the compression ratio, the DMW length can be determined based on the desired compression ratio.

[0066] At box 306, a pointer in the lookahead buffer can be set to be equal to the DMW length. This pointer can be used in a lazy matching algorithm to point to a symbol to be matched with a symbol in the history buffer (e.g., a pointer symbol), and in response to being set to be equal to the DMW length, the pointer can skip several symbols in the lookahead buffer based on the DMW length. For example, with a DMW length of 2, the pointer can skip the first and second symbols in the lookahead buffer and point to the third symbol in the lookahead buffer (e.g., the pointer symbol). In some cases, this pointer can also be referred to as an iteration of the lazy matching algorithm (e.g., a DMW iteration). For example, the pointer might initially point to the third symbol in the first iteration (as described in the previous example), and in the second iteration, the pointer could be updated to point to the second symbol, and so on. Therefore, a DMW iteration (e.g., i) may be associated with the current pointer value.

[0067] At box 308, a match can be obtained between the pointer symbol and a symbol in the history buffer. For example, a substring (e.g., a substring starting with the pointer symbol) can be searched in the lookahead buffer within the history buffer. In some cases, to be considered a valid substring match, the substring match must satisfy a threshold length (e.g., length '). The threshold length for matching this substring may be greater than the minimum substring matching length (e.g., ...). Subtract the iteration (e.g., the current pointer value). In some cases, a match search between the lookahead buffer and the history buffer can be performed by using parallel lookups in hardware and / or a rolling hash algorithm in software and / or hardware implementations.

[0068] When using rolling hashing for searching, substring matches of increasing length can be located. In such cases (e.g., when an increasing-length substring match is found), the maximum length of the substring matching the substring in the lookahead buffer and the history buffer in the i-th DMW iteration can be updated. The maximum length in the i-th DMW iteration can be expressed as... And the length falls within Substring matches within the specified range may be retained. Substring matches whose length falls outside the specified range may be discarded.

[0069] At box 310, based on DMW iterations, the identified substring matches can be included in the matching set. For example, in the i-th DMW iteration, the substring of length ' l The substring matching set of ' can be represented as Alternatively, substrings can be preserved in other sets, where the positions of matching substrings can be preserved in individual substring matching sets (in the history buffer).

[0070] Therefore, ordered sets can be generated ( Furthermore, in some cases, substring matches in an ordered set can be sorted in descending order based on the length of the substring matches. For example, such an ordered set can be represented as... In some embodiments, the ordered set may include, but is not limited to, lists, hash tables, and / or tree data structures, all of which may be stored in the memory of the CPU and / or microcontroller. Alternatively, in the case where the algorithm is implemented in hardware, one or more hardware elements may be used as containers to store (one or more) ordered sets.

[0071] When using a static dictionary, the symbols at the beginning of the lookahead buffer can be compared with the prefixes of keys in the history buffer. For example, the prefixes of previously retained keys and / or the suffixes associated with keys can be used for comparison with symbols in the lookahead buffer. In some cases, the key with the longest match (e.g., the one that matches the most symbols in the match) can be selected. When more than one key is identified as the longest match, additional considerations can be used to determine which longest matching key to select. For example, the compression ratio associated with the key can determine which longest matching key to select. In another example, a specific key that was most recently selected as the longest match can be selected as the longest match relative to another key that may not have been selected as the longest match recently.

[0072] In some cases, no substring match may be found between the lookahead buffer and the history buffer, or the substring match may not meet the minimum substring match length. In one case where no substring match is found, the first symbol at the beginning of the lookahead buffer can be transmitted as a literal character to the subsequent encoding state, such as in Huffman coding. Alternatively, in a second case where no substring match is found, the first 'i' symbols in the lookahead buffer can be transmitted as literal characters to the subsequent encoding state, such as in Huffman coding. After one or more symbols are sent as literal characters to the encoding state, the lookahead buffer can be adjusted to slide through the first symbol (e.g., in the first case) or 'i' symbols (e.g., in the second case), and the first symbol or 'i' symbols can be incorporated as part of the history buffer, as further described in method 500.

[0073] Method 300 may be modified, added to, or omitted without departing from the scope of this disclosure. For example, the designation of different elements in the manner described is intended to help explain the concepts presented herein and is not restrictive. Furthermore, method 300 may include any number of other elements, or may be implemented in a system or context other than the system or context described herein.

[0074] Figure 4 At least one embodiment according to the present disclosure is shown. Figure 2The flowchart illustrates an exemplary method 400 of method 200. Specifically, method 400 may detail traversing the matches in the match set to determine whether the substring match is extensible.

[0075] At box 402, the pointer in the lookahead buffer can be adjusted for use in subsequent iterations to obtain matches between lookahead buffer symbols and history buffer symbols. For example, the first DMW iteration can be performed with the pointer in the lookahead buffer pointing to the third symbol (e.g., the pointer equal to the DMW length), and the pointer can be adjusted for the second iteration to be equal to the DMW length less than one, so that the second DMW iteration can be performed with the pointer pointing to the second symbol in the lookahead buffer (e.g., the pointer equal to the DMW length less than one). In such cases, the pointer may cause symbols in the lookahead buffer to be skipped, and the number of symbols skipped this time may be one less than the number of symbols skipped in the first iteration when searching for substring matches between lookahead buffer symbols and history buffer symbols. As described herein, a full scan of the history buffer and / or dictionary search can be performed once, regardless of the number of DMW iterations. Alternatively, as described herein, subsequent iterations can look for extensions of identified matches.

[0076] At box 404, the substring match can be obtained from the match set. This match set can be obtained from another part of the lazy matching algorithm, for example... Figure 3 Box 310 is described in the diagram. If the matching set is empty, the algorithm may continue to check the matching set of previous iterations (e.g., the current DMW iteration incremented by one) to determine whether one or more symbols can be sent as literal characters to subsequent encoding states, and subsequently removed from the lookahead buffer and / or added to the history buffer. When using a static dictionary, symbols may not be added to the history buffer (e.g., since the history buffer is a static dictionary). If a substring match is found, method 400 may continue to box 406.

[0077] At box 406, an adjacent lookahead buffer symbol set adjacent to a symbol in the substring match can be compared with an adjacent history buffer symbol. This comparison can be used to expand the substring match. Expanding the substring match can add at least one or more symbols to the symbols contained in the substring match, which may increase the number of symbols contained in the substring match.

[0078] At box 408, a determination can be made regarding whether the substring match can be expanded. For example, if an adjacent lookahead buffer symbol matches an adjacent history buffer symbol, the substring match can be expanded into an expanded substring match. If the substring match is expanded, the method can continue to box 410. If the substring match is not expanded, the method can continue to box 412.

[0079] At box 410, extended substring matches can be added to a new substring match set, where the length of the matches in the new substring match set can be at least one symbol longer than the match set (e.g., a match set containing substring matches that have not yet been extended). In some cases, once the new substring match set is generated, the match set may be discarded because the new substring match set may contain extended substring matches that contain more symbols than the substring matches in the match set. After generating the new substring match set, method 400 can continue at box 412.

[0080] At box 412, the process of expanding the substring match can be repeated. For example, if it is determined in box 408 that the first substring match has not been expanded, the second substring match in the match set can be obtained, and it can be determined whether the second substring match can be expanded. In another example, if it is determined in box 408 that the first substring match has been expanded, the second substring match in the match set can be obtained, and it can be determined whether the second substring match can be expanded. This process can be repeated for some or all of the substring matches contained in the match set.

[0081] Method 400 may be modified, added to, or omitted without departing from the scope of this disclosure. For example, the designation of different elements in the manner described is intended to help explain the concepts presented herein and is not restrictive. Furthermore, method 400 may include any number of other elements, or may be implemented in a system or context other than those described herein.

[0082] Figure 5 At least one embodiment according to the present disclosure is shown. Figure 2 The flowchart illustrates an exemplary method 500 of method 200. Specifically, method 500 may detail the compression operation of substring matching symbols and / or literal characters in the lookahead buffer and the subsequent update of the lookahead buffer and / or history buffer.

[0083] At box 502, a substring match can be obtained from the match set. This substring match can be a substring match in the match set and / or an extended substring match in a new match set, both as described herein. In some cases, the obtained substring match may contain the longest length relative to other substring matches in the match set and / or other match sets.

[0084] At box 504, the length of the acquired substring match can be compared with a minimum length to determine if the substring match length meets the minimum length. The minimum length may be based on user input, the characteristics of the processing device executing the lazy matching algorithm, and / or other factors. If the substring match length meets the minimum length, method 500 may continue to box 506. If the substring match length does not meet the minimum length, method 500 may continue to box 508.

[0085] At box 506, symbols and / or leading symbols in the substring match can be sent to the compression module (e.g., encoding state, such as Huffman coding) for compression. If the substring match does not contain one or more symbols at the beginning of the lookahead buffer (e.g., the substring match begins with a later symbol, such as the second or third symbol in the lookahead buffer), the leading symbol (or additional literal characters) can be output as individual literal characters to the compression module, along with the symbols in the substring match, all of which can be compressed by the compression module.

[0086] At box 508, symbols in the substring match can be sent individually to the compression module as literal characters and compressed through the compression operation.

[0087] At box 510, the pointer in the lookahead buffer can be adjusted in response to a symbol being emitted from the lookahead buffer (e.g., as a substring match, as a literal character matching a substring, and / or as a literal character). For example, the lookahead buffer can be adjusted so that symbols emitted for compression can be removed and new symbols can be placed at the beginning of the lookahead buffer. Accordingly, the pointer can be adjusted relative to the new symbol and based on the DMW length. After symbols are emitted from the lookahead buffer, if there are no more symbols (or the number of remaining symbols in the lookahead buffer may not meet the minimum length for substring matching), the lazy matching algorithm can complete and stop.

[0088] At box 512, symbols emitted from the lookahead buffer as part of a compression operation can be added to the history buffer. These emitted symbols can be used in subsequent matching operations between symbols in the lookahead buffer and symbols in the history buffer to obtain substring matches as part of a lazy matching algorithm. In some cases, emitted symbols can be added to the beginning of the history buffer (e.g., the first symbol used for matching comparison with symbols in the lookahead buffer). Alternatively, or furthermore, the history buffer can be a static buffer (e.g., a dictionary-like buffer) that may not be updated with emitted symbols, and emitted symbols may be discarded. In the case of a static history buffer, the symbols and / or keys in the history buffer can be predefined, as described herein.

[0089] Method 500 may be modified, added to, or omitted without departing from the scope of this disclosure. For example, the designation of different elements in the manner described is intended to help explain the concepts presented herein and is not restrictive. Furthermore, method 500 may include any number of other elements, or may be implemented in a system or context other than that described.

[0090] Figure 6A block diagram 600 according to at least one embodiment of the present disclosure is shown. The block diagram 600 includes an exemplary look-ahead buffer 605 and an exemplary history buffer 610, which can be used for data compression employing a lazy matching algorithm. As shown, the look-ahead buffer 605 may contain multiple symbols, and the history buffer 610 may contain multiple symbols, with identical symbols (e.g., matching symbols) illustrated as having the same pattern. For example, the leftmost symbol in the look-ahead buffer 605 (e.g., the first symbol in the look-ahead buffer 605), illustrated with a first pattern, may be the same as the leftmost symbol in the history buffer 610 (e.g., the last symbol in the history buffer 610), which is also illustrated with the first pattern.

[0091] When the lazy matching algorithm identifies one or more substring matches between symbols in the lookahead buffer 605 and the history buffer 610, it can obtain the initial position 624 of the first symbol in the substring match in the history buffer 610. This initial position 624 reflects the number of symbols between the first symbol in the history buffer 610 and the first symbol of the substring match in the history buffer 610.

[0092] When the lazy matching algorithm attempts to expand the substring match between the symbols in lookahead buffer 605 and the symbols in history buffer 610, update position 626 can be obtained. This update position 626 can reflect the count of additional symbols and the number of symbols between the first symbol in history buffer 610 and the first symbol of the substring match in history buffer 610 (e.g., additional symbols besides the initial position 624).

[0093] When one or more matches are found between a substring in lookahead buffer 605 and one or more substrings in history buffer 610, the position and / or length of the matched substring can be obtained. For example, a first substring match 628 in history buffer 610 can be determined to match a substring in lookahead buffer 605, wherein the first substring match 628 may be at a first distance from the substring in lookahead buffer 605, and the first substring match 628 may include the first distance (e.g., measured from the first symbol in history buffer 610) and / or the associated length. Alternatively, or furthermore, a second substring match 630 in history buffer 610 can be determined to match a substring in lookahead buffer 605, wherein the second substring match 630 may be at a second distance from the substring in lookahead buffer 605, and the second substring match 630 may include the second distance and / or the same length as the first substring match 628.

[0094] If the lazy matching algorithm attempts to expand the substring match (e.g., by comparing the input adjacent symbols in lookahead buffer 605 with the historical adjacent symbols in history buffer 610), for example, the second substring match 630, and determines that the second substring match 630 can be expanded (e.g., by matching the input adjacent symbols with the historical adjacent symbols), then the expanded substring match 632 may be a second distance away from the substring in lookahead buffer 605, and the expanded substring match 632 may include the second distance and / or a second length, wherein the second length may be at least one more symbol than the second substring match 630 (and / or at least one more symbol longer than the first substring match 628).

[0095] Block diagram 600 may be modified, added to, or omitted without departing from the scope of this disclosure. For example, the designation of different elements in the manner described is intended to help explain the concepts presented herein and is not restrictive. Furthermore, block diagram 600 may include any number of other elements, or may be implemented in a system or context other than those described herein.

[0096] Figure 7 A flowchart of an exemplary method 700 for data compression using a static dictionary and a lazy matching algorithm, according to at least one embodiment of the present disclosure, is shown. At block 702, input data may be stored in a first buffer. The input data may include a first set of symbols.

[0097] At box 704, one or more keys can be retrieved from the second buffer. Each of these keys may comprise a buffer substring. In some cases, the second buffer may be a static dictionary. The buffer substring may be a predefined symbol stored in the second buffer.

[0098] In some cases, the buffer substring can be one or more suffixes associated with each of one or more keys. In some cases, one or more suffixes can be arranged in a suffix tree. Alternatively, or furthermore, one or more suffixes can be arranged in a suffix array. In some cases, the length of one or more suffixes can be less than the length of the keys associated with one or more keys.

[0099] At box 706, a portion of the first set of symbols can be compared with a buffer substring to identify one or more substring matches. A portion of the first set of symbols can be compared with one or more suffixes to identify one or more substring matches. In some cases, each of the one or more substring matches may include a substring length, which can be defined by the amount of substring matching symbols in the one or more substring matches. In some cases, the substring length may be greater than or equal to a minimum length. The minimum length can be a combination of the substring length and the length of the delayed matching window.

[0100] At box 708, the longest substring match among one or more substring matches can be identified. This longest substring match may include a substring match symbol. In some cases, the first length of the first longest substring match and the second length of the second longest substring match may be equal. Furthermore, in some cases, the first longest substring match may have a better compressibility ratio relative to the second longest substring match. In such cases, the first longest substring match may be selected as the longest substring match.

[0101] At box 710, keys from one or more keys associated with the longest substring match can be retained.

[0102] At box 712, compression can be performed on the substring matching symbols. At box 714, a portion of the first set of symbols can be removed from the first buffer.

[0103] Method 700 may be modified, added to, or omitted without departing from the scope of this disclosure. For example, compressed data may be output, such as to an output buffer, wherein the compressed data may include at least a portion of the input data that may have been compressed by a compression operation.

[0104] In another example, the designation of the different elements in the manner described is intended to help explain the concepts presented herein and is not restrictive. Furthermore, method 700 may include any number of other elements, or may be implemented in a system or context other than those described.

[0105] Figure 8 An example computing device 800 is illustrated, in which a set of instructions can be executed to cause a machine to perform any or more methods discussed herein. The computing device 800 may include a mobile phone, smartphone, netbook, rack server, router computer, server computer, personal computer, mainframe computer, laptop computer, tablet computer, desktop computer, or any computing device having at least one processor, etc., in which a set of instructions can be executed to cause a machine to perform any or more methods discussed herein. In alternative implementations, the machine may be connected (e.g., networked) to other machines in a local area network, intranet, extranet, or the Internet. The machine may operate as a server machine in a client-server network environment. The machine may include a personal computer (PC), set-top box (STB), server, network router, switch, or bridge, or any machine capable of executing a set of instructions (orderly or otherwise) to specify the operations to be performed by the machine. Furthermore, although only a single machine is shown, the term "machine" may also include any collection of machines that individually or collectively execute a set (or more) of instructions to perform any or more methods discussed herein.

[0106] The computing device 800 includes a processing device 802 (e.g., a processor), a main memory 804 (e.g., a read-only memory (ROM), flash memory, dynamic random access memory (DRAM), such as synchronous DRAM (SDRAM)), a static memory 806 (e.g., flash memory, static random access memory (SRAM)), and a data storage device 816, which communicate with each other via a bus 808.

[0107] Processing device 802 represents one or more general-purpose processing devices (e.g., microprocessors, central processing units, etc.). More specifically, processing device 802 may include a Complex Instruction Set Computing (CISC) microprocessor, a Reduced Instruction Set Computing (RISC) microprocessor, a Very Long Instruction Word (VLIW) microprocessor, or a processor implementing other instruction sets or combinations thereof. Processing device 802 may also include one or more special-purpose processing devices, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), network processors, etc. Processing device 802 is configured to execute instructions 826 to perform the operations and steps discussed herein.

[0108] The computing device 800 may also include a network interface device 822 that can communicate with a network 818. The computing device 800 may also include a display device 810 (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), an alphanumeric input device 812 (e.g., a keyboard), a cursor control device 814 (e.g., a mouse), and a signal generation device 820 (e.g., a speaker). In at least one implementation, the display device 810, the alphanumeric input device 812, and the cursor control device 814 may be combined into a single component or device (e.g., an LCD touchscreen).

[0109] Data storage device 816 may include computer-readable storage medium 824 on which one or more instruction sets 826 embody any one or more methods or functions described herein. During execution of the instructions 826 by computing device 800, the instructions 826 may reside wholly or at least partially in main memory 804 and / or processing device 802, which also constitute computer-readable media. The instructions may also be transmitted or received on network 818 via network interface device 822.

[0110] Although computer-readable storage medium 824 is shown as a single medium in the example implementation, the term "computer-readable storage medium" can include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) storing one or more sets of instructions. The term "computer-readable storage medium" can also include any medium capable of storing, encoding, or carrying sets of instructions for machine execution and causing the machine to perform any one or more methods of this disclosure. Therefore, the term "computer-readable storage medium" should be understood to include, but is not limited to, solid-state memory, optical media, and magnetic media.

[0111] The terms used in this disclosure, especially in the appended claims (e.g., the text of the appended claims), are generally intended to be “open-ended terms” (e.g., the term “comprising” should be interpreted as “including but not limited to”).

[0112] Furthermore, if there is an intent to describe a specific number of claims in an introduced claim, such intent will be explicitly stated in the claim, and where no such statement exists, such intent does not exist. For example, to aid understanding, the appended claims may include the use of the introductory phrases “at least one” and “one or more” to introduce a claim reference. However, the use of such phrases should not be construed as implying that a claim reference introduced by the indefinite article “a” or “a (a)” would limit any particular claim containing such an introduced claim reference to containing only one implementation of such a claim, even if the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “a (a)” (e.g., “a” and / or “a (a)” should be interpreted as meaning “at least one” or “one or more”); the same applies to the use of definite articles used to introduce claim references.

[0113] Furthermore, even when a specific number of claims are explicitly cited, those skilled in the art will recognize that such citations should be interpreted as referring to at least the number cited (e.g., a simple citation of "two citations," without further embellishment, implies at least two citations, or two or more citations). Additionally, in the use of conventions such as "at least one of A, B, and C" or "one or more of A, B, and C," such constructions are generally intended to include A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together, etc.

[0114] Furthermore, any extractive word or phrase preceding two or more alternative terms in the description, claims, or drawings should be understood to include the possibility of including one, any, or both terms. For example, the phrase "A or B" should be understood to include the possibility of including "A" or "B" or "A and B".

[0115] All examples and conditional language cited in this disclosure are intended for pedagogical purposes to aid the reader in understanding this disclosure and the concepts contributed by the inventors to advance the technology, and should be construed as not being limited to these specifically cited examples and conditions. Although implementations of this disclosure have been described in detail, various changes, substitutions, and alterations may be made without departing from the spirit and scope of this disclosure.

Claims

1. A method comprising: The input data is stored in a first buffer, and the input data includes a first set of symbols; Retrieve one or more keys from the second buffer, each of the one or more keys comprising a buffer substring; The portion of the first set of symbols is compared with the buffer substring to identify one or more substring matches; Identify the longest substring match among the one or more substring matches, the longest substring match including the substring matching symbol; Retain the keys from one or more keys associated with the longest substring match; Perform a compression operation on the matching symbols of the substring; as well as Remove a portion of the first group of symbols from the first buffer.

2. The method of claim 1, wherein the buffer substring is one or more suffixes associated with each of the one or more keys, and a portion of the first set of symbols is compared with the one or more suffixes to identify a match of the one or more substrings.

3. The method of claim 2, wherein the one or more suffixes are arranged in a suffix tree.

4. The method of claim 2, wherein the one or more suffixes are arranged in a suffix array.

5. The method of claim 2, wherein the length of the one or more suffixes is less than the length of the associated key in the one or more keys.

6. The method of claim 1, further comprising outputting compressed data, the compressed data including at least one symbol of the input data compressed by the compression operation.

7. The method of claim 4, wherein the buffer substring is a predetermined symbol stored in the second buffer.

8. The method of claim 1, wherein the second buffer is a static dictionary.

9. The method of claim 1, wherein each of the one or more substring matches includes a substring length defined by the amount of substring matching symbols in the one or more substring matches.

10. The method of claim 8, wherein the length of the substring is greater than or equal to a minimum length.

11. The method of claim 9, wherein the minimum length is a combination of the substring length and the delayed matching window length.

12. The method of claim 1, wherein the first length of the first longest substring match and the second length of the second longest substring match are equal, and the first longest substring match has a better compressibility ratio relative to the second longest substring match, wherein the first longest substring match is selected as the longest substring match.

13. A compression device, comprising: First buffer zone; Second buffer zone; as well as The data compression module can operate as follows: The input data is stored in the first buffer, and the input data includes a first set of symbols; Retrieve one or more keys from the second buffer, each of the one or more keys comprising a buffer substring; A portion of the first set of symbols is compared with the buffer substring to identify one or more substring matches; Identify the longest substring match among the one or more substring matches, the longest substring match including the substring matching symbol; Retain the keys from the one or more keys associated with the longest substring match; Perform a compression operation on the matching symbols of the substring; as well as Remove a portion of the first group of symbols from the first buffer.

14. The compression apparatus of claim 13, wherein the buffer substring is one or more suffixes associated with each of the one or more keys, and a portion of the first set of symbols is compared with the one or more suffixes to identify one or more substring matches.

15. The compression device of claim 14, wherein the one or more suffixes are arranged in a suffix tree.

16. The compression apparatus of claim 14, wherein the one or more suffixes are arranged in a suffix array.

17. The compression apparatus of claim 14, wherein the length of the one or more suffixes is less than the length of the associated key in the one or more keys.

18. The compression apparatus of claim 13, wherein the data compression module is further operable to output compressed data, the compressed data including at least a portion of the input data compressed by the compression operation.

19. The compression apparatus of claim 13, wherein the buffer substring is a predetermined symbol stored in the second buffer.

20. The compression apparatus of claim 13, wherein the first length of the first longest substring match is equal to the second length of the second longest substring match, and the first longest substring match has a better compressibility ratio relative to the second longest substring match, wherein the first longest substring match is selected as the longest substring match.