Parallel decompression of compressed data streams

By generating metadata to indicate the parallelism of the compressed data stream, the problem of low decompression efficiency of existing compression algorithms on parallel processing units is solved, realizing efficient parallel decompression in traditional systems while keeping the file size unchanged and improving the decompression speed.

CN120973755APending Publication Date: 2025-11-18NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511052152.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-08-25
Filing Date
2021-08-25
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing compression algorithms have low decompression efficiency on parallel processing units, and the widespread use of traditional file formats leads to high system reconfiguration costs, making it difficult to achieve fine-grained parallel decompression.

Method used

Metadata is generated to expose the degree of parallelism in the compressed data stream. The metadata indicates the boundaries and output positions in the data stream, allowing the decompressor to decompress the compressed data stream in parallel on a parallel processor without modifying the compressed data stream itself.

Benefits of technology

It achieves improved decompression speed without increasing the size of compressed data stream files, while maintaining compatibility with traditional systems and reducing the impact of system bandwidth and storage requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973755A_ABST
    Figure CN120973755A_ABST
Patent Text Reader

Abstract

The invention discloses parallel decompression of compressed data streams. In various examples, metadata corresponding to a compressed data stream compressed according to a serial compression algorithm, such as arithmetic coding, entropy coding, etc., may be generated in order to allow parallel decompression of the compressed data. Thus, the compressed data stream itself may not need to be modified, and the bandwidth and storage requirements of the system may be minimally affected. Furthermore, through parallel decompression, the system may benefit from faster decompression times while also reducing or completely removing the employment cycles of the system using metadata for parallel decompression.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese Patent Application No. 202110979990.8, filed August 25, 2021. BACKGROUND

[0002] For a long time, lossless compression algorithms have been used to reduce the size of data sets for storage and transmission. Many traditional compression algorithms rely on the Lempel-Ziv (LZ) algorithm, Huffman encoding, or a combination thereof. As an example, the DEFLATE compression format (Internet Standard RFC 1951) combines the LZ algorithm and Huffman encoding for use with email communications, downloading web pages, generating ZIP files for storage on hard drives, etc. Algorithms like DEFLATE can save bandwidth in data transmission and / or can conserve disk space by storing data with fewer bits. However, due to the strong dependency on previous input for reconstructing later input, traditional compression algorithms are inherently serial in nature, which makes these compression techniques less ideal for decompression on parallel processing units, such as graphics processing units (GPUs). As a result, fine-grained parallel decompression algorithms for processing compressed data are rare.

[0003] Most conventional approaches for parallel decompression rely on modifying the compression algorithm itself in order to remove the data hazard of the LZ algorithm and / or remove or limit the Huffman encoding step. Examples of existing methods for parallel decompression include LZ4 and LZ Sort and Set Empty (LZSSE). These and similar methods are able to achieve some benefits from parallel processing architectures - e.g., reduced run time - at the expense of some of the compression benefits of the LZ algorithm and / or Huffman encoding. For example, these parallel decompression algorithms typically result in a 10-15% increase in file size compared to the same file compressed under a traditional sequential implementation of the DEFLATE compression format.

[0004] Another drawback of these parallel decompression algorithms is the widespread use of traditional file formats presents a significant barrier to widespread adoption of any newly proposed format. For example, for a system that has stored data according to a more traditional compression format, like using the LZ algorithm, Huffman encoding, or a combination thereof, the system can need to be reconfigured to work with the new compression algorithm type. This reconfiguration can be costly because the bandwidth and storage requirements of the system can have been optimized for the lower bandwidth and reduced file size of the serial compression algorithm, and the increase in bandwidth and storage requirements of the parallel decompression algorithm can require additional resources. Furthermore, stored data from the existing compression format can have to be reformatted and / or new copies of the data can have to be stored in the updated format before removing the existing copies - thereby further increasing the time of the adoption period and potentially requiring additional resources. SUMMARY

[0005] Embodiments of the present disclosure relate to techniques for performing parallel decompression of compressed data streams. Systems and methods are disclosed that generate metadata for a data stream compressed according to more traditional compression algorithms (such as Lempel-Ziv (LZ), Huffman encoding, combinations thereof, and / or other compression algorithms) in order to expose different types of parallelism in the data stream for parallel decompression of the compressed data. For example, the metadata can indicate demarcations in the compressed data corresponding to individual data portions or chunks of the compressed data, demarcations of data segments within each content portion, and / or demarcations of dictionary segments within each data portion or chunk. Further, the metadata can indicate output locations in a data output stream so that a decompressor, especially when parallel decompression, can identify where to fit the decompressed data within the output stream. As such, and in contrast to conventional systems such as those described above, the metadata associated with the compressed stream results in a more minor (e.g., 1-2%) increase in the overall file size of the compressed data stream without requiring any modification to the compressed data stream itself. Thus, in contrast to conventional parallel decompression algorithms, the bandwidth and storage requirements of the system can be minimally impacted while also realizing the benefit of faster decompression times due to the parallel processing of the compressed data. Additionally, since the compressed stream is not affected (e.g., in the case of using DEFLATE format, the compressed stream still corresponds to DEFLATE format), issues with compatibility with older systems and files can be avoided since systems that employ central processing unit (CPU) decompression can ignore the metadata and serially decompress the compressed data according to conventional techniques, while systems that decompress using parallel processors such as GPUs can use the metadata to parallel decompress the data. BRIEF DESCRIPTION OF DRAWINGS

[0006] The present systems and methods for parallel decompression of compressed data streams are described in detail below with reference to the accompanying drawings, wherein:

[0007] Figure 1 depicted is an example data flow diagram showing a process 100 for parallel decompression of a compressed data stream, in accordance with some embodiments of the present disclosure;

[0008] Figure 2A depicted is an example table corresponding to metadata for parallel decompression of a compressed data stream, in accordance with some embodiments of the present disclosure;

[0009] Figure 2B depicted is an example table corresponding to metadata for parallel decompression of a compressed data stream, in accordance with some embodiments of the present disclosure;

[0010] Figure 2Cdepicts an example table corresponding to dictionaries and metadata associated with dictionaries, in accordance with some embodiments of the present disclosure;

[0011] Figure 2D depicts an example table corresponding to metadata for parallel decompression of chunks of a compressed data stream, in accordance with some embodiments of the present disclosure;

[0012] Figure 2E depicts an example table corresponding to replication of compressed data streams that are not suitable for parallel processing, in accordance with some embodiments of the present disclosure;

[0013] Figure 2F depicts an example table corresponding to replication of compressed data streams that are suitable for parallel processing, in accordance with some embodiments of the present disclosure;

[0014] Figure 3 depicts a flowchart corresponding to a method for generating metadata for a compressed data stream for parallel decompression of the compressed data stream, in accordance with some embodiments of the present disclosure;

[0015] Figure 4 depicts a flowchart corresponding to a method for parallel decompression of a compressed data stream, in accordance with some embodiments of the present disclosure;

[0016] Figure 5 depicts a block diagram of an example computing device suitable for implementing some embodiments of the present disclosure; and

[0017] Figure 6 is a block diagram of an example data center suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0018] Systems and methods are disclosed regarding parallel decompression of compressed data streams. Although primarily described herein with respect to data streams compressed using Lempel-Ziv (LZ) algorithms and / or Huffman encoding (e.g., DEFLATE, LZ4, LZ classification and set empty (LZSSE), PKZIP, LZ Jaccard distance (LZJD), LZWelch (LZW), BZIP2, finite state entropy, etc.), this is not intended to be limiting. As such, other compression algorithms and / or techniques can be used without departing from the scope of the invention. For example, Fibonacci encoding, Shannon-Fano encoding, arithmetic encoding, artificial bee colony algorithm, Bentley, Sleator, Tarjan, and Wei (BSTW) algorithm, prediction by partial matching (PPM), run-length encoding (RLE), entropy encoding, Rice encoding, Golomb encoding, dictionary type encoding, etc. As another example, the metadata generation and parallel decompression techniques described herein can be suitable for any compressed data format that includes variable bit lengths for encoding symbols and / or variable output sizes for copies (e.g., a copy can correspond to one symbol, two symbols, five symbols, etc.).

[0019] The metadata generation and decompression techniques described herein can be used in any technical space that implements data compression and decompression— especially for lossless compression and decompression. For example, but not limited to, the techniques described herein can be implemented for audio data, raster graphics, three-dimensional (3D) graphics, video data, cryptography, genetics and genomics, medical imaging (e.g., for compressing Digital Imaging and Communications in Medicine (DICOM) data in medicine), executable files, moving data to and from network servers, sending data between and among central processing units (CPUs) and graphics processing units (GPUs) (e.g., for increasing input / output (I / O) bandwidth between CPUs and GPUs), data storage (e.g., to reduce data footprint), email, text, messaging, compressed files (e.g., ZIP files, GZIP files, etc.), and / or other technical spaces. The systems and methods described herein can be particularly suitable for amplifying storage and increasing PCIe bandwidth for I / O intensive use cases— such as communicating data between CPUs and GPUs.

[0020] Reference is made to Figure 1 , Figure 1is an example data flow diagram illustrating a process 100 for parallel decompression of a compressed data stream, in accordance with some embodiments of the present disclosure. It should be understood that this and other arrangements described herein are set forth only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) can be used in addition to or instead of those shown, and some elements can be omitted altogether. Further, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components, and in any suitable combination and location of hardware, firmware, and / or software. Different functions described herein as being performed by entities can be carried out by hardware, firmware, and / or software associated with those entities. For example, different functions can be performed by a processor executing stored instructions.

[0021] The process 100 can include receiving and / or generating data 102. For example, the data 102 can correspond to any type of technical space, such as but not limited to those technical spaces described herein. For example, the data 102 can correspond to textual data, image data, video data, audio data, genomic sequencing data, and / or other data types or combinations thereof. In some embodiments, the data 102 can correspond to data stored and / or transmitted using lossless compression techniques.

[0022] The process 100 can include a compressor 104 compressing the data 102 to generate compressed data 106. The data 102 can be compressed according to any compression format or algorithm, such as but not limited to those compression formats or algorithms described herein. For example, but not by way of limitation, the data 102 can be compressed according to a Lempel-Ziv algorithm, Huffman encoding, DEFLATE format, and / or another compression format or technique.

[0023] The compressed data analyzer 108 can analyze the compressed data 106 to determine opportunities for parallelization therein. For example, the compressed data analyzer 108 can identify segments (or portions) within the compressed data 132 that correspond to portions of the data stream that can be processed at least partially in parallel without affecting the processing of other segments. In some embodiments, the number of segments can be the same for each block of data, or can be different (e.g., dynamically determined). The number of segments is not limited to any particular number; however, in some non-limiting embodiments, each block of compressed data can be split into 32 different segments, such that 32 threads (or co-processors) in a thread bundle on a GPU can process the 32 segments in parallel. As other non-limiting examples, the compressed data 106 or blocks thereof can be split into 4 segments, 12 segments, 15 segments, 64 segments, etc. The number of segments can correspond to each block of data and / or to each portion of the data structure for dictionary decoding corresponding to each block, as described herein. As such, the data structure (dictionary) can be split into multiple segments for parallel decoding, and the data can be split into (in embodiments, equal) multiple segments for parallel decoding— e.g., using a dictionary that has already been decoded.

[0024] To determine which portions of the compressed data 106 are associated with each segment, the compressed data analyzer 108 can perform a first pass on the compressed data 106 to determine the number of symbols or tokens within the compressed data 106. In a second pass, the number of symbols can then be used to determine how many and which symbols will be included in each segment. In some embodiments, the number of symbols can be divided equally or as equally as possible among the segments. For example, where there are 320 symbols and 32 segments, each segment can include 10 symbols. In other examples, the number of symbols can be adjusted— e.g., by adding or subtracting one or more symbols for one or more of the segments— in order to simplify decompression. For example, instead of selecting 10 symbols per segment in the above example, one or more of the segments can include 11 symbols (while other segments can include 9 symbols) in order to cause the segment boundaries to correspond to certain byte intervals— e.g., 4-byte intervals— that the decompressor 114 can more easily handle (e.g., by avoiding splitting the output between bytes of the compressed data 106).

[0025] Subsequently, the segments can be analyzed by the metadata generator 110 to generate metadata 112 corresponding to the compressed data 106 that provides information to the decompressor 114 to decompress the compressed data 106 in parallel. For example, within each segment, the metadata 112 can identify three pieces of information. First, the number of bits in the compressed data where to start decoding the segment; second, the location in the output buffer where the results of the decoding will be inserted; and third, the location or place within the copy list (or match) to start outputting the delayed copy - e.g., the copy index. For example, for the third type of metadata 112, because the decoding can be performed in parallel, where LZ algorithms are used, the decompressor 114 can not decode the copy serially, so the copy can be batched for later execution. As such, the copy index can be included in the metadata 112 to indicate to the decompressor 114 to save space in the output buffer for each copy, and can also store the copy index in a separate data array, such that once the decompressor 114 performs the first pass, the copy can be performed by the decompressor 114 to fill the data into the output buffer. In some embodiments, the copy window can be a set length - e.g., a sliding window. For example, in the case of using LZ77, the sliding window for the copy can be 32kb, while in other algorithms, the sliding window can be different (e.g., 16kb, 64kb, 128kb, etc.) or variable size. As such, the compressed data 106 can be generated based on the sliding window size. As a result of the metadata 112, parallelism on the GPU can be performed such that each thread of the GPU can start decoding a portion of the compressed data 106 independently from each other. In the above example using 32 segments, this process 100 can result in 32-way parallelism, and each thread can decode 1 / 32 of the compressed data 106 or block thereof.

[0026] In some embodiments, the metadata can correspond to the number of bits per segment, the number of output bytes per segment, and / or the number of copies in each segment. However, in other embodiments, prefix-sum operations can be performed on this data (e.g., the number of bits, the number of output bytes, and / or the number of copies) to generate the metadata 112 in prefix-sum format. As such, the metadata 112 can correspond to the input (bits, nibbles, bytes, etc.) position per segment (e.g., as determined using the number of bits, nibbles, or bytes of each previous segment), the output (bits, nibbles, bytes, etc.) position per segment (e.g., as determined using the number of output bits, nibbles, or bytes from the previous segment), and the number of copies included in each segment prior to the current segment for which the metadata 112 is being generated. An example of the difference between these two formats of metadata is illustrated in FIGS. 6A and 6B. Figure 2A and Figure 2BAs shown in FIG. 1, the metadata 112 can be generated by the metadata generator 110 and stored in the memory 104. In some embodiments, the metadata 112 can be generated by the metadata generator 110 based on the compressed data 106 and the dictionary 116. In some embodiments, the metadata 112 can be generated by the metadata generator 110 based on the compressed data 106 and the dictionary 116 as described in further detail herein. In some embodiments, because the values of the input bits, the output positions, and / or the copy indices of each segment are monotonically increasing, the metadata 112 can be compressed by storing a common offset (shared by all segments) and the difference between the input bits, the output positions, and the copy indices in each segment.

[0027] As described herein, the compressed data analyzer 108 can analyze the compressed data 106 to determine the metadata 112 corresponding to the content portion of the compressed data 106, but can also analyze the compressed data 106 to determine the metadata 112 corresponding to the dictionary portion (if present) of the compressed data 106 and / or to determine the metadata 112 corresponding to identifying the blocks within the larger stream of the compressed data 106. As an example, the content portion of the compressed data 106 can require a dictionary in order to be properly decoded by the decompressor 114. In embodiments using Huffman encoding, the dictionary can include a representation of a Huffman tree (or match tree). In some embodiments, such as in the case where both LZ algorithms and Huffman encoding are used (e.g., in DEFLATE format), a first Huffman encoding operation can be performed on the copied literals and lengths, and a second Huffman encoding operation can be performed on the distances. As such, two or more Huffman trees can be included within the dictionary for decoding each of the copied literals and lengths and distances.

[0028] In other embodiments, the dictionary can provide an indication as to what symbol or bit values of the symbol the compressed data 106 corresponds to, such that the decompressor 114 can use the dictionary to decompress the content portion of the compressed data 106. In some embodiments, the dictionary can be Huffman encoded and can also correspond to a Huffman tree used to decompress the compressed data 106. In the case where a dictionary is used for each block of the compressed data 106, such as in DEFLATE format, the metadata generator 110 can generate metadata 112 corresponding to the start input bit of each segment of the dictionary and the number of bits for each symbol in the content portion of the block of the compressed data 106 that the dictionary corresponds to. As such, the dictionary can be divided into segments based on the metadata 112 and processed in parallel using threads of a GPU. As described herein, the number of segments can be similar to the number of segments of the data or content portion of the blocks of the compressed data 106, or can be different depending on the embodiment. Furthermore, the dictionary can include padding or repetition, similar to the padding or repetition of the copy or match of the data segments of the compressed data 106, and the padding or repetition can be used to further compress the dictionary.

[0029] Compressed data 106 can be separated into any number of chunks based on any number of criteria as determined by compressor 104 and / or according to the compression format or algorithm used. For example, in the case of a change in frequency or priority in compressed data 106, a first chunk and a second chunk can be created. As a non-limiting example, the letters A, e, and i can be the most frequent for a first portion of compressed data 106, while the letters g, F, and k can be the most frequent for a second portion of compressed data 106. As such, depending on the particular compression algorithm used, the first portion can be separated into a first chunk and the second portion can be separated into a second chunk. Compressor 104 can determine any number of chunks for compressed data 106. Compressed data analyzer 108 can analyze the chunks to determine the location of the chunks within the larger stream of compressed data 106. As such, metadata generator 110 can generate metadata 112 that identifies the start input bits and output bytes (e.g., the first output byte position of the decoded data) of each chunk of compressed data 106, which can include uncompressed chunks. Because the chunks are separate from one another and are separately identified by metadata 112, the chunks can also be processed in parallel - e.g., in addition to the compressed data 106 within each chunk being processed in parallel. For example, in the case where each chunk includes 32 segments, a first thread bundle of a GPU can be used to execute a first chunk and a second thread bundle of the GPU, in parallel with the first chunk, can be used to execute a second chunk. In examples where one or more of the chunks are uncompressed, the uncompressed chunks can be sent without a dictionary and the input bits and output bytes of the uncompressed chunks can be used by decompressor 114 to copy the data directly to the output.

[0030] As a result, the metadata 112 can correspond to the input and output locations of each chunk within the larger stream, the input location of the dictionary within each chunk, and the bit values of each symbol of the dictionary, as well as the input location, output location, and copy index of each segment within each chunk. This metadata 112 can be used by the decompressor 114 to decode or decompress the compressed data 106 in various forms of parallelism. For example, as described herein, individual chunks can be decoded in parallel - e.g., using different GPU resources and / or parallel processing units. Additionally, within each (parallel decompressed) chunk, the dictionary (where present) can be divided into segments, and the segments can be decoded or decompressed in parallel (e.g., where there are 64 segments of the dictionary, all 64 segments can be decoded in parallel, e.g., by using 64 different threads or two thread warps of a GPU). Further, within each (parallel decompressed) chunk, the content portion of the chunk can be divided into segments, and the segments can be decoded or decompressed in parallel. Further, as defined herein, one or more of these copy or match operations can be performed in parallel by the decompressor 114 - e.g., where a copy depends on data that has already been decoded into the output stream, that copy can be performed in parallel with one or more other copies. Additionally, each individual copy operation can be performed in parallel. For example, where a copy has a length greater than 1, the copy of each symbol or character of the full copy can be performed in parallel by the decompressor 114 - e.g., with respect to the example of "issi", each character of "issi" can be performed in parallel (e.g., copy "i" on a first thread of a GPU, copy "s" on a second thread, copy "s" on a third thread, and copy "i" on a fourth thread, in order to generate the corresponding output bytes of the output stream). Figure 2F

[0031] ​The decompressor 114 can receive the compressed data 106 and metadata 112 associated therewith. The decompressor 114 can use the metadata 112 to separate the compressed data 106 into individual chunks (where there is more than one chunk). For example, the decompressor 114 can analyze the metadata 112 corresponding to the chunk level of the compressed data 106 and can determine the input (bit, nibble, byte, etc.) location of each chunk (e.g., the first bit corresponding to that chunk or the compressed data 106) and the output (bit, nibble, byte, etc.) location of each chunk (e.g., the first output location in the output stream where the data from the chunk (after decompression) is located). After identifying each chunk, the decompressor 114 can process each chunk serially (e.g., can process a first chunk, then a second chunk, etc.), can assign two or more of the chunks for parallel decompression by different GPU resources (e.g., by assigning a first chunk to a first GPU or a first set of threads thereof and a second chunk to a second GPU or a second set of threads of the first GPU, etc.), or a combination thereof. In some embodiments, each chunk can correspond to a different type or mode, such as an uncompressed mode chunk, a fixed code table mode chunk, a generated code table mode chunk, and / or other types. The decompressor 114 can decompress the compressed data 106 (and / or decode uncompressed data when in uncompressed mode) based on the mode and the metadata 112 can be different based on the mode. For example, in uncompressed mode, there can be no dictionary since there is no need to decompress the data and / or there can be no copying or matching. As such, the metadata can only indicate the input and output locations of the data such that the input data stream corresponding to the uncompressed chunk is directly copied to the output stream.

[0032] The decompressor 114 can decompress each data block using the metadata 112 associated with the dictionary and the content portion of the block. For example, for each block, the metadata 112 can identify the input (bit, nibble, byte, etc.) position of the dictionary and the bit value (or number of bits) of each symbol of each segment of data in the block. As described herein, the dictionary can be used by the decompressor 114 to accurately decompress the content portion of the block. The dictionary can be generated using Huffman encoding of the content portion of the block, and in some embodiments, the compressed data corresponding to the dictionary can also be Huffman encoded. As a result, in embodiments, the dictionary portion of the compressed data can be compressed using Huffman encoding, and the content portion of the compressed data can be Huffman encoded. The metadata 112 corresponding to the dictionary portion of the compressed data 106 within each block can indicate the input position of the segments of the dictionary. For example, where the dictionary is divided into 32 segments, the metadata 112 can indicate the starting input bit (and / or output byte or other position) of each segment of the dictionary. As such, the decompressor 114 can use the metadata 112 to decompress or decode the dictionary portion of the compressed data 106 in parallel (e.g., one segment per thread of a GPU). The dictionary can be compressed according to the LZ algorithm (in embodiments, in addition to using Huffman encoding), and thus, decompression of the dictionary portion of the compressed data 106 can include copying or padding. As such, where parallel decompression of the dictionary is performed, the first pass by the decompressor 114 can decode the actual bit values (e.g., corresponding to the bit length of each symbol in the dictionary), and leave placeholders for the bit values to be copied or to be padded. During a second pass, the decompressor 114 can perform the padding or copying operations to fill in the missing bit values corresponding to the symbols of the dictionary (e.g., as described herein with respect to FIG. 2). Figure 2C In more detail,

[0033] The decompressor 114 can use the metadata 112 corresponding to the content portion of the compressed data 106 for each block to identify a first input location (e.g., bit, nibble, byte, etc.) of each segment of the compressed data 106, an output location in an output stream of each segment of the compressed data 106 after decompression, and / or a replication index or number of replications for each segment of the compressed data 106. The decompressor 114 can perform prefix-sum operations to determine the input location, output location, and replication number for each segment. However, in other embodiments, the metadata 112 can instead indicate a number of bits in each segment, a number of output bytes in each segment, and a number of replications in each segment, as described herein, rather than using a prefix-sum format to identify the input location, output location, and replication index. The decompressor 114 can decompress the identified segments of the compressed data 106 in parallel. For example, using the identifiers from the metadata 112, the decompressor 114 can assign blocks or portions of the compressed data 106 corresponding to the segments to different threads of a GPU. A first pass through each segment of the compressed data 106 by the decompressor 114 can be performed to output decompressed literals (e.g., actual symbols) from the compressed data 106 directly to the output stream (e.g., at the locations identified by the metadata) and to store replication or match information in a separate queue for later processing (e.g., in a second pass by the decompressor 114), while reserving space in the output stream for the replications. The amount of space reserved in the output stream can be determined using the metadata 112. These queued replications or matches can be referred to herein as delayed replications.

[0034] After the delayed copies are queued and placeholders in the output stream are created, decompressor 114 can perform a second pass through the delayed copies. Depending on whether each copy is determined to be safe to copy (e.g., if the data to be copied has already been decompressed, or does not depend on another copy that has not yet been copied, then the copy can be determined to be safe), one or more of the copies can be performed in parallel. For example, decompressor 114 can look ahead in the copy sequence to find additional copies that can be performed in parallel. The ability to parallelize copies can be determined using metadata 112 and / or information corresponding to the copies. For example, the output location of the copy within the output stream (as determined from metadata 112), the source location from which the copy is made (as determined from encoding distance information corresponding to the copy), and / or the length of the copy (as determined from encoding length information corresponding to the copy) can be used to determine whether a copy is safe for parallel processing with one or more other copies. A copy can be safe to perform in parallel with another copy when the source ends before the current output cursor and the copy itself does not overlap. As an example, and based on experimentation, the number of bytes that can be copied in parallel can increase from 3-4 to 90-100 or more. This process provides significant additional opportunities for parallelism across threads, as well as for memory system parallelism within a single thread. As such, one or more of the copies (e.g., intra-block copies or inter-block copies) can be performed in parallel with one or more other copies. With respect to Figure 2E- Figure 2F Examples of safe and unsafe copies for parallel execution are described. Further, in some embodiments, symbols within a single copy can be performed in parallel. For example, where a copy has a length greater than 1, two or more threads (or co-processors) of a GPU can be used to copy individual symbols within the copy to the output stream (of bytes) in parallel.

[0035] As a result, decompressor 114 can output each of the symbols to the output stream by performing a first pass through compressed data 106 to output literals and performing a second pass through the copies to output symbols from the copies. The result can be an output stream corresponding to data 102 as originally compressed by compressor 104. In examples using lossless compression techniques, the data 102 output can be the same or substantially the same as the data 102 input to compressor 104.

[0036] In some embodiments, a binary tree search algorithm with a shared memory table can be performed on compressed data 106 to avoid divergence across threads that would occur in the case of typical fast path / slow path implementations found in CPU-based decoders or decompressors. For example, in a conventional implementation on a CPU, a large data array can be used to decode a certain number of bits at a time. With respect to the DEFLATE format, each symbol can range from 1 to 15 bits long, so when data is being decoded, it can not be immediately apparent to the decompressor how long each symbol is. As a result, a CPU decompressor takes one bit to see if it is a length 1 symbol, then another bit to see if it is a length 2 symbol, and so on until the actual number of bits corresponding to the symbol is determined. This task can be time consuming and can slow down the decompression process, even for CPU implementations. Thus, some approaches have implemented a method of analyzing multiple bits (e.g., 15 bits) at a time. In such embodiments, 15 bits can be pulled from the compressed data stream and a lookup table can be used to determine what symbol the data corresponds to. However, this process is wasteful because the sliding window can only be 32 kb, but the system must store 15 bits for analysis, even in the case where the symbol can only be compressed to 2 bits. Thus, in some implementations, a fast path / slow path method can be used, where 8 bits are extracted, a symbol lookup is performed on the 8 bits, and when the symbol is shorter than 8 bits, a fast path is used, and when the symbol is greater than 8 bits, a slow path is used to determine what symbol the data represents. This process is also time consuming and reduces the run time of the system used to decompress compressed data 106.

[0037] Instead of using a fast pass / slow path approach where a certain number of threads (e.g., 32) are executing on a certain number of symbols (e.g., 32), some of which will hit the fast path and some of which will hit the slow path, mixed together in a warp (e.g., in the presence of 32 segments), this is inefficient. To solve this problem, a binary search algorithm can be used to improve efficiency. For example, a binary search can be performed on a small table (e.g., a table 15 entries long) to determine which symbols the table belongs to. Due to the reduced size of the array, the array can be stored in shared memory on the chip, which can result in fast lookups on the GPU. Additionally, using a binary search algorithm can allow all threads to execute the same code even though they are looking at different parts of the array in shared memory. Thus, memory traffic can be reduced because the binary search can look at a symbol of length 8 to see if the symbol is longer than 8 bits or shorter than 8 bits. Furthermore, one or more (e.g., two) of the top levels of the binary tree can be cached in data registers to reduce the number of shared memory accesses per lookup (e.g., from 5 to 3). Thus, the first of the four accesses can always be the same access instead of loading it out of memory every time, the register can stay alive on the GPU. The next can be 4 or 12 instead of having another level of memory access, the system can choose to look at the symbol 4 register or the symbol 12 register, and this can reduce the total number of accesses by 2 or more (e.g., typically 4 for a binary search to get the length, and 1 more to get the actual symbol, so this process reduces from 4 plus 1 to 2 plus 1). As such, instead of loading the entry and then shifting the symbol to compare, the symbol itself is pre-shifted.

[0038] Further, in some embodiments, the input stream of compressed data 106 can be swizzled or interleaved. For example, because the chunk of compressed data 106 can be divided into some number of segments (e.g., 32) by the compressed data analyzer 108, each thread can read from a distant portion of the stream. Thus, the input stream can be interleaved at the segment boundaries in pre-processing (e.g., using the metadata 112) to improve data read locality. For example, in the case where the data 102 corresponds to an actual dictionary of all words comprising a particular language, one thread can read from words starting with the letter "A," another from words starting with the letter "D," another from words starting with the letter "P," and so on. To remedy this problem, the data can be reformatted so that all threads can read from adjacent memory. For example, the compressed data 106 can be interleaved using information from the index so that each thread can read from a similar cache line. As such, the data can be shuffled together so that when the threads are processing the data, they can have some similarity in the data even though the data is different. For the playing card example, the swizzle or interleaving of the data can allow each thread to process a card with the same number or character even though it is a different suit.

[0039] As another example, in the case where threads in a warp of a GPU are used to process segments, a warp synchronization data parallel loop can be executed to load and process the dictionary. For example, using an index and data parallel algorithm, the system can instruct the dictionary entries in parallel. When processed serially, the system can look at how many symbols are length 2, length 3, etc. However, instead of performing these calculations serially, the system can perform a data algorithm to compute or assign threads to each symbol in parallel, then report whether the symbol has a particular length, and then perform a warp reduction to the total number of warps. For example, in the case where 286 symbols are to be analyzed (e.g., 0-255 bytes, 256 ends of the chunk, 257-286 for different lengths), each of the 286 symbols can be analyzed in parallel.

[0040] Referring now to Figure 2A-2F Each of the described examples can correspond to data compressed according to the DEFLATE compression format and metadata 112 corresponding thereto. However, this is for example purposes only, and as described herein, the techniques of the present disclosure can be implemented for or applied to any type of data compression format, such as but not limited to those described herein.

[0041] Figure 2AAn example table 200A corresponding to metadata 112 for parallel decompression of compressed data stream is depicted in accordance with some embodiments of the present disclosure. For example, data 102 (or a portion thereof, such as a chunk thereof) can correspond to the word "Mississippi." Compressor 104 can compress data 102 according to a DEFLATE compression algorithm to generate a compressed version of data 102 (e.g., compressed data 106), which is represented as "Miss <copy length 4, distance 3> ppi." Moreover, compressed data 106 can be Huffman encoded, and thus, different symbols can be represented by a number of bits corresponding to a certain priority or frequency assessment of compressor 104. For a non-limiting example, "M" can be represented by 3 bits, copy can be represented by 4 bits (e.g., 3 bits for length and 1 bit for distance), and "i," "s," and "p" can each be represented by 2 bits in compressed data 106. For this example, assuming a chunk of compressed data 106 is broken into four segments (e.g., 4-way indexing), compressed data analyzer 108 can analyze compressed data 106 to determine a first segment including "Mi," a second segment including "ss," a third segment including copy and "p," and a fourth segment including "p." For example, eleven characters or symbols "Mississippi" can be broken into eight symbols (e.g., seven literal and one copy), and these segments can be generated to have substantially equal sizes. However, due to an odd number of symbols, the fourth segment can include only one symbol. Compressed data analyzer 108 can then determine a number of outputs (or output bytes) for each segment, a number of inputs (or input bits) for each segment, and / or a number of copies in each segment. In some examples, metadata generator 110 can directly use this information to generate metadata 112. However, in other examples, a prefix and operation can be performed on this data to generate metadata 112 according to table 200B.

[0042] With respect to Figure 2B , Figure 2BExample table 200B is depicted according to some embodiments of the present disclosure, corresponding to metadata 112 in a prefix and sum format, which is used for parallel decompression of a compressed data stream. For example, instead of multiple outputs, each segment may be identified by an output position within the output stream to indicate to the decompressor 114 where the output of the decompression symbol from the segment should begin. Instead of multiple inputs, input positions within the compressed data stream can be identified to indicate to the decompressor 114 where the decompression segment should begin, allowing a single thread of the GPU to be allocated to the segment for parallel processing. Furthermore, instead of the number of copies in each segment, the total number of runs of copies from previous segments of the block can be identified in the metadata 112 to indicate to the decompressor which copy corresponds to each delayed copy in the queue. Ultimately, in this example, the prefix and format of metadata 112 can indicate to decompressor 114 that within the content portion (or data portion) of the current block of compressed data, there are 11 bytes of output, 19 bits of input, and a copy, and can indicate where each segment begins in compressed data 106, where each segment and / or copy index is output.

[0043] See Figure 2C , Figure 2C Example table 200C, corresponding to a dictionary and metadata 112 associated with the dictionary, is depicted according to some embodiments of this disclosure. For example, using as described herein... Figure 2A and Figure 2BThe same number of bits of the described symbols (e.g., as determined using Huffman encoding), a dictionary can be generated to indicate the values. In this example, the dictionary can correspond to the lowercase and uppercase letters of the English alphabet. However, this is not intended to be limiting, and the dictionary can correspond to any type of symbol, including characters from any language, numbers, symbols (e.g.,! $, *, ^, and / or other symbol types), etc. As such, because the compressed data 106 can only correspond to M, i, s, and p, the dictionary portion of the compressed data 106 can be compressed to indicate these values. In such an example, the data string 202 can represent the data 102 corresponding to the dictionary, where each of the 52 characters (e.g., A-Z and a-z) is represented by a value corresponding to a number of bits. To further compress the dictionary, the compressor 104 can generate fill or copy symbols corresponding to repeated values from the data string 202. In this case, the repeated value is 0, so the compressed data 106 corresponding to the dictionary can be represented by “<fill 12x>3 <fill 21x>2 <fill 6x>2002 <fill 7x>”. The compressed data analyzer 108 can analyze the compressed data 106 corresponding to the dictionary and determine segment breaks (e.g., in the example using four segments, the compressed data 106 can be separated into 4 segments). The separation of the four segments is indicated by the dashed lines. The metadata generator 110 can then analyze the segment information to generate the metadata 112 corresponding to the dictionary portion of the block of compressed data 106— e.g., to indicate the starting input position and symbol number or index of each segment in the dictionary.

[0044] Referring now to Figure 2D , Figure 2DAn example table 200D corresponding to metadata 112 for parallel decompression of blocks of compressed data stream is depicted in accordance with some embodiments of the present disclosure. For example, assuming data 102 is "MississippiMississippiMiss", compressor 104 can separate data 102 into two blocks for compression: a first block corresponding to "Mississippi"; a 2nd block corresponding to "MississippiMiss". As such, to identify the location of the different blocks within compressed data 106 and the dictionary corresponding thereto, compressed data analyzer 108 can analyze compressed data 106 to determine the initial input location (e.g., first input bit, nibble, byte, etc.) and / or the initial output location (e.g., first bit, nibble, byte, etc.) of each block of compressed data 106. Accordingly, metadata 112 corresponding to the stream of compressed data 106 can indicate the number of inputs (e.g., bits, nibbles, bytes, etc.) and the number of outputs (e.g., bits, nibbles, bytes, etc.) of each block of compressed data 106, the number of inputs (e.g., bits, nibbles, bytes, etc.) and the number of symbols of each segment within each block, and / or the number of inputs (e.g., bits, nibbles, bytes, etc.), the number of outputs (e.g., bits, nibbles, bytes, etc.), and the number of copies of each segment within each block. In the case of performing prefix-sum operations, metadata 112 can instead include the initial input location and the initial output location of each block of compressed data, the initial input location and the symbol index of each segment of the dictionary portion of each block, and / or the initial input location, the initial output location, and the copy index of each segment of the content portion of each block (or data portion). In further embodiments, some combination of the two different metadata formats can be used such that the metadata of one or more of the blocks, dictionaries, or data is in prefix-sum format while one or more of the blocks, dictionaries, or data is not in prefix-sum format.

[0045] Metadata 112 can then be used by decompressor 114 to decompress the compressed data 106. For example, metadata 112 can be used to identify each block of compressed data 106, such that two or more blocks of compressed data 106— e.g., block A and block B— can be decompressed in parallel. For each block, metadata 112 can be used to determine the segments of the dictionary, such that the dictionary can be decompressed in parallel— e.g., one segment per thread or co-processor. The dictionary can then be used to decompress the content portion of the compressed stream. For example, metadata 112 can indicate the segments of the content portion of compressed data 106, and decompressor 114 can use the dictionary to decode literal quantities from compressed data 106 and output the literal quantities to an output stream. Decompressor 114 can further use metadata 112 and the copy information encoded in compressed data 106 to reserve portions of the output stream for copies, and populate a queue or data structure with information about each copy— e.g., source location, distance, length, etc. As described herein, the segments of the content portion of compressed data 106 can be decompressed in parallel. After decompression, decompressor 114 can perform copy operations on the delayed copies in the queue to fill the reserved placeholders in the output stream with the corresponding copied symbols. As an example, and with respect to Figure 2A The copy of "issi" indicated by source location 1, copy length 4, and distance 3 can be used to copy "i" to location 4, "s" to location 5, and "i" to location 6. The "i" at location 6 can be referred to as an overlapping copy, as the "i" at location 6 is copied from the "i" at location 4, which does not exist until the copy begins. As described herein, in some embodiments, individual copy operations can be performed in parallel, such that two or more of the "issi" copy can be performed in parallel using different threads of a GPU.

[0046] Further, in some embodiments, separate copies can be performed in parallel when the copy is determined to be safe. For example, with reference to Figure 2E , Figure 2EExample table 400E depicts copies of compressed data streams that are unsuitable for parallel processing, according to some embodiments of this disclosure. For example, in the case where compressed data 106 corresponds to "MississippiMississippi", compressed data 106 may include two copies (e.g., copies #1 and #2 as indicated in table 200E). In this example, decompressor 114 may determine, either before or during the execution of the first copy, whether one or more additional copies—e.g., a second copy—can be executed in parallel. Decompressor 114 may examine the source location of the second copy and the output location of the first copy to determine if there is any overlap. In this case, because the second copy depends on the output from the first copy, it may not be safe to execute the second copy in parallel with the first copy. Accordingly, the first and second copies can be executed sequentially.

[0047] As another example, and see Figure 2F , Figure 2F Example table 400F depicts copies of compressed data streams suitable for parallel processing, according to some embodiments of this disclosure. For example, in the case where compressed data 106 corresponds to "MississippiMiss", compressed data 106 may include two copies (e.g., copies #1 and #2 as indicated in table 200F). In this example, decompressor 114 may determine, either before or during the execution of the first copy, whether one or more additional copies—e.g., a second copy—can be executed in parallel. Decompressor 114 may examine the source location of the second copy and the output location of the first copy to determine if there is any overlap. In this case, it may be safe to execute the second copy in parallel with the first copy because the second copy does not depend on the output from the first copy (e.g., because the second copy can be executed without filling the output buffer with the result from the first copy). Accordingly, the first and second copies can be executed in parallel, thus providing an output of 8 symbols at a time, instead of providing outputs of 4 and 4 symbols sequentially.

[0048] See now Figure 3-4 Each block of methods 300 and 400 described herein includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be executed by a processor that executes instructions stored in memory. Methods 300 and 400 can also be embodied as computer-usable instructions stored on a computer storage medium. Methods 300 and 400 can be provided by a standalone application, service, or managed service (alone or in combination with another managed service) or plug-in to another product, to name a few. Furthermore, by way of example, regarding... Figure 1The processes 100 describe the methods 300 and 400. However, these methods 300 and 400 can additionally or alternatively be performed within any one process or by any one system or any combination of systems, including but not limited to those described herein.

[0049] Referring now to Figure 3 , Figure 3 A flowchart corresponding to a method 300 for generating metadata for a compressed data stream for parallel decompression of the compressed data stream is depicted in accordance with some embodiments of the present disclosure. At block B302, the method 300 includes analyzing the compressed data. For example, the compressed data analyzer 108 can analyze the compressed data 106.

[0050] At block B304, the method 300 includes determining a demarcation between a plurality of segments of the compressed data. For example, the compressed data analyzer 108 can determine a demarcation between segments of the compressed data 106.

[0051] At block B306, the method 300 includes generating, based at least in part on the demarcation and for at least two segments of the plurality of segments, metadata indicating an initial input location within the compressed data and an initial output location in output data corresponding to each of the at least two data segments. For example, the metadata generator 110 can generate metadata 112 corresponding to the segments to identify the initial input location, the initial output location, and / or the copy index for some or all of the segments of the content portion of each block of the compressed data 106.

[0052] At block B308, the method 300 includes sending the compressed data and the metadata to a decompressor. For example, the compressed data 106 and the metadata 112 can be used by the decompressor 114 to at least partially parallel decompress the compressed data 106.

[0053] Referring now to Figure 4 , Figure 4 A flowchart corresponding to a method 400 for parallel decompression of a compressed data stream is depicted in accordance with some embodiments of the present disclosure. At block B402, the method 400 includes receiving compressed data and metadata corresponding thereto. For example, the decompressor 114 can receive the compressed data 106 and the metadata 112.

[0054] At block B404, the method 400 includes determining, based on the metadata, an initial input location and an initial output location corresponding to the compressed data. For example, the metadata 112 can indicate the initial input location in the compressed data 106 and the initial output location in the output data stream corresponding to each block of the compressed data 106.

[0055] At block B406, the method 400 includes determining, based on the initial input position and the initial output position, input dictionary positions and symbol indices for two or more dictionary segments of a dictionary of the compressed data. For example, the metadata 112 can indicate the initial input position and symbol indices for the segments of the dictionary corresponding to the compressed data 106.

[0056] At block B408, the method 400 includes decompressing the dictionary based on the input dictionary positions at least in part in parallel. For example, the metadata 112 can indicate the segments of the dictionary, and this information can be used by the decompressor 114 to process each segment of the dictionary in parallel using threads of a GPU.

[0057] At block B410, the method 400 includes determining, based on the initial input position and the initial output position, input segment positions, output segment positions, and copy index values for at least two segments of a plurality of segments of the compressed data. For example, the decompressor 114 can use the metadata 112 to determine, for each segment of the compressed data 106 in a block or portion of data, an initial input position in the compressed data 106, an initial output position in an output stream, and a copy index (e.g., a number of copies in a segment prior to the current segment) of the compressed data 106.

[0058] At block B412, the method 400 includes decompressing the at least two segments in parallel according to the input segment positions and the output segment positions to generate a decompressed output. For example, the decompressor 114 can use the metadata 112 and the dictionary to generate the data 102 from the compressed data 106. As such, once the data 102 has been recovered, the data 102 can be used on a receiving end to perform one or more operations. For example, in the case that the data 102 was compressed and passed from a CPU to a GPU for parallel processing, the data can then be passed back to the CPU. In the case that the data 102 corresponds to text, messaging, or email, the data can be displayed on a device (e.g., a user device or client device). In the case that the data 102 corresponds to video, audio, images, etc., the data can be output using a display, speaker, headphones, earpiece, etc. In the case that the data 102 corresponds to a website, the website can be displayed within a browser on a receiving device (e.g., a user device or client device). As such, the decompressed data can be used in any of a variety of ways, and due to the parallel decompression, the data can be made available more quickly while using less memory resources than conventional approaches.

[0059] Example Computing Device

[0060] Figure 5is a block diagram of an example computing device 500 suitable for implementing some embodiments of the present disclosure. The computing device 500 can include an interconnection system 502 coupling the following components: a memory 504, one or more central processing units (CPUs) 506, one or more graphics processing units (GPUs) 508, a communication interface 510, input / output (I / O) ports 512, input / output components 514, a power supply 516, one or more presentation components 518 (e.g., one or more displays), and one or more logical units 520. In at least one embodiment, the one or more computing devices 500 can include one or more virtual machines (VMs), and / or any of the components thereof can include virtual components (e.g., virtual hardware components). For a non-limiting example, the one or more GPUs 508 can include one or more vGPUs, the one or more CPUs 506 can include one or more vCPUs, and / or the one or more logical units 520 can include one or more virtual logical units. As such, the one or more computing devices 500 can include discrete components (e.g., a full GPU dedicated to the computing device 500), virtual components (e.g., a portion of a GPU dedicated to the computing device 500), or a combination thereof.

[0061] Although Figure 5 The various blocks shown in the computing device 500 are meant to be illustrative only. For example, in some embodiments, a presentation component 518 such as a display device can be considered an I / O component 514 (e.g., if the display is a touchscreen). As another example, a CPU 506 and / or GPU 508 can include memory (e.g., the memory 504 can represent a storage device in addition to the memory of the GPU 508, CPU 506, and / or other components). In other words, Figure 5 The computing device of the computing device 500 is merely illustrative. Distinction is not made between a “workstation,” “server,” “laptop,” “desktop,” “tablet,” “client device,” “mobile device,” “hand-held device,” “game console,” “electronic control unit (ECU),” “virtual reality system,” and / or other device or system types, as all are contemplated within the scope of the computing device 500. Figure 5 The computing device of the computing device 500 is merely illustrative. Distinction is not made between a “workstation,” “server,” “laptop,” “desktop,” “tablet,” “client device,” “mobile device,” “hand-held device,” “game console,” “electronic control unit (ECU),” “virtual reality system,” and / or other device or system types, as all are contemplated within the scope of the computing device 500.

[0062] The interconnection system 502 can represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnection system 502 can include one or more bus types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 506 can be directly connected to the memory 504. Also, the CPU 506 can be directly connected to the GPU 508. Where there are direct or point-to-point connections between components, the interconnection system 502 can include a PCIe link to perform the connection. In these examples, a PCI bus need not be included in the computing device 500.

[0063] The memory 504 can include any of a wide variety of computer-readable media. Computer-readable media can be any available media that can be accessed by the computing device 500. Computer-readable media can include both volatile and nonvolatile media, and removable and non-removable media. By way of example, and not limitation, computer-readable media can comprise computer storage media and communication media.

[0064] Computer storage media can include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules, and / or other data types. For example, the memory 504 can store computer readable instructions (e.g., which represent programs and / or program elements, such as an operating system). Computer storage media can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the computing device 500. Computer storage media, when used in this text, does not encompass signals per se.

[0065] Computer storage media can include computer-readable instructions, data structures, program modules, and / or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and includes any information delivery media. The term "modulated data signal" can refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, computer storage media can include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media. Combinations of the above should also be included within the scope of computer readable media.

[0066] The CPUs 506 can be configured to execute at least some computer-readable instructions in order to control one or more components of the computing device 500 to perform one or more of the methods and / or processes described herein. Each of the CPUs 506 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of handling a large number of software threads concurrently. The CPUs 506 can include any type of processors and can include different types of processors depending on the type of computing device 500 being implemented (e.g., processors with fewer cores for mobile devices and processors with more cores for servers). For example, depending on the type of computing device 500, the processors can be Advanced RISC Machines (ARM) processors implemented using reduced instruction set computing (RISC) or x86 processors implemented using complex instruction set computing (CISC). The computing device 500 can include one or more CPUs 506 in addition to one or more microprocessors or supplemental co-processors such as math co-processors.

[0067] In addition to or in place of CPU(s) 506, GPU(s) 508 can be configured to execute at least some computer-readable instructions to control one or more components of computing device 500 to perform one or more methods and / or processes described herein. GPU(s) 508 can be integrated GPUs (e.g., with CPU(s) 506) and / or GPU(s) 508 can be discrete GPUs. In embodiments, GPU(s) 508 can be co-processors to CPU(s) 506. Computing device 500 can use GPU(s) 508 to render graphics (e.g., 3D graphics) or perform general purpose computing. For example, GPU(s) 508 can be used for general purpose computing on GPUs (GPGPU). GPU(s) 508 can include hundreds or thousands of cores capable of processing hundreds or thousands of software threads concurrently. GPU(s) 508 can generate pixel data for an output image in response to rendering commands (e.g., received from CPU(s) 506 over a host interface). GPU(s) 508 can include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. Display memory can be included as part of memory 504. GPU(s) 508 can include two or more GPUs operating in parallel (e.g., via a link). The link can connect the GPUs directly (e.g., using NVLINK) or through a switch (e.g., using NVSwitch). When combined together, each GPU 508 can generate pixel data or GPGPU data for a different portion of an output or a different output (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can include its own memory or can share memory with other GPUs.

[0068] In addition to or in place of CPU(s) 506 and / or GPU(s) 508, logic unit(s) 520 can be configured to execute at least some computer-readable instructions to control one or more components of computing device 500 to perform one or more methods and / or processes described herein. In embodiments, CPU(s) 506, GPU(s) 508, and / or logic unit(s) 520 can perform any combination of methods, processes, and / or portions thereof, discretely or jointly. Logic unit(s) 520 can be part of and / or integrated with one or more of CPU(s) 506 and / or GPU(s) 508, and / or logic unit(s) 520 can be discrete components or otherwise external to CPU(s) 506 and / or GPU(s) 508. In embodiments, logic unit(s) 520 can be co-processors to CPU(s) 506 and / or GPU(s) 508.

[0069] Examples of logic 520 include one or more processing cores and / or components thereof, such as tensor cores (TCs), tensor processing units (TPUs), pixel visual cores (PVCs), visual processing units (VPUs), graphics processing clusters (GPCs), texture processing clusters (TPCs), streaming multi-processors (SMs), tree traversal units (TTUs), artificial intelligence accelerators (AIAs), deep learning accelerators (DLAs), arithmetic logic units (ALUs), application specific integrated circuits (ASICs), floating point units (FPUs), input / output (I / O) elements, peripheral component interconnects (PCI) or peripheral component interconnect express (PCIe) elements, and the like.

[0070] Communication interface 510 can include one or more receivers, transmitters, and / or transceivers that enable computing device 500 to communicate with other computing devices via electronic communication networks, including wired and / or wireless communications. Communication interface 510 can include components and functionality enabling communication over any of a number of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet or InfiniBand), low power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.

[0071] I / O ports 512 can enable computing device 500 to be logically coupled to other devices including I / O components 514, presentation components 518, and / or other components, some of which can be built into (e.g., integrated with) computing device 500. Illustrative I / O components 514 include a microphone, mouse, keyboard, joystick, game pad, game controller, dish satellite antenna, scanner, printer, wireless device, etc. I / O components 514 can provide a natural user interface (NUI) that processes audio, speech, or other physiological inputs generated by a user. In some instances, the inputs can be transmitted to an appropriate network element for further processing. The NUI can implement any combination of speech recognition, stylus recognition, facial recognition, biometric recognition, gesture recognition both on screen and adjacent to the screen, air gestures, head and eye tracking, and touch recognition (as described in more detail below) associated with a display of the computing device 500. Computing device 500 can include depth cameras, infrared cameras, RGB cameras, touch screen technology, and combinations of these, such as a stereoscopic camera system to capture depth camera images. Additionally, computing device 500 can include accelerometers or gyroscopes (e.g., as part of an inertial measurement unit (IMU)) to enable location detection by considering motion, acceleration, and / or orientation. In some examples, the output of the accelerometers or gyroscopes can be used by computing device 500 to render immersive augmented reality or virtual reality.

[0072] Power supply 516 can include a hard-wired power supply, a battery power supply, or a combination thereof. Power supply 516 can supply power to computing device 500 to enable operation of components of computing device 500.

[0073] Presentation component 518 can include a display (e.g., a monitor, a touchscreen, a television screen, a heads-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. Presentation component 518 can receive data from other components (e.g., GPU 508, CPU 506, etc.) and output the data (e.g., as images, video, sound, etc.).

[0074] Example data center

[0075] Figure 6 An example data center 600 that can be used in at least one embodiment of the present disclosure is shown. Data center 600 can include a data center infrastructure layer 610, a framework layer 620, a software layer 630, and / or an application layer 640.

[0076] As Figure 6 shown, data center infrastructure layer 610 can include resource orchestrator 612, grouped computing resources 614, and node computing resources (“node C.R.”) 616(1)-616(N), where “N” represents any whole, positive integer. In at least one embodiment, node C.R.s 616(1)-616(N) can include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid state drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power supply modules, and / or cooling modules, etc. In some embodiments, one or more of node C.R.s 616(1)-616(N) can correspond to a server having one or more of the above-described computing resources. Further, in some embodiments, node C.R.s 616(1)-616(N) can include one or more virtual components, such as a vGPU, a vCPU, etc., and / or one or more of node C.R.s 616(1)-616(N) can correspond to a virtual machine (VM).

[0077] In at least one embodiment, the grouped computing resources 614 may include individual groups (not shown) of node CR616 housed in one or more racks, or a plurality of racks (also not shown) housed in data centers in various geographic locations. Individual groups of node CRs within the grouped computing resources 614 may include computing, networking, memory, or storage resources that can be configured or allocated to support groups of one or more workloads. In at least one embodiment, several node CR616, including CPUs, GPUs, and / or other processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. One or more racks may also include any number of power modules, cooling modules, and / or network switches, in any combination.

[0078] Resource coordinator 622 may be configured or otherwise control one or more nodes CR616(1)-616(N) and / or grouped computing resources 614. In at least one embodiment, resource coordinator 622 may include a Software Design Infrastructure (“SDI”) management entity for data center 600. Resource coordinator 622 may include hardware, software, or some combination thereof.

[0079] In at least one embodiment, such as Figure 6 As shown, framework layer 620 may include a job scheduler 632, a configuration manager 634, a resource manager 636, and / or a distributed file system 638. Framework layer 620 may include a framework of software 632 supporting software layer 630 and / or one or more applications 642 of application layer 640. Software 632 or application 642 may respectively include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. Framework layer 620 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark, which can utilize distributed file system 638 for large-scale data processing (e.g., "big data"). TM(“Spark”). In at least one embodiment, job scheduler 632 can include a Spark driver to facilitate scheduling of workloads supported by various tiers of data center 600. Configuration manager 634 can be capable of configuring different tiers, such as software tier 630 and framework tier 620 including Spark and a distributed file system 638 for supporting large-scale data processing. Resource manager 636 can manage clustered or grouped computing resources mapped to or allocated for supporting distributed file system 638 and job scheduler 632. In at least one embodiment, clustered or grouped computing resources can include grouped computing resources 614 on data center infrastructure layer 610. Resource manager 636 can coordinate with resource orchestrator 612 to manage these mapped or allocated computing resources.

[0080] In at least one embodiment, software 632 included in software tier 630 can include software used by at least portions of node C.R.s 616(1)-616(N), grouped computing resources 614, and / or distributed file system 638 of framework tier 620. One or more types of software can include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.

[0081] In at least one embodiment, one or more application programs 642 included in application tier 640 can include one or more types of application programs used by at least portions of node C.R.s 616(1)-616(N), grouped computing resources 614, and / or distributed file system 638 of framework tier 620. One or more types of application programs can include, but are not limited to, any number of genomics application programs, cognitive computing and machine learning application programs including training or inferencing software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning application programs used in conjunction with one or more embodiments.

[0082] In at least one embodiment, any of configuration manager 634, resource manager 636, and resource orchestrator 612 can implement any number and type of self-modifying actions based on any number and type of data acquired in any technically feasible fashion. Self-modifying actions can relieve data center operators of data center 600 from making possibly poor configuration decisions and can avoid underutilized and / or poorly performing portions of a data center.

[0083] Data center 600 may include tools, services, software, or other resources to train one or more machine learning models or to use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, a machine learning model can be trained by calculating weight parameters based on a neural network architecture using the software and / or computing resources described above with respect to data center 600. In at least one embodiment, information can be inferred or predicted using trained or deployed machine learning models corresponding to one or more neural networks using the resources described above with respect to data center 600 by using weight parameters calculated through one or more training techniques such as, but not limited to, those described herein.

[0084] In at least one embodiment, the data center 600 may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, and / or other hardware (or corresponding virtual computing resources) to perform training and / or inference. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.

[0085] Example network environment

[0086] A network environment suitable for implementing embodiments of this disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may be... Figure 5 This is implemented on one or more instances of computing devices 500—for example, each device may include similar components, features, and / or functions of one or more computing devices 500. Furthermore, in the case of implementing backend devices (e.g., servers, NAS, etc.), the backend devices may be included as part of a data center 600, examples of which are described in this document. Figure 6 To describe in more detail.

[0087] Components of a network environment can communicate with each other via one or more networks, which can be wired, wireless, or both. A network may include multiple networks or one of multiple networks. For example, a network may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks such as the Internet and / or the Public Switched Telephone Network (PSTN), and / or one or more private networks. In cases where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connectivity.

[0088] A compatible network environment can include one or more peer-to-peer network environments (in which case, a server can not be included in the network environment) and one or more client-server network environments (in which case, one or more servers can be included in the network environment). In a peer-to-peer network environment, functionality described herein with respect to a server can be implemented on any number of client devices.

[0089] In at least one embodiment, a network environment can include one or more cloud-based network environments, distributed computing environments, combinations thereof, and the like. A cloud-based network environment can include a framework layer, a work scheduler, a resource manager, and a distributed file system implemented on one or more servers, which can include one or more core network servers and / or edge servers. The framework layer can include a framework that supports one or more applications of a software layer and / or an application layer. The software or applications can include web-based service software or applications, respectively. In an embodiment, one or more client devices can use the web-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer can be, without limitation, a type of free and open-source software web application framework, such as can be used for large-scale data processing (e.g., “big data”) using the distributed file system.

[0090] A cloud-based network environment can provide cloud computing and / or cloud storage that performs any combination of the computing and / or data storage functionality described herein (or one or more portions thereof). Any of these different functionalities can be distributed across multiple locations from central or core servers (e.g., one or more data centers that can be distributed across states, regions, countries, globally, and the like). If a connection to a user (e.g., a client device) is relatively close to an edge server, a core server can designate at least a portion of functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), can be public (e.g., available to many organizations), and / or combinations thereof (e.g., a hybrid cloud environment).

[0091] One or more client devices can include those described herein with respect to Figure 5At least some of the components, features, and functionalities of the described one or more example computing devices 500. By way of example, and not limitation, a client device can be embodied as a personal computer (PC), a laptop computer, a mobile device, a smartphone, a tablet computer, a smartwatch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or device, a video player, a camera, a surveillance device or system, a vehicle, a ship, a spacecraft, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or system, an entertainment system, a vehicle computer system, an embedded system controller, a remote control, an appliance, a consumer electronic device, a workstation, an edge device, any combination of these depicted devices, or any other suitable device.

[0092] The present disclosure can be described in the general context of machine-usable instructions or computer code, including computer-executable instructions such as program modules, being executed by a computer or other machine, such as a personal data assistant or other handheld device. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types. The present disclosure can be practiced in a variety of system configurations, including hand-held devices, consumer electronics, general- purpose computers, more specialty computing devices, etc. The present disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network.

[0093] As used in this document, the recitation of “and / or” with respect to two or more elements should be interpreted to mean only one element, or a combination of elements. For example, “element A, element B, and / or element C” can include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or element A, B, and C. In addition, “at least one of element A or element B” can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, “at least one of element A and element B” can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0094] The subject matter of the present disclosure is described with specificity herein to meet statutory requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the inventors have contemplated that the claimed subject matter might also be embodied in other ways, to include different steps or combinations of steps similar to the ones described in this document, in conjunction with other present or future technologies. Moreover, although the terms "step" and / or "block" might be used herein to connote different elements of methods employed, the terms should not be interpreted as implying any particular order among or between various steps herein disclosed unless and except when the order of individual steps is explicitly described.

Claims

1. A method comprising: Analyze the compressed data to determine the data segments of the compressed data; Metadata is generated for the data segments, the metadata indicating the position of the input segment within the compressed data and the position of the output segment in the output data corresponding to each data segment in the data segments; as well as Associate the metadata with the compressed data.

2. The method of claim 1, wherein the metadata further indicates a replication index corresponding to the data segment.

3. The method of claim 1, wherein the metadata further indicates an initial input position and an initial output position corresponding to the compressed data.

4. The method according to claim 1, further comprising: The data segments are decompressed in parallel according to the input segment position and the output segment position to generate a decompressed output.

5. The method of claim 1, wherein the input segment position indicates the block input position of each of at least two blocks of the compressed data, and the output segment position indicates the block output position of each of the at least two blocks of the compressed data, wherein each data segment corresponds to a single block of the at least two blocks.

6. The method according to claim 1, further comprising: Receive input data; as well as The compressed data is generated by at least compressing the input data.

7. The method of claim 1, wherein analyzing the compressed data to determine the data segments comprises: The compressed data is analyzed to determine the boundaries between the data segments of the compressed data.

8. The method according to claim 1, further comprising: Identify dictionary segments of the dictionary associated with the compressed data; Generate second metadata for the dictionary segments, the second metadata indicating the initial dictionary position corresponding to each dictionary segment in the dictionary segments; as well as Associate the second metadata with the compressed data.

9. The method according to claim 8, further comprising: The dictionary segments are decompressed in parallel according to the initial dictionary position of each dictionary segment to generate the dictionary; as well as Using the dictionary, the data segments are decompressed in parallel according to the input segment position and the output segment position to generate a decompressed output.

10. A method comprising: Identify dictionary segments that are associated with compressed data; Generate metadata for the dictionary segments, the metadata indicating the initial dictionary position corresponding to each dictionary segment in the dictionary segments; as well as Associate the metadata with the compressed data.

11. The method of claim 10, wherein the metadata further indicates a symbol index for each dictionary segment in the dictionary segments.

12. The method of claim 10, further comprising: The dictionary segments are decompressed in parallel based on their initial dictionary positions to generate the dictionary.

13. The method of claim 10, further comprising: The compressed data is analyzed to identify data segments of the compressed data, wherein the dictionary segments correspond to the data segments; A second metadata is generated at least in part based on the data segments, the second metadata indicating the input segment position of each data segment in the data segments; as well as Associate the second metadata with the compressed data.

14. The method of claim 13, wherein identifying the dictionary segments comprises: The dictionary is divided into dictionary segments, at least in part, based on the data segments.

15. The method of claim 13, further comprising: The dictionary segments are decompressed in parallel according to the initial dictionary position of each dictionary segment to generate the dictionary; as well as Using the dictionary, the data segments are decompressed in parallel according to the input segment position of each data segment to generate a decompressed output.

16. The method of claim 10, further comprising: Receive input data; as well as The compressed data is generated at least by compressing the input data.

17. A system comprising: One or more processing units are configured to generate metadata for data segments identified by analyzing compressed data, the metadata indicating the location of input segments within the compressed data and the location of output segments in output data corresponding to each data segment in the compressed data.

18. The system of claim 17, wherein the one or more processing units are further configured to: generate second metadata, the second metadata indicating an initial dictionary position corresponding to each dictionary segment of the dictionary associated with the compressed data.

19. The system of claim 17, wherein the one or more processing units are further configured to: decompress the data segments in parallel according to the input segment position and the output segment position to generate a decompressed output.

20. The system of claim 17, wherein the system comprises at least one of the following: Control systems for autonomous or semi-autonomous machines; Sensing systems for autonomous or semi-autonomous machines; A system used to perform simulation operations; A system used to perform deep learning operations; A system for performing real-time streaming broadcasts; Systems used to perform video surveillance services; Systems for performing intelligent video analytics; Systems implemented using edge devices; A system for generating graphical output using ray tracing; A system containing one or more virtual machines (VMs); A system that is at least partially implemented in a data center; or A system that utilizes cloud computing resources at least in part.