A decompression engine for decompressing compressed input data comprising a plurality of data streams

By using a hardware decompression engine for parallel decoding and decompression, the problem of low decompression efficiency in electronic devices is solved, achieving more efficient data decompression and improving the overall performance of electronic devices.

CN114222973BActive Publication Date: 2025-12-09ADVANCED MICRO DEVICES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080056956.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-19
Filing Date
2020-08-17
Publication Date
2025-12-09
Estimated Expiration
2040-08-17

AI Technical Summary

Technical Problem

Existing electronic devices are inefficient at decompressing compressed data, requiring a large number of computational operations and memory accesses, which leads to performance degradation.

Method used

It employs a hardware decompression engine, which includes N decoders and one decompressor. Through parallel decoding and decompression technology, it can quickly decompress compressed data consisting of N data streams.

Benefits of technology

Improving decompression efficiency reduces memory access and power consumption, freeing up other functional blocks in the electronic device to perform other operations, thus enhancing overall performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114222973B_ABST
    Figure CN114222973B_ABST
Patent Text Reader

Abstract

An electronic device having a decompression engine including N decoders and a decompressor decompresses compressed input data including N data streams. Upon receiving a command to decompress the compressed input data, the decompression engine causes each of the N decoders to individually and substantially in parallel with the other decoders of the N decoders to decode a respective one of the N streams from the compressed input data. Each decoder outputs a respective type of decoded data stream used to generate a command associated with a compression standard for decompressing the compressed input data. The decompressor next generates a command from the decoded data streams output by the N decoders to decompress the data using the compression standard to reconstruct the original data. The decompressor next executes the command to reconstruct the original data and store the original data in a memory or provide the original data to another entity.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Related Art

[0002] Some electronic devices perform operations for compressing data, such as user or system files, data streams or sequences, and the like. For example, electronic devices can compress data to reduce the size of the data, to enable more efficient storage of the data in memory, transmission of the data between electronic devices over a network, and the like. Many of these electronic devices use compression standards, such as compression standards based on Lempel-Ziv 77 (LZ77) or LZ78, and / or encoding standards, such as Huffman coding, to generate compressed data from original data.

[0003] Although compressing data can reduce the effort and energy involved in storing and handling the data, compressed data must be decompressed before it can be used for many operations. This means that the electronic device must perform operations to reverse the effects of the compression standard and / or encoding standard's operations and thus recover the original data before such operations can be performed. In many electronic devices, data is decompressed in software (i.e., using software routines or applications), which typically requires a general purpose processor to perform a large number of computational operations and corresponding memory accesses for decompressing the compressed data. Decompressing data in software is inefficient due to the large number of computational operations and memory accesses. BRIEF DESCRIPTION OF DRAWINGS

[0004] Figure 1 A block diagram is presented showing compressed data including data streams, according to some embodiments.

[0005] Figure 2 A block diagram is presented showing an electronic device, according to some embodiments.

[0006] Figure 3 A block diagram is presented showing a decompression subsystem, according to some embodiments.

[0007] Figure 4 A block diagram is presented showing a decoder subsystem, according to some embodiments.

[0008] Figure 5 A block diagram is presented showing a decompressor subsystem, according to some embodiments.

[0009] Figure 6 A flow diagram is presented showing a process for decompressing compressed input data including multiple streams, according to some embodiments.

[0010] Throughout the drawings and the description, same drawing reference numerals refer to same elements. DETAILED DESCRIPTION

[0011] The following description presents the embodiments so that one skilled in the art can make and use the described embodiments and in the context of the particular applications and requirements. Those skilled in the art will readily understand various modifications to the described embodiments and the generic principles defined herein can be applied to other embodiments and applications without departing from the scope of the described embodiments. Thus, the described embodiments are not intended to be limited to the described embodiments but are to be accorded the widest scope consistent with the principles and features disclosed herein.

[0012] Terminology

[0013] In the following description, various terms are used to describe embodiments. The following is a simplified and general description of one of these terms. It should be noted that this term can have a large number of additional aspects, which are not listed here for the sake of clarity and conciseness and therefore the description is not intended to limit the term.

[0014] Functional block: A functional block refers to a group, collection, and / or set of one or more interrelated circuit elements such as integrated circuit elements, discrete circuit elements, and the like. Circuit elements are "interrelated" in that they share at least one property. For example, interrelated circuit elements can be included in, fabricated on, or otherwise coupled to a particular integrated circuit chip or portion thereof, can participate in the execution of a given function (computing or processing function, memory function, and the like), can be controlled by common control elements and / or common blocks, and the like. A functional block can include any number of circuit elements, from a single circuit element (e.g., a single integrated circuit logic gate) to millions or billions of circuit elements (e.g., an integrated circuit memory).

[0015] Compressed data

[0016] In described embodiments, operations are performed on compressed data and using compressed data. Generally, compressed data is output from one or more compression, encoding, and / or other operations on original data that result in at least some of the original data being replaced with data that can be used to reconstruct the original data (to perform lossless compression), with data that is similar to the original data (to perform lossy compression), and / or with other values. In described embodiments, various types of data can be compressed, including user or system files (e.g., audio and / or video files, text files, documents, executable files, operating system files, spreadsheets, etc.), data streams or sequences (e.g., audio and / or video data streams, sequences of data received through a network interface, etc.), data captured from sensors (e.g., cameras and / or speakers, thermometers, vibration sensors, etc.), etc. In described embodiments, numerous compression, encoding, and / or other operation standards, algorithms, or formats, or combinations thereof, can be used to compress data, including prefix coding standards, dictionary coding compression standards, delta encoding, autoregressive model compression, etc.

[0017] As used herein, the terms "compressed data" and "compression" are used broadly to refer to operations on original data that result in at least some of the original data being replaced with other data that can be used to fully or approximately reconstruct the original data. As noted above, these operations include various compression, encoding, and / or other operation standards, algorithms, or formats, or combinations thereof. These data should therefore not be interpreted as being limited to only operations such as dictionary coding compression and / or other operations that can sometimes be considered "compression" operations.

[0018] In some embodiments, when compressing data, the electronic device creates compressed data that includes N separate data streams, where N is greater than or equal to one. The N data streams are independent chunks or portions of the compressed data from one another and include various types of data that can be decoded, for example, in a decoder during a decompression operation and then used to generate commands associated with the compression standard for decompressing the compressed input data. In some embodiments, each stream is self-contained in that each stream can be decoded without involving data other than the data of the stream. Figure 1 A block diagram showing compressed data including data streams according to some embodiments is presented. As can be seen in Figure 1 In particular, the compressed data 100 (which can be or include a file, a data stream or sequence, etc.) includes three data streams: streams 102-106, each of which is in Figure 1The data streams in the compressed data 100 are labeled using different padding patterns (i.e., diagonal padding for stream 102, etc.). Each data stream includes a plurality of data bytes or blocks of a respective type. For example, assume that the compressed data 100 is compressed using a dictionary coding compression standard, each stream includes elements of the dictionary coding compression standard to be used to generate commands associated with the dictionary coding compression standard for decompressing compressed input data such as literals, command tags, distances, lengths, and / or other values. For this example, literals are strings that have not been detected as being repeated within a specified number of bytes of the original data (i.e., the data is compressed), command tags are values that identify specified commands for performing decompression operations (e.g., adding / appending one or more literals to reconstructed original data, etc.), and lengths and distances are used to indicate repeated strings to be added / appended to reconstructed original data. For example, stream 102 can include (and possibly only include) literals, stream 104 can include (and possibly only include) command tags, and stream 106 can include (and possibly only include) distances or lengths.

[0019] In some embodiments, Figure 1 Each data stream in the compressed data 100 begins with a header, i.e., headers (HD) 108, 112, and 116, respectively, that includes a description of the content, formatting, source electronic device, routing values, and / or length of the stream, etc. In some embodiments, each data stream includes metadata, i.e., metadata (MD) 110, 114, and 118, respectively, that is or includes information such as a decoding reference for decoding the respective stream from the compressed data, information about the stream from the compressed data, etc. The header and metadata of each stream are followed by the body of the stream, which includes compressed data of the respective type. Each of streams 102, 104, and 106 begins at a given location in the compressed data, e.g., at a given byte, frame, or block, and follows the previous stream except for stream 102. For example, the header 108 of stream 102 can be at byte 0 of the compressed data (e.g., at a starting address in memory, a first location in a total data stream or sequence, etc.), and the length of the combination of header 108, metadata 110, and stream 102 can be 50 kB, and thus the header 112 of stream 104 can be at 50 kB, and so on. Figure 1 The streams in the compressed data 100 span Figure 1 The lines in the compressed data 100, so stream 102 includes a portion in the first line of compressed data 100 and a portion in the second line of compressed data 100; this is for illustration only. As noted above, a stream is a sequence of data, such as bytes or blocks retrieved from a file, a data stream or sequence received over a network interface, etc.

[0020] In some embodiments, the electronic device generates compressed data (e.g., compressed data 100) from raw data by first compressing the raw data using a compression standard. For example, in some embodiments, the electronic device compresses the raw data using a dictionary-based compression standard such as a compression standard based on LZ77 or LZ78 to generate the compressed data. During or after the compression operation, the electronic device divides the compressed data into N streams (e.g., streams 102-106), each of the N streams including data of a respective type associated with the compression standard used to generate the compressed input data as well as headers and possibly metadata. For example, in some embodiments, for the LZ77 or LZ78 based compression standard described above, the data types include some or all of literals, command tags, distances, and lengths, and thus the individual streams can include and possibly only include literals, command tags, distances, and lengths. The electronic device next encodes each of the N streams using an encoding standard to further compress each stream. For example, in some embodiments, the electronic device uses a prefix-based encoding standard such as Huffman coding for encoding the streams.

[0021] Although a particular sequence of compression and encoding is described as being used to create compressed data and specific instances of compressed data are shown in Figure 1 FIGS. 1-3, the described embodiments are not limited to the sequence and / or instances of compressed data. In general, any combination of compression and / or encoding can be used to produce any version of compressed data from raw data, so long as the compressed data can be decompressed in a decompression engine as described herein.

[0022] SUMMARY

[0023] In the described embodiments, the electronic device performs operations for decompressing compressed data that includes N data streams. The electronic device includes a hardware decompression engine that has N decoders and one decompressor. In other words, the decompression engine (which is itself a functional block) includes functional blocks for the N decoders and one decompressor. The N decoders and one decompressor in the hardware decompression engine are used to decompress the compressed data that includes the N data streams.

[0024] In some embodiments, when decompressing compressed data that includes N data streams, the decompression engine first receives a command or request. The command or request identifies the compressed data (e.g., identifies a starting address in memory of the compressed data, a location of a stream or sequence from which the compressed data can be fetched, etc.) and requests decompression of the compressed data. The command or request can also identify configurations and settings for compressing data used by the decompression engine when decompressing the compressed data. The command or request can also include or identify certain values, such as initial decompression data or other values used by the decompression engine when decompressing the compressed data.

[0025] Upon receiving the command or request, the decompression engine causes each of the N decoders to individually decode a respective data stream of the N data streams substantially in parallel using the coding standard (e.g., the prefix coding standard and / or another coding standard). For the decoding operation, the decompression engine communicates a position of the first stream to a stream header decoder in a first decoder of the N decoders (i.e., the decoder of the N decoders selected to decode the first stream). The stream header decoder decodes a header in the first stream, determines a length of the first stream, communicates the length back to the decompression engine, and the first decoder of the N decoders continues decoding the first stream. The decompression engine uses the length to determine a starting position of the second stream in the compressed data and communicates a position of the second stream to a stream header decoder in a second decoder of the N decoders, which continues as the stream header decoder in the first decoder. For the remaining decoders, the decompression engine continues in this manner, determining a starting position of the respective stream and beginning decoding for each of the decoders. As used herein, “substantially in parallel” thus means that the N decoders decode the respective streams at approximately the same time, as offset by overhead for beginning decoding in each of the N decoders, etc. It is noted that in some embodiments, rather than the above per-stream header instance, a header of a first stream of the N streams (or generally a header of the compressed data) includes length and / or starting position information for all N streams. In these embodiments, the decompression engine obtains the length and / or starting position information for all streams from the header of the first stream (or the header of the compressed data), and can use the respective length and / or starting position information obtained from the header of the first stream (or the header of the compressed data) to start all remaining decoders (or all decoders) at approximately the same time.

[0026] It is recalled that each of the N data streams includes respective types of data for generating commands associated with the compression standard used to compress the compressed input data, and thus the output from each of the N decoders is a respective type of decoded data stream. For example, in embodiments in which the compression standard used to compress the data is a dictionary coding compression standard, the data types can include literals, command tags, distances, lengths, and / or other data types. For this example, in some embodiments, one or more of the N streams includes literals, one or more of the N streams includes command tags, etc. In some embodiments, a stream includes only one particular data type rather than a mix of data types.

[0027] In some embodiments, at least one of the decoders generates, creates, and / or initializes a decoding reference (e.g., one or more tables, lists, and / or other records) used by the decoder for decoding operations. For example, in some embodiments, a given decoder obtains information for generating a decoding reference from a specified location in a stream to be decoded by the given decoder (e.g., from a metadata byte in the stream immediately after a header of the stream, etc.). For example, in embodiments in which a given decoder decodes a stream using Huffman coding, the given decoder can obtain information for initializing / populating a decoding table, which includes a mapping between Huffman codes and symbols (e.g., patterns of bits or bytes) to be used for decoding operations.

[0028] In some embodiments, some or all of the decoders include substream decoders. Each of the substream decoders in the decoders is a different decoding functional block that can be used by the decoders to decode a separate portion of data from a stream being decoded. For example, in embodiments in which a decoder has M substream decoders, each of the M substream decoders obtains and decodes a respective next chunk of data, symbols, etc. from a stream in a round-robin manner. In some of these embodiments, the M substream decoders each decode a next chunk of data, encoded symbols, etc. substantially in parallel. For example, in some embodiments, in a given cycle, time period, etc. of a control clock, each of the M substream decoders decodes at least one chunk of data, symbols, etc. such that all M substream decoders produce one or more decoded symbols or other decoded data segments per cycle of the control clock. In some embodiments, a stream combiner combines outputs from the substream decoders for forwarding to a decompressor.

[0029] After decoding data in each of the N streams, the N decoders forward a stream of decoded data of a respective data type to a decompressor in a decompression engine. The decompressor generates commands for decompressing data using a compression standard to reconstruct the original data from the stream of decoded data output by the N decoders. For this operation, the decompressor obtains data from some or all of the streams for generating each command, and then assembles a command from the data. Continuing the above example, in some embodiments, for each command, the decompressor obtains some or all of a next literal, command tag, distance, length, and / or other data type from a respective stream and assembles a command therefrom. At the end of this operation, the decompressor has a single command that, when executed, will cause the decompressor to obtain and / or generate a next chunk of data (e.g., a P-byte chunk of data) for reconstructing the original data.

[0030] In some embodiments, the decompressor includes an operation combiner that performs operations for combining groups of two or more commands into an aggregate command. In these embodiments, when a command is generated in the decompressor, two or more commands are buffered (i.e., temporarily stored) and then an attempt is made to combine the two or more commands into an aggregate command. For example, commands that access adjacent memory locations, the same cache line, etc. can be combined into an aggregate command. By combining commands, operations for executing the commands (e.g., memory accesses, etc.) and associated with executing the commands can be performed more efficiently.

[0031] The decompressor next executes the command to reconstruct the original data. For this operation, the execution function block in the decompressor executes the command that causes the execution function block to generate and / or fetch the next block of data for reconstructing the original data. For example, and continuing the dictionary coding compression standard example from above, the command can cause the execution function block to fetch the next word or words and add / append the next word to the reconstruction of the original data (i.e., to the output file or stream generated by the decompressor). As another example, and again continuing the dictionary coding compression standard example, the command can cause the execution function block to read a pre-fix of a specified length from an identified location in the original data that has already been reconstructed (i.e., from data output earlier by the execution function block) and add / append the pre-fix to the reconstruction of the original data.

[0032] In some embodiments, the decompressor includes a read manager function block that fetches data from memory for executing a command before the command is executed in the execution function block. In some of these embodiments, the command is buffered to wait for the data to return from memory and is executed when the data returns. In some embodiments, the decompressor includes a history buffer that is a first-in-first-out (FIFO) buffer, a specified number of words and / or pre-fixes are stored in the FIFO buffer and fed back to a related subsequent command. In these embodiments, the words and / or pre-fixes are eventually read from the FIFO buffer and appended to the reconstruction of the original data, i.e., stored as decompressed data in memory, output as a stream or sequence of decompressed data, etc.

[0033] By decompressing compressed data including N data streams using a hardware decompression engine having N decoders and one decompressor, the described implementations perform operations in hardware that existing devices perform using software. The hardware decompression engine is faster and more efficient (e.g., requires fewer memory accesses, uses less power, etc.) than using a software entity for performing the same operations. In addition, using the decompression engine frees other functional blocks (e.g., processing subsystems, etc.) in the electronic device for performing other operations. The decompression engine thus improves the overall performance of the electronic device, which in turn improves user satisfaction.

[0034] Electronic device

[0035] Figure 2 A block diagram showing electronic device 200 is presented in accordance with some embodiments. As can be seen in Figure 2 electronic device 200 includes processor 202 and memory 204. Processor 202 is a functional block that performs computations, decompressions, and other operations in electronic device 200. Processor 202 includes processing subsystem 206 and decompression subsystem 208. Processing subsystem 206 includes one or more functional blocks that perform general-purpose computations, decompressions, and other operations, such as central processing unit (CPU) cores, graphics processing unit (GPU) cores, embedded processors, and / or application-specific integrated circuits (ASICs).

[0036] Decompression subsystem 208 is a functional block that performs operations for decompressing compressed input data including individual data streams. Generally, decompression subsystem 208 takes as input a command or request to decompress a specified compressed input data, such as a file, data stream, or sequence, etc., and returns uncompressed data that fully or approximately reconstructs the original data from which the compressed input data was generated. Decompression subsystem 208 is described in more detail below.

[0037] Memory 204 is a functional block that performs operations of a memory (e.g., main memory) in electronic device 200. Memory 204 includes memory circuitry (i.e., storage elements, access elements, etc.) for storing data and instructions used by functional blocks in electronic device 200, and control circuitry for handling access (e.g., read, write, check, delete, invalidate, etc.) to the data and instructions in the memory circuitry. The memory circuitry in memory 204 includes computer-readable memory circuitry such as fourth generation double data rate synchronous dynamic random access memory (DDR4 SDRAM), static random access memory (SRAM), or a combination thereof.

[0038] The electronic device 200 is shown using a particular number and arrangement of elements (e.g., functional blocks and devices such as the processor 202, the memory 204, etc.). However, the electronic device 200 is simplified for illustrative purposes. In some embodiments, a different number or arrangement of elements is present in the electronic device 200. For example, the electronic device 200 can include a power subsystem, a display, etc. Generally, the electronic device 200 includes sufficient elements to perform the operations described herein.

[0039] Although the decompression subsystem 208 is shown as included in the processor 202, in some embodiments, the decompression subsystem 208 is a separate and / or independent functional block. For example, the decompression subsystem 208 can be implemented (by itself or with supporting circuit elements and functional blocks) on a separate integrated circuit chip, etc. Generally, in the described embodiments, the decompression subsystem 208 is suitably located in the electronic device 200 to enable performance of the operations described herein. Figure 2

[0040] The electronic device 200 can be or can be included in any electronic device that performs data decompression or other operations. For example, the electronic device 200 can be or can be included in an electronic device such as a desktop computer, a laptop computer, a wearable electronic device, a tablet computer, a smartphone, a server, an artificial intelligence device, a virtual or augmented reality device, a network appliance, a toy, an audiovisual device, a household appliance, a controller, a vehicle, etc., and / or combinations thereof.

[0041] Decompression engine

[0042] In the described embodiments, a decompression subsystem in an electronic device performs operations for decompressing compressed input data. Figure 3 A block diagram showing the decompression subsystem 208 is presented in accordance with some embodiments. As can be seen in Figure 3 The decompression subsystem 208 includes a decompression engine 300 and a memory interface 302. The decompression engine 300 is a functional block that performs operations for decompressing compressed input data and associated with decompressing compressed input data. The decompression engine 300 includes a decoder subsystem 304 and a decompressor subsystem 306 that are functional blocks each performing respective portions of operations for decoding encoded data and decompressing compressed data and associated with decoding encoded data and decompressing compressed data. The decoder subsystem 304 and the decompressor subsystem 306 are described in greater detail below.

[0043] ​Memory interface 302 is a functional block that performs operations for and associated with memory accesses (e.g., memory reads, writes, invalidations, deletions, etc.) by decompression engine 300. For example, in some embodiments, original data reconstructed from compressed input data is stored in memory (e.g., memory 204) by memory interface 302. As another example, in some embodiments, compressed input data is retrieved or fetched from memory by memory interface 302 in preparation for decompressing the compressed input data. Although memory interface 302 is shown in Figure 3 FIG. 1, in some embodiments, decompression subsystem 208 includes and uses one or more additional or different interfaces, such as network interfaces, IO device interfaces, inter-processor communication interfaces, etc. Generally, in the described embodiments, decompression subsystem 208 includes one or more functional blocks or devices for retrieving compressed input data, for writing out reconstructed original data, and / or for exchanging communications (e.g., commands, etc.) with other functional blocks in an electronic device (e.g., electronic device 200) in which decompression subsystem 208 is located.

[0044] It should be noted that the following discussion for the example in Figures 3-5 assumes that compressed input data is fetched from memory (e.g., memory 204) for decompression and original data reconstructed from the compressed input data is stored in memory 204. This is not a requirement. In some embodiments, compressed data is fetched from other sources and / or stored in or provided to other destinations in electronic device 200, etc. (e.g., via network interfaces, IO devices, etc.). Additionally, for the example in Figures 3-5 , it is assumed that original data is processed as described above for generating compressed input data, i.e., first compressed using a dictionary coding compression standard (e.g., a standard based on LZ77 or LZ78), then partitioned into N streams by data type (e.g., literal, command tag, distance, length, etc.) of the dictionary coding compression standard, and finally individual streams are separately encoded using a prefix coding standard (e.g., Huffman coding). This too is not a requirement. In some embodiments, different compression, encoding, and / or other operation standards, algorithms, or formats or combinations thereof can be used for generating compressed input data, and decompression and / or operations otherwise reversed from those described similarly. Figures 3-5

[0045] Figure 4 ​A block diagram showing decoder subsystem 304 is presented in accordance with some embodiments. Decoder subsystem 304 includes a plurality of individual decoders 400-404. Each of decoders 400-404 performs operations for decoding a given type of data to generate commands associated with a compression standard for decompressing compressed input data (e.g., literals, command tags, distances, etc.). In some embodiments, each decoder 400-404 is dedicated to decoding a given type of data, such as by having circuit elements designed for decoding the given type of data included in the decoder. In some embodiments, each of decoders 402-404 includes internal elements similar to those shown in decoder 400 (which are not shown for brevity), although this is not a requirement. In general, each decoder includes sufficient internal elements to decode one or more types of data.

[0046] As can be seen in Figure 4 Decoder 400 includes a read manager 406, which is a functional block that performs operations associated with fetching compressed input data 432 from memory and directing portions of compressed input data 432 to substream decoders 408-412. Substream decoders 408-412 are each a functional block that decodes and otherwise processes a respective portion of a stream being decoded in decoder 400. Decoder 400 also includes a stream header decoder (SHD) 414, which is a functional block that performs operations for processing information from a header of a stream being decoded in decoder 400. Decoder 400 also includes buffers 416-420, which are functional blocks that perform operations for buffering (i.e., temporarily storing) portions of a stream being decoded in decoder 400 and providing the portions of the stream to a corresponding substream decoder (e.g., buffer 416 provides portions of a stream to substream decoder 408, etc.). Decoder 400 also includes a stream combiner 422, which is a functional block that performs operations for combining individual decoded data portions output from substream decoders 408-412 into a single decoded data stream for output from decoder 400 to decompressor subsystem 306. Decoder 400 also includes a decoded reference builder (DRB) 424, which is a functional block that performs operations for initializing and populating a decoded reference (DREF) 426 from which information is fetched and used by substream decoders 408-412 to decode portions of a stream.

[0047] During operation of the decoder subsystem 304, a command header decoder (CHD) 428 receives a command 430 (which can be a message, packet, request, and / or other form of communication or included in the communication) that identifies and requests decompression of compressed input data 432. For example, in some embodiments, the command 430 identifies the compressed input data 432 by an absolute or relative address, location, and / or reference in a memory in which the compressed input data 432 is stored. In some embodiments, the command 430 also includes information that identifies configurations and settings used to compress the compressed input data 432, information that identifies configurations and settings to be used by the decoder subsystem 304 and / or decompressor subsystem 306 to decode and / or decompress the compressed input data 432. For example, in some embodiments, when configurable or selectable encoding and / or compression options are used when compressing the compressed input data 432, the command 430 can include an indication (e.g., in units of specified bits, etc.) of how the options were configured or selected. In some embodiments, the command 430 also includes or identifies information to be used by the decoder subsystem 304 and / or decompressor subsystem 306 to decode and / or decompress the compressed input data 432. For example, in some embodiments, the command 430 includes a dictionary code map (e.g., included in a packet or message with the dictionary code map) that is used to initialize a decompression reference, such as a table, etc., in the decompressor subsystem 306. In some embodiments, the command header decoder 428 communicates information (e.g., the dictionary code map, etc.) from or associated with the command 430 to the decompressor subsystem 306.

[0048] Upon receiving the command 430, in some embodiments, the command header decoder 428 communicates to the read manager (MGR) 406 in the decoder 400 (which is the first / initial decoder to start decoding the streams from the compressed input data 432) the address or location in memory where the first stream (e.g., stream 102 in compressed data 100) among the N streams in the compressed input data 432 begins. Using the address or location in memory, the read manager 406 starts fetching the first stream in the compressed input data 432 from memory. As the first stream begins to return from memory, the read manager 406 forwards the initial portion to the buffer 416 (i.e., stores a specified number of bytes from the first stream in memory elements in the buffer 416). The stream header decoder 414 gets at least some of the initial portion of the first stream from the buffer 416 and processes the information from the header (e.g., header 108 from stream 102) of the first stream included in the initial portion. Among the information from the header of the first stream is the length of the first stream, which the stream header decoder 406 gets and sends to the command header decoder 428. The command header decoder 428 computes or determines the address or location in memory of the second stream (e.g., stream 104) among the N streams in the compressed input data 432 based on the length of the first stream and communicates to the read manager in the decoder 402 (which is the second decoder to start decoding the streams from the compressed input data 432) the address or location in memory where the second stream in the compressed input data begins. The decoder 402 then starts decoding the second stream using similar operations as those performed by the decoder 400. In this way, the command header decoder 404 initiates the decoders 400-404 in a cascading or daisy chain manner, where each decoder after the first decoder is initiated based on information obtained by the command header decoder 428 from the previous decoder. In some embodiments, the decoders 400-404 decode the respective streams “substantially in parallel” rather than completely in parallel due to the offset in the start time of each of the decoders 400-404 in determining the length of each of the N streams.

[0049] In some embodiments, unlike the above example where the header of each stream includes length information for determining the starting position of the next stream in the compressed data, the header of the first stream (or generally the header of the compressed data) among the N streams includes length and / or starting position information for all N streams. In these embodiments, the command header decoder 428 obtains the length and / or starting position information for all streams from the header of the first stream (or generally the header of the compressed data) and can use the respective length and / or starting position information obtained from the header of the first stream (or the header of the compressed data) to initiate all the remaining decoders (i.e., decoders 402-404) (or all the decoders) at approximately the same time.

[0050] In some embodiments, the stream header decoder 414 also obtains information about the configuration and / or settings used to encode and / or compress the first stream from the stream header information and / or metadata in the initial portion of the first stream (e.g., the header 108 and / or metadata 110). For example, in some embodiments, the stream header decoder 414 obtains information about the format or arrangement of the data in the first stream, which can not be the same format or arrangement as the data in other streams in the compressed input data 432. The stream header decoder 414 then communicates the information about the configuration and / or settings to the decompressor subsystem 306 for use therein, i.e., so that the decompressor subsystem 306 can take the configuration and / or settings into account when processing data from the first stream.

[0051] In some embodiments, the decoding reference builder 424 in the decoder 400 also obtains at least some of the initial portion of the first stream from the buffer 416 and processes the metadata (e.g., the metadata 110) of the first stream included in the initial portion to obtain information for initializing and populating the decoding reference 426. For example, in some embodiments, the metadata in the first portion of the first stream includes a mapping between codes and symbols (e.g., patterns of bits or bytes) to be used for decoding the stream in each of the sub-stream decoders 408-412.

[0052] Along with forwarding the initial portion of the first stream to buffer 416, read manager 406 forwards a corresponding portion of the first stream to buffers 418 and 420. Respective blocks or portions of data (e.g., 32-bit data blocks per clock cycle, etc.) are fed from buffers 416-420 into each of the corresponding sub-stream decoders 408-412. Each of sub-stream decoders 408-412 decodes the blocks or portions of data received using a corresponding decoding technique using information from the decode reference 426 to generate one or more symbols. In some embodiments, sub-stream decoders 408-412 operate substantially in parallel when decoding respective blocks or portions of data. For example, in some embodiments, each sub-stream decoder decodes a respective block or portion of data every control clock cycle, and thus produces at least one decoded symbol (i.e., decoded data) for every control clock cycle. Each of sub-stream decoders 408-412 then outputs the respective decoded data to stream combiner 422. Stream combiner 422 receives a stream of symbols from each of sub-stream decoders 408-412 and combines the streams of symbols from all decoders 408-412 to generate a decoded data stream that is output from decoder 400. For example, in some embodiments, stream combiner 422 takes the next symbol from each of decoders 408-412 in a round-robin manner and appends the next symbol to an aggregated or combined decoded data stream. Stream combiner 422, and more broadly decoder 400, then outputs the decoded data stream to decompressor subsystem 306. Additionally, each of decoders 402-404 (e.g., a stream combiner or another functional block therein) outputs a respective decoded data stream to decompressor subsystem 306 (as indicated by the arrows from decoders 402-404 on the right side of FIG. 4). Figure 4

[0053] In some embodiments, one or more of decoders 402-404 are secondary decoders that decode secondary streams from among the N streams in the compressed input data. Decoding of secondary streams is less complex than that of primary streams, and thus secondary decoders are simplified and can include fewer internal elements. For example, in some embodiments, secondary streams include additional values having a fixed length (i.e., a fixed number of bits or bytes) of information such as words, command tags, lengths, and / or distances and / or portions thereof (with the remaining portions included in primary streams), and thus secondary streams can be decoded simply by retrieving the fixed length data groupings or blocks from the secondary streams.

[0054] ​In some embodiments, some or all of the streams include raw or unencoded stream data. For example, in some cases, an encoded stream makes the size of the stream larger due to the effects of encoding, and thus the stream is not encoded by the electronic device that compressed the raw data. In some of these embodiments, some or all of the buffers 416-420 can bypass the corresponding substream decoder (i.e., not perform a decoding operation in the corresponding substream decoder) and pass the raw or unencoded stream elements directly to the stream combiner 422 and / or the decompressor subsystem 306.

[0055] Figure 5 A block diagram showing the decompressor subsystem 306 is presented in accordance with some embodiments. The decompressor subsystem 306 includes internal elements for receiving decoded data output by the decoder subsystem 304 and using the decoded data to generate commands for decompressing compressed input data. The decompressor subsystem 306 also includes internal elements for executing the commands to reconstruct the raw data from which the compressed input data was generated and storing the reconstructed raw data in memory.

[0056] As in Figure 5As can be seen, the elements in the decompressor subsystem 306 include a command assembler 500, which is a functional block that performs operations for assembling (i.e., generating, creating, etc.) commands from information obtained from buffers 502-504, which are functional blocks that each buffer respective values such as command tags, lengths, etc. for assembling commands. The decompressor subsystem 306 also includes an operation combiner (OP COMB) 508, which is a functional block that performs operations for combining commands (or operations resulting therefrom) into new aggregate commands for subsequent execution in the decompressor subsystem 306 where possible. The decompressor subsystem 306 also includes a buffer 506 and a literal buffer 510, which are functional blocks that perform operations for buffering literals in decoded data output by the decoder subsystem 304 for use in the decompressor subsystem 306. The decompressor subsystem 306 also includes a read scheduler (SCH) 512 and a read return (RETRN) 514, which are functional blocks that perform operations for obtaining data from memory to be used in reconstructing original data when executing commands (e.g., strings from the reconstructed original data that were previously written to memory). The decompressor subsystem 306 also includes an operation controller (OP CTRLR) 516 and a command executor (CMD EXE) 518, which are functional blocks that control when commands are executed and execute commands, respectively. The decompressor subsystem 306 also includes an operation (OP) cache 520 and a history buffer 522, which are functional blocks that perform operations for caching blocks (e.g., 32 byte blocks, etc.) of commands and reconstructed original data to be used in executing commands and / or feeding back to subsequent commands as well as for the history buffer 522 from which the reconstructed original data is written out to memory.

[0057] During operation of decompressor subsystem 306, decoded data is received by each of buffers 502-506 (e.g., command tags, lengths, and literals, respectively). Command tags and lengths are retrieved (e.g., in a first-in, first-out order) from buffers 502-504 as needed by command assembler 500 and used by command assembler 500 to generate commands for reconstructing the original data. For example, based solely on a command tag (i.e., based on information or values included in the command tag), command assembler 500 can generate a command to fetch the next literal or literals and add / append them to the reconstructed original data. As another example, based on a command tag and a length and / or distance, command assembler 500 can generate a command to fetch a chunk or block of a specified size (e.g., number of bits or bytes) of data from an identified location in previously reconstructed original data and add / append the fetched chunk or block of data to the reconstructed original data. It should be noted that in some embodiments, there are two or more command assemblers (as shown by the additional blocks following command assembler 500) and the command assemblers operate substantially in parallel for generating commands as described above.

[0058] The generated commands are forwarded from command assembler 500 to operation combiner 506, which collects groups of two or more individual commands and attempts to combine the two or more commands into one or more aggregate commands. Each aggregate command performs all of the operations of the two or more individual commands, but can do so with less computational effort (from command executor 518), less memory and / or cache memory accesses, etc. In some embodiments, when an aggregate command cannot be generated from two or more commands, the two or more commands themselves are output separately from operation combiner 506.

[0059] The commands or aggregate commands (collectively referred to as "commands" for the remainder of this example for clarity and brevity) are forwarded from operation combiner 506 to read scheduler 512. Read scheduler 512 analyzes the commands to determine whether a memory access is needed to fetch a chunk or block of data from reconstructed data that has already been written to memory. For example, is the command to copy a chunk or block of data from previously reconstructed original data to be used to perform the current command. When a memory read is to be performed to fetch a block or chunk of data, read scheduler 512 sends a read request to memory for the block and chunk of data.

[0060] Commands pass from the read scheduler 512 to the operation controller 516, which determines whether the command can be executed or the command is to be held pending data return from memory. When the command can be executed, the operation controller 516 forwards the command to the command executor 518 for execution, or stores the command in the operation cache 520 and causes the command executor 518 to fetch the command from the operation cache 520 for execution. Otherwise, when the command is to be held pending data return from memory, the operation controller 516 forwards the command to the operation cache 520 to be stored while awaiting return of the data. When the data subsequently returns, the read return 514 forwards the data to the operation controller 516, which stores the data in the operation cache 520 and causes the command executor 518 to fetch the command and data from the operation cache 520 for execution.

[0061] Executing the command causes the command executor 518 to perform the operation indicated by the command. For example, in some embodiments, the command is to add / append one or more literals to the end of the reconstructed original data, and thus the command executor 518 fetches the one or more literals from the literal buffer 510 (or the operation cache 520 or the history buffer 522) and adds / appends the one or more literals to the end of the reconstructed original data. In other words, the command executor 518 outputs the data (e.g., bits or bytes) of the literals fetched from the literal buffer 510 to the operation cache 520 in order, and the data of the literals is eventually written from the operation cache 520 to the file, stream, etc. in memory of the reconstructed original data. As another example, in some embodiments, a given command is to fetch a string of a certain length from a distance back into the original data that has been reconstructed, and add / append the string to the end of the reconstructed original data, and thus the command executor 518 performs a memory access to fetch the string from the original data that has been reconstructed. Since the read scheduler 512 fetched the string from memory earlier, the memory access should hit the operation cache 520 (or the history buffer 522, as described below), and the command executor 518 fetches the string and adds / appends the string to the end of the reconstructed original data.

[0062] In some embodiments, the history buffer 522 is used to temporarily store and feed back recently output data (e.g., literals and strings) output by the command executor 518. In these embodiments, as chunks of data (e.g., T-bit literals, K-bit or byte strings) are output from the command executor 518, the chunks of data are written to the history buffer (perhaps through the operation cache 520) in a first-in-first-out order. The chunks of data are held in the history buffer 522 for a given amount of time or until the history buffer 522 is full, and then written out to memory in a first-in-first-out order. As noted above, while resident in the history buffer 522, the chunks of data are available for feeding back to subsequent commands.

[0063] Although Figure 3 Although only one decompression engine in the decompression subsystem 208 is shown, in some embodiments, the decompression subsystem 208 includes multiple decompression engines. In these embodiments, the multiple decompression engines can be used (e.g., substantially in parallel) to decompress portions of the compressed input data or can be used to decompress different / separate compressed input data. For example, in some embodiments, each of the multiple decompression engines can be used to decompress a respective segment or section of the same compressed input data.

[0064] Although Figure 4 Although three decoders 400-404 are shown, in some embodiments, a different number of decoders are used. Additionally, in some embodiments, a different number of other internal elements are present in the decompression engine, such as the command assembler 500, the command executor 518, etc. Also, in some embodiments, a different number of substream decoders and buffers are used in the decoder subsystem 304 and the decompressor subsystem 306. This is illustrated in Figures 4-5 by the ellipses between the substream decoders and the buffers. Generally, in the described embodiments, the decompression engine includes sufficient internal elements to perform the operations described herein.

[0065] Process for decompressing compressed input data

[0066] In the described embodiments, a decompression engine (e.g., the decompression engine 300) that includes a decoder subsystem and a decompressor subsystem (e.g., the decoder subsystem 304 and the decompressor subsystem 306) performs operations for decompressing compressed input data that includes one or more separate data streams (e.g., the streams 102-106 in the compressed data 100). Figure 6 A flowchart showing a process for decompressing compressed input data that includes multiple streams is presented in accordance with some embodiments. It should be noted that, Figure 6The illustrated operations are presented in general examples of operations performed by some embodiments. Operations performed by other embodiments include different operations, operations performed in different orders, and / or operations performed by different entities or functional blocks.

[0067] For the example in Figure 6 , it is assumed that the input data has been compressed using a combination of dictionary coding compression (e.g., using a compression standard based on LZ77 or LZ78) and prefix coding (e.g., using a Huffman coding standard). However, this is not a requirement. For example, in some embodiments, the data is first encoded, then split into streams, and finally compressed, as opposed to the example in Figure 6 . In general, the described embodiments are operable with any combination of compression and / or encoding that can be decompressed using a decompression engine as described herein. Additionally, during the compression process, for example, after or during compression and before encoding, it is assumed that the compressed data has been split into N data streams. As described herein, each of the N data streams includes and possibly only includes data of a respective type used to generate commands associated with the compression standard for decompressing the compressed input data. For example, in some embodiments, the data types include some or all of literals, command tags, distances, lengths, and / or other data associated with a dictionary coding compression standard, and thus each of the N streams includes at least and possibly only one of these data types.

[0068] As can be seen in Figure 6 , the process begins when the decompression engine receives a command to decompress the compressed input data that includes N data streams (step 600). For this operation, the decompression engine receives a packet, message, or signal that includes the command and / or an identification of the compressed input data over a signal line, bus, or other communication interface. For example, in some embodiments, the identification of the command includes one or more bits or bytes organized in a pattern that identifies the command. As another example, in some embodiments, the identification of the compressed input data includes a location from which to obtain the compressed input data, such as an address in a memory, an identifier for a network interface device, etc.

[0069] The decompression engine then causes each of N decoders (e.g., decoders 408-412) in the decompression engine to individually and substantially in parallel decode a respective one of N streams of compressed input data (step 602). This operation involves the decompression engine (e.g., command header decoder, stream header decoder, etc.) determining a location in the compressed input data for each of the N streams and causing a different decoder among the N decoders to undertake decoding each stream. Decoding each stream includes each decoder (or another entity) performing a lookup in a respective decoding reference (e.g., decoding reference 426) to determine a symbol associated with a block of decoded data in each stream and outputting the symbol.

[0070] The decoders each output a stream of decoded data (i.e., symbols) of a respective type of compression standard to the decompressor (e.g., decompressor subsystem 306) (step 604). This operation involves the decoders outputting decoded data of a respective type for use in generating commands associated with the compression standard for decompressing the compressed input data. For example, one of the decoders can output decoded data including a command tag to be used, alone or in combination with other information (e.g., length, distance, etc.) output by one or more other decoders, for generating a command for decompressing the compressed input data.

[0071] The decompressor then generates commands from the streams of decoded data for decompressing the data using the compression standard to reconstruct the original data (step 606). As described above, for this operation, the decompressor takes chunks or portions of data (e.g., W-bit portions) from some or all of the streams and uses and / or combines the chunks or portions of data to create the commands. For example, the decompressor can take a command tag from one of the streams of decoded data that identifies a copy command that will cause the enactor (e.g., command enactor 518) to add / append a copy of a previous string from early in the reconstruction of the original data to the end of the reconstruction of the original data. In this case, the decompressor also takes a distance (e.g., number of bytes back into the original data) from a second one of the streams and / or a length (e.g., number of bytes in the string) from a third one of the streams to combine with the command tag.

[0072] The decompressor then executes the commands to reconstruct the original data (step 608). As described above, this operation involves the enactor (e.g., command enactor 518) in the decompressor executing the commands to generate and / or take strings and / or literals for adding / appending to the reconstruction of the original data. The decompressor then stores the reconstructed original data in memory (step 610).

[0073] In some embodiments, at least one electronic device (e.g., electronic device 200) uses code and / or data stored on a non-transitory computer-readable storage medium to perform some or all of the operations described herein. More specifically, at least one electronic device reads code and / or data from the computer-readable storage medium when performing the described operations and executes the described code and / or uses the described data. The computer-readable storage medium can be any device, medium, or combination thereof that stores code and / or data for use by an electronic device. For example, the computer-readable storage medium can include, but is not limited to, volatile and / or non-volatile memory, including flash memory, random access memory (e.g., eDRAM, RAM, SRAM, DRAM, DDR4 SDRAM, etc.), non-volatile RAM (e.g., phase change memory, ferroelectric random access memory, spin-transfer torque random access memory, magnetoresistive random access memory, etc.), read-only memory (ROM), and / or magnetic or optical storage mediums (e.g., disk drives, magnetic tape, CDs, DVDs, etc.).

[0074] In some embodiments, one or more hardware modules perform the operations described herein. For example, a hardware module can include, but is not limited to, one or more processors / cores / central processing units (CPUs), application-specific integrated circuit (ASIC) chips, neural network processors or accelerators, field-programmable gate arrays (FPGAs), decompression engines, compute units, embedded processors, graphics processing units (GPUs) / graphics cores, pipelines, accelerated processing units (APUs), functional units, controllers, accelerators, and / or other programmable logic devices. When such hardware modules are activated, the hardware modules perform some or all of the operations. In some embodiments, a hardware module includes one or more general-purpose circuits that are configured by executing instructions (program code, firmware, etc.) to perform the operations.

[0075] In some embodiments, data structures representing some or all of the structures and mechanisms described herein (e.g., electronic device 200 or some portion thereof) are stored on a non-transitory computer-readable storage medium, which includes a database or other data storage mechanism that can be read by a computer or other electronic device and that directly or indirectly stores data that can be used to manufacture hardware that includes the structures and mechanisms described herein. For example, the data structure can be an behavioral-level description or register-transfer level (RTL) description of hardware functionality in a high level design language (HDL) such as Verilog or VHDL. The description can be read by a synthesis tool, which can synthesize the description to produce a hardware description in a netlist format that includes a list of gates or other circuit elements that can be used to fabricate hardware that includes the structures and mechanisms described herein. The netlist can then be placed and routed to produce a data set that describes geometric shapes to be applied to masks that are used to manufacture the hardware that includes the structures and mechanisms described herein. The data set can then be used to fabricate the hardware that includes the structures and mechanisms described herein. Alternatively, the database on the computer accessible storage medium can be a netlist (with or without synthesis libraries) or a data set (as appropriate) or a graphical data system (GDS) II data.

[0076] In this specification, variables or unspecified values (i.e., general descriptions of values in the absence of specific instances of the values) are represented by letters such as N, M, and X. As used herein, although similar letters can be used at different locations in this specification, the variables and unspecified values are not necessarily the same in each case, i.e., some or all of the general variables and unspecified values can be expected to be different variable and values. In other words, in this specification, N and any other letter used to represent variables and unspecified values are not necessarily related to each other.

[0077] The expression “and / or” or “or” as used herein is intended to present one and / or the case of an equivalent of “at least one” of the elements in the list associated with the or. For example, in the sentence “the electronic device performs a first operation, a second operation, and / or the like,” the electronic device performs at least one of the first operation, the second operation, and other operations. In addition, the elements associated with the “and / or” in the list are merely instances from among a group of instances, and at least some of the instances can not occur in some embodiments.

[0078] The foregoing description of implementations has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the implementations to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art. Additionally, the above disclosure is intended to be illustrative, and not restrictive. It is therefore intended to cover any modifications and variations of the implementations.

Claims

1. An electronic device for decompressing compressed input data comprising N data streams, the N data streams having been generated from original data by compressing the original data using a compression standard to create compressed data prior to separating the original data; The compressed data is divided into N streams, each of the N streams including data of a corresponding type for generating commands associated with the compression standard for decompressing the compressed input data; The electronic device includes: and encodes each of the N streams using an encoding standard. Memory; and A decompression engine, comprising N decoders and one decompressor, is configured to: Receive a command to decompress the compressed input data; In each of the N decoders, a corresponding one of the N streams from the compressed input data is decoded individually and in parallel with the other N decoders, and each decoder outputs a decoded data stream of a corresponding type for generating a command associated with the compression standard for decompressing the compressed input data; In the decompressor, the decoded data stream output by the N decoders generates commands for decompressing the compressed input data using the compression standard to reconstruct the original data; The command is executed in the decompressor to reconstruct the original data; and The original data is stored in the memory.

2. The electronic device of claim 1, wherein decoding the corresponding one of the N streams in each of one or more of the N decoders comprises: Information for generating a decoding reference is obtained from a specified position in one of the N streams, the decoding reference including information used by the decoder to decode the corresponding stream from the compressed input data.

3. The electronic device of claim 1, wherein at least one of the N decoders comprises two or more sub-stream decoders and a stream combiner, wherein: Each of the two or more substream decoders is configured to: A single portion of data is obtained from one of the N streams decoded by the decoder containing the sub-stream decoder; Decode the individual portions of the data individually and in parallel with the other substream decoders of the two or more substream decoders; and Output the decoded data portion associated with the individual portion of the data; and The stream combiner is configured to: Receive the decoded data portion from each substream decoder; Combine the decoded data portions to generate the decoded data stream; and Output the decoded data stream.

4. The electronic device of claim 1, wherein at least one of the N decoders is a secondary decoder, the secondary decoder comprising fewer internal elements than the primary decoder among the N decoders used to decode the corresponding one of the N streams.

5. The electronic device of claim 1, wherein the decompressor comprises one or more buffers and at least one command assembler, wherein: Each of the one or more buffers stores data from a single decoded data stream derived from the outputs of the N decoders; and The command assembler: Data is retrieved from the one or more buffers; and The command for decompressing the data is generated from the data.

6. The electronic device of claim 5, wherein the decompressor comprises at least two command assemblers, wherein each of the at least two command assemblers is configured to: obtaining a first portion of data from the one or more buffers; and obtain a second portion of data from a decoded data stream output by the N decoders; and generate, from the first and second portions of data, some of the commands used to decompress the data.

7. The electronic device of claim 5, wherein the decompressor comprises an operation combiner configured to combine two or more commands into an aggregate command used to decompress the data.

8. The electronic device of claim 1, wherein the decoded data stream output by the N decoders comprises some or all of the following: literals, command tags, distances, and lengths.

9. The electronic device of claim 1, wherein the decompression engine comprises a command header decoder configured to: determine a starting position in the compressed input data for each of the N streams by at least one of processing the commands to decompress the compressed input data and communicating with a stream header decoder in some or all of the N decoders; and communicate the starting position in the compressed input data for each of the N streams to the respective one of the N decoders.

10. The electronic device of claim 1, wherein when the commands are executed in the decompressor to reconstruct the original data, the decompressor is configured to: pre-fetch data from memory when the data is used to execute a command; buffer the commands while the data is pre-fetched from memory; and execute the commands while the data is returned from memory.

11. The electronic device of claim 1, wherein: the decompressor reconstructs the original data in chunks of a specified size, and a command can depend on data in chunks reconstructed by a previous command; and the decompressor stores last M reconstructed chunks in a history buffer, the reconstructed chunks in the history buffer can be used to feed back to subsequent commands and the reconstructed chunks are written from the history buffer to the memory in a first-in-first-out order.

12. The electronic device of claim 1, wherein the encoding standard is a prefix coding standard and the compression standard is a dictionary coding compression standard.

13. A method for decompressing compressed input data comprising N data streams in an electronic device, the electronic device comprising a memory and a decompression engine having N decoders and one decompressor, the N data streams having been generated from original data by compressing the original data using a compression standard to create compressed data prior to separating the original data; divide the compressed data into N streams, each of the N streams comprising a respective type of data used to generate commands associated with the compression standard used to decompress the compressed input data; and encode each of the N streams using an encoding standard, the method comprising: receiving, by the decompression engine, commands to decompress the compressed input data; decoding, in each of the N decoders individually and in parallel with other decoders of the N decoders, a respective one of the N streams from the compressed input data, each decoder outputting a respective type of decoded data stream used to generate commands associated with the compression standard used to decompress the compressed input data; generating, in the decompressor, a command for decompressing the compressed input data using the compression standard to reconstruct the original data from the decoded data stream output by the N decoders; executing, in the decompressor, the command to reconstruct the original data; and storing the original data in the memory.

14. The method of claim 13, wherein decoding the respective one of the N streams in each of one or more of the N decoders comprises: obtaining, from a designated location in the respective one of the N streams, information for generating a decoding reference, the decoding reference comprising information used by the decoder to decode a respective stream from the compressed input data.

15. The method of claim 13, wherein at least one of the N decoders comprises two or more sub-stream decoders and a stream combiner, and wherein the method further comprises: obtaining, by each of the two or more sub-stream decoders, a separate portion of data from the one of the N streams decoded by the decoder in which the sub-stream decoder resides; decoding, by each of the two or more sub-stream decoders, separately and in parallel with the other sub-stream decoders of the two or more sub-stream decoders, the separate portion of data; outputting, by each of the two or more sub-stream decoders, a decoded data portion associated with the separate portion of data; receiving, by the stream combiner, the decoded data portion from each of the two or more sub-stream decoders; combining, by the stream combiner, the decoded data portions to generate the decoded data stream; and outputting, by the stream combiner, the decoded data stream.

16. The method of claim 13, wherein at least one of the N decoders is a secondary decoder, the secondary decoder comprising fewer internal elements relative to a primary decoder of the N decoders for decoding the respective one of the N streams.

17. The method of claim 13, wherein the decompressor comprises one or more buffers and at least one command assembler, wherein the method further comprises: storing, by each of the one or more buffers, data from a separate one of the decoded data streams output by the N decoders; obtaining, by the command assembler, data from the one or more buffers; and generating, by the command assembler, the command for decompressing the data from the data.

18. The method of claim 17, wherein the decompressor comprises an operation combiner, and the method further comprises: combining, by the operation combiner, two or more commands into an aggregate command for decompressing the data.

19. The method of claim 13, wherein the decoded data stream output by the N decoders comprises some or all of the following: literals, command tags, distances, and lengths.

20. The method of claim 13, wherein the decompression engine comprises a command header decoder, and the method further comprises: ​ ​ determined by the command header decoder by at least one of processing the command to decompress the compressed input data and communicating with stream header decoders in some or all of the N decoders to determine a starting location in the compressed input data for each of the N streams; and communicating by the command header decoder to the respective one of the N decoders each of the N streams at the starting location in the compressed input data.

21. The method of claim 13, wherein executing the command in the decompressor to reconstruct the original data comprises: prefetching by the decompressor the data from memory when the data in memory is used to execute a command; buffering by the decompressor the command while the data is prefetched from memory; and executing by the decompressor the command when the data is returned from memory.

22. The method of claim 13, wherein the decompressor reconstructs the original data in chunks of a specified size, and a command can depend on data in chunks reconstructed by a previous command, and the method further comprises: storing by the decompressor the last M reconstructed chunks in a history buffer, the reconstructed chunks in the history buffer available for feedback to a subsequent command and the reconstructed chunks written from the history buffer to the memory in a first-in-first-out order.

23. The method of claim 13, wherein the encoding standard is a prefix coding standard and the compression standard is a dictionary coding compression standard.