Data compression using conditional entropy model

By using neural network technology in the data compression system, the latent representation of data and the latent representation of entropy model are generated, and the neural network parameters are jointly trained, the problem of inefficient data compression in the existing technology is solved, and more efficient data compression and decompression effects are achieved.

CN119990190APending Publication Date: 2025-05-13GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411898542.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-07-20
Filing Date
2019-07-22
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has problems of inefficiency in the data compression and decompression process, especially when processing complex data such as images, it is difficult for conventional systems to achieve both high compression rates and low distortion.

Method used

A neural network-based system is adopted to generate a potential representation of data through an encoder neural network, and a potential representation of an entropy model is generated using a hyperencoder neural network. In combination with an autoregressive context neural network, neural network parameters are jointly trained to optimize the performance metric of rate distortion.

Benefits of technology

It realizes more efficient data compression and decompression than conventional systems, which can improve compression rates and reduce communication network bandwidth and storage requirements without compromising data quality and authenticity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990190A_ABST
    Figure CN119990190A_ABST
Patent Text Reader

Abstract

The invention relates to data compression using conditional entropy models. Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for compressing and decompressing data. In one aspect, a method includes processing data using an encoder neural network to generate a potential representation of the data; processing the potential representation of the data using a hyper-encoder neural network to generate a potential representation of an entropy model; generating an entropy encoded representation of the potential representation of the entropy model; generating an entropy encoded representation of the potential representation of the data using the potential representation of the entropy model; and determining a compressed representation of the data from an entropy encoded representation of (i) a potential representation of the data and (ii) a potential representation of an entropy model used to entropy encode the potential representation of the data.
Need to check novelty before this filing date? Find Prior Art

Description

Description of the case This application is a divisional application of Chinese invention patent application No. 201980020216.5, filed on July 22, 2019. Technical Field

[0001] This specification relates to data compression. Background Art

[0002] Compressing data refers to determining a representation of the data that takes up less space in memory and / or requires less bandwidth, such as for transmission over a network. Compressed data may be stored (e.g., in a logical data storage area or a physical data storage device), sent to a destination over a communications network (e.g., the Internet), or used in any other manner. Typically, the data can be reconstructed (approximately or exactly) from a compressed representation of the data. Summary of the invention

[0003] This specification describes a system implemented as a computer program on one or more computers in one or more locations that performs data compression and data decompression.

[0004] According to a first aspect, a method implemented by a data processing device is provided, the method comprising processing data using an encoder neural network to generate a potential representation of the data. Processing the potential representation of the data using a super encoder neural network to generate a potential representation of an entropy model, wherein the entropy model is defined by one or more probability distribution parameters characterizing a probability distribution of one or more code symbols. An entropy coded representation of the potential representation of the entropy model is generated. Generating an entropy coded representation of the potential representation of the data using the potential representation of the entropy model comprises determining the probability distribution parameters defining the entropy model using the potential representation of the entropy model. Quantized code symbols of the potential representation of the data are autoregressively processed by one or more neural network layers to generate a context output set. Quantization of the context output and the potential representation of the entropy model are processed by one or more neural network layers to generate the probability distribution parameters defining the entropy model. Generating an entropy coded representation of the potential representation of the data using the entropy model. Determining a compressed representation of the data from the entropy coded representations of: (i) the potential representation of the data and (ii) the potential representation of the entropy model used to entropy encode the potential representation of the data.

[0005] In some implementations, the data includes images and the encoder neural network is a convolutional neural network.

[0006] In some implementations, processing the data using an encoder neural network to generate a latent representation of the data includes determining an ordered collection of code symbols representing the data by quantizing the latent representation of the data.

[0007] In some implementations, generating an entropy encoded representation of a latent representation of the entropy model includes quantizing the latent representation of the entropy model.

[0008] In some implementations, the entropy model is defined by a corresponding code symbol probability distribution for each code symbol in an ordered collection of code symbols representing the data.

[0009] In some implementations, each code symbol probability distribution is a Gaussian distribution convolved with a uniform distribution, and the respective probability distribution parameters defining each code symbol probability distribution include respective mean and standard deviation parameters of the Gaussian distribution.

[0010] In some implementations, generating an entropy-encoded representation of a latent representation of the entropy model includes entropy encoding a quantized latent representation of the entropy model using one or more predetermined probability distributions.

[0011] In some implementations, using an entropy model to generate an entropy encoded representation of a latent representation of the data includes arithmetically encoding an ordered collection of code symbols representing the data using a probability distribution of code symbols defining the entropy model.

[0012] In some implementations, code symbols of the quantized latent representation of the data are associated with an ordering. Autoregressively processing the code symbols of the quantized latent representation of the data by one or more neural network layers to generate a plurality of context outputs includes generating a respective context output for each code symbol of the quantized latent representation of the data. Generating a context output for a given code symbol of the quantized latent representation of the data includes processing, using the one or more neural network layers, inputs of one or more code symbols of the quantized latent representation of the data preceding the given code symbol of the quantized latent representation of the data to generate a context output for the given code symbol of the quantized latent representation of the data.

[0013] In some implementations, the input does not include: (i) a given code symbol of a quantized latent representation of the data, or (ii) any code symbol of a quantized latent representation of the data that immediately follows the given code symbol of the quantized latent representation of the data.

[0014] In some implementations, one or more neural network layers are masked convolutional neural network layers.

[0015] In some implementations, processing the context output and the quantized latent representation of the entropy model by one or more neural network layers to generate probability distribution parameters defining the entropy model includes generating corresponding probability distribution parameters representing a corresponding code symbol probability distribution for each code symbol of the quantized latent representation of the data. For each code symbol of the quantized latent representation of the data, processing inputs including: (i) the context output of the code symbol of the quantized latent representation of the data, and (ii) the quantized latent representation of the entropy model using one or more neural network layers to generate probability distribution parameters representing the code symbol probability distribution for the code symbol of the quantized latent representation of the data.

[0016] In some implementations, parameters of a neural network layer used to: (i) generate a compressed representation of the data, and (ii) generate a reconstruction of the data from the compressed representation of the data are jointly trained using machine learning training techniques to optimize a rate-distortion performance metric.

[0017] In some implementations, parameters of (i) a superencoder neural network and (ii) a neural network layer for autoregressively processing code symbols of a quantized latent representation of data to generate a set of context outputs are jointly trained using machine learning training techniques to optimize a rate-distortion performance metric.

[0018] In some implementations, the rate-distortion performance metric includes: (i) a first rate term based on the size of an entropy encoded representation of a latent representation of the data, (ii) a second rate term based on the size of an entropy encoded representation of the latent representation of an entropy model, and (iii) a distortion term based on a difference between the data and a reconstruction of the data.

[0019] In some implementations, the distortion term is scaled by a hyperparameter that determines the rate-distortion tradeoff.

[0020] In some implementations, the encoder neural network includes one or more generalized division normalization (GDN) nonlinearities.

[0021] In some implementations, the compressed representation of the data includes a bitstream.

[0022] In some implementations, the method further includes sending or storing a compressed representation of the data.

[0023] According to a second aspect, a method is provided, the method comprising obtaining an entropy coded representation of: (i) a latent representation of a data set, and (ii) a latent representation of an entropy model for entropy encoding the latent representation of the data. The latent representation of the entropy model is determined by processing the latent representation of the data using a super encoder neural network. The entropy model is defined by one or more probability distribution parameters characterizing a probability distribution of one or more code symbols. Entropy decoding the latent representation of the data comprises determining the probability distribution parameters defining the entropy model using the latent representation of the entropy model. Entropy decoding the latent representation of the data using the entropy model comprises, for one or more code symbols of quantization of the latent representation of the data, processing: (i) one or more preceding code symbols of the quantized latent representation of the data preceding the code symbol in an ordering of code symbols of the quantized latent representation of the data, and (ii) quantization of the latent representation of the entropy model to generate a code symbol probability distribution corresponding to the code symbol of the quantized latent representation of the data. The code symbol probability distribution corresponding to the code symbol of the quantized latent representation of the data is used to entropy decode the code symbol of the quantized latent representation of the data. Determining a reconstruction of the data from the latent representation of the data comprises processing the latent representation of the data using a decoder neural network.

[0024] According to a third aspect, a method implemented by a data processing device is provided, comprising processing data using an encoder neural network to generate a potential representation of the data. The potential representation of the data includes an ordered collection of code symbols representing the data. The potential representation of the data is processed using a super encoder neural network to generate a potential representation of an entropy model. The entropy model is defined by one or more probability distribution parameters that characterize the probability distribution of one or more code symbols. An entropy coded representation of the potential representation of the entropy model is generated. An entropy coded representation of the potential representation of the data is generated using the potential representation of the entropy model, comprising determining the probability distribution parameters that define the entropy model using the potential representation of the entropy model. An entropy coded representation of the potential representation of the data is generated using the entropy model. A compressed representation of the data is determined from the entropy coded representations of: (i) the potential representation of the data and (ii) the potential representation of the entropy model used to entropy encode the potential representation of the data.

[0025] In some implementations, processing the data using the encoder neural network to generate a latent representation of the data includes determining an ordered collection of code symbols representing the data by quantizing outputs of the encoder neural network.

[0026] In some implementations, generating an entropy encoded representation of a latent representation of the entropy model includes quantizing the latent representation of the entropy model.

[0027] In some implementations, the entropy model is defined by a corresponding code symbol probability distribution for each code symbol in a latent representation of the data.

[0028] In some implementations, each code symbol probability distribution is a Gaussian distribution convolved with a uniform distribution, and the respective probability distribution parameters defining each code symbol probability distribution include respective mean and standard deviation parameters of the Gaussian distribution.

[0029] In some implementations, generating an entropy-encoded representation of a latent representation of the entropy model includes entropy encoding the latent representation of the entropy model using one or more predetermined probability distributions.

[0030] In some implementations, generating an entropy encoded representation of a latent representation of the data using an entropy model includes arithmetically encoding the latent representation of the data using a code symbol probability distribution that defines the entropy model.

[0031] In some implementations, determining parameters of a probability distribution defining the entropy model using the latent representation of the entropy model includes processing the latent representation of the entropy model by one or more neural network layers.

[0032] In some implementations, the method further includes autoregressively processing, by one or more neural network layers, code symbols of a latent representation of the data to generate a set of context outputs. The context outputs and the latent representation of the entropy model are processed by one or more neural network layers to generate probability distribution parameters defining the entropy model.

[0033] In some implementations, parameters of a neural network used to: (i) generate a compressed representation of the data and (ii) generate a reconstruction of the data from the compressed representation of the data are jointly trained using machine learning training techniques to optimize a rate-distortion performance metric.

[0034] In some implementations, the rate-distortion performance metric includes: (i) a first rate term based on the size of an entropy encoded representation of a latent representation of the data, (ii) a second rate term based on the size of an entropy encoded representation of the latent representation of an entropy model, and (iii) a distortion term based on a difference between the data and a reconstruction of the data.

[0035] In some implementations, the distortion term is scaled by a hyperparameter that determines the rate-distortion tradeoff.

[0036] In some implementations, the encoder neural network includes one or more generalized division normalization (GDN) nonlinearities.

[0037] In some implementations, the compressed representation of the data includes a bitstream.

[0038] In some implementations, the method further includes sending or storing a compressed representation of the data.

[0039] According to a fourth aspect, a method implemented by a data processing device is provided, comprising obtaining an entropy coded representation of each of the following: (i) a potential representation of the data including an ordered collection of code symbols representing the data, and (ii) a potential representation of an entropy model for entropy encoding the potential representation of the data. The potential representation of the entropy model is determined by processing the potential representation of the data using a super encoder neural network. The entropy model is defined by one or more probability distribution parameters that characterize the probability distribution of one or more code symbols. Entropy decoding the potential representation of the data includes determining the probability distribution parameters that define the entropy model using the potential representation of the entropy model. Entropy decoding the potential representation of the data using the entropy model. Determining (reconstructing) data from the potential representation of the data includes processing the potential representation of the data using a decoder neural network.

[0040] As described above, the compression system entropy encodes a latent representation of the data determined by processing the data using an encoder neural network. However, in some cases, the compression system does not include an encoder neural network and directly entropy encodes components of the input data. For example, if the input data is an image, the compression system may directly entropy encode pixel intensities / color values ​​of the image. When the compression system directly entropy encodes components of the input data (i.e., without using an encoder neural network), the decompression system similarly directly entropy decodes components of the input data (i.e., without using a decoder neural network). An example of such a method may include a method implemented by a data processing device, the method comprising: processing data using a super encoder neural network to generate a latent representation of an entropy model, wherein the entropy model is defined by one or more probability distribution parameters characterizing a probability distribution of one or more code symbols; generating an entropy encoded representation of the latent representation of the entropy model; generating an entropy encoded representation of the data using the latent representation of the entropy model, comprising: determining probability distribution parameters defining the entropy model using the latent representation of the entropy model; and generating an entropy encoded representation of the data using the entropy model; and determining a compressed representation of the data from the entropy encoded representations of: (i) the data and (ii) a latent representation of the entropy model used to entropy encode the data.

[0041] According to a fifth aspect, a system is provided, the system comprising one or more computers and one or more storage devices storing instructions, which, when executed by the one or more computers, cause the one or more computers to perform the operations of the corresponding methods of any previous aspects.

[0042] According to a sixth aspect, one or more computer storage media are provided, wherein the one or more computer storage media store instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the corresponding methods of any previous aspect.

[0043] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.

[0044] The compression system described in this specification compresses input data using a conditional entropy model determined based on the input data (using a neural network) rather than using, for example, a static predetermined entropy model. Determining the entropy model based on the input data makes the entropy model richer and more accurate, for example by capturing spatial dependencies in the input data, thereby enabling the input data to be compressed at a higher rate that can be achieved in some conventional systems.

[0045] The system described in this specification uses a collection of neural networks that are jointly trained to optimize a rate-distortion objective function to determine a conditional entropy model. In order to determine the conditional entropy model, the system can use a "super encoder" neural network to process the quantized potential representation of the data to generate a "hyper-prior" that implicitly characterizes the conditional entropy model. The hyper-prior is then compressed and included in the compressed representation of the input data as side information. Typically, a more complex hyper-prior can specify a more accurate conditional entropy model that enables the input data to be compressed at a higher rate. However, increasing the complexity of the hyper-prior can cause the hyper-prior itself to be compressed at a lower rate. Machine learning techniques are used to train the system described in this specification to adaptively determine the complexity of the hyper-prior for each input data set in order to optimize the overall compression rate. This allows the system described in this specification to achieve a higher compression rate than some conventional systems.

[0046] The system described in this specification is capable of determining a conditional entropy model for compressing input data using an autoregressive "context" neural network that enables the system to learn a richer entropy model without increasing the size of the compressed representation of the data. The system is capable of jointly training a superencoder neural network and an autoregressive context neural network so that the super prior can store information needed to reduce uncertainty in the context neural network while avoiding information that can be accurately predicted by the autoregressive context neural network.

[0047] By compressing and decompressing data more efficiently than conventional systems, the systems described herein may enable more efficient data transmission (e.g., by reducing the communication network bandwidth required to send data) and more efficient data storage (e.g., by reducing the amount of memory required to store data). Furthermore, through the disclosed methods, improved efficiency may be achieved without compromising data quality and / or authenticity.

[0048] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a block diagram of an example compression system.

[0050] Figure 2 is a block diagram of an example decompression system.

[0051] Figure 3 A table describing an example architecture of a neural network used by a compression / decompression system is shown.

[0052] Figure 4 Illustrate an example of the dependency of a context output generated by a context neural network on an ordered collection of code symbols representing input data.

[0053] Figure 5 Figure 2 is a rate-distortion graph comparing the performance of the compression / decompression system described in this specification with the performance of other compression / decompression systems.

[0054] Figure 6 The diagram illustrates a qualitative example of the performance gain that can be achieved by using the compression / decompression system described in this specification.

[0055] Figure 7 is a flow chart of an example process for compressing data.

[0056] Figure 8 is a flow chart of an example process for decompressing data.

[0057] Like reference numbers and designations throughout the various drawings refer to like elements. DETAILED DESCRIPTION

[0058] This specification describes a data compression system and a data decompression system. The compression system is configured to process input data (e.g., image data, audio data, video data, text data, or any other suitable type of data) to generate a compressed representation of the input data. The decompression system can process the compressed data to generate a (approximate or exact) reconstruction of the input data.

[0059] Typically, the compression system and the decompression system may be co-located or remotely located, and the compressed data generated by the compression system may be provided to the decompression system in any of a variety of ways. For example, the compressed data may be stored (e.g., in a physical data storage device or a logical data storage area) and then subsequently retrieved from storage and provided to the decompression system. As another example, the compressed data may be sent to a destination via a communication network (e.g., the Internet), where the compressed data is subsequently retrieved and provided to the decompression system.

[0060] To compress input data, a compression system maps the input data to a quantized potential representation as an ordered collection of "code symbols", such as a vector or matrix of code symbols. Each code symbol is obtained from a discrete set of possible code symbols, such as a set of integer values. The compression system compresses the code symbols by entropy encoding the code symbols using a conditional entropy model, i.e., an entropy model that depends on the input data. The conditional entropy model defines a corresponding code symbol probability distribution (i.e., a probability distribution over the set of possible code symbols) corresponding to each code symbol in the ordered collection of code symbols representing the input data. The compression system then generates a compressed representation of the input data based on: (i) the compressed code symbols, and (ii) "side information" that characterizes the conditional entropy model used to compress the code symbols.

[0061] The decompression system can decompress the data by recovering the conditional entropy model from the compressed data and using the conditional entropy model to decompress (ie, entropy decode) the compressed code symbols. The decompression system can then reconstruct the original input data by mapping the code symbols back to a reconstruction of the input data.

[0062] Using an adaptive and input data dependent conditional entropy model (rather than, for example, a static predetermined entropy model) can enable the input data to be compressed more efficiently. These and other features are described in more detail below.

[0063] Figure 1 is a block diagram of an example compression system 100. Compression system 100 is an example system implemented as a computer program on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.

[0064] Compression system 100 processes input data 102 to generate compressed data 104 representing input data 102 using: (1) an encoder neural network 106, (2) a super encoder neural network 108, (3) a super decoder neural network 110, (4) an entropy model neural network 112, and optionally (5) a context neural network 114. As will be described with reference to Figure 2 As described in more detail, a rate-distortion objective function is used to jointly train the neural network used by the compression system (as well as the neural network used by the decompression system). In general, each neural network described in this document can have any suitable neural network architecture that enables it to perform the functions described therein. Figure 3 Example architectures of neural networks used by the compression system and the decompression system are described in more detail.

[0065] The encoder neural network 106 is configured to process the input data 102 ( ) to generate a latent representation 116 of the input data 102 ( ). As used throughout this document, a "latent representation" of data refers to a representation of the data as an ordered collection of values, such as a vector or matrix of values. In one example, the input data may be an image, the encoder neural network 106 may be a convolutional neural network, and the latent representation 116 of the input data may be a multi-channel feature map output by the final layer of the encoder neural network 106. Typically, the latent representation 116 of the input data may be more compressible than the input data itself, and in some cases may have a lower dimensionality than the input data.

[0066] To facilitate compression of the latent representation 116 of the input data using entropy coding techniques, the compression system 100 uses a quantizer Q 118 to quantize the latent representation 116 of the input data to generate an ordered set of code symbols 120 A quantized value refers to a member of a discrete set of possible code symbols that maps that value. For example, the set of possible code symbols may be integer values, and compression system 100 may perform quantization by rounding real-valued numbers to integer values.

[0067] As will be described in greater detail below, the compression system 100 uses a superencoder neural network 108, a superdecoder neural network 110, and an entropy model neural network 112 to generate a conditional entropy model for entropy encoding code symbols 120 representing input data.

[0068] The super encoder neural network 108 is configured to process the latent representation 116 of the input data to generate a “super prior” 122 ( ) (sometimes referred to as "hyperparameters"), i.e., a potential representation of the conditional entropy model. In one example, the superencoder neural network 108 can be a convolutional neural network, and the hyper-prior 122 can be a multi-channel feature map output by the final layer of the superencoder neural network 108. The hyper-prior implicitly characterizes an entropy model associated with the input data that will enable the code symbols 120 representing the input data to be efficiently compressed.

[0069] The compressed data 104 typically includes a compressed representation of the hyper-prior 122 to enable the decompression system to recover the conditional entropy model. To this end, the compression system 100 uses a quantizer Q 136 to quantize the hyper-prior 122 to generate a quantized hyper-prior 124. , and generates a compressed representation 126 of the quantized hyper-prior 124, for example as a bit string, i.e., a binary digit string. In one example, the compression system 100 compresses the quantized hyper-prior 124 using an entropy coding engine 138 in accordance with a predetermined entropy model that specifies one or more predetermined code symbol probability distributions.

[0070] The super-decoder neural network 110 is configured to process the quantized super-prior 124 to generate a super-decoder output 128 , and the entropy model neural network 112 is configured to process the super decoder output 128 to generate a conditional entropy model. That is, the super decoder 110 and the entropy model neural network 112 jointly decode the quantized super prior to generate an output that explicitly defines the conditional entropy model.

[0071] The conditional entropy model specifies a corresponding code symbol probability distribution corresponding to each code symbol 116 representing the input data. Typically, the output of the entropy model neural network 112 includes distribution parameters that define each code symbol probability distribution of the conditional entropy model. In one example, each code symbol probability distribution of the conditional entropy model can be a Gaussian distribution (parameterized by mean and standard deviation parameters) convolved with a unit uniform distribution. In this example, the output of the entropy model neural network 112 can specify corresponding values ​​of the Gaussian mean and standard deviation parameters for each code symbol probability distribution of the conditional entropy model.

[0072] Optionally, the compression system 100 can additionally use a context neural network 114 in determining the conditional entropy model. The context neural network 114 is configured to autoregressively process code symbols 120 representing input data (i.e., in accordance with the ordering of the code symbols) to generate a corresponding "context output" 130 for each code symbol. The context output of each code symbol depends only on the code symbols that precede it in the ordered set of code symbols representing the input data, and not on the code symbol itself or on the code symbol that immediately follows it. The context output 130 of a code symbol can be understood as causal context information that can be used by the entropy model neural network 112 to generate a more accurate code symbol probability distribution for the code symbol.

[0073] The entropy model neural network 112 is capable of processing the context outputs 130 generated by the context neural network 114 (i.e., in addition to the super decoder outputs 128) to generate a conditional entropy model. In general, the code symbol probability distribution for each code symbol depends on the context outputs of the code symbol, and optionally, on the context outputs of the code symbols preceding the code symbol, but not on the context outputs of the code symbols immediately following the code symbol. As will be seen in reference Figure 2 Described in more detail, this results in a causal dependency of the conditional entropy model on the code symbols representing the input data, which ensures that the decompression system can recover the conditional entropy model from the compressed data.

[0074] In contrast to the super-prior 122, which must be included as side information in the compressed data 104 (thus increasing the overall compressed file size), the autoregressive context neural network 114 provides a "free" source of information (discounting computational cost) because it does not require the addition of any side information. Jointly training the context neural network 114 and the super encoder 108 enables the super-prior 122 to store information that is complementary to the context output 130, while avoiding information that can be accurately predicted using the context output 130.

[0075] The entropy coding engine 132 is configured to compress the code symbols 120 by entropy coding the code symbols 120 representing the input data according to the conditional entropy model. The entropy coding engine 132 can implement any suitable entropy coding technique, such as arithmetic coding technique, range coding technique, or Huffman coding technique. The compressed code symbols 134 can be represented in any of a variety of ways, such as, for example, as a bit string.

[0076] The compression system 100 generates compressed data 104 based on: (i) compressed code symbols 134 and (ii) compressed hyper-prior 126. For example, the compression system may generate the compressed data by concatenating respective bit strings representing the compressed code symbols 134 and the compressed hyper-prior 126.

[0077] In some cases, the compression system 100 uses the context neural network 114 instead of the super prior 122 to determine an entropy model for entropy encoding code symbols representing data. In these cases, the compression system 100 does not use the super encoder network 108 or the super decoder network 110. Instead, the system generates an entropy model by autoregressively processing the code symbols representing data 120 to generate a context output 130 and then processing the context output 130 using the entropy model neural network 112.

[0078] Figure 2 is a block diagram of an example decompression system 200. Decompression system 200 is an example system implemented as a computer program on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.

[0079] The decompression system 200 processes the compressed data 104 generated by the compression system to generate a reconstruction 202 that approximates the original input data using: (1) a super-decoder neural network 110, (2) an entropy model neural network 112, (3) a decoder neural network 204, and optionally (4) a context neural network 114. The super-decoder neural network 110, the entropy model neural network 112, and the context neural network 114 used by the decompression system share the same parameter values ​​as the corresponding neural networks used by the compression system.

[0080] To recover the conditional entropy model, the decompression system 200 obtains the quantized hyper-prior 124 from the compressed data 104. For example, the decompression system 200 may obtain the quantized hyper-prior 124 by entropy decoding a compressed representation 126 of the quantized hyper-prior 124 included in the compressed data 104 using the entropy decoding engine 206. In this example, the entropy decoding engine 206 may entropy decode the compressed representation 126 of the quantized hyper-prior 124 using the same (e.g., predetermined) entropy model used to entropy encode the compressed representation 126 of the quantized hyper-prior 124.

[0081] The super-decoder neural network 110 is configured to process the quantized super-prior 124 to generate a super-decoder output 128 , and the entropy model neural network 112 is configured to process the super decoder output 128 to generate a conditional entropy model, i.e., in a manner similar to the compression system. The entropy decoding engine 208 is configured to entropy decode the compressed code symbols 134 included in the compressed data 104 according to the conditional entropy model to recover the code symbols 120.

[0082] In the case where the compression system uses the context neural network 114 to determine the conditional entropy model, the decompression system 200 also uses the context neural network 114 to recover the conditional entropy model. Figure 1 As described, the context neural network 114 is configured to autoregressively process the code symbols 120 representing the input data to generate a corresponding context output 130 for each code symbol. After initially receiving the compressed data 104, the decompression system 200 does not have access to the complete set of decompressed code symbols 120 provided as input to the context neural network 114. As will be described in more detail below, the decompression system 200 accounts for this by decompressing the code symbols 120 sequentially according to the order of the code symbols. The context output 130 generated by the context neural network 114 is provided to the entropy model neural network 112, which processes the context output 130 along with the super decoder output 128 to generate a conditional entropy model.

[0083] To illustrate that the decompression system 200 does not initially have access to the complete set of decompressed code symbols 120 provided as input to the context neural network 114, the decompression system 200 sequentially decompresses the code symbols 120 in accordance with the order of the code symbols. In particular, the decompression system may decompress the first code symbol using, for example, a predetermined code symbol probability distribution. To decompress each subsequent code symbol, the context neural network 114 processes one or more previous code symbols (i.e., which have already been decompressed) to generate a corresponding context output 130. The entropy model neural network 112 then processes (i) the context output 130, (ii) the super decoder output 128, and optionally (iii) one or more previous context outputs 130 to generate a corresponding code symbol probability distribution, which is then used to decompress the code symbol.

[0084] The decoder neural network 204 is configured to process the ordered collection of code symbols 120 to generate a reconstruction 202 that approximates the input data. That is, the operations performed by the decoder neural network 204 approximately cause the decoder neural network 204 to approximate the input data. Figure 1 The described encoder neural network performs the inversion of operations.

[0085] In some cases, the decompression system 200 determines an entropy model for entropy decoding code symbols representing data using the context neural network 114 instead of the super prior. In these cases, the decompression system 200 does not use the super decoder network 110. Instead, the system generates an entropy model by autoregressively processing the code symbols representing data 120 to generate a context output 130 and then processing the context output 130 using the entropy model neural network 112.

[0086] The compression system and the decompression system can be jointly trained using machine learning training techniques (e.g., stochastic gradient descent) to optimize a rate-distortion objective function. More specifically, the encoder neural network, the super-encoder neural network, the super-decoder neural network, the context neural network, the entropy model neural network, and the decoder neural network can be jointly trained to optimize the rate-distortion objective function. In one example, the rate-distortion objective function (“performance metric”) It can be given by the following formula: in It means that the input data The code symbols are entropy encoded in the conditional entropy model The probability of refers to the quantized hyper-prior The entropy model used to entropy encode the quantized hyper-prior The probability of is a parameter that determines the rate-distortion tradeoff, and Refers to input data Reconstruction of input data In the rate-distortion objective function described with reference to equations (1) to (4), represents the size of the compressed code symbol representing the input data (e.g., in bits), characterizes the size of the compressed hyper-prior (e.g., in bits), and Characterizes the difference ("distortion") between the input data and a reconstruction of the input data.

[0087] Typically, a more complex hyper-prior can specify a more accurate conditional entropy model that enables the code symbols representing the input data to be compressed at a higher rate. However, increasing the complexity of the hyper-prior can cause the hyper-prior itself to be compressed at a lower rate. By jointly training the compression system and the decompression system, the balance between (i) the size of the compressed hyper-prior, and (ii) the increased compression rate from the more accurate entropy model can be learned directly from the training data.

[0088] In some implementations, the compression system and the decompression system do not use an encoder neural network or a decoder neural network. In these implementations, as described above, the compression system can generate code symbols representing the input data by directly quantizing the input data, and the decompression system can generate a reconstruction of the input data as a result of decompressing the code symbols.

[0089] Figure 3 A table 300 is shown that describes an example architecture of a neural network used by a compression / decompression system in the specific case where the input data consists of an image. More specifically, table 300 describes example architectures of an encoder neural network 302, a decoder neural network 304, a super encoder neural network 306, a super decoder neural network 308, a context neural network 310, and an entropy model neural network 312.

[0090] Each row of table 300 corresponds to a corresponding layer. Convolutional layers are designated with a "Conv" prefix followed by the kernel size, number of channels, and downsampling stride. For example, the first layer of encoder neural network 302 uses a kernel with 192 channels and a stride of 2. kernel. The prefix “Deconv” corresponds to upsampling convolution, whereas “Masked” corresponds to masked convolution. GDN stands for generalized divisive normalization, while IGDN is the inverse of GDN.

[0091] In reference Figure 3 In the example architecture described, the entropy model neural network uses Kernel. This architecture enables the entropy model neural network to generate a conditional entropy model with the following property: the code symbol probability distribution corresponding to each code symbol does not depend on the context output corresponding to the subsequent code symbol (as described earlier). As another example, the same effect can be achieved by using a masked convolution kernel.

[0092] Figure 4 An example of the dependency of a context output generated by a context neural network on an ordered collection of code symbols 402 representing input data is illustrated. To generate a context output for code symbol 404, the context neural network processes one or more code symbols preceding code symbol 404, which are illustrated as shaded.

[0093] Figure 5 A rate-distortion graph 500 is shown comparing the performance of the compression / decompression system described in this specification with the performance of other compression / decompression systems. For each compression / decompression system, the rate-distortion graph 500 shows the peak signal-to-noise ratio (PSNR) achieved by the system (on the vertical axis) for various bits per pixel (BPP) compression rates (on the horizontal axis).

[0094] Line 502 on graph 500 corresponds to the compression / decompression system described in this specification, line 504 corresponds to the Better Portable Graphics (BPG) system, and line 506 corresponds to the system described in reference to: J. Balle, D. Minnen, S. Singh, S. J. Hwang, N. Johnston, “Variational image compression with a scale hyperprior,” 6 th Int. Conf. on Learning Representations, 2018, line 508 corresponds to the system described in reference to D. Minnen, G. Toderici, S. Singh, SJ Hwang, M. Covell, “Image-dependent local entropy models for image compression with deep networks,” Int. Conf. on Image Processing, 2018, line 510 corresponds to the JPEG 2000 system, and line 512 corresponds to the JPEG system. It can be appreciated that the compression / decompression system described in this specification generally outperforms other compression / decompression systems.

[0095] Figure 6A qualitative example of the performance gain that can be achieved by using the compression / decompression system described in this specification is illustrated. Image 602 is reconstructed after compression using the JPEG system at 0.2309 BPP, and image 604 is the same image reconstructed after compression using the compression / decompression system described in this specification at 0.2149 BPP. It can be appreciated that, despite the use of a larger BPP to compress image 602 than image 604, image 604 (corresponding to the system described in this specification) is of substantially higher quality (e.g., has fewer artifacts) than image 602 (corresponding to JPEG).

[0096] Figure 7 700 is a flow chart of an example process 700 for compressing data. For convenience, process 700 is described as being performed by a system of one or more computers located in one or more locations. For example, a compression system appropriately programmed in accordance with the present specification, such as reference Figure 1 The described compression system is capable of performing process 700.

[0097] The system receives data to be compressed (702). The data may be in any suitable form, such as image data, audio data, or text data.

[0098] The system processes the data using an encoder neural network to generate a latent representation of the data (704). In one example, the data is image data, the encoder neural network is a convolutional neural network, and the latent representation of the data is an ordered collection of feature maps output by a final layer of the encoder neural network.

[0099] The system processes the latent representation of the data using a superencoder neural network to generate a latent representation of the conditional entropy model, i.e., a "super-prior" (706). In one example, the superencoder neural network is a convolutional neural network and the super-prior is a multi-channel feature map output by the final layer of the superencoder neural network.

[0100] The system quantizes and entropy encodes the hyper-prior (708). The system can entropy encode the quantized hyper-prior using, for example, a predetermined entropy model defined by one or more predetermined code symbol probability distributions. In one example, the predetermined entropy model can specify a corresponding predetermined code symbol probability distribution for each code symbol of the quantized hyper-prior. In this example, the system can entropy encode each code symbol of the quantized hyper-prior using the corresponding predetermined code symbol probability distribution. The system can use any suitable entropy encoding technique, such as a Huffman coding technique or an arithmetic coding technique.

[0101] The system uses the super prior to determine a conditional entropy model (710). In particular, the system processes the quantized super prior using a super decoder neural network and then processes the super decoder neural network output using an entropy model neural network to generate an output that defines the conditional entropy model. For example, the entropy model neural network can generate an output that specifies corresponding distribution parameters for each code symbol probability distribution that defines the conditional entropy model.

[0102] In some implementations, to determine the entropy model, the system additionally uses a context neural network to regressively process the code symbols of the quantized latent representation of the data to generate a corresponding context output corresponding to each code symbol. In these implementations, the entropy model neural network generates a conditional entropy model by processing the context output together with the super decoder neural network output.

[0103] The context neural network can have any suitable neural network architecture that enables it to autoregressively process the code symbols of the quantized latent representation of the data. In one example, the context neural network is a neural network with The masked convolutional layer is a masked convolutional neural network with a masked convolutional layer having convolution kernels. In this example, the context neural network generates a context output for each code symbol by processing an appropriate subset of the previous code symbols, i.e., rather than processing every previous code symbol. In some cases, the masked convolutional layer can generate a context output for the current code symbol by dynamically zeroing components of the convolution kernels of the layer operating on (i) the current code symbol and (ii) the code symbol immediately following the current code symbol.

[0104] The system entropy encodes the quantized latent representation of the data using the conditional entropy model (712). For example, the system can entropy encode each code symbol of the quantized latent representation of the data using the corresponding code symbol probability distribution defined by the conditional entropy model. The system can use any suitable entropy encoding technique, such as a Huffman coding technique or an arithmetic coding technique.

[0105] The system determines a compressed representation of the data based on: (i) a compressed (i.e., entropy coded) quantized latent representation of the data and (ii) a compressed (i.e., entropy coded) quantized hyper-prior, e.g., by concatenating them (714). After determining the compressed representation of the data, the system can store the compressed representation of the data (e.g., in a logical data storage area or a physical data storage device) or transmit the compressed representation of the data (e.g., over a data communication network).

[0106] Figure 8 is used to describe the example process 700 (see Figure 7Flowchart of an example process 800 for decompressing data compressed as described herein. For convenience, process 800 is described as being performed by a system of one or more computers located in one or more locations. For example, a decompression system appropriately programmed in accordance with the present specification, such as reference Figure 2 The decompression system described is capable of performing process 800.

[0107] The system obtains compressed data (802). As described above, the compressed data includes (i) a compressed (i.e., entropy encoded) quantized latent representation of the data and (ii) a compressed (i.e., entropy encoded) quantized hyper-prior. The system can obtain the compressed data, for example, from a data store (e.g., a logical data store area or a physical data store device) or as a transmission over a communication network (e.g., the Internet).

[0108] The system entropy decodes the quantization super prior (804). For example, the system can entropy decode the quantization super prior using a predetermined entropy model defined by a set of predetermined code symbol probability distributions used to entropy encode the quantization super prior. In this example, the system can entropy decode each code symbol of the quantization super prior using a corresponding predetermined code symbol probability distribution defined by the predetermined entropy model.

[0109] The system determines a conditional entropy model for entropy encoding a quantized latent representation of the data (806). To determine the conditional entropy model, the system processes the quantized hyper-prior using a super-decoder neural network and then processes the super-decoder neural network output using an entropy model neural network to generate distribution parameters that define the conditional entropy model. As will be described in more detail below, in some implementations, the system determines the conditional entropy model using an autoregressive context neural network.

[0110] The system entropy decodes the quantized latent representation of the data using the conditional entropy model (808). In particular, the system entropy decodes each code symbol of the quantized latent representation of the data using the corresponding code symbol probability distribution defined by the conditional entropy model.

[0111] In some implementations, the system determines the conditional entropy model using a context neural network. As described above, the context neural network is configured to autoregressively process the code symbols of the latent representation of the data to generate a corresponding context output for each code symbol. The entropy model neural network generates the conditional entropy model by processing the context output in addition to the super decoder neural network output.

[0112] In these implementations, the system sequentially entropy decodes the code symbols of the quantized potential representation of the data. For example, to entropy decode the first code symbol, the context neural network may process the placeholder input to generate a corresponding context output. The entropy model neural network may process the context output together with the super decoder neural network output to generate a corresponding code symbol probability distribution for entropy decoding the first code symbol. For each subsequent code symbol, the context neural network may process an input including one or more previous code symbols (i.e., which have been entropy decoded) to generate a corresponding context output. Then, as before, the entropy model neural network may process the context output of the code symbol together with the super decoder neural network output to generate a corresponding code symbol probability distribution for entropy decoding the code symbol.

[0113] The system determines a reconstruction of the data by processing the quantized latent representation of the data using a decoder neural network (810). In one example, the decoder neural network can be a deconvolutional neural network that processes the quantized latent representation of the image data to generate a (approximate or exact) reconstruction of the image data.

[0114] This specification uses the term "configured" in conjunction with system and computer program components. For a system of one or more computers to be configured to perform a particular operation or action, it is meant that the system has installed thereon software, firmware, hardware, or a combination of software, firmware, hardware that, in operation, causes the system to perform those operations or actions. For one or more computer programs to be configured to perform a particular operation or action, it is meant that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform the operation or action.

[0115] Embodiments of the subject matter and functional operations described in this specification may be implemented with digital electronic circuits, with tangibly implemented computer software or firmware, with computer hardware including the structures disclosed in this specification and their structural equivalents, or with a combination of one or more of them. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by a data processing device or for controlling the operation of a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access storage device, or a combination of one or more of them. Alternatively or in addition, the program instructions may be encoded on an artificially generated propagation signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device for execution by a data processing device.

[0116] The term "data processing apparatus" refers to data processing hardware and includes all kinds of apparatus, devices and machines for processing data, including by way of example a programmable processor, a computer or a plurality of processors or computers. An apparatus may also be or further include special-purpose logic circuitry, for example, an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). An apparatus may optionally include, in addition to hardware, code that creates an execution environment for a computer program, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0117] A computer program, which may also be referred to or described as a program, software, software application, app, module, software module, script or code, may be written in any form of programming language including compiled or interpreted languages ​​or declarative or procedural languages; and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data, such as one or more scripts stored in a markup language document; in a single file dedicated to the program or in multiple coordinated files, such as files storing one or more modules, subroutines or portions of code. A computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communications network.

[0118] Similarly, the term "engine" is used broadly in this specification to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a specific engine; in other cases, multiple engines may be installed and run on the same computer or multiple computers.

[0119] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by a special purpose logic circuit such as an FPGA or ASIC, or by a combination of a special purpose logic circuit and one or more programmed computers.

[0120] A computer suitable for executing a computer program may be based on a general-purpose microprocessor or a special-purpose microprocessor or both, or any other kind of central processing unit. Typically, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for executing or implementing instructions and one or more storage devices for storing instructions and data. The central processing unit and the memory may be supplemented by a dedicated logic circuit or incorporated in the dedicated logic circuit. Typically, a computer will also include one or more large-capacity storage devices for storing data, such as a disk, a magneto-optical disk, or an optical disk, or may be operationally coupled to receive data from the one or more large-capacity storage devices or to transfer data to the one or more large-capacity storage devices, or both for storing data. However, a computer need not have such a device. In addition, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game controller, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive, etc.

[0121] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including by way of example semiconductor storage devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks.

[0122] To provide interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and pointing device, such as a mouse or trackball, that the user can use to provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, voice, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on the user's device in response to receiving a request from the web browser. In addition, the computer may interact with the user by sending a text message or other form of message to a personal device, such as a smart phone running a messaging application, and then receiving a response message from the user.

[0123] The data processing apparatus for implementing a machine learning model may also include, for example, dedicated hardware accelerator units for processing common and computationally intensive parts of machine learning training or production, i.e., inference, workloads.

[0124] Machine learning models can be implemented and deployed using a machine learning framework, such as the TensorFlow framework, the Microsoft Cognitive Toolkit framework, the Apache Singa framework, or the Apache MXNet framework.

[0125] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes a back-end component, such as a data server; or includes a middleware component, such as an application server; or includes a front-end component, such as a client computer with a graphical user interface, a web browser, or an app that a user can use to interact with an implementation of the subject matter described in this specification; or includes any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0126] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by means of computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server transmits data such as HTML pages to the user device, for example, for the purpose of displaying data to a user interacting with the device as a client and receiving user input from the user. Data generated at the user device, for example, the result of the user interaction, may be received from the device at the server.

[0127] Although this specification contains many specific implementation details, these should not be interpreted as limitations on the scope of any invention or that may be claimed, but rather as descriptions of features that may be specific to a particular embodiment of a particular invention. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable subcombination. In addition, although features may be described above as functioning in certain combinations and even initially claimed as such, one or more features from a claimed combination may be removed from the combination in some cases, and a claimed combination may be directed to a subcombination or variations of a subcombination.

[0128] Similarly, although operations are depicted in the drawings and described in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order, or that all illustrated operations be performed to achieve the desired results. In some cases, multitasking and parallel processing can be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0129] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the actions recited in the claims may be performed in a different order and still achieve the desired results. As an example, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A method implemented by a data processing device, the method comprising: receiving an entropy encoded version of a latent representation of data, the latent representation of the data comprising a collection of code symbols associated with an ordering; sequentially entropy decoding each code symbol in the latent representation of the data according to the ordering of the code symbols in the latent representation of the data, comprising, for each code symbol of a plurality of code symbols in the latent representation of the data: processing at least one or more preceding code symbols that precede the code symbol in the ordering of code symbols in the latent representation of the data and that have been previously entropy decoded to generate a code symbol probability distribution corresponding to the code symbol; as well as entropy decoding the code symbol using the code symbol probability distribution corresponding to the code symbol; as well as The latent representation of the data is processed using a decoder neural network to generate a reconstruction of the data from the latent representation of the data.

2. The method according to claim 1, wherein: For each code symbol of the plurality of code symbols in the latent representation of the data, processing at least one or more preceding code symbols that precede the code symbol in the ordering of code symbols in the latent representation of the data and that have been previously entropy decoded to generate a code symbol probability distribution corresponding to the code symbol comprises: At least one or more previous code symbols are processed using one or more neural networks to generate the code symbol probability distribution corresponding to the code symbol.

3. The method according to claim 2, wherein: The one or more neural networks that process the at least one or more previous code symbols to generate the code symbol probability distribution corresponding to the code symbol have been trained using a machine learning training technique for optimizing a rate-distortion performance metric.

4. The method according to any one of claims 2 to 3, wherein: For each code symbol of the plurality of code symbols in the latent representation of the data, processing at least one or more previous code symbols using one or more neural networks to generate the code symbol probability distribution corresponding to the code symbol comprises: Processing at least one or more previous code symbols using a contextual neural network to generate a context output for the code symbol; and At least the context output of the code symbol is processed using an entropy model neural network to generate the code symbol probability distribution corresponding to the code symbol.

5. The method according to claim 4, further comprising: receiving a latent representation of an entropy model used to entropy encode the latent representation of the data, wherein the latent representation of the entropy model is generated by processing the latent representation of the data using a superencoder neural network; and processing the latent representation of the entropy model using a super-decoder neural network to generate a super-decoder output; Wherein for each code symbol of the plurality of code symbols in the potential representation of the data, processing at least the context output of the code symbol using the entropy model neural network to generate the code symbol probability distribution corresponding to the code symbol comprises: The entropy model neural network is used to process both (i) the super decoder output and (ii) the context output of the code symbol to generate the code symbol probability distribution corresponding to the code symbol.

6. The method according to any one of claims 1 to 3, wherein: Each code symbol probability distribution is a Gaussian distribution convolved with a uniform distribution, and Each code symbol probability distribution is defined by probability distribution parameters, which include corresponding mean and standard deviation parameters of the Gaussian distribution.

7. The method according to any one of claims 1 to 3, wherein: The data includes one or more of the following: image data, audio data, and text data.

8. A system comprising one or more computers and one or more storage devices storing instructions, wherein the instructions, when executed by the one or more computers, cause the one or more computers to perform operations comprising: receiving an entropy encoded version of a latent representation of data, the latent representation of the data comprising a collection of code symbols associated with an ordering; sequentially entropy decoding each code symbol in the latent representation of the data according to the ordering of the code symbols in the latent representation of the data, comprising, for each code symbol of a plurality of code symbols in the latent representation of the data: processing at least one or more preceding code symbols that precede the code symbol in the ordering of code symbols in the latent representation of the data and that have been previously entropy decoded to generate a code symbol probability distribution corresponding to the code symbol; as well as entropy decoding the code symbol using the code symbol probability distribution corresponding to the code symbol; as well as The latent representation of the data is processed using a decoder neural network to generate a reconstruction of the data from the latent representation of the data.

9. The system according to claim 8, wherein: For each code symbol of the plurality of code symbols in the latent representation of the data, processing at least one or more preceding code symbols that precede the code symbol in the ordering of code symbols in the latent representation of the data and that have been previously entropy decoded to generate a code symbol probability distribution corresponding to the code symbol comprises: At least one or more previous code symbols are processed using one or more neural networks to generate the code symbol probability distribution corresponding to the code symbol.

10. The system according to claim 9, wherein: The one or more neural networks that process the at least one or more previous code symbols to generate the code symbol probability distribution corresponding to the code symbol have been trained using a machine learning training technique for optimizing a rate-distortion performance metric.

11. The system according to any one of claims 9 to 10, wherein: For each code symbol of the plurality of code symbols in the latent representation of the data, processing at least one or more previous code symbols using one or more neural networks to generate the code symbol probability distribution corresponding to the code symbol comprises: Processing at least one or more previous code symbols using a contextual neural network to generate a context output for the code symbol; and At least the context output of the code symbol is processed using an entropy model neural network to generate the code symbol probability distribution corresponding to the code symbol.

12. The system according to claim 11, wherein: The operations further include: receiving a latent representation of an entropy model used to entropy encode the latent representation of the data, wherein the latent representation of the entropy model is generated by processing the latent representation of the data using a superencoder neural network; and processing the latent representation of the entropy model using a super-decoder neural network to generate a super-decoder output; Wherein for each code symbol of the plurality of code symbols in the potential representation of the data, processing at least the context output of the code symbol using the entropy model neural network to generate the code symbol probability distribution corresponding to the code symbol comprises: The entropy model neural network is used to process both (i) the super decoder output and (ii) the context output of the code symbol to generate the code symbol probability distribution corresponding to the code symbol.

13. The system according to any one of claims 8 to 10, wherein: Each code symbol probability distribution is a Gaussian distribution convolved with a uniform distribution, and Each code symbol probability distribution is defined by probability distribution parameters, which include corresponding mean and standard deviation parameters of the Gaussian distribution.

14. The system according to any one of claims 8 to 10, wherein: The data includes one or more of the following: image data, audio data, and text data.

15. A computer-readable storage medium storing instructions which, when executed by one or more computers, cause the one or more computers to perform operations comprising: receiving an entropy encoded version of a latent representation of data, the latent representation of the data comprising a collection of code symbols associated with an ordering; sequentially entropy decoding each code symbol in the latent representation of the data according to the ordering of the code symbols in the latent representation of the data, comprising, for each code symbol of a plurality of code symbols in the latent representation of the data: processing at least one or more preceding code symbols that precede the code symbol in the ordering of code symbols in the latent representation of the data and that have been previously entropy decoded to generate a code symbol probability distribution corresponding to the code symbol; as well as entropy decoding the code symbol using the code symbol probability distribution corresponding to the code symbol; as well as The latent representation of the data is processed using a decoder neural network to generate a reconstruction of the data from the latent representation of the data.

16. The computer-readable storage medium of claim 15, wherein: For each code symbol of the plurality of code symbols in the latent representation of the data, processing at least one or more preceding code symbols that precede the code symbol in the ordering of code symbols in the latent representation of the data and that have been previously entropy decoded to generate a code symbol probability distribution corresponding to the code symbol comprises: At least one or more previous code symbols are processed using one or more neural networks to generate the code symbol probability distribution corresponding to the code symbol.

17. The computer-readable storage medium of claim 16, wherein: The one or more neural networks that process the at least one or more previous code symbols to generate the code symbol probability distribution corresponding to the code symbol have been trained using a machine learning training technique for optimizing a rate-distortion performance metric.

18. The computer-readable storage medium according to any one of claims 16 to 17, wherein: For each code symbol of the plurality of code symbols in the latent representation of the data, processing at least one or more previous code symbols using one or more neural networks to generate the code symbol probability distribution corresponding to the code symbol comprises: Processing at least one or more previous code symbols using a contextual neural network to generate a context output for the code symbol; and At least the context output of the code symbol is processed using an entropy model neural network to generate the code symbol probability distribution corresponding to the code symbol.

19. The computer-readable storage medium of claim 18, wherein: The operations further include: receiving a latent representation of an entropy model used to entropy encode the latent representation of the data, wherein the latent representation of the entropy model is generated by processing the latent representation of the data using a superencoder neural network; and processing the latent representation of the entropy model using a super-decoder neural network to generate a super-decoder output; Wherein for each code symbol of the plurality of code symbols in the potential representation of the data, processing at least the context output of the code symbol using the entropy model neural network to generate the code symbol probability distribution corresponding to the code symbol comprises: The entropy model neural network is used to process both (i) the super decoder output and (ii) the context output of the code symbol to generate the code symbol probability distribution corresponding to the code symbol.

20. The computer-readable storage medium according to any one of claims 15 to 17, wherein: Each code symbol probability distribution is a Gaussian distribution convolved with a uniform distribution, and Each code symbol probability distribution is defined by probability distribution parameters, which include corresponding mean and standard deviation parameters of the Gaussian distribution.