Data compression using integer neural networks
By using integer neural network to calculate the entropy model in the data compression system, the decoding failure problem during operation on different platforms is solved, and the system reliability and performance stability are improved.
Patent Information
- Application Number
- CN202510110089.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-09-27
- Filing Date
- 2019-09-18
- Publication Date
- 2025-06-06
AI Technical Summary
When existing data compression systems operate on different hardware and software platforms, they may lead to decoding failures due to differences in entropy models, and the implementation differences in floating-point arithmetic and numerical rounding operations increase the uncertainty of the system.
Integer neural networks are used to calculate the entropy model, and all operations are realized through integer arithmetic and lookup tables, ensuring that the system can operate deterministically on different platforms and generate the same entropy model.
It improves the reliability of the data compression system on different hardware and software platforms, reduces decoding failures due to differences in entropy model, and has a small impact on the performance of the compression system and the decompression system.
Smart Images

Figure CN120106134A_ABST
Abstract
Description
[0001] Description of the case
[0002] This application is a divisional application of a patent application with application number 201980062275.9 (PCT / US2019 / 051624), which was submitted to the China Patent Office on March 23, 2021, with an international application date of September 18, 2019, and the invention name is “Data Compression Using Integer Neural Networks”. Background Art
[0003] This specification relates to data compression.
[0004] Compressing data refers to determining a representation of the data that takes up less space in memory. The compressed data may be stored (e.g., in a logical data storage area or a physical data storage device), transmitted to a destination via a communication network (e.g., the Internet), or used in any other manner. Typically, the data can be reconstructed (approximately or exactly) from the compressed representation of the data. Summary of the invention
[0005] This specification describes a system implemented as a computer program on one or more computers in one or more locations that can reliably perform data compression and data decompression on a variety of hardware and software platforms by using integer neural networks.
[0006] According to one aspect, a method for entropy encoding data performed by one or more data processing devices is provided, the method defining a sequence including a plurality of components, wherein each component specifies a corresponding code symbol from a discrete set of predetermined possible code symbols, the method comprising: for each of the plurality of components: processing inputs including (i) a corresponding integer representation of each of one or more components of data preceding the component in the sequence, (ii) an integer representation of one or more corresponding latent variables characterizing the data, or (iii) both using an integer neural network to generate data defining a probability distribution of the component of the data over a predetermined set of possible code symbols, wherein: the integer neural network has a plurality of integer neural network parameter values, and each of the plurality of integer neural network parameter values is an integer; the integer neural network comprises a plurality of integer neural network layers, each integer neural network layer being configured to process a corresponding integer neural network layer input to generate a corresponding integer neural network layer output, and processing the integer neural network layer input to generate the integer neural network layer output comprises: generating an intermediate result by processing the integer neural network layer input according to the plurality of integer neural network layer parameters using integer valued operations; and generating the integer neural network layer output by applying an integer valued activation function to the intermediate result; and generating the entropy encoded representation of the data using the corresponding probability distribution determined for each of the plurality of components.
[0007] In some implementations, the predetermined set of possible code symbols may include a set of integer values.
[0008] In some implementations, the data may represent an image.
[0009] In some implementations, the data may represent latent representations of an image generated by processing the image using different integer neural networks.
[0010] In some implementations, generating an entropy coded representation of the data using the respective probability distribution determined for each of the plurality of components may include using an arithmetic coding process to generate an entropy coded representation of the data using the respective probability distribution determined for each of the plurality of components.
[0011] In some implementations, generating an intermediate result by processing an integer neural network layer input according to an integer neural network layer parameter set using integer-valued operations may include generating a first intermediate result by multiplying the integer neural network layer input by an integer-valued parameter matrix or convolving the integer neural network layer input by an integer-valued convolution filter.
[0012] In some implementations, the method may further include generating a second intermediate result by adding an integer-valued bias vector to the first intermediate result.
[0013] In some implementations, the method may further include generating a third intermediate result by dividing each component of the second intermediate result by an integer-valued rescaling factor, wherein the division is performed using a rounded division operation.
[0014] In some implementations, an integer-valued activation function may be defined by a lookup table that defines a mapping from each integer value in a predetermined set of integer values to a corresponding integer output.
[0015] In some implementations, a plurality of integer neural network parameter values for the integer neural network may be determined by a training process; and during training the integer neural network, the integer neural network parameter values may be stored as floating point values, and the integer neural network parameter values stored as floating point values may be rounded to integer values before being used in calculations. Rounding the integer neural network parameter values stored as floating point values to integer values may include: scaling the floating point values; and rounding the scaled floating point values to the nearest integer value. The floating point values may be transformed by a parameterized mapping before being scaled, wherein the parameterized mapping r(·) is defined by: The integer neural network parameter values may define the parameters of the convolutional filters, and the floating point values may be scaled by a factor s defined as follows: s = (max((-2 K-1 ) -1 L, (2K-1 -1) - 1 H),∈) -1 , where K is the bit width of the kernel, L is the minimum value of the set of floating-point values defining the parameters of the convolution filter, H is the maximum value of the set of floating-point values defining the parameters of the convolution filter, and ∈ is a positive constant. The floating-point values can be scaled based on the bit width of the convolution kernel.
[0016] In some implementations, an integer neural network can include an integer neural network layer configured to generate an integer neural network layer output by applying an integer-valued activation to an intermediate result, wherein the integer-valued activation function performs a clipping operation and during training the integer neural network, a gradient of the activation function is replaced with a scaled generalized Gaussian probability density.
[0017] In some implementations, the data may be processed using a neural network to generate one or more corresponding latent variables that characterize the data.
[0018] In some embodiments, for each of the multiple components, the probability distribution of the component on a predetermined set of code symbols may be a Gaussian distribution convolved with a uniform distribution, and the data defining the probability distribution of the component on the predetermined set of code symbols may include corresponding mean and standard deviation parameters of the Gaussian distribution.
[0019] According to another aspect, a method for entropy decoding data defining a sequence comprising a set of components is provided, wherein each component specifies a corresponding code symbol from a discrete set of predetermined possible code symbols. The method includes obtaining an entropy coded representation of the data; generating a corresponding reconstruction of each component of the data from the entropy coded representation of the data, for each component of the data, including: determining a corresponding probability distribution over the predetermined set of possible code symbols; and entropy decoding the component of the data using the corresponding probability distribution over the predetermined set of possible code symbols.
[0020] For each component, determining a corresponding probability distribution over a predetermined set of possible code symbols includes: processing inputs including (i) corresponding integer representations of each of one or more previously determined components of data preceding the component in the sequence of components, (ii) integer representations of one or more latent variables characterizing the data, or (iii) both, using an integer neural network to generate data defining a corresponding probability distribution over the predetermined set of possible code symbols. The integer neural network has a set of integer neural network parameter values, and each integer neural network parameter value is an integer. The integer neural network includes a set of integer neural network layers, each integer neural network layer being configured to process a corresponding integer neural network layer input to generate a corresponding integer neural network layer output, and processing the integer neural network layer input to generate the integer neural network layer output includes: generating intermediate results by processing the integer neural network layer input according to the integer neural network layer parameter set using integer valued operations; and generating the integer neural network layer output by applying an integer valued activation function to the intermediate results.
[0021] In some implementations, the data may represent latent representations of an image generated by processing the image using different integer neural networks.
[0022] In some implementations, entropy decoding the components of the data using corresponding probability distributions over a predetermined set of possible code symbols may include using an arithmetic decoding process to entropy decode the components of the data using corresponding probability distributions over a predetermined set of possible code symbols.
[0023] According to yet another aspect, a system is provided that includes one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the above-described method aspects.
[0024] According to another aspect, one or more computer storage media are provided that store instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the above-described method aspects.
[0025] It will be readily appreciated that features described in the context of one aspect may be combined with other aspects.
[0026] Particular embodiments of the subject matter described in this specification can be implemented to realize one or more of the following advantages.
[0027] This specification describes a data compression system that compresses data by processing the data using one or more neural networks to generate an entropy model, and then entropy encoding the data using the entropy model. The "entropy model" specifies the corresponding code symbol probability distribution of each code symbol in an ordered set of code symbols representing the data (i.e., a probability distribution over a set of possible code symbols, e.g., integers), as described in more detail below. A decompression system can decompress the data by reproducing the entropy model using one or more neural networks, and then entropy decoding the data using the entropy model.
[0028] Often, the compression system and the decompression system may operate on different hardware or software platforms that use, for example, different implementations of floating-point arithmetic and numerical rounding operations. Thus, performing operations using floating-point arithmetic may cause the decompression system to compute an entropy model that is slightly different from the entropy model computed by the compression system. However, in order for data compressed by the compression system to be reliably reconstructed by the decompression system, the compression system and the decompression system must use the same entropy model. Even small differences between the respective entropy models used by the compression system and the decompression system may result in catastrophic decoding failures, for example, where the data reconstructed by the decompression system is substantially different from the data compressed by the compression system.
[0029] The compression system and decompression system described in this specification both use integer neural networks to compute entropy models, i.e., neural networks that implement all operations using integer arithmetic, lookup tables, or both. The integer neural network can operate deterministically across different hardware and software platforms, i.e., independent of how different platforms implement floating point arithmetic and numerical rounding operations. Thus, the use of integer neural networks enables the compression system and the decompression system to compute the same entropy model independent of the hardware and software platforms, and thereby enables the compression system and the decompression system to operate reliably on different hardware and software platforms, i.e., by reducing the likelihood of catastrophic decoding failures.
[0030] In addition to increasing the reliability of the compression and decompression systems, the use of integer neural networks may have little impact on the performance (e.g., rate-distortion performance) of the compression and decompression systems relative to the use of floating-point neural networks.
[0031] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 An example integer neural network is shown.
[0033] Figure 2Illustration of examples of nonlinearities (activation functions) that can be applied to the intermediate results of an integer neural network layer to generate the layer output.
[0034] Figure 3 The figure shows the results of entropy decoding of the same image using: (i) an entropy model calculated using an integer neural network, and (ii) an entropy model using floating-point arithmetic.
[0035] Figure 4 is a block diagram of an example compression system for entropy encoding a data set using an entropy model computed using an integer neural network.
[0036] Figure 5 is a block diagram of an example decompression system for entropy decoding a data set using an entropy model computed using an integer neural network.
[0037] Figure 6 A table describing an example architecture of an integer neural network that may be used by a compression / decompression system is shown.
[0038] Figure 7 Figure 2 compares the rate-distortion performance of compression / decompression systems using: (i) an integer neural network, and (ii) a neural network using floating-point arithmetic.
[0039] Figure 8 is an example of comparing the frequency of decompression failure rates due to floating point rounding errors on a dataset of RGB images when using (i) a neural network implementing floating point arithmetic and (ii) an integer neural network.
[0040] Fig. 9 is a flow chart of an example process for processing integer layer inputs to generate integer layer outputs.
[0041] Fig.10 is a flow chart of an example process for processing compressed data using an entropy model computed using an integer neural network.
[0042] Fig.11 is a flow chart of an example process for reconstructing compressed data using an entropy model computed using an integer neural network.
[0043] Like reference numbers and designations throughout the various drawings refer to like elements. DETAILED DESCRIPTION
[0044] This specification describes a compression system and a decompression system that can operate reliably across a variety of hardware and software platforms by using integer neural networks (i.e., neural networks that implement all operations using integer arithmetic, lookup tables, or both).
[0045] A compression system compresses data (e.g., image data, video data, audio data, text data, or any other suitable kind of data) represented by an ordered set of code symbols (i.e., integer values) by entropy encoding the code symbols using an entropy model (e.g., arithmetic coding or Huffman coding). As used throughout this document, a "code symbol" refers to an element extracted from a discrete set of possible elements, e.g., an integer value. The entropy model defines a corresponding code symbol probability distribution (i.e., a probability distribution over the set of possible code symbols) corresponding to each code symbol in the ordered set of code symbols representing the data. A decompression system may entropy decode the data using the same entropy model as the compression system.
[0046] Typically, a compression system can compress data more efficiently by using a "conditional" entropy model, that is, an entropy model that is dynamically calculated based on the compressed data, rather than, for example, a predefined entropy model. The decompression system must also dynamically calculate the entropy model in order to decompress the data. Typically, the compression and decompression systems can operate on different hardware or software platforms that, for example, implement floating-point arithmetic and numerical rounding operations differently, which can cause the compression and decompression systems to calculate different entropy models. In one example, the compression system can be implemented in a data center, while the decompression system can be implemented on a user device. However, in order for data compressed by the compression system to be reliably reconstructed by the decompression system, the compression system and the decompression system must use the same entropy model. Even small differences between the respective entropy models used by the compression system and the decompression system can result in catastrophic decoding failures.
[0047] The compression and decompression system described in this specification uses an integer neural network that operates deterministically across hardware and software platforms to calculate an entropy model for entropy encoding / decoding data. Therefore, the compression and decompression system described in this specification can operate reliably across different hardware and software platforms while greatly reducing the risk of decoding failures due to differences between entropy models used to entropy encode / decode data.
[0048] As used throughout this document, the term integer arithmetic refers to the basic arithmetic operations (e.g., addition, subtraction, multiplication, and division) applied to integer inputs to produce integer outputs. In the case of division, any fractional portion of the result of the division may be ignored or rounded to ensure that the output is an integer value. As used throughout this document, a lookup table refers to a data structure that stores data that defines a mapping from each input value in a predetermined set of input values to a pre-computed output value.
[0049] Typically, the compression system and the decompression system may be co-located or remotely located, and the compressed data generated by the compression system may be provided to the decompression system in any of a variety of ways. For example, the compressed data may be stored (e.g., in a physical data storage device or a logical data storage area) and then subsequently retrieved from the storage and provided to the decompression system. As another example, the compressed data may be transmitted over a communication network (e.g., the Internet) to a destination where the compressed data is subsequently retrieved and provided to the decompression system.
[0050] The compression system may be used to compress any suitable kind of data, for example image data, audio data, video data or text data.
[0051] These and other features are described in greater detail below.
[0052] Figure 1 An example integer neural network 100 is shown, which may be implemented as a computer program on one or more computers in one or more locations. The integer neural network 100 is configured to process integer-valued network inputs 102 using operations implemented by integer arithmetic, lookup tables, or both to generate integer-valued network outputs 104. A compression system (e.g., as described in reference Figure 4 ) and a decompression system (e.g., as described in reference Figure 5 As described herein, an integer neural network (e.g., integer neural network 100) can be used to reliably compute the same entropy model despite operating on different hardware or software platforms.
[0053] Integer neural network 100 processes network input 102 using one or more integer neural network layers (e.g., integer neural network layer 106) to generate network output 104. In general, network input 102 and network output 104 may be represented as ordered collections of integer values, such as vectors or matrices of integer values.
[0054] Each integer neural network layer 106 is configured to process integer-valued layer inputs 108 using operations implemented by integer arithmetic, lookup tables, or both to generate integer-valued layer outputs 110. For example, an integer neural network layer 106 may process a layer input u to generate a layer output w according to the following operation:
[0055]
[0056] w=g(v) (2)
[0057] Where, in order: a linear transformation H is applied to the layer input u, a bias vector b is added, the result is element-wise divided by a vector c to produce an intermediate result v, and then an element-wise nonlinearity (activation function) g(·) is applied to the intermediate result to generate the layer output w. The integer neural network layer parameters H, b, and c and all intermediate results are defined as integers. To define the intermediate result v as an integer, the element-wise division operation can be implemented as a round-down division operation (equivalent to division followed by rounding to the nearest integer). In a programming language such as C, this can be implemented with integer operands m and n:
[0058]
[0059] The \\ is a floor operation.
[0060] The form of the linear transformation H mentioned in equation (1) depends on the type of integer neural network layer. In one example, the integer neural network layer can be a fully connected layer, and the linear transformation H can be implemented as a matrix multiplication, i.e., it transforms the layer input u by multiplying the layer input u with an integer-valued parameter matrix. In another example, the integer neural network layer can be a convolutional layer, and the linear transformation H can be implemented as a convolution operation, i.e., it transforms the layer input u by convolving the layer input u with one or more integer-valued convolution kernels. Typically, applying the linear transformation H to the layer input u involves a matrix multiplication or a convolution operation, which increases the possibility of integer overflow of the intermediate result v. Division by the parameter c (which is optional) can reduce the possibility of integer overflow of the intermediate result v.
[0061] The nonlinearity g(·) mentioned in equation (2) can be implemented in any of a variety of ways. In one example, the nonlinearity g(·) can be a rectified linear unit (ReLU) nonlinearity that clips the value of the intermediate result v within a predetermined range, for example:
[0062] g QReLU (v)=max(min(v,255),0) (4)
[0063] Wherein, in this example, g(.) clips the value of the intermediate result to the range [0, 255]. In another example, the nonlinearity g(·) can be a hyperbolic tangent nonlinearity, such as:
[0064]
[0065] where Q(·) represents a quantization operator that rounds its input to the nearest integer value. In this example, the hyperbolic tangent nonlinearity can be represented by a lookup table to ensure that its output is independent of the implementation of the hyperbolic tangent function on any particular platform.
[0066] In some cases, the nonlinearity g(·) may be scaled so that its range matches the bit width used to represent the integer layer output. The bit width of an integer refers to the number of bits (e.g., binary digits) used to represent the integer. In one example, the integer layer output can be a signed integer with a bit width of 4 binary digits, which means that it can take integer values in the range of -7 to 7, and the nonlinearity g(·) can be scaled to generate integer outputs in the range of -7 to 7, for example, as in equation (5). Scaling the nonlinearity g(·) to match its range with the bit width of the integer layer output can enable the integer layer to generate richer layer outputs by maximizing the use of the dynamic range of the layer output.
[0067] Any suitable integer format may be used to represent the learning parameters and (intermediate) outputs of the integer neural network layer 106. For example, referring to equations (1)-(2), the learning parameters and (intermediate) outputs of the integer neural network layer 106 may have the following format:
[0068] H: 8-bit signed integer
[0069] b, v: 32-bit signed integer
[0070] c: 32-bit unsigned integer
[0071] W: 8-bit unsigned integer
[0072] Integer neural network 100 may have any suitable neural network architecture. Figure 6 A more detailed description of the compression system (ref. Figure 4 Description) and decompression system (reference Figure 5 Description) using an integer neural network.
[0073] During the training process, the parameter values of the integer neural network are iteratively adjusted to optimize an objective function, such as a rate-distortion objective function, as described below with reference to equations (15)-(18). More specifically, in each of a plurality of training iterations, the parameter values of the integer neural network may be adjusted based on a gradient descent optimization process (e.g., Adam) using the gradient of the objective function with respect to the parameters of the integer neural network. In order to efficiently accumulate small gradient signals during training, the integer neural network parameters are stored as floating point values, but are mapped (rounded) to integer values before being used for calculations. That is, in each training iteration, the floating point representations of the integer neural network parameters are mapped to integer values before being used to calculate the integer neural network parameter value updates for the training iteration. Next, some examples of mapping the floating point representations of the integer neural network parameters H, b, and c (described with reference to equation (1)) to integer representations are described.
[0074] In one example, the integer-valued bias vector parameter b may be obtained by mapping the floating-point-valued bias vector parameter b′ as follows:
[0075] b=Q(2 K b′)(6)
[0076] where K is the bit width b of the integer representation, and Q(·) represents a quantization operator that rounds its input to the nearest integer value. The floating-point bias vector parameter b′ is scaled by 2 K Integer layers can be enabled to generate richer layer outputs by maximizing the use of the dynamic range of the integer-valued bias vector parameter b.
[0077] In another example, an integer-valued division parameter C may be obtained by mapping a floating-point-valued division parameter map c′ as follows:
[0078] c=Q(2 K r(c′)) (7)
[0079]
[0080] where K is the bit width of the integer representation c, and Q(·) represents a quantization operator that rounds its input to the nearest integer value. In this example, the parameterized mapping r(.) ensures that the value of c is always positive while gently scaling down the gradient magnitude at c′ close to zero, which can reduce the possibility of perturbations that cause large fluctuations in the intermediate result v, especially when c′ is small.
[0081] In another example, the linear transformation H can be given by H = [h 1 ,h 2 , ..., h N ](that is, each h i is a convolution transform defined by, for example, a corresponding convolution filter defined by a two-dimensional (2D) or three-dimensional (3D) array of integer values, and a linear transformation parameter H′=[h′ 1 , h′ 2 , ..., h′ N ] can be mapped to an integer-valued linear transformation parameter H as:
[0082] h i =Q(s(h′ i )·h′ i ), i=1,...,N (9)
[0083] s(h′)=(max((-2 K-1 ) -1 L, (2 K-1 -1) -1 H),∈)-1 (10)
[0084] where Q(·) represents a quantization operator that rounds its input to the nearest integer value, K is the bit width of each component of the integer representation of the convolution filter, L is the minimum floating point value in the convolution filter h′, H is the maximum floating point value in the convolution filter h′, and ∈ is a positive constant. In this example, the scaling parameter S rescales each convolution filter so that at least one of its minimum and maximum parameter values hits the dynamic range boundary (-2 K-1 and 2 K-1 -1) while keeping zero at zero. Given the floating point representation of the convolution filter, this represents the best possible quantization of the convolution filter and thus maximizes accuracy. The positive constant ∈ reduces instabilities and errors caused by division by zero.
[0085] During training, the gradient of the objective function is calculated with respect to the floating point valued parameters of the integer neural network. However, in some cases, the gradient cannot be calculated directly because some operations performed by the integer neural network are non-differentiable. In these cases, an approximation to the gradient of the non-differentiable operation can be applied to train the integer neural network. The following are some examples of such approximations.
[0086] In one example, the mapping from floating point values to integer values of integer neural network parameters uses a non-differentiable quantization operator, for example, as described with reference to equations (6), (7), and (9). In this example, the gradient of the quantization operator can be replaced by an identity function. In particular, with reference to equations (6), (7), and (9), the gradient can be calculated as:
[0087]
[0088] in Denotes the partial derivative operator. In the example illustrated in equation (11), the scaling parameter s(h′) is i ) -1 is treated as if it is constant (i.e., although it depends on h′ i ).
[0089] In another example, the rounding division operation (described with reference to equation (1)) is not differentiable. In this example, the gradient of the rounded division operation can be replaced by the gradient of the floating point division operation.
[0090] In another example, the nonlinearity (activation function) g(·) (please refer to equation (2)) can be non-differentiable, and the gradient of the nonlinearity can be replaced by the gradient of a continuous function that approximates the nonlinearity. In a specific example, the nonlinearity can be a quantized ReLU, as described with reference to equation (4), and the gradient of the nonlinearity can be replaced by a scaled generalized Gaussian probability density with a shape parameter β, for example, given by:
[0091]
[0092] in, represents the partial derivative operator, And K is the bit width of the integer layer output.
[0093] After training is complete, integer-valued network parameters (e.g., H, b, and C, as described with reference to equation (1)) are calculated more than once from corresponding floating-point-valued network parameters (e.g., according to equations (6)-(10)). Thereafter, the integer-valued network parameters can be used for inference, e.g., for calculating an entropy model used by a compression / decompression system, which will be described in more detail below.
[0094] Figure 2 An example of a graph illustrating a nonlinearity (activation function) that may be applied to an intermediate result of an integer neural network layer to generate a layer output, e.g., as described with reference to equation (2). In particular, graph 202 uses circles (e.g., circle 204) to illustrate a quantized ReLU nonlinearity g that clips integer values to the range [0, 15]. QReLU (v) = max(min(v, 15), 0). This nonlinearity can be implemented deterministically using a lookup table or using a digital clipping operation. Examples of scaled generalized Gaussian probability densities with different values of shape parameter β are plotted together with the quantized ReLU, for example, as illustrated by line 206. During training, the gradient of the quantized ReLU can be replaced by the gradient of the scaled generalized Gaussian probability density, as described earlier. Graph 208 uses circles (e.g., circle 210) to illustrate the quantized hyperbolic tangent nonlinearity This nonlinearity can be implemented deterministically using a lookup table.The corresponding continuous hyperbolic tangent nonlinearity used to compute the gradient is plotted along with the quantized hyperbolic tangent nonlinearity, for example, as illustrated in line 212 .
[0095] Figure 3The results of entropy decoding the same image using: (i) an entropy model calculated using an integer neural network - 302, and (ii) an entropy model calculated using floating point arithmetic - 304. When the entropy model is calculated using floating point arithmetic, the image is initially decoded correctly (starting from the upper left corner) until floating point rounding errors cause slight differences between the corresponding entropy models calculated by the compression and decompression systems, at which point the errors propagate catastrophically, causing the image to be decoded incorrectly.
[0096] In general, a compression system may use one or more integer neural networks to determine an entropy model for entropy encoding an ordered collection of code symbols representing a data set (e.g., image data, video data, or audio data). For example, an ordered collection of code symbols representing the data may be obtained by quantizing a representation of the data into an ordered collection of floating point values by rounding each floating point value to the nearest integer value. As previously described, the entropy model specifies a corresponding code symbol probability distribution corresponding to each component (code symbol) of the ordered collection of code symbols representing the data. The compression system may generate a code symbol probability distribution for each component of the ordered collection of code symbols by processing, using one or more integer neural networks: (i) integer representations of one or more previous code symbols, (ii) integer representations of one or more latent variables representing the data, or (iii) both. A "latent variable" representing the data refers to an alternative representation of the data, which is generated, for example, by processing the data using one or more neural networks.
[0097] Figure 4 is a block diagram of an example compression system 400 for entropy encoding a data set using an entropy model computed by an integer neural network. Compression system 400 is an example system implemented as a computer program on one or more computers in one or more locations that implement the systems, components, and techniques described below. Compression system 400 is provided for purposes of illustration only, and in general, various components of compression system 400 are optional, and other architectures of the compression system are possible.
[0098] The compression system 400 processes the input data 402 to generate compressed data 404 representing the input data 402 using: (1) an encoder integer neural network 406, (2) a super encoder integer neural network 408, (3) a super decoder integer neural network 410, (4) an entropy model integer neural network 412, and optionally (5) a context integer neural network 414. As will be described with reference to Figure 5 As described in more detail, the integer neural network used by the compression system (and the integer neural network used by the decompression system) is jointly trained using a rate-distortion objective function. In general, each integer neural network described in this document can have any suitable integer neural network architecture that enables it to perform its described function. Figure 6 An example architecture of an integer neural network used by the compression system and the decompression system is described in more detail.
[0099] The encoder integer neural network 406 is configured to process the quantized (integer) representation (x) of the input data 402 to generate an ordered collection of code symbols (integer) 420 representing the input data. In one example, the input data may be an image, the encoder integer neural network 406 may be a convolutional integer neural network, and the code symbols 420 may be a multi-channel feature map output by the final layer of the encoder integer neural network 406.
[0100] The compression system 400 uses a super encoder integer neural network 408, a super decoder integer neural network 410, and an entropy model integer neural network 412 to generate a conditional entropy model for entropy encoding code symbols 420 representing input data, as will be described in more detail below.
[0101] The superencoder integer neural network 408 is configured to process the code symbols 420 to generate a set of latent variables (characterizing the code symbols) referred to as "super-priors" 422(z) (sometimes referred to as "hyper-parameters"). In one example, the superencoder integer neural network 408 can be a convolutional integer neural network, and the super-prior 422 can be a multi-channel feature map output by the final layer of the superencoder integer neural network 408. The super-prior implicitly characterizes an entropy model associated with the input data that will enable the code symbols 420 representing the input data to be efficiently compressed.
[0102] The compressed data 404 typically includes a compressed representation of the hyper-prior 422 to enable the decompression system to regenerate the conditional entropy model. To this end, the compression system 400 generates a compressed representation 426 of the hyper-prior 422, for example, as a bit string, i.e., a binary digit string. In one example, the compression system 400 compresses the hyper-prior 422 using an entropy encoding engine 438 according to a predetermined entropy model that specifies one or more predetermined code symbol probability distributions.
[0103] The super-decoder integer neural network 410 is configured to process the super-prior 422 to generate a super-decoder output 428 (ψ), and the entropy model integer neural network 412 is configured to process the super-decoder output 428 to generate a conditional entropy model. That is, the super-decoder 410 and the entropy model integer neural network 412 jointly decode the super-prior to generate an output that explicitly defines the conditional entropy model.
[0104] The conditional entropy model specifies a corresponding code symbol probability distribution corresponding to each code symbol 420 representing the input data. In general, the output of the entropy model integer neural network 412 specifies the distribution parameters that define each code symbol probability distribution of the conditional entropy model. In one example, each code symbol probability distribution of the conditional entropy model can be a Gaussian distribution (parameterized by mean and standard deviation parameters) convolved with a unit uniform distribution. In this example, the output of the entropy model integer neural network 412 can be the mean parameter of the Gaussian distribution and standard deviation parameters Specified as:
[0105]
[0106] Where N is the number of code symbols in the ordered set of code symbols 420 representing the input data, μ min is the minimum allowable mean value, μ max is the maximum permissible average value, is the integer value output by the last layer of the entropy model integer neural network with L possible values in the range [0, L-1], σ min is the minimum allowable standard deviation, σ max is the maximum allowed standard deviation, and is the integer value output by the last layer of the entropy model integer neural network with L possible values in the range [0, L-1]. In this example, during training, the gradient can be determined by the reformulated back propagation provided by equations (13) and (14). After training, the code symbol probability distribution can be represented by a lookup table by pre-calculating all possible code symbol probability values based on (i) code symbols and (ii) mean and standard deviation parameters.
[0107] Optionally, the compression system 400 can additionally use a context integer neural network 414 in determining the conditional entropy model. The context integer neural network 414 is configured to autoregressively process the code symbols 420 (i.e., in accordance with the ordering of the code symbols) to generate a corresponding integer "context output" 430 (Φ) for each code symbol. The context output for each code symbol depends only on the code symbols that precede it in the ordered set of code symbols representing the input data, and not on the code symbol itself or on the code symbols that immediately follow it. The context output 430 for a code symbol can be understood as causal context information that can be used by the entropy model integer neural network 412 to generate a more accurate code symbol probability distribution for the code symbol.
[0108] The entropy model integer neural network 412 is capable of processing the context outputs 430 (i.e., in addition to the super decoder outputs 428) generated by the context integer neural network 414 to generate a conditional entropy model. In general, the code symbol probability distribution for each code symbol depends on the context outputs of the code symbol, and optionally, on the context outputs of the code symbols preceding the code symbol, but not on the context outputs of the code symbols immediately following the code symbol. As will be described with reference to Figure 5 Described in more detail, this results in a causal dependency of the conditional entropy model on the code symbols representing the input data, which ensures that the decompression system can regenerate the conditional entropy model from the compressed data.
[0109] In contrast to the super-prior 422 which must be included as side information in the compressed data 404 (thus increasing the overall compressed file size), the autoregressive context integer neural network 414 provides a "free" source of information (discounting computational cost) because it does not require the addition of any side information. Jointly training the context integer neural network 414 and the superencoder integer neural network 408 enables the super-prior 422 to store information that is complementary to the context output 430, while avoiding information that can be accurately predicted using the context output 430.
[0110] The entropy coding engine 432 is configured to compress the code symbols 420 by entropy coding the code symbols 420 representing the input data according to the conditional entropy model. The entropy coding engine 432 can implement any suitable entropy coding technique, such as arithmetic coding technique, range coding technique, or Huffman coding technique. The compressed code symbols 434 can be represented in any of a variety of ways, such as, for example, as a bit string.
[0111] The compression system 400 generates compressed data 404 based on: (i) compressed code symbols 434 and (ii) compressed hyper-prior 426. For example, the compression system can generate compressed data by concatenating respective bit strings representing the compressed code symbols 434 and the compressed hyper-prior 426.
[0112] Alternatively, the compression system 400 may use the context integer neural network 414 instead of the super prior 422 to determine an entropy model for entropy encoding code symbols representing data. In these cases, the compression system 400 does not use the super encoder integer neural network 408 or the super decoder integer neural network 410. Instead, the compression system 400 generates an entropy model by autoregressively processing the code symbols representing data 420 using the context integer neural network 100 to generate a context output 430 and then processing the context output 430 using the entropy model integer neural network 412.
[0113] Alternatively, the encoder integer neural network 406 may be implemented using floating point arithmetic instead of integer arithmetic and a lookup table, i.e., the encoder neural network may be implemented as a conventional floating point neural network instead of an integer neural network. The encoder neural network 406 is not used to calculate the entropy model, and therefore the implementation of the encoder neural network 406 using floating point arithmetic does not affect the reproducibility of the decompression system to the entropy model. However, the implementation of the encoder neural network using floating point arithmetic may still result in the decompression system being unable to accurately reconstruct the original input data, whereas the implementation of the encoder neural network as an integer neural network may enable the decompression system to accurately reconstruct the original input data.
[0114] In general, the decompression system can use one or more integer neural networks to reproduce the same entropy model calculated by the compression system, and then use the entropy model to entropy decode the code symbols. More specifically, the decompression system can generate a code symbol probability distribution for each component of an ordered collection of code symbols by using one or more integer neural networks to process: (i) integer representations of one or more previous code symbols, (ii) integer representations of one or more latent variables that characterize the data, or (iii) both. As a result of using integer neural networks, the decompression system can calculate an entropy model that exactly matches the entropy model calculated by the compression system even if the compression system and the decompression system are running on different hardware or software platforms (e.g., using different implementations of floating-point operations).
[0115] Figure 5 is a block diagram of an example decompression system 500 for entropy decoding a data set using an entropy model calculated by an integer neural network. Decompression system 500 is an example system that is implemented as a computer program on one or more computers in one or more locations that implement the systems, components, and techniques described below. Decompression system 500 is described here for illustration purposes only, and in general, various components of decompression system 500 are optional, and other architectures of the decompression system are possible.
[0116] The decompression system 500 processes the compressed data 404 generated by the compression system to generate a reconstruction 502 that approximates the original input data using: (1) a super decoder neural network 510, (2) an entropy model integer neural network 412, (3) a decoder integer neural network 504, and optionally (4) a context integer neural network 114. The super decoder integer neural network 410, the entropy model integer neural network 412, and the context integer neural network 414 used by the decompression system share the same parameter values as the corresponding integer neural networks used by the compression system.
[0117] To regenerate the conditional entropy model, the decompression system 500 obtains the super-prior 422 from the compressed data 404. For example, the decompression system 500 may obtain the super-prior 422 by entropy decoding a compressed representation 426 of the super-prior 422 included in the compressed data 404 using the entropy decoding engine 506. In this example, the entropy decoding engine 506 may entropy decode the compressed representation 426 of the super-prior 422 using the same (e.g., predetermined) entropy model used to entropy encode the compressed representation 426 of the super-prior 422.
[0118] The super decoder integer neural network 410 is configured to process the quantized super prior 422 to generate a super decoder output 428 (Ψ), and the entropy model integer neural network 412 is configured to process the super decoder output 428 to generate a conditional entropy model, i.e., in a manner similar to the compression system. The entropy decoding engine 208 is configured to entropy decode the compressed code symbols 434 included in the compressed data 404 according to the conditional entropy model to recover the code symbols 420.
[0119] In the case where the compression system uses the context integer neural network 414 to determine the conditional entropy model, the decompression system 500 also uses the context integer neural network 114 to regenerate the conditional entropy model. Figure 4 As described, the context integer neural network 414 is configured to autoregressively process the code symbols 420 representing the input data to generate a corresponding context output 430 for each code symbol. After initially receiving the compressed data 404, the decompression system 500 does not have access to the complete set of decompressed code symbols 420 provided as input to the context integer neural network 414. As will be described in more detail below, the decompression system 500 accounts for this by sequentially decompressing the code symbols 420 according to the ordering of the code symbols. The context output 430 generated by the context integer neural network 414 is provided to the entropy model integer neural network 412, which processes the context output 430 along with the super decoder output 428 to generate a conditional entropy model.
[0120] To illustrate that the decompression system 500 does not initially have access to the complete set of decompressed code symbols 420 provided as input to the context integer neural network 414, the decompression system 500 decompresses the code symbols 420 in a sorted order according to the code symbols. In particular, the decompression system may decompress the first code symbol using, for example, a predetermined code symbol probability distribution. To decompress each subsequent code symbol, the context integer neural network 414 processes one or more previous code symbols (i.e., which have already been decompressed) to generate a corresponding context output 430. The entropy model integer neural network 412 then processes (i) the context output 430, (ii) the super decoder output 428, and optionally (iii) one or more previous context outputs 430 to generate a corresponding code symbol probability distribution, which is then used to decompress the code symbol.
[0121] The decoder integer neural network 504 is configured to process the ordered collection of code symbols 420 to generate a reconstruction 502 that approximates the input data. That is, the operations performed by the decoder integer neural network 504 approximately cause the code symbols 420 to be represented by the reference 420. Figure 4 Describes the inverse of the operation performed by the encoder integer neural network.
[0122] The compression system and the decompression system can be jointly trained using machine learning training techniques (e.g., stochastic gradient descent) to optimize a rate-distortion objective function. More specifically, the encoder integer neural network, the super encoder integer neural network, the super decoder integer neural network, the context integer neural network, the entropy model integer neural network, and the decoder integer neural network can be jointly trained to optimize the rate-distortion objective function. In one example, the rate-distortion objective function (“performance metric”) It can be given by the following formula:
[0123]
[0124] in It means that the input data The code symbols are entropy encoded in the conditional entropy model The probability of Super prior In the entropy model used to entropy encode the super prior , λ is a parameter that determines the rate-distortion tradeoff, and refers to the reconstruction of the input data x and the input data In the rate-distortion objective functions described in reference equations (15) to (18), R latent represents the size of the compressed code symbol representing the input data (e.g., in bits), R hyper-priorcharacterizes the size of the compressed hyper-prior (e.g., in bits), and E reconstruction Characterizes the difference ("distortion") between the input data and a reconstruction of the input data.
[0125] Typically, a more complex hyper-prior can specify a more accurate conditional entropy model that enables the code symbols representing the input data to be compressed at a higher rate. However, increasing the complexity of the hyper-prior can cause the hyper-prior itself to be compressed at a lower rate. By jointly training the compression system and the decompression system, the balance between (i) the size of the compressed hyper-prior, and (ii) the increased compression rate from the more accurate entropy model can be learned directly from the training data.
[0126] In some implementations, the compression system and the decompression system do not use an encoder neural network or a decoder neural network. In these implementations, as described above, the compression system can generate code symbols representing the input data by directly quantizing the input data, and the decompression system can generate a reconstruction of the input data as a result of decompressing the code symbols.
[0127] Figure 6 A table 600 is shown that describes an example architecture of an integer neural network used by a compression / decompression system in the particular case where the input data consists of an image. More specifically, table 600 describes example architectures of an encoder integer neural network 602, a decoder integer neural network 604, a super encoder integer neural network 606, a super decoder integer neural network 608, a context integer neural network 610, and an entropy model integer neural network 612.
[0128] Each row of table 600 corresponds to a corresponding layer. Convolutional layers are designated with a "Conv" prefix followed by the kernel size, number of channels, and downsampling stride. For example, the first layer of encoder integer neural network 602 uses a 5×5 kernel with 192 channels and a stride of 2. The "Deconv" prefix corresponds to upsampling convolution, while "Masked" corresponds to masked convolution. GDN stands for generalized divisive normalization, and IGDN is the inverse of GDN.
[0129] In reference Figure 6 In the example architecture described, the entropy model integer neural network uses a 1×1 kernel. This architecture enables the entropy model integer neural network to generate a conditional entropy model with the following property: the code symbol probability distribution corresponding to each code symbol does not depend on the context output corresponding to the subsequent code symbol (as described earlier). As another example, the same effect can be achieved by using a masked convolution kernel.
[0130] Figure 7Graph 700 shows a comparison of the rate-distortion performance of compression / decompression systems using: (i) an integer neural network (702) and (ii) a neural network implementing floating point arithmetic (704). The horizontal axis of graph 700 indicates the number of bits per pixel of the compressed data, and the vertical axis indicates the peak signal-to-noise ratio (PSNR) of the reconstructed data (to the left and up is better). It can be appreciated that the use of an integer neural network does little to change the rate-distortion performance of the compression / decompression system, but enables the system to be reliably deployed on different hardware and software platforms.
[0131] Figure 8 804). An example of comparing the frequency of decompression failure rates due to floating point rounding errors on an RGB data set when using (i) a neural network (802) that implements floating point arithmetic, and (ii) an integer neural network (804). When the compression and decompression systems are implemented on the same platform (e.g., 806), no decompression failures occur when using either system. When the compression and decompression systems are implemented on different platforms (e.g., different CPUs, different GPUs, or one on a CPU and one on a GPU), a large number of decompression failures occur when using a neural network that implements floating point arithmetic, while no decompression failures occur when using an integer neural network. It can be appreciated that the use of integer neural networks can greatly improve the reliability of compression / decompression systems implemented on different hardware or software platforms.
[0132] Fig. 9 is a flow chart of an example process 900 for processing integer layer inputs to generate integer layer outputs. For convenience, process 900 will be described as being performed by a system of one or more computers located in one or more locations. For example, an integer neural network, e.g., Figure 1 The integer neural network 100 is appropriately programmed according to the present specification and can perform the process 900.
[0133] The system receives integer layer input (902). The integer layer input may be represented as an ordered collection of integer values, such as a vector or matrix of integer values.
[0134] The system generates intermediate results (904) by processing integer neural network layer inputs according to values of integer neural network layer parameters using integer valued operations. For example, the system can generate a first intermediate output by multiplying the layer input by an integer valued parameter matrix or by convolving the layer input with an integer convolution filter. The system can generate a second intermediate output by adding an integer valued bias vector to the first intermediate result. The system can generate a third intermediate result by dividing each component of the second intermediate result by an integer valued scaling factor, wherein the division is performed using a rounding division operation. An example of generating intermediate results is described with reference to equation (1). The integer valued operations can be implemented using integer arithmetic or using a precomputed lookup table.
[0135] The system generates a layer output by applying an integer-valued activation function to the intermediate result (906). The integer-valued activation function can be, for example, a quantized ReLU activation function (e.g., as described with reference to equation (4)) or a quantized hyperbolic tangent activation function (e.g., as described with reference to equation (5)). The activation function can be implemented, for example, using integer arithmetic or using a lookup table that defines a mapping from each integer value in a predetermined set of integer values to a corresponding pre-computed integer output.
[0136] Fig.10 1 is a flow chart of an example process 1000 for compressing data using an entropy model calculated using an integer neural network. For convenience, process 1000 is described as being performed by a system of one or more computers located in one or more locations. For example, a compression system, such as Figure 4 The compression system 400, appropriately programmed according to the present specification, can perform process 4000.
[0137] The system receives data to be compressed (1002). The data may be any suitable kind of data, for example, image data, audio data, video data, or text data.
[0138] The system generates a representation of the data as an ordered collection (e.g., a sequence) of code symbols (1004). In one example, the system quantizes the representation of the data into an ordered collection of floating point values to generate a representation of the data as an ordered collection of integer values (code symbols). In another example, after quantizing the data, the system can process the data using an encoder integer neural network and identify the output of the encoder integer neural network as code symbols representing the data. In this example, the output of the encoder integer neural network can be referred to as a "latent representation" of the data.
[0139] The system generates an entropy model using one or more integer neural networks that specifies a corresponding code symbol probability distribution for each component (code symbol) of an ordered collection of code symbols representing data (1006). For example, the system can generate a code symbol probability distribution for each code symbol by processing, using one or more integer neural networks, inputs including (i) corresponding integer representations of each of one or more components (code symbols) preceding the component (code symbol), (ii) integer representations of one or more latent variables representing the data, or (iii) both. The latent variables representing the data can be generated, for example, by processing the code symbols representing the data using one or more other integer neural networks. Reference Figure 4 An example of using an integer neural network to generate an entropy model is described in more detail. Figure 4 In the examples described, the latent variables that characterize the data are called "hyper-priors."
[0140] The system entropy encodes code symbols representing the data using an entropy model (1008). That is, the system generates an entropy-encoded representation of the data using a corresponding code symbol probability distribution determined for each component (code symbol) of the ordered set of code symbols representing the data. The system can use any suitable entropy encoding technique for entropy encoding the code symbols, such as an arithmetic coding process.
[0141] The system uses the entropy coded code symbols to determine a compressed representation of the data (1010). Figure 4 An example of determining a compressed representation of data from entropy encoded code symbols is described.
[0142] Fig.11 1 is a flow chart of an example process 1100 for reconstructing compressed data using an entropy model calculated using an integer neural network. For convenience, process 1100 is described as being performed by a system of one or more computers located in one or more locations. For example, a decompression system, such as Figure 5 The decompression system 500 can perform process 1100 when appropriately programmed according to the present specification.
[0143] The system obtains compressed data, for example, from a data storage (e.g., a logical data storage area or a physical data storage device), or as a transmission over a data communication network (e.g., the Internet) (1102). Typically, the compressed data includes an entropy coded representation of an ordered collection of code symbols representing the original data. Fig.10 Describes an example process for generating compressed data.
[0144] The system uses one or more integer neural networks to reproduce an entropy model (1104) for entropy encoding code symbols representing data. The entropy model specifies a corresponding code symbol probability distribution for each component (code symbol) of an ordered set of code symbols representing the data. The system can generate a code symbol probability distribution for each component (code symbol) by processing inputs including (i) a corresponding integer representation of each of one or more components (code symbols) preceding the component (code symbol), (ii) an integer representation of one or more latent variables representing the data, or (iii) both using one or more integer neural networks. Reference Figure 5 An example of using an integer neural network to reproduce the entropy model is described in more detail. Figure 5 In the described examples, latent variables characterizing the data, called "super-priors", are included in the compressed data, ie, in addition to the code symbols representing the entropy encoding of the data.
[0145] The system entropy decodes each component (code symbol) of the ordered collection of code symbols representing the data using the corresponding code symbol probability distribution specified by the entropy model 1106. The system can entropy decode the code symbols using, for example, an arithmetic decoding or Huffman decoding process.
[0146] The system generates a reconstruction of the original data (1108). For example, the entropy decoded code symbols themselves can represent a reconstruction of the input data. As another example, the system can generate a reconstruction of the original data by using a decoder integer neural network (e.g., as described in reference Figure 5 ) processes the entropy decoded code symbols to generate a reconstruction of the original data.
[0147] This specification uses the term "configured" in conjunction with system and computer program components. For a system of one or more computers to be configured to perform a particular operation or action, it is meant that the system has installed thereon software, firmware, hardware, or a combination of software, firmware, hardware that, in operation, causes the system to perform those operations or actions. For one or more computer programs to be configured to perform a particular operation or action, it is meant that the one or more programs include instructions that, when executed by a data processing device, cause the device to perform the operation or action.
[0148] Embodiments of the subject matter and functional operations described in this specification may be implemented with digital electronic circuits, with tangibly implemented computer software or firmware, with computer hardware including the structures disclosed in this specification and their structural equivalents, or with a combination of one or more of them. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by a data processing device or for controlling the operation of a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access storage device, or a combination of one or more of them. Alternatively or in addition, the program instructions may be encoded on an artificially generated propagation signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device for execution by a data processing device.
[0149] The term "data processing apparatus" refers to data processing hardware and includes all kinds of apparatus, devices and machines for processing data, including by way of example a programmable processor, a computer or a plurality of processors or computers. An apparatus may also be or further include special-purpose logic circuitry, for example, an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). An apparatus may optionally include, in addition to hardware, code that creates an execution environment for a computer program, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0150] A computer program, which may also be referred to or described as a program, software, software application, app, module, software module, script or code, may be written in any form of programming language including compiled or interpreted languages or declarative or procedural languages; and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data, such as one or more scripts stored in a markup language document; in a single file dedicated to the program or in multiple coordinated files, such as files storing one or more modules, subroutines or portions of code. A computer program may be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a data communications network.
[0151] Similarly, the term "engine" is used broadly in this specification to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Typically, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a specific engine; in other cases, multiple engines may be installed and run on the same computer or multiple computers.
[0152] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by a special purpose logic circuit such as an FPGA or ASIC, or by a combination of a special purpose logic circuit and one or more programmed computers.
[0153] A computer suitable for executing a computer program may be based on a general-purpose microprocessor or a special-purpose microprocessor or both, or any other kind of central processing unit. Typically, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for executing or implementing instructions and one or more storage devices for storing instructions and data. The central processing unit and the memory may be supplemented by a dedicated logic circuit or incorporated in the dedicated logic circuit. Typically, a computer will also include one or more large-capacity storage devices for storing data, such as a disk, a magneto-optical disk, or an optical disk, or may be operationally coupled to receive data from the one or more large-capacity storage devices or to transfer data to the one or more large-capacity storage devices, or both for storing data. However, a computer need not have such a device. In addition, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game controller, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive, etc.
[0154] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including by way of example semiconductor storage devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks.
[0155] To provide interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and pointing device, such as a mouse or trackball, that the user can use to provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, voice, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on the user's device in response to receiving a request from the web browser. In addition, the computer may interact with the user by sending a text message or other form of message to a personal device, such as a smart phone running a messaging application, and then receiving a response message from the user.
[0156] The data processing apparatus for implementing a machine learning model may also include, for example, dedicated hardware accelerator units for processing common and computationally intensive parts of machine learning training or production, i.e., inference, workloads.
[0157] The machine learning model can be implemented and deployed using a machine learning framework, such as the TensorFlow framework, the Microsoft Cognitive Toolkit framework, the Apache Singa framework, or the Apache MXNet framework.
[0158] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes a back-end component, such as a data server; or includes a middleware component, such as an application server; or includes a front-end component, such as a client computer with a graphical user interface, a web browser, or an app that a user can use to interact with an implementation of the subject matter described in this specification; or includes any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0159] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by means of computer programs running on respective computers and having a client-server relationship with each other. In some embodiments, the server transmits data such as HTML pages to the user device, for example, for the purpose of displaying data to a user interacting with the device as a client and receiving user input from the user. Data generated at the user device, for example, the result of the user interaction, may be received from the device at the server.
[0160] Although this specification contains many specific implementation details, these should not be interpreted as limitations on the scope of any invention or that may be claimed, but rather as descriptions of features that may be specific to a particular embodiment of a particular invention. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable subcombination. In addition, although features may be described above as functioning in certain combinations and even initially claimed as such, one or more features from a claimed combination may be removed from the combination in some cases, and a claimed combination may be directed to a subcombination or variations of a subcombination.
[0161] Similarly, although operations are depicted in the drawings and described in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order, or that all illustrated operations be performed to achieve the desired results. In some cases, multitasking and parallel processing can be advantageous. In addition, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0162] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the actions recited in the claims may be performed in a different order and still achieve the desired results. As an example, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A method for decompressing data executed by one or more computers, the method include: obtaining: (i) a compressed representation of the data, wherein the data comprises a sequence of code symbols from a set of possible code symbols, and wherein the data has been compressed by entropy encoding the data, and (ii) a set of latent variables derived from the data; processing an input including the set of latent variables derived from the data using an integer neural network to generate data defining an entropy model for decompressing the data, wherein: The entropy model defines, for each position in the code symbol sequence, a corresponding probability distribution over the set of possible code symbols; The integer neural network has a plurality of integer neural network parameter values, and each of the plurality of integer neural network parameter values is an integer; and The data is decompressed by entropy decoding the compressed representation of the data using the entropy model generated by the integer neural network.
2. The method of claim 1, wherein the set of possible code symbols comprises a plurality of integer values. The method of claim 1 , wherein the data represents an image.
4. The method of claim 1, wherein the data represents a latent representation of the image generated by processing the image using a second integer neural network.
5. The method of claim 1 , further comprising, after decompressing the data: The code symbol sequence is processed using a third integer neural network to generate a reconstructed image.
6. The method of claim 1 , wherein decompressing the data by entropy decoding the compressed representation of the data using the entropy model generated by the integer neural network comprises, for each position in the code symbol sequence: An arithmetic decoding process is used to entropy decode the code symbol at the position using the corresponding probability distribution over the set of possible code symbols.
7. The method of claim 1 , wherein the integer neural network comprises a plurality of integer neural network layers, each integer neural network layer being configured to process a corresponding integer neural network layer input to generate a corresponding integer neural network layer output, and processing the integer neural network layer input to generate the integer neural network layer output include: generating an intermediate result by processing the integer neural network layer input according to a plurality of integer neural network layer parameters using integer-valued operations; as well as The integer neural network layer outputs are generated by applying an integer-valued activation function to the intermediate results.
8. The method of claim 6, wherein intermediate results are generated by processing the integer neural network layer input according to a plurality of integer neural network layer parameters using integer value operations include: A first intermediate result is generated by multiplying the integer neural network layer input by an integer-valued parameter matrix or convolving the integer neural network layer input by an integer-valued convolution filter.
9. The method according to claim 7, further comprising: include: A second intermediate result is generated by adding an integer-valued bias vector to the first intermediate result.
10. The method according to claim 8, further comprising: include: A third intermediate result is generated by dividing each component of the second intermediate result by a rescaling vector of integer values, wherein the division is performed using a rounded division operation.
11. The method of claim 6, wherein the integer-valued activation function is defined by a lookup table that defines a mapping from each integer value in the predetermined set of integer values to a corresponding integer output.
12. The method according to claim 1, in: The plurality of integer neural network parameter values of the integer neural network are determined through a training process; and During training of the integer neural network: The integer neural network parameter values are stored as floating point values, and The integer neural network parameter values stored as floating point values are rounded to integer values before being used in calculations.
13. The method of claim 1, wherein the set of latent variables derived from the data is generated by processing the data using a fourth integer neural network.
14. The method of claim 1, wherein for each position in the code symbol sequence, the corresponding probability distribution over the set of possible code symbols is a Gaussian distribution convolved with a uniform distribution.
15. A method according to claim 13, wherein for each position in the code symbol sequence, the data defining the corresponding probability distribution over the set of possible code symbols includes corresponding mean and standard deviation parameters of the Gaussian distribution.
16. The method of claim 1, wherein the set of latent variables derived from the data is obtained include: obtaining a compressed representation of the set of latent variables derived from the data; as well as The set of latent variables derived from the data is decompressed by entropy decoding the compressed representation of the set of latent variables derived from the data using a predetermined entropy model.
17. The method of claim 1, wherein the integer neural network is configured to process a representation of the set of latent variables as a collection of integer values.
18. The method of claim 1, wherein the integer neural network is configured to generate the entropy model over a sequence of processing steps, wherein each processing step corresponds to a respective position in the sequence of code symbols, and wherein at each processing step, the integer neural network performs an operation, wherein the operation include: processing inputs comprising: (i) the set of latent variables derived from the data, and (ii) corresponding integers representing each code symbol at any position in the code symbol sequence preceding the position associated with the processing step, so as to generate the probability distribution over the set of possible code symbols corresponding to the position associated with the processing step.
19. A system comprising one or more computers and one or more storage devices storing instructions, the instructions, when executed by the one or more computers, causing the one or more computers to perform operations for decompressing data, the operations include: obtaining: (i) a compressed representation of the data, wherein the data comprises a sequence of code symbols from a set of possible code symbols, and wherein the data has been compressed by entropy encoding the data, and (ii) a set of latent variables derived from the data; processing an input including the set of latent variables derived from the data using an integer neural network to generate data defining an entropy model for decompressing the data, wherein: The entropy model defines, for each position in the code symbol sequence, a corresponding probability distribution over the set of possible code symbols; The integer neural network has a plurality of integer neural network parameter values, and each of the plurality of integer neural network parameter values is an integer; and The data is decompressed by entropy decoding the compressed representation of the data using the entropy model generated by the integer neural network.
20. One or more non-volatile computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations for decompressing data, the operations include: obtaining: (i) a compressed representation of the data, wherein the data comprises a sequence of code symbols from a set of possible code symbols, and wherein the data has been compressed by entropy encoding the data, and (ii) a set of latent variables derived from the data; processing an input including the set of latent variables derived from the data using an integer neural network to generate data defining an entropy model for decompressing the data, wherein: The entropy model defines, for each position in the code symbol sequence, a corresponding probability distribution over the set of possible code symbols; The integer neural network has a plurality of integer neural network parameter values, and each of the plurality of integer neural network parameter values is an integer; and The data is decompressed by entropy decoding the compressed representation of the data using the entropy model generated by the integer neural network.
21. A method for compressing data executed by one or more computers, the method include: obtaining data to be compressed, wherein the data comprises a sequence of code symbols from a set of possible code symbols; An input comprising integer representations of a set of latent variables derived from the data is processed using an integer neural network to generate data defining an entropy model for compressing the data, wherein: The entropy model defines, for each position in the code symbol sequence, a corresponding probability distribution over the set of possible code symbols; The integer neural network has a plurality of integer neural network parameter values, and each of the plurality of integer neural network parameter values is an integer; as well as The data is compressed by entropy encoding the data using the entropy model generated by the integer neural network.
22. The method of claim 21, wherein the set of possible code symbols comprises a plurality of integer values.
23. The method of claim 21, wherein the data to be compressed represents an image.
24. The method of claim 21, wherein the data to be compressed comprises a latent representation of the image generated by processing the image using a second integer neural network.
25. The method of claim 21 , wherein compressing the data by entropy encoding the data using the entropy model generated by the integer neural network comprises, for each position in the code symbol sequence: An arithmetic coding process is used to entropy encode the code symbol at the position using the corresponding probability distribution over the set of possible code symbols.
26. The method of claim 21, wherein the integer neural network comprises a plurality of integer neural network layers, each integer neural network layer being configured to process a respective integer neural network layer input to generate a respective integer neural network layer output, and processing the integer neural network layer input to generate the integer neural network layer output include: generating an intermediate result by processing the integer neural network layer input according to a plurality of integer neural network layer parameters using integer-valued operations; as well as The integer neural network layer outputs are generated by applying an integer-valued activation function to the intermediate results.
27. The method of claim 26, wherein intermediate results are generated by processing the integer neural network layer input according to a plurality of integer neural network layer parameters using integer valued operations include: A first intermediate result is generated by multiplying the integer neural network layer input by an integer-valued parameter matrix or convolving the integer neural network layer input by an integer-valued convolution filter.
28. The method according to claim 27, further comprising: include: A second intermediate result is generated by adding an integer-valued bias vector to the first intermediate result.
29. The method according to claim 28, further comprising: include: A third intermediate result is generated by dividing each component of the second intermediate result by a rescaling vector of integer values, wherein the division is performed using a rounded division operation.
30. The method of claim 26, wherein the integer-valued activation function is defined by a lookup table that defines a mapping from each integer value in the predetermined set of integer values to a corresponding integer output.
31. The method according to claim 21, in: The plurality of integer neural network parameter values of the integer neural network are determined through a training process; and During training of the integer neural network: The integer neural network parameter values are stored as floating point values, and The integer neural network parameter values stored as floating point values are rounded to integer values before being used in calculations.
32. The method of claim 21, wherein the set of latent variables derived from the data is generated by processing the data using a fourth integer neural network.
33. The method of claim 21, wherein for each position in the code symbol sequence, the corresponding probability distribution over the set of possible code symbols is a Gaussian distribution convolved with a uniform distribution.
34. A method according to claim 33, wherein for each position in the code symbol sequence, the data defining the corresponding probability distribution over the set of possible code symbols includes corresponding mean and standard deviation parameters of the Gaussian distribution.
35. The method of claim 21, wherein the integer neural network is configured to generate the entropy model over a sequence of processing steps, wherein each processing step corresponds to a respective position in the sequence of code symbols, and wherein at each processing step, the integer neural network performs an operation, wherein the operation include: processing inputs comprising: (i) the set of latent variables derived from the data, and (ii) corresponding integers representing each code symbol at any position in the code symbol sequence preceding the position associated with the processing step, so as to generate the probability distribution over the set of possible code symbols corresponding to the position associated with the processing step.
36. A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for compressing data, the operations include: obtaining data to be compressed, wherein the data comprises a sequence of code symbols from a set of possible code symbols; An input comprising integer representations of a set of latent variables derived from the data is processed using an integer neural network to generate data defining an entropy model for compressing the data, wherein: The entropy model defines, for each position in the code symbol sequence, a corresponding probability distribution over the set of possible code symbols; The integer neural network has a plurality of integer neural network parameter values, and each of the plurality of integer neural network parameter values is an integer; as well as The data is compressed by entropy encoding the data using the entropy model generated by the integer neural network.
37. One or more non-volatile computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations for compressing data, the operations include: obtaining data to be compressed, wherein the data comprises a sequence of code symbols from a set of possible code symbols; An input comprising integer representations of a set of latent variables derived from the data is processed using an integer neural network to generate data defining an entropy model for compressing the data, wherein: The entropy model defines, for each position in the code symbol sequence, a corresponding probability distribution over the set of possible code symbols; The integer neural network has a plurality of integer neural network parameter values, and each of the plurality of integer neural network parameter values is an integer; as well as The data is compressed by entropy encoding the data using the entropy model generated by the integer neural network.
38. The non-volatile computer storage medium of claim 37, wherein the set of possible code symbols comprises a plurality of integer values.
39. The non-volatile computer storage medium of claim 37, wherein the data to be compressed represents an image.
40. The non-volatile computer storage medium of claim 37, wherein the data to be compressed comprises a latent representation of the image generated by processing the image using a second integer neural network.