Method and data processing system for encoding, transmission and decoding of lost images or video - Patents.com
Patent Information
- Application Number
- JP2024525029
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-26
- Filing Date
- 2022-10-26
- Publication Date
- 2025-10-31
AI Technical Summary
Existing image and video compression techniques using artificial intelligence (AI) often result in poor compression quality and require significant data transmission, leading to increased energy consumption in communication networks.
A method involving training neural networks for encoding and decoding lossy images or videos, where the difference between input and output images is evaluated to update network parameters, using functions based on square, absolute values, and variance of multiple outputs, and incorporating clustering algorithms to specialize models for specific data clusters.
Improves compression quality and reduces data transmission requirements, resulting in more efficient energy use in communication networks.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method and system for encoding, transmitting and decoding lost images or lost videos, a method, apparatus, computer program and computer-readable storage medium for encoding and transmitting lost images or lost videos, and a method, apparatus, computer program and computer-readable storage medium for receiving and decoding lost images or lost videos.
[0002] Demand for image and video content from users of communication networks is increasing, which is driving demand for higher resolution image and video content from communication networks and computers, which is increasing the load on communication networks and the amount of data being transmitted, which in turn is increasing the energy usage of communication networks.
[0003] To mitigate the effects of these problems, image and video content is compressed before being transmitted over networks. Image and video content can be compressed using either lossless or lossy compression. Lossless compression compresses images and videos in a way that allows all of the original information in the content to be recovered. However, there is a limit to the amount of data reduction that can be achieved when using lossless compression. Lossy compression involves the loss of information from the image or video during the compression process. Known compression techniques attempt to minimize the apparent loss of information by removing information that changes the image or video after decompression in a way that is not particularly noticeable to the human visual system.
[0004] Artificial intelligence (AI)-based compression techniques achieve image and video compression and decompression by using trained neural networks in the compression and decompression process. Typically, during neural network training, differences between the original and decompressed images and videos are analyzed, and the neural network parameters are modified to reduce these differences while minimizing the data required to transmit the content. However, AI-based compression methods can sometimes produce poor compression results in terms of the appearance of the compressed images and videos and the amount of information required for transmission.
[0005] According to the present invention, there is provided a method of training one or more neural networks, the one or more neural networks being for use in encoding, transmitting, and decoding lossy images or lossy videos, the method comprising the steps of: receiving an input image at a first computer system; encoding an input image using a first neural network to generate a latent representation, and transmitting the latent representation to a second computer system; decoding the latent representation using a second neural network to generate an output image; generating an output image, the output image being an approximation of the input image; The first neural network and the second neural network each consist of multiple layers. at least one of the layers of the first or second neural network includes a transformation; The method further comprises the steps of: evaluating the difference between the output image and the input image and evaluating a function based on the output of the transformation; updating parameters of the first neural network and the second neural network based on the estimated difference and the estimated function; and Repeating the steps above using the first set of input images to generate a first trained neural network and a second trained neural network.
[0006] The function may be based on multiple outputs of multiple transformations, each transformation being in a different one of the multiple layers of the first or second neural network.
[0007] The contribution of each of the multiple outputs to the evaluation of the function may be scaled by a predetermined value.
[0008] The output of the transform may consist of multiple values, and the function may be based on the square of each of the multiple values.
[0009] The output of the transform consists of multiple values, and the function may be based on the absolute value of each of the multiple values.
[0010] The function may be based on the square and absolute value of each of the multiple values.
[0011] It may be based on the average value of each of the multiple values.
[0012] The squared and absolute value contributions in the evaluation of the function may be scaled by a predetermined value.
[0013] The function may also be based on the variance of multiple values.
[0014] The contribution of the evaluated function to the update to the parameters of the first neural network and the second neural network may be scaled by a predetermined value.
[0015] The contribution of the evaluation function to updates to the parameters of the first and second neural networks may be scaled by a value, which is incrementally updated in at least one iteration of the method steps.
[0016] At least one of the multiple layers of the first neural network or the second neural network may include an additional transform applied after the transform, and parameters of the linear transform may be additionally updated based on the evaluated difference and the evaluated function, where the function is not based on the output of the additional transform.
[0017] According to the present invention, there is provided a method for encoding, transmitting and decoding a lossy image or video, the method comprising the following steps: receiving an input image at a first computer system; encoding an input image using a first trained neural network to generate a latent representation; generating latent representations; transmitting the latent representation to a second computer system; decoding the latent representation using a second trained neural network to generate an output image, the output image being an approximation of the input image; Here, the first trained neural network and the second trained neural network are trained according to the above method.
[0018] According to the present invention, there is provided a method for lossy image or video encoding and transmission, the method comprising the following steps: receiving an input image at a first computer system; encoding an input image using a first trained neural network to generate a latent representation; transmitting the latent representation; Here, the first trained neural network is trained according to the method described above.
[0019] According to the present invention, there is provided a method for receiving and decoding a lost image or video, the method comprising the following steps: receiving at a second computer system the latent representations transmitted according to the method; decoding the latent representation using a second trained neural network to generate an output image, the output image being an approximation of the input image; Here, a second trained neural network is trained according to the method described above.
[0020] According to the present invention, there is provided a method for encoding, transmitting and decoding lossy images and videos, the method comprising the following steps: receiving an input image at a first computer system; encoding an input image using a first trained neural network to generate latent representations; performing a first operation on the latent representation to obtain a residual latent; and transmitting the residual latent to a second computer system; performing a second operation on the latent residual to obtain a retrieved latent representation (wherein the second operation includes performing an operation on previously obtained pixels of the retrieved latent); and Decoding the obtained latent representation using a second trained neural network to generate an output image, the output image being an approximation of the input image.
[0021] The first operation may be a linear operation.
[0022] The first operation may be a non-linear operation.
[0023] The first operation may be a differentiable operation.
[0024] The first operation may be a non-differentiable operation.
[0025] The operations performed on the retrieved potential previously retrieved pixels may include matrix operations.
[0026] At least one of the operations may include an explicit iterative operation.
[0027] At least one operation may include an implicit iterative operation.
[0028] An implicit iterative operation may involve computing a rounding function, and an approximation may be used to determine the resolution of the rounding function.
[0029] The second operation may be an iterative operation that includes operator decomposition.
[0030] The iterative operation may include defining at least two variables based on the latent representation and the obtained latent, and performing an operation on each of the at least two variables in turn.
[0031] The iterative operation may include defining at least two further variables each corresponding to a range of one of the at least two variables, and further performing the operation on each of the at least two further variables in turn.
[0032] The iterative operation may further include a step in which, during each iteration step, one of the at least two variables is changed based on a difference between the value of the variable and the value of the variable in the previous iteration step.
[0033] If the matrix operations are negative definite, each operation may be a descending step.
[0034] If the matrix operation is singular, the iterative method may include additional terms to obtain a unique solution for the second operation.
[0035] The second operation may consist of performing a conditioning operation on the matrix operation.
[0036] The first operation may include a quantization operation.
[0037] The first operation may be further based on a latent corresponding to the obtained latent expression. The latent expression and the latent expression corresponding to the obtained latent expression may be equal.
[0038] The first operation may be additionally based on a latent representation corresponding to the obtained latent representation, and the latent representation and the latent representation corresponding to the obtained latent representation may be equal. The latent representation and the latent representation corresponding to the obtained latent representation may differ from each other by a predetermined error value.
[0039] The method may further include the following steps: encoding the latent representation using a fourth trained neural network to generate a super-latent representation; transmitting the quantized hyperlatent representation to a second computer system; Decoding the quantized hyper-latencies using a fifth trained neural network, wherein the operations performed on previously obtained pixels of the obtained latents are based on the output of the fifth trained neural network.
[0040] According to the present invention, there is provided a method of training one or more neural networks, the one or more neural networks being for use in encoding, transmitting and decoding lossy images or lossy videos, the method comprising the steps of: receiving an input image at a first computer system; encoding an input image using a first neural network to generate a latent representation; and performing a first operation on the latent representation to obtain a residual latent; performing a second operation on the latent residual to obtain a obtained latent representation; decoding the quantized latents using a second neural network to generate an output image, the output image being an approximation of the input image; evaluating the difference between the output image and the input image; updating parameters of the first neural network and the second neural network based on the estimated difference; Repeating the above steps using the first set of input images to generate a first trained neural network and a second trained neural network.
[0041] The method may further include the steps of: encoding the latent representation using a fourth neural network to generate a hyper-latent representation; A step of performing a quantization process on the hyper-latent representation to generate a quantized hyper-latent representation. generating quantized hyperlatencies; transmitting the quantized hyperlatent representation to a second computer system; decoding the quantized hyper-latencies using a fifth neural network, wherein the operations performed on the previously obtained pixels of the obtained latents are based on the output of the fifth trained neural network; Here, the parameters of the fourth neural network and the fifth neural network are additionally updated based on the evaluated difference to generate a fourth trained neural network and a fifth trained neural network.
[0042] According to the present invention, there is provided a method for encoding and transmitting a lossy image or video, the method comprising the following steps: receiving an input image at a first computer system; encoding an input image using a first trained neural network to generate a latent representation; performing a first operation on the latent representation to obtain a residual latent; Sending the residual latent.
[0043] According to the present invention, there is provided a method for receiving and decoding a lost image or video, the method comprising the steps of: receiving, at a second computer system, the residual latencies transmitted according to the method; performing a second operation on the latent residual to obtain a retrieved latent representation, wherein the second operation consists of performing an operation on previously obtained pixels of the retrieved latent; Decoding the obtained latent representation using a second trained neural network to generate an output image, where the output image is an approximation of the input image.
[0044] According to the present invention, there is provided a method for encoding, transmitting and decoding lossy images and lossy videos, the method comprising the steps of: receiving an input image at a first computer system; encoding an input image using the first trained neural network to generate a latent representation; performing a first operation on the latent representation to obtain a residual latent, where the first operation includes performing a quantization operation, and the size of the bins used in the quantization operation is based on the input image; transmitting the residual latent to a second computer system; performing a second operation on the latent residual to obtain a retrieved latent representation, where the second operation includes performing an operation on previously obtained pixels of the retrieved latent; Decoding the obtained latent representation using a second trained neural network, where the output image is an approximation of the input image.
[0045] At least one of the first operation and the second operation may include solving an implicit system of equations.
[0046] The bin size may differ between at least two of the pixels or channels of the latent representation.
[0047] The quantization process may involve performing an operation on the value of each pixel of the latent representation that corresponds to the size of the bin assigned to that pixel.
[0048] The quantization process may involve subtracting the mean value of the latent representation from each pixel of the latent representation.
[0049] The quantization process may include a rounding function.
[0050] The operations performed on the retrieved potential previously obtained pixels may consist of matrix operations.
[0051] The matrices that define the matrix operations may have at least one of the following properties: upper triangular, lower triangular, banded, monotonic, symmetric, and bounded by eigenvalues.
[0052] The quantization process may include a third trained neural network.
[0053] The method may further include the following steps: encoding the latent representation using a fourth trained neural network to generate a hyper-latent representation; performing a quantization process on the hyper-latent representation to generate a quantized hyper-latent representation; transmitting the quantized hyperlatent representation to a second computer system; decoding the quantized hyperlatency using a fifth trained neural network to obtain bin sizes; Here, the decoding of the quantized latent uses the obtained bin size; The operations performed on the previously obtained pixels of the obtained potential are based on the output of the fifth trained neural network.
[0054] The output of the fifth trained neural network may be processed by a further function to obtain the bin size.
[0055] A further function may be a sixth trained neural network.
[0056] The size of the bins used in the quantization process of the hyper-latent representation may be based on the input image.
[0057] According to the present invention, there is provided a method of training one or more neural networks, the one or more neural networks being for use in encoding, transmitting, and decoding lossy images or videos, the method comprising the steps of: receiving an input image at a first computer system; encoding an input image using a first neural network to generate a latent representation; performing a first operation on the latent representation to obtain a residual latent representation, where the first operation includes performing a quantization operation, and the size of the bins used in the quantization operation is based on the input image; transmitting the residual latent to a second computer system; performing a second operation on the residual latent to obtain a obtained latent, wherein the second operation performing operations on previously obtained pixels of the obtained latent; decoding the obtained latent representation using a second neural network to generate an output image; evaluating the difference between the output image and the input image; updating parameters of the first neural network and the second neural network based on the estimated difference; updating parameters of the first neural network and the second neural network based on the estimated difference; Repeating the above steps using the first set of input images to generate a first trained neural network and a second trained neural network.
[0058] The method may further include the steps of: encoding the latent representation using a third neural network to generate a hyper-latent representation; performing a quantization process on the hyper-latent representation to generate a quantized hyper-latent representation; transmitting the quantized hyperlatent representation to a second computer system; decoding the quantized hyperlatency using a fourth neural network to obtain bin sizes; Here, the decoding of the quantized latent uses the obtained bin size; The operations performed on the previously obtained pixels of the obtained potential are based on the output of the fourth neural network; The parameters of the third neural network and the fourth neural network are additionally updated based on the estimated difference to obtain a third trained neural network and a fourth trained neural network.
[0059] The quantization process may include a first quantization approximation.
[0060] The method may further comprise the step of determining a rate for the encoding, transmission and decoding process of the lossy image or video; The parameters of the neural network are incrementally updated based on the determined rate; Based on the rate, the neural network parameters are incrementally updated; A second quantization approximation is used to determine the rate of the compression process; The second quantized approximation is different from the first quantized approximation.
[0061] Updating the parameters of the neural network may include the following steps: Evaluating a loss function based on the difference between the output image and the input image; Evaluating the gradient of the loss function; Backpropagating the gradient of the loss function through the neural network; Here, a third quantization approximation is used during backpropagation of the gradient of the loss function; The third quantification approximation is the same approximation as the first quantification approximation.
[0062] According to the present invention, there is provided a method for lossy image or lossy video encoding and transmission, the method comprising the following steps: receiving an input image at a first computer system; encoding an input image using the first trained neural network to generate a latent representation; performing a first operation on the latent representation to obtain a residual latent, wherein the first operation includes performing a quantization process, and the size of the bins used in the quantization process is based on the input image; Sending the residual latent.
[0063] According to the present invention, there is provided a method for receiving and decoding a lost image or video, the method comprising the following steps: receiving at a second computer system the residual latencies transmitted according to the method; performing a second operation on the residual latent to obtain a retrieved latent representation, where the second operation includes performing an operation on previously obtained pixels of the retrieved latent; Decoding the obtained latent representation using a second trained neural network, where the output image is an approximation of the input image.
[0064] According to the present invention, there is provided a method for encoding, transmitting and decoding lossy images and lossy videos, the method comprising the steps of: receiving an input image at a first computer system; assigning input images to clusters using a clustering algorithm; encoding an input image using the first trained neural network to generate a latent representation; transmitting the latent representation to a second computer system; decoding the latent representation using a second trained neural network to generate an output image, where the output image is an approximation of the input image; wherein at least one of the first trained neural network and the second trained neural network is selected from the first plurality of trained neural networks and the second plurality of trained neural networks, respectively, based on the assigned cluster.
[0065] The first network and the second network may each be selected based on the assigned cluster.
[0066] The input image may be a subimage obtained from the initial image.
[0067] The initial image may be divided into multiple input images of the same size; Each of the multiple input images may be encoded, transmitted, and decoded according to the method for lossy image and video encoding, transmission, and decoding.
[0068] The multiple latent representations may be combined to form a combined latent representation before being transmitted to the second computer system, may be split to obtain the multiple received latent representations, and may be split to obtain the multiple received latent representations before decoding with the second trained neural network.
[0069] One of the plurality of first neural networks and one of the plurality of second neural networks may not be associated with any cluster of the clustering algorithm.
[0070] The clustering algorithm may include at least one of the following algorithms: averaging, mini-batch averaging, similarity propagation, mean shift, spectral clustering, Ward hierarchical clustering, agglomerative clustering, DBSCAN, OPTICS, Gaussian mixtures, BIRCH, deep clustering, and trained neural networks.
[0071] The clustering performed by the clustering algorithm may be based on the following properties of at least one channel of the input image: mean, variance, histogram probability and power spectral density.
[0072] The clustering performed by the clustering algorithm may be based on the output of a third trained neural network that receives the input image as input.
[0073] Each of the first plurality of trained neural networks may have the same architecture. Each of the second plurality of trained neural networks may have the same architecture.
[0074] The clusters of the clustering algorithm may be based on the media categories of the input images.
[0075] The clusters of the T-clustering algorithm may be based on subcategories of the media category of the input images.
[0076] In accordance with the present invention, there is provided a method for training one or more algorithms, The one or more algorithms are for use in encoding, transmitting, and decoding lossy images or lossy videos, and the method includes the steps of: receiving an input image at a first computer system; assigning input images to clusters using a clustering algorithm; encoding an input image using a first neural network to generate a latent representation; decoding the latent representation using a second neural network to generate an output image, where the output image is an approximation of the input image; at least one of a first trained neural network and a second trained neural network is selected from a first plurality of neural networks and a second plurality of neural networks, respectively, based on the assigned cluster; The method further comprises the steps of: evaluating the difference between the output image and the input image; updating parameters of the first neural network and the second neural network based on the evaluated difference; Repeating the above steps with the first set of input images to generate a first plurality of trained neural networks and a second plurality of trained neural networks.
[0077] The clustering performed by the clustering algorithm may be based on the output of a third neural network that receives the input image as input; The parameters of the third neural network are additionally updated based on the estimated difference.
[0078] According to the present invention, there is provided a method for encoding and transmitting a lossy image or video, the method comprising the following steps: receiving an input image at a first computer system; assigning input images to clusters using a clustering algorithm; encoding an input image using the first trained neural network to generate a latent representation; generating latent representations; transmitting the latent representation; Here, a first trained neural network is selected from the first plurality of trained neural networks based on the assigned cluster.
[0079] According to the present invention, there is provided a method for receiving and decoding a lost image or video, the method comprising the following steps: receiving at a second computer system the latency transmitted according to the method; decoding the latent representation using a second trained neural network to generate an output image; Here, a second trained neural network is selected from the second plurality of trained neural networks based on the assigned cluster.
[0080] According to the present invention there is provided a data processing system configured to perform any of the above methods.
[0081] According to the present invention, there is provided a data processing apparatus arranged to carry out any of the above methods.
[0082] According to the present invention there is provided a computer program comprising instructions which, when executed by a computer, cause the computer to carry out any of the above methods.
[0083] According to the present invention, there is provided a computer-readable storage medium containing instructions which, when executed by a computer, cause the computer to carry out any of the methods described above. [Brief explanation of the drawings]
[0084] Aspects of the present invention will now be described by way of example with reference to the following figures: [Figure 1] 1 illustrates an example of an image or video compression, transmission and decoding pipeline. [Figure 2]A further example of an image or video compression, transmission and decompression pipeline including a hypernetwork is shown. [Figure 3] 1 shows the loss curves for the AI compression system with and without second moment penalty. [Figure 4] We show the generalized and specialized models trained on dataset D. [Figure 5] An example of the training process for a specialized AI compression process is shown below. [Figure 6] An example of the encoding process of the specialized AI compression process is shown below. [Figure 7] An example of a decoding process for specialized AI compression processing is shown below. DETAILED DESCRIPTION OF THE INVENTION
[0085] Compression can be applied to any form of information to reduce the amount of data required to store it, i.e., file size. Images and video are examples of information that can be compressed. The file size required to store information is sometimes referred to as the rate during the compression process, particularly when referring to compressed files. Generally, compression can be classified as lossless or lossy. Both types of compression result in smaller file sizes. However, with lossless compression, no information is lost when information is compressed and subsequently decompressed. That is, the original file storing the information is completely reconstructed during the decompression process. In contrast, with lossy compression, information is lost during the compression and decompression process, and the reconstructed file may differ from the original file. Image and video files containing image and video data are common targets for compression. JPEG, JPEG2000, AVC, HEVC, and AVI are examples of compression processes for image and / or video files.
[0086] In compression processes involving images, the input image can be represented as x. Data representing the image can be stored in a tensor of dimensions H×W×C, where H is the image height, W is the image width, and C is the number of channels in the image. Each H×W data point in the image represents the pixel value of the image at the corresponding location. Each channel C of the image represents a different component of the image at each pixel, which are combined when the image file is displayed by a device. For example, an image file may have three channels, each representing the red, green, and blue components of the image. In this case, the image information is stored in an RGB color space, sometimes referred to as a model or format. Other examples of color spaces or formats include the CMK and YCbCr color models. However, the channels in an image file are not limited to storing color information; other information may also be represented in the channels. Because video is considered a continuous series of images, compression processes applied to images may also be applied to video. Each image that makes up a video is sometimes referred to as a frame of the video.
[0087] The output image may differ from the input image. JPEG2024540005000002.jpg119 (an x with a ^ symbol above it will hereinafter be referred to as "x (hat)"; similarly, a letter with a ^ symbol above it will hereinafter be referred to as the letter (hat)). The difference between an input image and an output image is sometimes referred to as distortion or image quality difference. Distortion can be measured using any distortion function that accepts an input image and an output image and provides an output that numerically represents the difference between the input and output images. An example of such a method is using the mean squared error (MSE) between the pixels of the input and output images, although many other methods for measuring distortion are known to those skilled in the art. The distortion function may consist of a trained neural network.
[0088] In general, the rate and distortion of a lossy compression process are related: an increase in rate results in a decrease in distortion, and a decrease in rate results in an increase in distortion. Changes in distortion can correspondingly affect the rate. The relationship between these quantities for a given compression technique can be defined by the rate-distortion formula:
[0089] AI-based compression processes often use neural networks. A neural network is an operation that can be performed on an input to produce an output. A neural network may consist of multiple layers. The first layer of the network receives the input. The layer performs one or more operations on the input to produce the output of the first layer. The output of the first layer is then passed to the next layer of the network, which performs one or more operations in a similar manner. The output of the final layer is the output of the neural network.
[0090] Each layer of a neural network may be divided into nodes. Each node may receive at least a portion of the inputs from the previous layer and provide outputs to one or more nodes in the subsequent layer. Each node of a layer may perform one or more operations of the layer on at least a portion of the inputs to the layer. For example, a node may receive inputs from one or more nodes in a previous layer. The one or more operations include convolutions, weights, biases, and activation functions. Convolution operations are used in convolutional neural networks. If a convolution operation is present, the convolution may be performed across the inputs to the layer. Alternatively, the convolution may be performed across at least a portion of the inputs to the layer.
[0091] Each of the one or more operations may be defined by one or more parameters associated with the operation. For example, a weight operation may be defined by a weight matrix that defines the weights applied to each input from each node in the previous layer to each node in the current layer. In this example, each value in the weight matrix is a parameter of the neural network. A convolution may be defined by a convolution matrix, also known as a kernel. In this example, one or more of the values in the convolution matrix may be a parameter of the neural network. An activation function may also be defined by values that may be parameters of the neural network. The parameters of the network may be varied during training of the network.
[0092] Other characteristics of a neural network may be predetermined and therefore do not change during training of the network. For example, the number of layers of the network, the number of nodes of the network, one or more operations performed in each layer, and the connections between layers may be predetermined and therefore fixed before the learning process occurs. These predetermined characteristics may be referred to as hyperparameters of the network. These characteristics may also be referred to as the network's architecture.
[0093] To train a neural network, a training set of inputs with known expected outputs, sometimes called ground truth, can be used. The initial parameters of the neural network are randomized, and the first training inputs are provided to the network. The network's output is compared to the expected output, and based on the difference between the output and the expected output, the network's parameters are changed to reduce the difference between the network's output and the expected output. This process is repeated for multiple training inputs to train the network. The difference between the network's output and the expected output can be defined by a loss function. The result of the loss function can be calculated using the difference between the network's output and the expected output to determine the gradient of the loss function. Gradient descent backpropagation of the loss function can be used to update the neural network's parameters using the gradient of the loss function, dL / dy. Multiple neural networks in a system can be trained simultaneously through backpropagation of the gradient of the loss function to each network.
[0094] For AI-based image or video compression, the loss function may be expressed as a rate-distortion equation: Loss = D + λ * R, where D is the distortion function, λ is a weighting factor, and R is the rate loss. λ is sometimes called the Lagrangian multiplier. Lagrangian multipliers serve as weights for specific terms in the loss function relative to each other term, and can be used to control which terms in the loss function are prioritized when training the network.
[0095] For AI-based image or video compression, a training set of input images can be used. An example of a training set of input images is the Kodak image set (e.g., www.cs.albany.edu / xypan / research / snr / Kodak.html). An example of a training set of input images is the IMAX image set. An example of a training set of input images is the Imagenet dataset (e.g., www.image-net.org / download). An example of a training set of input images is the CLIC Training Dataset P ("professional") and M ("mobile") (e.g., http: / / challenge.compression.cc / tasks / ).
[0096] An example of an AI-based compression process 100 is shown in Figure 1. As a first step in the AI-based compression process, an input image 5 is provided. The input image 5 is then subjected to a function f θ The input image is provided to a trained neural network 110 characterized by the following: The encoder neural network 110 generates an output based on the input image. This output is called a latent representation of the input image 5. In a second step, the latent representation is quantized in a quantization process 140 characterized by an operation Q, resulting in a quantized latent representation. The quantization process converts the continuous latent representation into a discrete quantized latent representation. An example of a quantization process is a rounding function.
[0097] In a third step, the quantized latent quantities are entropy coded in an entropy coding process 150 to generate a bitstream 130. The entropy coding process may be, for example, range coding or arithmetic coding. In a fourth step, the bitstream 130 may be transmitted over a communication network.
[0098] In the fifth step, the bitstream is entropy decoded in an entropy decoding process 160. The quantized latent quantities are then decoded using the function g θThe neural network 120 acts as a decoder and decodes the quantized latents. The trained neural network 120 generates an output based on the quantized latents. The output may be an output image of the AI-based compression process 100. The encoder-decoder system may be called an autoencoder.
[0099] The above-described system may be distributed across multiple locations and / or devices. For example, the encoder 110 may be located on a device such as a laptop computer, a desktop computer, a smartphone, or a server. The decoder 120 may be located on another device, called a recipient device. The system used to encode, transmit, and decode the input image 5 to obtain the output image 6 is sometimes called a compression pipeline.
[0100] The AI-based compression process may further include a hypernetwork 105 for the transmission of meta-information that improves the compression process. h θ and a trained neural network 115 that acts as a hyperdecoder g h θ An example of such a system is shown in Figure 2. Components of the system not further described may be assumed to be the same as those described above. The neural network 115, acting as a hyper-decoder, receives the latents that are the output of the encoder 110. The hyper-encoder 115 generates an output based on the latent representation, sometimes called the hyper-latent representation. The hyper-latent is then computed as Q h The hyperlatencies are quantized in a quantization process 145 characterized by Q h The quantization process 145 characterized by Q may be the same as the quantization process 140 characterized by Q described above.
[0101] In a manner similar to that described above for the quantized latent, the quantized hyper-latent is then entropy coded in an entropy coding process 155 to produce a bitstream 135. The bitstream 135 may be entropy decoded in an entropy decoding process 165 to retrieve the quantized hyper-latent. The quantized hyper-latent is then used as input to a trained neural network 125, which functions as a hyper-decoder. However, in contrast to the compression pipeline 100, the output of the hyper-decoder may not be an approximation of the input to the hyper-decoder 115. Instead, the output of the hyper-decoder is used to provide parameters for use in the entropy coding process 150 and the entropy decoding process 160 of the main compression process 100. For example, the output of the hyper-decoder 125 may include one or more of the mean, standard deviation, variance, or other parameters used to describe the probability model of the entropy coding process 150 and the entropy decoding process 160 of the latent representation. In the example shown in FIG. 2, only a single entropy decoding process 165 and hyper-decoder 125 are shown for simplicity. However, in practice, decompression processes are typically performed on a separate device, so there will be a duplication of these processes on the device used for encoding to provide the parameters used by the entropy encoding process 150.
[0102] Further transformations can be applied to the latents and / or hyperlatents at any stage of the AI-based compression process 100. For example, the latents and / or hyperlatents can be converted to residual values before the entropy encoding processes 150, 155 are performed. The residual values can be determined by subtracting the mean of the distribution of the latents or hyperlatents from each latent or hyperlatent. The residual values can also be normalized.
[0103] To perform training of the AI-based compression process described above, a training set of input images may be used, as described above. During the training process, the parameters of both the encoder 110 and the decoder 120 may be updated simultaneously at each training step. If a hypernetwork 105 is also present, the parameters of both the hyperencoder 115 and the hyperdecoder 125 may also be updated simultaneously at each training step.
[0104] The training process may further include a generative adversarial network (GAN). When applied to an AI-based compression process, in addition to the compression pipeline described above, the system includes an additional neural network that functions as a classifier. The classifier receives input and outputs a score based on the input to provide an indication of whether the classifier considers the input to be ground truth or fake. For example, this indication is a score, where a high score is associated with a true input and a low score is associated with a fake input. In such a case, training to distinguish between true and false inputs uses a loss function that maximizes the difference between output instructions that distinguish between true and false inputs.
[0105] When a GAN is incorporated into the training of the compression process, the output image 6 may be provided to a classifier. The output of the classifier is used in the loss function of the compression process as a measure of distortion of the compression process. Alternatively, the classifier may receive both the input image 5 and the output image 6, and the difference between the output representations may be used in the loss function of the compression process as a measure of distortion of the compression process. The training of the neural network acting as the classifier and other neural networks in the compression process may be performed simultaneously. When the trained compression pipeline is used for image or video compression and transmission, the classifier neural network is removed from the system, and the output of the compression pipeline becomes the output image 6.
[0106] Incorporating a GAN into the training process may cause the decoder 120 to perform hallucinations. Hallucination is a process that adds information to the output image 6 that was not present in the input image 5. In one example, the hallucination process may add fine details to the output image 6 that were not present in the input image 5 or that were not received by the decoder 120. The hallucinations that are performed may be based on quantized latent information received by the decoder 120.
[0107] As described above, a video is composed of a series of images arranged sequentially. The AI-based compression process 100 described above may be applied multiple times to compress, transmit, and decompress the video. For example, each frame of the video may be compressed, transmitted, and decompressed individually. The received frames may then be grouped together to obtain the original video.
[0108] A number of concepts related to the AI compression process described above will now be described. Although each concept will be described separately, one or more of the concepts described below may be applied in the AI-based compression process described above.
[0109] Second Moment Penalty In neural (AI-based) compression, a mathematical model ("network") is generally composed of a collection of tensors ("neurons") that apply mathematical transformations to input data (e.g., images / videos) and return an output, as described above. Groups of neurons in a network can be organized into successive "layers," with the property that the output tensor of a particular layer is passed as input to the next layer. The number of layers in a network is called the "depth" of the network. Given input data x, the transformation of a layer (the sum of the transformations of all neurons) is most commonly an affine linear transformation plus a nonlinear function er ("activation function"):
number
[0110] The transformed tensors y and z are sometimes called the "pre-activation feature map" and "feature map", respectively. The tensors A and b are examples of the "parameters" associated with a layer (there can be additional parameters). The transformed tensors y and z are therefore the output of each transformation.
[0111] As mentioned above, a network "learns" by sequentially propagating data (e.g., images / videos) through the network while incrementally adjusting parameters according to a differentiable loss measure called L (the "loss function"). The process of data propagation to adjust parameters is commonly called "training." This contrasts with "inference," which simply refers to propagating data (typically not used in the training process) through the network without adjusting parameters. The loss associated with the inference path (the "validation loss") is a direct measure of the network's general performance level for the task it is assigned to adapt to.
[0112] AI image and video compression commonly employs networks with many layers ("deep networks"). To speed up training, find local minima in the training loss, and ensure good gradient propagation, deep networks may need to adjust the inputs of their layers as the input distribution changes during training. This is particularly relevant when the loss function requires training two networks simultaneously ("adversarial training"). The necessary adjustments are often achieved using so-called "normalization layers," which explicitly adjust the inputs (most commonly the mean and variance) to pre-specified target values. For example, in a batch normalization scheme, the input for the training step consists of a collection of training samples (a "(mini)batch"). The elements of each batch are transformed so that the mean of the batch equals zero and the mean and standard deviation of the batch equals one.
number
[0113] We now describe an example of an alternative approach called second moment penalty (SMP). In contrast to the regularization approach described above, the second moment penalty can vary the loss function based on at least one of the pre-activation feature maps in the network. In one example, the additional term may be based on the element-wise mean square of at least one of the pre-activation feature maps in the network. The additional term may also be based on all subsequent pre-activation feature maps in the network. The size of the additional term is a positive real number λ. SMP can be controlled by
[0114] In an example with all these characteristics, the SMP loss function is defined as follows:
number
[0115] Another advantage of SMP is its ability to control the variance of feature maps throughout training. Deep neural network training frequently encounters problems where the variance of feature maps or parameter gradients explodes or vanishes. A common remediation strategy in the deep learning community is to select weight (parameter) initialization heuristics with favorable properties. However, because parameters are initialized before training, the impact of the initialization choice diminishes over time. Second-moment penalty can control the variance of feature weights throughout training, providing a more robust solution to problems where variance explodes or vanishes.
[0116] Other adjustments to losses include: Average of absolute values JPEG2024540005000006.jpg1129This is equivalent to imposing a zero-mean Laplacian prior, which promotes sparsity in the feature maps. Elastic net penalty JPEG2024540005000007.jpg1246Here α i and β i is a predetermined positive real number, which corresponds to a hierarchical elastic net prior. · It has properties similar to other positive functions of the elements of the feature map with properties similar to those above.
[0117] In addition to the additional term in the loss function, an affine linear transformation may be added to each layer of the network, allowing the network to adapt to the penalty or, optimally, avoid it altogether. A single layer transformation looks like this (see (1)):
number
[0118] An additional transformation may be performed after the pre-activation feature map y is taken, yielding two tensors, A SMP and b SMPmay be an additional adjustable parameter in a given layer. Two successive linear maps can be fused in a single nonlinear operation, thus leaving the execution time unaffected. Other possible transformations of pre-activation feature maps include (but are not limited to): A single convolutional layer, Dense layer, -Individual deep neural networks for each layer Depthwise convolution.
[0119] The distribution targeting property of SMP and its variants is that the additional term in the SMP loss function (3) is the negative log-likelihood isotropic Gaussian distribution of a zero-mean Gaussian. To explain this in more detail, consider one element of the preactivation feature map of layer i. An element is specified by four coordinates: batch, channel, height, and width. Therefore, we define this element as y i (b,c,h,w). The contribution of this element to the overall loss function is
number
[0120] The positive integer N is the result of averaging all pre-activation feature maps and is equal to the number of elements in the i-th feature map. i (b,c,h,w) 2 can be considered as the negative log-likelihood of a standard normal random variable with mean 0 and standard deviation 1. SMP λ / N is the standard deviation SMP Rescale to / N. N and λ SMP Since is a constant, this distribution is y i The same is true for any element of y. i Since the samples (b,c,h,w) are independent and identically distributed, the network i Train a normalizing transformation that pushes the distribution on the training dataset of (b,c,h,w) to a normal distribution with the aforementioned parameters.
[0121] Extending the second moment penalty to change the target distribution can be equivalent to replacing the loss contribution with the negative log-likelihood of the desired target distribution. For example, at location 0 and scale λ, SMP The loss contribution corresponding to the Laplace distribution of / N is
number
[0122] The scope of the approach is not limited to probability distributions with analytically tractable likelihoods: since it only requires sampling from the distribution, more complex distributions can be used by leveraging sampling algorithms such as Markov Chain Monte Carlo.
[0123] In addition to the above arrangements, the second moment penalty can be extended in the following ways, but not limited to: In addition to targeting a mean of 0, you can target a variance equal to 1. The penalty terms described above target the first and second moments of the probability distribution of the pre-actuation features. Different penalty terms targeting other moment combinations can also be implemented. Each layer of the network has its own hyperparameter λ SMP , which allows fine-grained control of the normalized intensity for each layer. ·Hyperparameter λ SMP can be varied during training and follows an annealing schedule that controls its size at different stages. Any combination of the above.
[0124] Figure 3 shows the loss curves for an AI compression system trained with and without second-order moment penalty as described above. The x-axis shows the number of training iterations, and the y-axis shows the system loss. The AI compression system consists of a 7-layer encoder and a 6-layer decoder. The top line in the figure corresponds to the system trained without SMP, and the bottom line corresponds to the same system trained with SMP. We can see that the system loss is improved by 5%. With SMP, the bits per pixel improved by 6% and the MSE by 3%.
[0125] The second-order moment penalty and transformation or augmentation can be used with any neural network layer architecture, including but not limited to dense layers, residual blocks, attention layers, and long short-term memory layers, and therefore can be used with a wide variety of neural network models, including but not limited to densely connected neural networks, convolutional neural networks, generative adversarial networks, transformers, and residual networks.
[0126] context In AI-based compression, the latent variable y is usually quantized to the variable y(hat) by the following formula:
number
[0127] An extension of this approach is quantization by nonlinear implicit equations.
number
[0128] If the latent is in d-dimensional space, where A is a dxd matrix, then the quantized latent variables appear on both sides of the equation here, which must be solved. The idea of this quantization scheme is that the quantized neighboring pixels may contain relevant information about the given pixel and can therefore be employed to make a better position prediction. In other words, JPEG2024540005000013.jpg717 can be considered as a "boosted" positional prediction, which allows for even lower compression ratios of the quantized latent.
[0129] The matrix A can be fixed (perhaps pre-trained), or predicted in some way, such as by a hypernetwork, or derived in some way for each image.
[0130] Unfortunately, solving (8) is generally very difficult. In the very special case where the A matrix is strictly lower (or upper) triangular with respect to the underlying permutation, then (8) can be easily solved by nonlinear forward (or inverse) substitution. Otherwise, it is not at all obvious how (8) can be solved.
[0131] We now describe some methods for solving the coding problem arising from (8), including the inverse problem of decoding (the problem of solving a system of linear equations).
[0132] One way to solve the coding problem (8) is to use fixed-point iteration. Given JPEG2024540005000014.jpg810,
number
[0133] The method converges when A is symmetric positive definite and the spectral norm is less than 1. This is undesirable because it over-constrains the structure of A.
[0134] Alternatively, the auxiliary variables Introducing JPEG2024540005000016.jpg835 leads to many algorithms for solving (8). The variable ξ is sometimes called the residual. We can rewrite this system slightly so that the unknown variable is on the left-hand side:
number
[0135] This is a nonlinear saddle point problem. An example of a linear saddle point problem and its algorithm is described in Michele Benzi, Gene H. Golub, and Jorg Liesen. Numerical solution of saddle point problems. Acta Numerica, 14:1-137, May 2005, which is incorporated herein by reference.
[0136] This problem belongs to the general class of nonlinear saddle point problems of the form
number
[0137] If F is maximally monotonic (as explained in the next paragraph) and C is symmetric and positive semidefinite, then the existence of a solution is guaranteed (this means that C+C T (This can also be relaxed to require positive semidefinite values of .) A saddle point problem written in this way (using these properties) is called a saddle point problem in standard form.
[0138] A set-value function T(·) is monotonic if 〈uv,xy〉 ≥ 0 for all u∈T(x) and v∈T(y). A set-value function T(x) is maximally monotonic if there is no other monotone operator B such that the graph of T is properly contained in B.
[0139] Note that round(ξ) is monotonic (but not maximally monotonic). This is problematic, and the existence and uniqueness results for nonlinear saddle point problems are only guaranteed if the operator is maximally monotonic.
[0140] Therefore, the rounding function can be extended by defining a maximal monotonic extension round(·) via a set-valued function.
number
[0141] where JPEG2024540005000020.jpg817 is the set of integers plus half, and floor(x) and ceil(x) are functions that return the largest integer less than or equal to x and the smallest integer greater than or equal to x, respectively. In what follows, unless otherwise noted, we use the maximal monotonic function Confusing JPEG2024540005000021.jpg911 with round(·).
[0142] For general real-valued A, (10) is not a standard form because, in general, A may have both positive and negative eigenvalues with real parts, in which case A is not positive definite. If A is positive definite, then (11) identifies C=A, and Set JPEG2024540005000022.jpg1032, B2=I, F=round(·), and specify x=ξ, z=y(hat). Obviously, a=-μ and b=y-μ.
[0143] (10) can be rewritten in standard form as follows: If A is an eigendecomposition, A=QAQ -1 (Other decompositions such as Schur decomposition may also be used.) Let Q be the linear equation. -1 Multiply by Let JPEG2024540005000023.jpg925. Equation (10) becomes as follows.
number
[0144] Spectral decomposition assumes that the eigenvalues are ordered in ascending order from smallest real value to largest real value.
number
[0145] The diagonal Λ contains eigenvalues with both positive and negative real parts. ·C is a submatrix Λ containing eigenvalues with positive real parts ≧0 Identify with. Similarly, let Z be the eigenvalues of A that have positive real parts. It is defined as corresponding to the partial vector of JPEG2024540005000026.jpg127 (the symbol ~ above z will be referred to as "z (tilde)" below). Similarly, B1 and B2 are the eigenvalues in A corresponding to the partial vectors of z (tilde) with positive real parts, Q(I+Λ) and Q -1 is obtained by taking a submatrix of ·Define x as the concatenation of g and the remaining elements of z (elements whose eigenvalues in A have negative real parts). ·F is defined as the rounded function of ξ plus all other linear terms. Set a and b to the appropriate components of the right-hand side.
[0146] Once the system has been converted to standard form, several algorithms are available for solving it. One example is an algorithm that relies on the heuristic of alternating updates of x and z. This is an example of an iterative method that uses operator decomposition. For example, the problem can be divided into two or more sets of variables and updates can be performed iteratively for each subset of variables.
[0147] A saddle point problem can be loosely thought of as arising from a minimax optimization problem in which some scalar function is minimized over x and maximized over z. A saddle point algorithm can therefore be thought of as alternating "descent steps in x" and "upward steps in z", which respectively decrease and increase the value of the scalar function.
[0148] The iterative update takes the form: initial iteration value x 0 and z 0 is set, and then the following updates are performed until the convergence criterion is met: JPEG2024540005000027.jpg1172 solution x k+l Let's say. JPEG2024540005000028.jpg1078 solution to z k+l Let's say.
[0149] where τ and σ are step size parameters. The magnitudes of these step size parameters can be appropriately chosen to ensure convergence and depend on the cocoercivity (Lipschitz constant) of F(·) and C, respectively. For example, the step size can be chosen to be less than or equal to half the reciprocal of the Lipschitz constant (cocoercivity modulus) of the operator.
[0150] where JPEG2024540005000029.jpg1312 is x k or x k+l and similarly JPEG2024540005000030.jpg1110 is z k or z k+lThe choice of using the current iterator (superscript k) or the updated iterator (superscript k+1) determines whether each of these update rules is explicit or implicit. If the current iterator is used, the update is explicit. If the updated iterator is used, the system is implicit. Note that it is possible to perform both implicit and explicit updates. For example, it is possible to perform implicit updates, which are simple and easy to implement. These can be thought of as the correspondence between updating x and explicit updates of z.
[0151] An explicit update is computationally an explicit Euler step (if we think of the system as discretizing ordinary differential equations). Another name for an explicit update step is a forward step or forward update rule.
[0152] However, despite their simplicity, explicit update rules may not lead to convergence, especially when F has a large (or infinite) Lipschitz constant, as is the case for the rounding function F. Therefore, implicit update rules can be used instead.
[0153] The implicit update involves the following iterations: JPEG2024540005000031.jpg1073 solution x k+l Let's say. The solution of JPEG2024540005000032.jpg1175 is z k+l Let's say.
[0154] The first (implicit) update rule can be rearranged
number
[0155] The second update rule can be rearranged for the implicit case as follows:
number
[0156] That is, in resolvent notation, JPEG2024540005000037.jpg865. Since C is linear, this amounts to inverting the matrix I + σC.
[0157] The operator F is the maximal monotonic rounding function JPEG2024540005000038.jpg1014 can be constructed. JPEG2024540005000039.jpg1120 can be determined using the following algorithm. Set it to JPEG2024540005000040.jpg1120 and check the following: 1. Let's set the file name to JPEG2024540005000041.jpg1027. If JPEG2024540005000042.jpg965 is executable in r, Returns JPEG2024540005000043.jpg729. 2. Otherwise, Let's set it to JPEG2024540005000044.jpg1129. If JPEG2024540005000045.jpg1071 is executable in r, then Returns JPEG2024540005000046.jpg729. 3. Otherwise, Returns JPEG2024540005000047.jpg1127.
[0158] If x is a vector, the algorithm is applied pointwise to the elements of x.
[0159] JPEG2024540005000048.jpg911 is a function JPEG2024540005000049.jpg1055 is a partial derivative of the quadratic function JPEG2024540005000050.jpg1018 Just note that this is a piecewise linear approximation of 1018, where the pieces are linear on the unit square of the integer lattice.
[0160] Operators can also be split further, for example: If JPEG2024540005000051.jpg1030 is defined by converting (10) to standard form (11), the update step is to first update x, then z < Update z ≧ (or update some other order of these variables). These further partitions can make their updates implicit or explicit. The problem can be further subdivided into other partitions of the system, allowing updates to be applied to the partition variables in any order.
[0161] Additionally, accelerated update rules can be implemented. For example, z k+1 ←z k+1 +θ(z k+1 -z k ), where θ is the acceleration parameter. Similarly, x k+1 ←x k+1 +θ(x k+1 -x k ) These substitutions can be performed after each set of iterations. They have the effect of accelerating the convergence of the algorithm.
[0162] When A is negative definite, the problem is no longer a saddle point problem, since there is no longer a direction in which the system (10) contracts in y. If (10) is maximally monotonic in both ξ and y, then the system can be viewed as a problem of finding a zero of a maximally monotone operator. Instead of a "descent step in x" and an "ascension step in z", we perform two descent steps in the split variables. In other words, all the algorithms from the previous section are applicable, except for replacing "+σ" with "-σ". The resolvent of the linear term is The result is JPEG2024540005000052.jpg921.
[0163] To ensure that the system is monotonic, Certain additional conditions may be required for JPEG2024540005000053.jpg119 and B2. Secondly, if the system is not monotonic, additional steps may be required to ensure the stability of the iterative algorithm.
[0164] If A is singular, it simply means that the solution to the problem is not unique. In this case, We can add additional terms to the system, such as JPEG2024540005000054.jpg1113, in which case the algorithm simply chooses a unique solution from among many others.
[0165] Here, we will explain the decoding process. ξ and y (hat) encode Once y(hat) is found by one of the above algorithms, the variable ξ is then quantized by ξ(hat) = round(ξ), coded and sent to the bitstream. Then, at decoding time, ξ(hat) is recovered from the bitstream. decode is recovered by solving:
number
[0166] This is the latent expression. y(hat) encodeand the decoded y(hatt) which corresponds to, but may not be identical to, the latent representation. decode The error between and depends on: I+A condition number The number of elements of ξ with half-integer values. The error introduced here is the maximum monotonic extension of round( This is due to the use of JPEG2024540005000056.jpg78. When the maximal monotonic function outputs a set (when ξ is a half integer), it differs from the naive round function by up to 1 in absolute value. This introduces some error between encoding and decoding, but can be improved by properly conditioning I+A as well.
[0167] L context and learned bins As mentioned above, compression algorithms can be divided into two phases: encoding and decoding. During the encoding stage, input data is transformed into latent variables that have a smaller representation (in bits) than the original input variables. During the decoding phase, an inverse transformation is applied to the latent variables to recover the original data (or an approximation of the original data).
[0168] AI-based compression systems must also undergo training, which is the process of selecting the parameters of the AI-based compression system that will achieve good compression results (small file size and minimal distortion). Training involves running parts of the encoding and decoding algorithms to determine how to adjust the parameters of the AI-based compression system.
[0169] To be precise, for Al-based compression, the encoding usually looks like this:
number
[0170] where x is the data to be compressed (image or video) and f eneis an encoder, typically a neural network with parameters θ that are trained. The encoder converts input data x into a latent representation y, which is a lower-dimensional, refined form for further compression.
[0171] To further compress y and transmit it as a stream of bits, established lossless coding algorithms such as arithmetic coding are used. These lossless coding algorithms require y to be discrete rather than continuous, and require knowledge of the probability distribution of the latent representation. To achieve this, a quantization function Q (usually round-to-nearest) is used to convert the continuous data into a discrete value y (hat).
[0172] The desired probability distribution P(y(hat)) is found by fitting a probability distribution to the latent space. The probability distribution can be learned directly, but is often a parametric distribution with parameters determined by a hypernetwork consisting of a hyperencoder and a hyperdecoder. When using such a hypernetwork, an additional bitstream z(hat) (also called "side information") may be coded, transmitted, and decoded:
number
[0173] Decryption proceeds as follows:
number
[0174] Summary: The distribution of the latent P(y(hat)) is used in an arithmetic decoder (or other lossless decoding algorithm) to convert the bitstream into quantized latent y(hat). Then, the function f dectransforms the quantified potential into a lossy reconstruction of the input data (denoted as x). In AI-based compression, f dec is typically a neural network and depends on the learned parameters θ.
[0175] When using a hypernetwork, the side information bitstream is first decoded and then used to obtain the parameters needed to construct P(y) needed to decode the main bitstream.
[0176] A key step in typical AI-based image and video compression pipelines is "quantization," where the residual of the latent representation is typically rounded to the nearest integer. This is necessary for algorithms to losslessly encode the bitstream. However, the quantization step introduces its own information loss that impacts reconstruction quality.
[0177] The quantization function can be improved by training a neural network to predict the size of the quantization bin that should be used for each latent pixel. Typically, the latent y is rounded to the nearest integer, which corresponds to a "bin size" of 1. That is, all possible values of y in an interval of length 1 are mapped to the same y (hat):
number
[0178] However, this is not an optimal choice of information loss: for some latent pixels, more information can be ignored (equivalently: use bins larger than 1) without significantly affecting the reconstruction quality; for other latent pixels, the optimal bin size is smaller than 1.
[0179] Learned quantization solves this problem by predicting the quantization bin size for each pixel in the image, which is a tensor You can do this with JPEG2024540005000061.jpg841 and modify the quantization function as follows:
number
[0180] ξ(Hat) y is called the "quantized latent residual." Therefore, Equation 21 becomes:
number
[0181] Note that the learned quantization bin size is incorporated into the modification of the quantization function Q, so that the data we want to encode and transmit can take advantage of the learned quantization bin size. For example, instead of encoding the latent y, we can encode the mean-subtracted latent y-μ y If you want to encode this, you can do this:
number
[0182] Similarly, hyperlatencies, hyperhyperlatencies, and any other object we wish to quantize can all be quantized using a modified quantization function Q for a properly trained Δ. Δ can be used.
[0183] We will now detail several architectures for predicting the quantization bin size. An example architecture is to use a hypernetwork to predict the quantization bin size Δ. The bitstream is encoded as follows:
number
[0184] At decoding time, the bitstream is losslessly decoded as usual. Then, Δ is used to multiply the two elements to obtain ξ y The result of this transformation is y (hat) and is passed to the decoder network as usual:
number
[0185] We will now describe in detail some variants of the above architecture: The size of the quantization bins of the hyperlatencies can be a learned parameter, or we can include a hyperlatency network that predicts these variables. Quantization bin size Δ from hyper decoder y After obtaining , this tensor can be further processed with a nonlinear function, which will typically be a neural network (but is not limited to this choice).
[0186] We also emphasize that our method for learning quantization bins is compatible with all methods that convey meta-information, including hyperpliers, autoregressive models, and implicit models.
[0187] In AI-based compression, it is advantageous to transform the latent variables before quantization using a "context model." A context model improves performance by transforming each residual by a function of its neighboring residuals, thus encoding context information. Upon decoding, this transformation is reversed.
[0188] Many types of context models are computationally expensive, making them difficult or even impossible to use for AI-based compression for end-user devices. However, if the context model is a linear function of neighborhood residuals and is defined through an implicit equation, fast implicit equation-solving algorithms exist, making context modeling feasible from a runtime perspective. An example of such an implicit equation is augmenting the mean-subtracted latent y-μ with a linear transformation of y:
number
[0189] where L is a lower triangular matrix that can be output from the hyperdecoder, or can be a trained model parameter, or can be predicted by another neural network. y appears on both sides of the above equation, so it is defined implicitly.
[0190] The learned quantization and the implicit context model can naturally be combined, in which case the implicit equation becomes:
number
[0191] One challenge of implicit context modeling is the need to train a neural network (running an optimization procedure to find the network parameters that yield the best possible performance). The standard approach is to use the backpropagation algorithm with an optimization procedure such as stochastic gradient descent (or one of its close variants). This presents several challenges: Backpropagation with stochastic gradient descent requires computing the gradient of every transformation in the neural network, but the rounding function is not differentiable. The backpropagation algorithm must use the chain rule of multivariable calculus to "propagate" gradients down each layer of the network. Most automatic differentiation software packages cannot propagate gradients through implicit equations.
[0192] By selecting a specific propagation gradient with a rounding operation, we can compute all gradients of the implicit Equation 28, and these gradients take a form that allows us to design an algorithm that propagates gradients through an AI-based compression pipeline that uses this implicit context model with learned quantization.
[0193] The rounding operation can be replaced by a "rounding by straight-through estimator" that defines the derivative of the rounding function and ignores the rounding operation:
number
[0194] Now we can calculate all the derivatives of the implicit equation 28. The easiest way to do this is to define the function:
number
[0195] Then the implicit equation becomes
number
number
[0196] Calculate the necessary partial derivatives ( Please remember the definition of JPEG2024540005000073.jpg1131 :)
number
number
[0197] The rate loss required for encoding and decoding the bitstream incurs additional gradient computations. Suppose the quantities fed into the entropy model are noise quantized. Then, using learned quantization, the typical rate loss for AI-based compression with an implicit context model is:
number
number
number
[0198] Then the total derivative becomes:
number
[0199] The following partial derivatives are needed (introducing the use of the Kronecker delta):
number
number
[0200] An example algorithm using the above is as follows: 1. Pass image or video x through the encoder to get y. 2. Pass y through the hyperencoder to get z. 3. Quantify z to get z(hat). 4. Pass z (hat) through the hyperdecoder to obtain μ, σ, L, and Δ. 5. Up to this point, all object and network parameters are tracked by an automatic differentiation software package. 6. Use an implicit equation solver to Find y (hat) from JPEG2024540005000082.jpg1381. 7.y(hat) is not currently tracked by automatic differentiation software packages. 8. y(hat) as function y(hat) with grads Surrounded by y (hat) with grads We define the partial derivative of as follows:
number
[0201] The algorithm in the previous section follows from the structure of the implicit equations, so many variations are possible: No learned quantification (remove derivatives of Δ and set Δ=1 everywhere Δ appears) Learned quantization where Δ is processed in an additional neural network layer Implicit equations where L has a banded structure, is upper or lower triangular, or has other structure required for a solution to exist, such as monotonicity, symmetry, or eigenvalue bounds. L does not come from a hyperdecoder, but is instead a learned parameter or predicted from another neural network. -Extension to implicit context modeling such as hyperhyperlatent and hyperhyperlatent. When RoundWithSTE quantization is used instead of noise quantization in rate loss
[0202] Specialization in AI compression As mentioned above, in AI-based image and video compression, a pipeline may consist of a neural network (sometimes called a model) composed of multiple convolutional layers, nonlinear activation functions, and downsampling / upsampling operations. The pipeline can be divided into two components: an encoder that takes in original data samples, such as input images or videos, and returns a latent representation (or latent), and a decoder that takes the latent as input and outputs a (lossy) approximation of the original data samples.
[0203] AI-based compression pipelines train on large sets of data samples. Traditionally, the training set JPEG2024540005000084.jpg1831 should be sufficiently diverse for the model to capture multiple modes in the data space. However, if the data space is extremely high-dimensional and complex, the multimodality of the data space may force over-generalization, resulting in poor performance.
[0204] This concept is illustrated in Figure 4. On the right is an idealized representation of a generalized model trained on the full dataset D, and on the left is an idealized representation of specialized models trained on datasets into which the full dataset D has been partitioned. This concept of separating the data points of the full dataset into different clusters of data and training separate models for each cluster is sometimes called specialization.
[0205] A specialization in AI-based image and video compression is pipelines that exploit the inherent multimodality of the data. Models (cluster models) that can identify the relevant modes (clusters) of each data sample, and training regimes that fit one model to each cluster separately, are expected to significantly improve compression performance. In other words, each D k We can find a partition of the dataset D into K clusters such that K are mutually exclusive and collectively form an exhaustive set of clusters. Then, we can train at least one encoder, decoder, or one model consisting of both an encoder and decoder for each cluster to model the clusters separately, as shown on the right side of Figure 4.
[0206] The model architecture of the encoder and / or decoder may be fixed, in which case the weights of the specialized models may simply be hot-swapped from memory based on the cluster index of each input, which can be cheaply coded into the bitstream.
[0207] Training an AI-based compression pipeline based on the above specialization concept consists of two stages: 1. Cluster the dataset into a set of K small clusters with similar features. 2. Initialize K models, assign each to a cluster set, and train them.
[0208] A schematic of this training process is shown in Figure 5. The first step in specialization is to obtain a cluster model or algorithm 24 that can categorically assign each training and validation instance in the training dataset 20 to a cluster, based on which a compression model can be adapted.
[0209] One way to use a cluster algorithm is to use a model with predefined parameters. Examples of such clustering algorithms include k-means, mini-batch k-means, affinity propagation, mean-shift, spectral clustering, Ward hierarchical clustering, agglomerative clustering, DBSCAN, OPTICS, Gaussian mixtures, BIRCH, and deep clustering. An example of deep clustering is described in Caron, M., Bojanowski, P., Joulin, A. & Douze, M. (2018) Deep Clustering for Unsupervised Learning of Visual Learning, which is incorporated herein by reference.
[0210] The clustering algorithm 24 receives each instance of the training dataset 20 and assigns each instance to a cluster based on one or more features of the instance. For input images, the clustering algorithm can perform the assignment based on one or more of the following features: RGB channel statistics (mean, variance, higher moments), RGB color histogram probabilities, channel statistics and histogram probabilities in other color spaces, 1D power spectral density, neural network-based classification probabilities, file metadata, and deep features. The number of clusters to which an instance is assigned may be predetermined. If neural network-based classification probabilities are used, the neural network used may be trained prior to training the compression model described below. In this case, the neural network is pre-trained on the training dataset 20.
[0211] Alternatively, clustering may be performed by a neural network trained end-to-end in conjunction with the compression model described below. Examples of back-probable clustering techniques include: hard cluster assignment in the forward pass and soft cluster assignment in the backward pass, Gumbel-softmax reparameterization as described in Jang, E., Gu, S. & Poole, B. (2016) Categorical Reparameterization with Gumbel-Softmax, which is incorporated herein by reference, and deep clustering.
[0212] Once the clustering algorithm is performed, the next step is training the specialized network. Using an appropriate clustering model or algorithm, each data sample in the training dataset 20 can be associated with an assigned cluster to derive K cluster sets 25 for training multiple specialized compression pipelines 200. Each specialized compression pipeline 200 consists of a separate encoder 250 and decoder 260. Each specialized compression pipeline 200 can be initialized from scratch rather than restarting from a pre-trained model. When training with an adversarial loss, a separate classifier must be used for each cluster set 25. Each specialized pipeline can then be trained on the set of instances assigned to the associated cluster, as described above.
[0213] Once the pipelines are trained, in operation, a clustering algorithm 24 receives an input image and assigns the input image to a particular cluster. Based on the assigned cluster, a particular dedicated pipeline is selected for encoding, transmission, and decoding of the input image. The cluster index selected by the clustering algorithm may be transmitted as metadata in addition to the latent representation of the image so that the correct decoder 260 can be selected at the receiving end.
[0214] The network architecture is not restricted to being the same across the compression pipeline associated with each cluster set and can be treated as a separate hyperparameter optimization step after cluster analysis. However, from an operational perspective, it may be more efficient for the encoder and / or decoder architecture to be similar or identical across one or more of the cluster sets, allowing for more efficient loading in and out of network parameters into the computational graph on mobile hardware.
[0215] A generalized compression pipeline can also be combined to work in conjunction with multiple specialized compression pipelines. Such a model serves as a fallback option if a specialized pipeline proves to compress a particular data sample worse than a generalized pipeline. To obtain this model, training of the compression pipeline can be performed on the training data set 20 without using a clustering algorithm.
[0216] The concept of specialization allows for the creation of pipelines that are specialized for domains or media categories. For example, one compression model might be specialized for movies, one for sports, one for video conferencing, one for a general fallback model, etc. A hierarchical structure is also possible, with one level of clustering performed on media subcategories, e.g., one model per movie genre (action, anime, horror, sci-fi), one model per individual sport (soccer, skiing, boxing), etc.
[0217] An alternative approach to the above is to be domain-agnostic and let a deep clustering model or feature extractor, such as a neural network, determine the data partitioning. In this approach, the neural network receives input images as input, and a further clustering algorithm can be applied to the neural network's output to obtain appropriate clusters. The input images may then be passed through an associated compression pipeline, as described above. In this approach, the neural network may be trained to produce outputs that result in better clustering of the input images by the clustering algorithm. For example, the loss function used to train the neural network may be based on the variance between the neural network's outputs to encourage variance in the outputs. The neural network may be the encoder of an autoencoder system, and the output is subsequently used as the input to a decoder, whose output is an approximation of the input images.
[0218] An image may be divided into several sub-images that are used as input to the system. Each sub-image may be the same size. For example, each image may be macroblocked into small, fixed-resolution contiguous blocks of 32x32 residuals, and K specialized models, each taking the 32x32 residual input, may be identified and trained. This approach can be implemented if a neural network determines the data division before clustering, as described above. In this approach, each block is processed with a specialized transform coding algorithm that is best suited to that block.
[0219] A visualization of such an encoding system is shown in Figure 6. An initial image 34 is divided into multiple sub-images or blocks 35. Each sub-image 35 is assigned to a cluster by a clustering algorithm 24 and encoded by the cluster's associated encoder 250 to generate multiple latent representations associated with the multiple sub-images 35.
[0220] For such systems, it may be beneficial to combine or combine the latent components from each block into a latent space for efficient entropy modeling during transmission via the bitstream 250. To do this, the latent blocks output from each encoder can be strung together in a predetermined arrangement. Techniques such as those described in Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenbor, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J. & Houlsby, N. (2020) can also be utilized. A picture is worth 16x16 words: Scale and Wu, B., Xu, C., Dai, X., Wan, A., Zhang, P., Yan, Z., Tomizuka, M., Gonzalez, J., Keutzer, K. & Vajda, P. (2020) Visual Transformers: Token-Based Image Representation and Processing for Computer Vision, which is incorporated herein by reference, treat each latent block as a token to model dependencies between blocks.
[0221] The cluster index used for each instance may be included in the bitstream at a specific step 270 from the clustering algorithm. Fortunately, this is barely 1 byte per patch, a fixed number of bits per residual (bpp) that depends only on the patch size and the number of specialized models (for K=100 and a 32x32 patch, a constant bpp for the cluster index encoding might be 0.0068).
[0222] A visualization example of a decoding system corresponding to the example of Figure 6 is shown in Figure 7. As shown in Figure 7, in the decoding system, the index is obtained from the bitstream in step 280. Each block is processed through a corresponding dedicated decoder, after which the images are stitched together and may be post-processed with a boundary removal filter.
[0223] An example of results using the specialized pipeline is shown below. 16M preprocessed 256x256 image crops were processed using the clustering algorithm MiniBatchK Means to form K=100 clusters, each assigned a cluster index between 0 and 99. For each cluster, a dedicated compression pipeline consisting of a 7E6D128 architecture was trained. A generalized model was trained on the entire uncropped image dataset. Sample cluster results are shown below. [Table 1]
[0224] As shown, the specialized model outperforms the generalized model in each cluster in both rate and distortion performance.
Claims
1. 1. A method of training one or more neural networks, the one or more neural networks being for use in encoding, transmitting, and decoding lossy images or lossy videos, the method comprising: receiving an input image at a first computer system; encoding the input image using a first neural network to generate a latent representation; decoding the latent representation using a second neural network to generate an output image, the output image being an approximation of the input image; Including, wherein the first neural network and the second neural network each include a plurality of layers; and at least one of the plurality of layers of the first or second neural network includes a transformation; The method comprising: evaluating the difference between the output image and the input image and evaluating a function based on the output of the transformation; updating parameters of the first neural network and the second neural network based on the estimated difference and the estimated function; repeating the steps above with a first set of input images to generate a first trained neural network and a second trained neural network.
2. 2. The method of claim 1, wherein the function is based on multiple outputs of multiple transforms, each transform being in a different one of multiple layers of the first or second neural network.
3. The method of claim 2 , wherein the contribution of each of the plurality of outputs to the evaluation of the function is scaled by a predetermined value.
4. the output of the transformation comprises a plurality of values; The method of any one of claims 1 to 3, wherein the function is based on the square of each of the plurality of values.
5. the output of the transformation comprises a plurality of values; The method of any one of claims 1 to 3, wherein the function is based on the absolute value of each of the plurality of values.
6. The method of claim 4 , wherein the function is based on the square and absolute value of each of the plurality of values.
7. The method of claim 4 , wherein the function is based on an average of each of the plurality of values.
8. 7. The method of claim 6, wherein the squared and absolute value contributions to the evaluation of the function are scaled by a predetermined value.
9. The method of claim 4 , wherein the function is further based on a variance of the plurality of values.
10. 2. The method of claim 1, wherein the contribution of a cost function to the updates to the parameters of the first neural network and the second neural network is scaled by a predetermined value.
11. the contribution of a cost function to the update to the parameters of the first neural network and the second neural network is scaled by a value; The method of claim 1 , wherein the value is updated incrementally in at least one iteration of the steps of the method.
12. at least one of the layers of the first neural network or the second neural network includes an additional transformation applied after the transformation; The parameters of the linear transformation are additionally updated based on the estimated difference and the estimated function. The method of claim 1 , wherein the function is not based on the output of the additional transform.
13. 1. A method for encoding, transmitting and decoding lossy images or video, said method comprising: receiving an input image at a first computer system; encoding the input image using a first trained neural network to generate a latent representation; transmitting the latent representation to a second computer system; decoding the latent representation using a second trained neural network to generate an output image, the output image being an approximation of the input image; 10. The method of claim 1, wherein the first trained neural network and the second trained neural network are trained according to the method of claim 1.
14. 1. A method for encoding and transmitting lossy images or video, said method comprising: receiving an input image at a first computer system; encoding the input image using a first trained neural network to generate a latent representation; transmitting the latent representation; 10. The method of claim 1, wherein the first trained neural network is trained according to the method of claim 1.
15. 1. A method for receiving and decoding a lost image or video, said method comprising: receiving at a second computer system the latent representation transmitted according to the method of claim 14; decoding the latent representation using a second trained neural network to generate an output image, the output image being an approximation of the input image; 10. The method of claim 1, wherein the second trained neural network is trained according to the method of claim 1.
16. 10. A data processing system configured to perform the method of claim 1.
17. 16. A data processing apparatus configured to carry out the method according to claim 14 or 15.
18. A computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of claim 14 or 15.
19. A computer-readable storage medium containing instructions that, when executed by a computer, cause the computer to perform the method of claim 14 or 15.