Method and data processing system for encoding, transmitting and decoding a lossy image or video - Patents.com

JP2024528208A5Active Publication Date: 2025-08-06INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024506650
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-08-03
Filing Date
2022-08-03
Publication Date
2025-08-06
Estimated Expiration
2042-08-03

AI Technical Summary

Technical Problem

Existing image and video compression technologies face challenges in achieving efficient data reduction while minimizing perceptible information loss, particularly in Al-based methods that struggle with poor compression results and high transmission requirements.

Method used

A method involving neural networks for encoding and decoding lossy images and videos, utilizing quantization processes with variable bin sizes based on input images, and training networks to minimize differences through iterative parameter updates, with region-of-interest adjustments and hyper-latent representations.

Benefits of technology

Improves compression efficiency by up to 1.5% in rate and 1.9% in distortion reduction, enhancing the quality of reconstructed images and videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

1. A method for lossy image and video encoding, transmission, and decoding, comprising: receiving an input image at a first computer system; encoding the input image with a first trained neural network to generate a latent representation; performing a quantization process on the latent representation to generate a quantized latent representation, where a bin size used in the quantization process is based on the input image; and transmitting the quantized latent to a second computing system and decoding the quantized latent with a second trained neural network to generate an output image, where the output image is an approximation of the input image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a method and system for encoding, transmitting and decoding a lost image or video, a method, apparatus, computer program and computer readable storage medium for encoding and transmitting a lost image or video, and a method, apparatus, computer program and computer readable storage medium for receiving and decoding a lost image or video.

[0002] There is an increasing demand for image and video content from users of communication networks. The demand is increasing not only for the number of images and the duration of videos, but also for higher resolution content. This increases the load on communication networks and the amount of data transmitted, which in turn increases the energy usage of communication networks.

[0003] To reduce the impact of these problems, image and video content is compressed for transmission over networks. Image and video content compression can be lossless or lossy. Lossless compression compresses images and videos in such a way that all of the original information contained in the content can be restored. However, there is a limit to the amount of data reduction that can be achieved when using lossless compression. With lossy compression, information is lost from images and videos during the compression process. Known compression techniques attempt to minimize the apparent loss of information by removing information that results in changes to the image and video after decompression that are not particularly noticeable to the human visual system.

[0004] Artificial intelligence (Al)-based compression techniques achieve image and video compression and decompression by using trained neural networks in the compression and decompression process. Typically, during training of the neural network, the differences between the original images and videos and the compressed and decompressed images and videos are analyzed, and the parameters of the neural network are modified to reduce the differences while minimizing the data required to transmit the content. However, Al-based compression methods can sometimes produce poor compression results in terms of the appearance of the compressed images and videos and the amount of information required for transmission.

[0005] According to the present invention, there is provided a method for lossy image and video encoding, transmission and decoding, the method including the steps of receiving an input image at a first computer system, encoding the input image using a first trained neural network to generate a latent representation, performing a quantization process on the latent representation to generate a quantized latent, where a bin size used in the quantization process is based on the input image, transmitting the quantized latent to a second computer system, and decoding the quantized latent with a second trained neural network to generate an output image, where the output image is an approximation of the input image.

[0006] The size of the bin may vary between at least two pixels of the latent representation.

[0007] The bin sizes may differ between at least two channels of the latent representation.

[0008] A bin size may be assigned to each pixel in the latent representation.

[0009] The quantization process may involve performing an operation on the value of each pixel of the latent representation that corresponds to the bin size assigned to that pixel.

[0010] The quantization process may involve subtracting the mean value of the latent representation from each pixel of the latent representation.

[0011] The quantization process may include a rounding function.

[0012] The size of the bins used to decode the quantized latent may be based on previously decoded pixels of the quantized latent.

[0013] The quantization process may be constructed from a third trained neural network.

[0014] A third trained neural network may receive as input at least one previously decoded pixel of the quantized latents.

[0015] The method may further include encoding the latent representation with a fourth trained neural network to generate a hyper-latent representation, performing a quantization process on the hyper-latent representation to generate a quantized hyper-latent, transmitting the quantized hyper-latent to a second computer system, and decoding the quantized hyper-latent with a fifth trained neural network to obtain a size of the bins, where the decoding of the quantized latent uses the obtained size of the bins.

[0016] The output of the fifth trained neural network may be processed by a further function to obtain the bin size.

[0017] The further function may be a sixth trained neural network.

[0018] The size of the bins used in the quantization process of the hyper-latent representation may be based on the input image.

[0019] The method may further include identifying at least one region of interest in the input image and reducing a size of a bin used in the quantization process for at least one corresponding pixel of the latent representation within the identified region of interest.

[0020] The method may further include identifying at least one region of interest in the input image and using different quantization processes for at least one corresponding pixel of the latent representation within the identified region of interest.

[0021] The at least one region of interest may be identified by a seventh trained neural network.

[0022] The locations of the one or more regions of interest may be stored in a binary mask, and the binary mask may be used to obtain the size of the bins.

[0023] According to the present invention, there is provided a method of training one or more neural networks for use in encoding, transmitting, and decoding lossy images or videos, the method including the steps of receiving an input image at a first computer system; encoding the input image with the first neural network to generate a latent representation; performing a quantization process on the latent representation to generate a quantized latent representation, where a size of a bin used in the quantization process is based on the input image; decoding the quantized latent representation with a second neural network to generate an output image, where the output image is an approximation of the input image; determining a quantity based on a difference between the output image and the input image; updating parameters of the first neural network and the second neural network based on the determined quantity; and repeating the above steps using a first set of input images to generate a first trained neural network and a second trained neural network.

[0024] The method may further include encoding the latent representation using a third neural network to generate a hyper-latent representation; performing a quantization process on the hyper-latent representation to generate a quantized hyper-latent representation; transmitting the quantized hyper-latent representation to a second computer system; decoding the quantized hyper-latent representation using a fourth neural network to obtain a bin size, wherein the decoding of the quantized latent uses the obtained bin size; and parameters of the third neural network and the fourth neural network are additionally updated based on the determined amount to obtain a third trained neural network and a fourth trained neural network.

[0025] The quantization process may include a first quantization approximation.

[0026] The determined amount may be additionally based on a rate associated with the quantized potential, and a second quantized approximation may be used to determine the rate associated with the quantized potential, and the second quantized approximation may differ from the first quantized approximation.

[0027] The determined quantity may include a loss function, and updating the parameters of the neural network may include evaluating a gradient of the loss function and backpropagating the gradient of the loss function through the neural network, wherein a third quantized approximation is used during the backpropagation of the gradient of the loss function, the third quantized approximation being the same approximation as the first quantized approximation.

[0028] The parameters of the neural network may additionally be updated based on the distribution of the bin sizes.

[0029] At least one parameter of the distribution may be learned.

[0030] The distribution may be an inverse gamma distribution.

[0031] The distribution may be determined by a fifth neural network.

[0032] According to the present invention, there is provided a method for lossy image or video encoding and transmission, the method including the steps of receiving an input image at a first computer system; encoding the input image using a first trained neural network to generate a latent representation; performing a quantization process on the latent representation to generate a quantized latent, where a bin size used in the quantization process is based on the input image; and transmitting the quantized latent.

[0033] According to the present invention, there is provided a method for receiving and decoding a lost image or video, the method comprising the steps of receiving, at a second computer system, a quantized latent transmitted according to the above method, and decoding the quantized latent using a second trained neural network to generate an output image, the output image being an approximation of the input image.

[0034] According to the present invention, there is provided a method of training one or more neural networks for use in encoding, transmitting, and decoding lossy images or videos, the method including the steps of: receiving an input image at a first computer system; encoding the input image using the first neural network to generate a latent representation; performing a quantization process on the latent representation to generate a quantized latent; decoding the quantized latent using a second neural network to generate an output image, the output image being an approximation of the input image; determining a quantity based on a difference between the output image and the input image; updating parameters of the first and second neural networks based on the determined quantity; and repeating the above steps using a plurality of sets of input images to generate one trained neural network and a second trained neural network, wherein at least one of the plurality of sets of input images includes a first proportion of images including a particular feature and at least one other of the plurality of sets of input images includes a second proportion of images including the particular feature, the second proportion being different from the first proportion.

[0035] The first proportion may be all of the images in the set of input images.

[0036] The particular feature may be any of a human face, an animal face, a letter, an eye, lips, a logo, a car, a flower, and a pattern.

[0037] Each of the multiple sets of input images may be used the same number of times during the repetitions of the method steps.

[0038] The difference between the output image and the input image may be determined at least in part by a neural network acting as a classifier.

[0039] A separate neural network acting as a classifier may be used for each set of multiple sets of input images.

[0040] One or more parameters of the neural network acting as a classifier may be updated for a first number of training steps, and one or more other parameters of the neural network acting as a classifier may be updated for a second number of training steps, the second number being smaller than the first number.

[0041] The determined amount may be additionally based on rates associated with the quantized potentials, and the parameter updates for at least one of the multiple sets of input images may use a first weighting for the rates associated with the quantized potentials and the parameter updates for at least one other of the multiple sets of input images may use a second weighting for the rates associated with the quantized potentials, the second weighting being different from the first weighting.

[0042] The difference between the output image and the input images may be determined at least in part using a plurality of perceptual metrics, and parameter updates for at least one of the plurality of sets of input images may use a first set of weightings for the plurality of perceptual metrics and parameter updates for at least one other of the plurality of sets of input images may use a second set of weightings for the plurality of perceptual metrics, the second set of weightings being different from the first set of weightings.

[0043] The input image may be a modified image in which one or more regions of interest have been identified by a third trained neural network and other regions of the image have been masked.

[0044] The region of interest may be a region that includes one or more features of a human face, an animal face, text, eyes, lips, a logo, a car, a flower, or a pattern.

[0045] The locations of the one or more areas of interest may be stored in a binary mask.

[0046] The binary mask may be an additional input to the first neural network.

[0047] According to the present invention, there is provided a method for encoding, transmitting and decoding lossy images and videos, the method comprising the steps of receiving an input image at a first computer system, encoding the input image using a first trained neural network to generate a latent representation, performing a quantization process on the latent representation to generate a quantized latent, transmitting the quantized latent to a second computer system, and decoding the quantized latent using a second trained neural network to generate an output image, the output image being an approximation of the input image, wherein the first trained neural network and the second trained neural network are trained according to the above method.

[0048] According to the present invention, there is provided a method for lossy image or video encoding and transmission, the method comprising the steps of receiving an input image at a first computer system, encoding the input image using a first trained neural network to generate a latent representation, performing a quantization process on the latent representation to generate a quantized latent, and transmitting the quantized latent, wherein the first trained neural network is trained according to the above method.

[0049] According to the present invention there is provided a method for receiving and decoding a lost image or video, the method comprising the steps of receiving at a second computer system latents quantized according to the method of claim 48 and decoding the quantized latents using a second trained neural network to generate an output image, the output image being an approximation of the input image, the second trained neural network having been trained according to the above method.

[0050] According to the present invention, there is provided a method of training one or more neural networks for use in encoding, transmitting and decoding lossy images or videos, the method comprising the steps of: receiving an input image at a first computer system; encoding the input image with the first neural network to generate a latent representation; performing a quantization process on the latent representation to generate a quantized latent representation; decoding the quantized latent representation using a second neural network to generate an output image, the output image being an approximation of the input image; determining a quantity based on a rate associated with the quantized latent, where evaluation of the rate includes a step of interpolating a discrete probability rate mass function; updating parameters of the first neural network and the second neural network based on the determined quantity; and repeating the above steps using a set of multiple input images to generate a first trained neural network and a second trained neural network.

[0051] At least one parameter of the discrete probability rate mass function may be additionally updated based on the estimated rate.

[0052] The method may further include encoding the latent representation using a third neural network to generate a hyper-latent representation; performing a quantization process on the hyper-latent representation to generate a quantized hyper-latent; and decoding the quantized hyper-latent using a fourth neural network to obtain at least one parameter of a discrete probability mass function, wherein parameters of the third neural network and the fourth neural network are additionally updated based on the determined amount to obtain a third trained neural network and a fourth trained neural network.

[0053] The interpolation may include at least one of piecewise constant interpolation, nearest neighbor interpolation, linear interpolation, polynomial interpolation, spline interpolation, piecewise cubic interpolation, Gaussian process, and kriging.

[0054] The discrete probability rate mass function may be a categorical distribution.

[0055] The categorical distribution may be parameterized by at least one vector.

[0056] The categorical distribution may be obtained by softmax projection of the vectors.

[0057] The discrete probability mass function may be parameterized by at least a mean parameter and a scale parameter.

[0058] The discrete probability mass function may be multivariate.

[0059] The discrete probability mass function may be constructed from a plurality of points, where a first set of adjacent points of the plurality of points may have a first spacing and a second set of adjacent points of the plurality of points may have a second spacing, where the second spacing is different from the first spacing.

[0060] The discrete probability mass function may be constructed from a plurality of points, where a first set of adjacent points of the plurality of points may have a first spacing and a second set of adjacent points of the plurality of points may have a second spacing, where the second spacing is equal to the first spacing.

[0061] At least one of the first interval and the second interval may be obtained using a fourth neural network.

[0062] At least one of the first interval and the second interval may be obtained based on a value of at least one pixel of the latent representation.

[0063] According to the present invention, there is provided a method for encoding, transmitting and decoding lossy images and videos, the method comprising the steps of receiving an input image at a first computer system, encoding the input image using a first trained neural network to generate a latent representation, performing a quantization process on the latent representation to generate a quantized latent, transmitting the quantized latent to a second computer system, and decoding the quantized latent using a second trained neural network to generate an output image, where the output image is an approximation of the input image, wherein the first trained neural network and the second trained neural network are trained according to the above method.

[0064] According to the present invention, there is provided a method for encoding and transmitting a lossy image or video, the method comprising the steps of receiving an input image at a first computer system, encoding the input image using a first trained neural network to generate a latent representation, performing a quantization process on the latent representation to generate a quantized latent, and transmitting the quantized latent, wherein the first trained neural network is trained according to the above method.

[0065] According to the present invention, there is provided a method for receiving and decoding a lost image or video, the method comprising the steps of receiving, in a second computer system, latents quantized according to the above method, and decoding the quantized latents using a second trained neural network to generate an output image, the output image being an approximation of the input image, the second trained neural network being trained according to the above method.

[0066] According to the present invention, there is provided a method for encoding, transmitting and decoding lossy images and videos, the method comprising the steps of receiving an input image at a first computer system, encoding the input image using a first trained neural network to generate a latent representation, performing a first operation on the latent representation to obtain a residual latent, transmitting the residual latent to a second computer system, performing a second operation on the residual latent to obtain a retrieved latent representation, the second operation including performing an operation on previously obtained pixels of the retrieved latent, and decoding the retrieved latent representation using the second trained neural network to generate an output image, the output image being an approximation of the input image.

[0067] The operations on the retrieved potential acquired pixels may be performed for each pixel for which a retrieved potential acquired pixel is available.

[0068] At least one of the first operation and the second operation may include solving an implicit system of equations.

[0069] The first operation may include a quantization operation.

[0070] The operations performed on the retrieved potential previously obtained pixels may include matrix operations.

[0071] The matrices that define the matrix operations may be sparse.

[0072] The matrices that define the matrix operation may have zero values ​​that correspond to captured potential pixels that are not captured at the time the matrix operation is performed.

[0073] The matrix that defines the matrix operation may be lower triangular.

[0074] The second operation may be constructed with a standard forward substitution.

[0075] The operations performed on the previously obtained pixels of the obtained latent may be constructed from a third trained neural network.

[0076] The method may further include encoding the latent representation using a fourth trained neural network to generate a hyper-latent representation, transmitting the quantized hyper-latent to a second computer system, and decoding the quantized hyper-latent using a fifth trained neural network, where operations performed on previously obtained pixels of the obtained latent are based on an output of the fifth trained neural network.

[0077] Decoding the quantized hyper-latencies with the fifth trained neural network may further generate mean parameters, and the implicit equation system may further include the mean parameters.

[0078] According to the present invention, there is provided a method of training one or more neural networks, the one or more neural networks being for use in encoding, transmitting, and decoding lossy images or videos, the method comprising the steps of: receiving an input image at a first computer system; encoding the input image using the first neural network to generate a latent representation; performing a first operation on the latent representation to obtain a latent residual; performing a second operation on the latent residual to obtain a obtained latent representation, the second operation including performing an operation on previously obtained pixels of the obtained latent; decoding the quantized latent using a second neural network to generate an image, the output image being an approximation of the input image; determining a quantity based on a difference between the output image and the input image; updating parameters of the first neural network and the second neural network based on the determined quantity; and repeating the above steps using a first set of input images to generate a first trained neural network and a second trained neural network.

[0079] The operations performed on the retrieved potential previously obtained pixels may include matrix operations.

[0080] The parameters of the matrices defining the matrix operation may additionally be updated based on the determined amount.

[0081] The operations performed on the previously obtained pixels of the obtained latent may include a third neural network, and parameters of the third neural network may be additionally updated based on the determined amount to generate a third trained neural network.

[0082] The method further includes the steps of encoding the latent representation using a fourth neural network to generate a hyper-latent representation, performing a quantization process on the hyper-latent representation to generate a quantized hyper-latent representation, transmitting the quantized hyper-latent representation to a second computer system, and decoding the quantized hyper-latent representation using a fifth neural network, wherein operations performed on previously obtained pixels of the obtained latent are based on the output of the fifth trained neural network, and parameters of the fourth neural network and the fifth neural network are additionally updated based on the determined amount to generate the fourth trained neural network and the fifth trained neural network.

[0083] According to the present invention, there is provided a method for lossy image or video encoding and transmission, the method including the steps of receiving an input image at a first computer system, encoding the input image using a first trained neural network to generate a latent representation, performing a first operation on the latent representation to obtain a residual latent, and transmitting the residual latent.

[0084] According to the present invention, there is provided a method for receiving and decoding a lost image or video, the method comprising the steps of: receiving, at a second computer system, a residual latent transmitted according to the above method; performing a second operation on the residual latent to obtain a obtained latent representation, the second operation comprising performing an operation on previously obtained pixels of the obtained latent; and decoding the obtained latent representation using a second trained neural network to generate an output image, the output image being an approximation of the input image.

[0085] In accordance with the present invention, there is provided a method of training one or more neural networks for use in encoding, transmitting, and decoding lossy images or videos, the method including the steps of receiving an input image at a first computer system, encoding the input image with the first neural network to generate a latent representation, entropy encoding the latent representation, transmitting the entropy encoded latent representation to a second computer system, entropy decoding the entropy encoded latent representation, and decoding the latent representation using the second neural network. the step of: decoding the current image to generate an output image, the output image being an approximation of the input image; determining a quantity based on a difference between the output image and the input image; updating parameters of the first and second neural networks based on the determined quantity; and repeating the above steps with a first set of input images to generate a first trained neural network and a second trained neural network, wherein entropy decoding of the entropy encoded latent representation is performed pixel by pixel, and an order of decoding pixel by pixel is incrementally updated based on the determined quantity.

[0086] The pixel-by-pixel decoding order may be based on the latent representation.

[0087] Entropy decoding of entropy encoded latent images may involve operations based on previously decoded pixels.

[0088] Determining the pixel-by-pixel decoding order may include ordering the pixels of the latent representation in a directed acyclic graph.

[0089] Determining the pixel-by-pixel decoding order may involve operating on the latent representation using multiple adjacency matrices.

[0090] Determining the pixel-by-pixel decoding order may include splitting the latent representation into multiple sub-images.

[0091] Multiple sub-images may be obtained by convolving the latent representation with multiple binary mask kernels.

[0092] Determining a pixel-by-pixel decoding order may include ranking multiple pixels of the latent representation based on a magnitude of a quantity associated with each pixel.

[0093] The quantity associated with each pixel may be a position parameter or a scale parameter associated with that pixel.

[0094] The quantity associated with each pixel may additionally be updated based on the estimated difference.

[0095] Determining the pixel-by-pixel decoding order may include a wavelet decomposition of multiple pixels of the latent representation.

[0096] The pixel-by-pixel decoding order may be based on the frequency content of the wavelet decomposition associated with the pixels.

[0097] The method may further include encoding the latent representation using a fourth trained neural network to generate a hyper-latent representation, transmitting the hyper-latent to a second computer system, and decoding the hyper-latent using a fifth trained neural network, where a pixel-by-pixel decoding order is based on an output of the fifth trained neural network.

[0098] According to the present invention, there is provided a method for lossy image and video encoding, transmission and decoding, the method comprising the steps of receiving an input image at a first computer system, encoding the input image using a first trained neural network to generate a latent, entropy encoding the latent, transmitting the entropy encoded latent to a second computer system, entropy decoding the entropy encoded latent, and decoding the latent using a second trained neural network to generate an output image, the output image being an approximation of the input image, wherein the first trained neural network and the second trained neural network are trained according to the above method.

[0099] According to the present invention, there is provided a method for lossy image or video encoding and transmission, the method comprising the steps of receiving an input image at a first computer system, encoding the input image using a first trained neural network to generate a latent representation, entropy encoding the latent representation, and transmitting the entropy encoded latent representation, wherein the first trained neural network is trained according to the above method.

[0100] According to the present invention, there is provided a method for receiving and decoding a lossy image or video, the method comprising the steps of receiving, in a second computer system, an entropy coded latent representation transmitted in accordance with the above method, and decoding the latent representation using a second trained neural network to generate an output image, the output image being an approximation of the input image, the second trained neural network being trained in accordance with the above method.

[0101] According to the present invention, there is provided a method of training one or more neural networks for use in encoding, transmitting and decoding lossy images or videos, the method comprising the steps of receiving an input image at a first computer system; encoding the input image with a first neural network to generate a latent representation; decoding the latent representation with a second neural network to generate an output image, the output image being an approximation of the input image; and determining a quantity based on a difference between the output image and the input image and a rate associated with the latent representation, wherein a first weighting is applied when determining the quantity. the step of applying a weighting to a difference between the force image and the input image, and a second weighting to a rate associated with the latent representation; updating parameters of the first and second neural networks based on the determined amount; and repeating the above steps with the first set of input images to generate a first trained neural network and a second trained neural network, wherein after at least one of the repetitions of the above steps, at least one of the first weighting and the second weighting is additionally updated based on a further amount, the further amount being based on at least one of the difference between the output image and the input image and the rate associated with the latent representation.

[0102] At least one of a difference between the output image and the input image and a rate associated with the latent representation may be recorded for each iteration of the steps, and the further amount may be based on a plurality of previously recorded differences between the output image and the input image and / or a plurality of previously recorded rates associated with the latent representation.

[0103] The further amount may be based on an average of multiple previously recorded differences or rates.

[0104] The average may be at least one of an arithmetic mean, a median, a geometric mean, a harmonic mean, an exponential moving average, a smoothed moving average, and a linear weighted moving average.

[0105] Outliers may be removed from previously recorded differences or rates before determining further quantities.

[0106] Outliers may be removed only for the first predetermined number of iterations of the step.

[0107] A rate associated with the latent expression may be calculated using a first method when determining the amount and may be calculated using a second method when determining the further amount, the first method being different from the second method.

[0108] At least one iteration of the steps may be performed using an input image from a second set of input images, and when an input image from the second set of input images is used, the parameters of the first neural network and the second neural network may not be further updated.

[0109] The determined quantity may further be based on the output of a neural network acting as a classifier.

[0110] In accordance with the present invention, there is provided a method of training one or more neural networks for use in encoding, transmitting, and decoding lossy video, the method comprising the steps of receiving an input video at a first computer system; encoding a plurality of frames of the input video with a first neural network to generate a plurality of latent representations; decoding the plurality of latent representations with a second neural network to generate a plurality of frames of an output video, the output video being an approximation of the input video; and determining an amount based on a difference between the output image and the input image and a rate associated with the plurality of latent representations, the first weighting being an approximation of the output image. the step of applying a weighting factor to a difference between the output image and the input image, and a second weighting factor to rates associated with the plurality of latent representations; further updating parameters of the first and second neural networks based on the determined amount; and repeating the above steps with the plurality of input images to generate a first trained neural network and a second trained neural network, wherein after at least one of the repetitions of the above steps, at least one of the first and second weighting factors is additionally further updated based on a further amount, the further amount being based on at least one of the difference between the output image and the input image and the rates associated with the plurality of latent representations.

[0111] The input video includes at least one I-frame and multiple P-frames.

[0112] The amount may be based on a plurality of first weightings or second weightings, each of the weightings corresponding to one of a plurality of frames of the input video.

[0113] After at least one of the iterations of the steps, at least one of the plurality of weightings may be additionally updated based on an additional amount associated with each weighting.

[0114] Each additional amount may be based on a predetermined target value of the difference between the output frame and the input frame, or a rate associated with the latent representation.

[0115] The additional amount associated with the I-frame may have a first target value and at least one additional amount associated with the P-frame may have a second target value, the second target value being different from the first target value.

[0116] Each additional amount associated with a P frame may have the same target value.

[0117] The plurality of first weightings or second weightings may be initialized to zero.

[0118] According to the present invention, there is provided a method for lossy image and video coding, transmission and decoding, the method comprising the steps of receiving an input image at a first computer system, encoding the input image using a first trained neural network to generate a latent representation, performing a quantization process on the latent representation to generate a quantized latent, transmitting the quantized latent to a second computer system, and decoding the quantized latent using a second trained neural network to generate an output image, where the output image is an approximation of the input image, wherein the first trained neural network and the second trained neural network are trained according to the above method.

[0119] According to the present invention, there is provided a method for lossy image or video encoding and transmission, the method comprising the steps of receiving an input image or video at a first computer system, encoding the input image or video using a first trained neural network to generate a latent representation, and transmitting the latent representation, the first trained neural network being trained according to the above method.

[0120] According to the present invention, there is provided a method for receiving and decoding a lost image or video, the method comprising the steps of receiving, at a second computer system, a latent representation according to the above method, and decoding the latent representation using a second trained neural network to generate an output image or video, the output image or video being an approximation of the input image or video, the second trained neural network being trained according to the above method.

[0121] According to the present invention, there is provided a method for lossy image or video encoding, transmission and decoding, the method comprising: receiving an input image at a first computer system; encoding the input image using a first trained neural network to generate a latent representation; performing a quantization process on the latent representation to generate a quantized latent; entropy encoding the quantized latent using a probability distribution, the probability distribution being defined using a tensor network; transmitting the entropy encoded quantized latent to a second computer system; entropy decoding the entropy encoded quantized latent using the probability distribution to obtain a quantized latent; and decoding the quantized latent using a second trained neural network to generate an output image, where the output image is an approximation of the input image.

[0122] The probability distribution may be defined by a Hermitian operator operating on the quantized latents, where the Hermitian operator is defined by a tensor network.

[0123] A tensor network may include a non-normal core tensor and one or more orthonormal tensors.

[0124] The method may further include encoding the latent representation using a third trained neural network to generate a hyper-latent representation, performing a quantization process on the hyper-latent representation to generate a quantized hyper-latent representation, transmitting the quantized hyper-latent representation to a second computer system, and decoding the quantized hyper-latent representation using a fourth trained neural network, where an output of the fourth trained neural network is one or more parameters of the tensor network.

[0125] The tensor network may be composed of an unnormalized core tensor and one or more orthonormal tensors, and the output of the fourth trained neural network may be one or more parameters of the unnormalized core tensor.

[0126] One or more parameters of the tensor network may be computed using one or more pixels of the latent representation.

[0127] The probability distribution may be associated with a subset of the pixels of the latent representation.

[0128] The probability distributions may be associated with the channels of the latent representation.

[0129] The tensor network may be at least one of a tensor tree, a locally refined state, a Borne machine, a matrix product state, and a projected entangled pair state factorization.

[0130] According to the present invention, there is provided a method for training one or more networks, the one or more networks being used for encoding, transmitting, and decoding lossy images or videos, comprising the steps of: receiving a first input image; encoding the first input image using a first neural network to generate a latent representation; performing a quantization process on the latent representation to generate a quantized latent; entropy encoding the quantized latent using a probability distribution, the probability distribution being defined using a tensor network; entropy decoding the entropy encoded quantized latent using the probability distribution to obtain a quantized latent; decoding the quantized latent using a second neural network to generate an output image, the output image being an approximation of the input image; determining an amount based on a difference between the output image and the input image; updating parameters of the first neural network and the second neural network based on the determined amount; and repeating the above steps using a plurality of input images to generate a first trained neural network and a second trained neural network.

[0131] One or more of the parameters of the tensor network may be additionally updated based on the determined quantity.

[0132] The tensor network may include an unnormalized core tensor and one or more orthonormal tensors, and parameters of all tensors in the tensor network except for the unnormalized core tensor may be updated based on the determined quantity.

[0133] A tensor network may be computed using the latent representation.

[0134] The tensor network may be computed based on linear interpolation of the latent representation.

[0135] The determined quantity may further be based on the entropy of the tensor network.

[0136] According to the present invention, there is provided a method for lossy image or video encoding and transmission, the method including the steps of receiving an input image at a first computer system, encoding the input image using a first trained neural network to generate a latent representation, performing a quantization process on the latent representation to generate a quantized latent, entropy encoding the quantized latent using a probability distribution, where the probability distribution is defined using a tensor network, and transmitting the entropy encoded quantized latent.

[0137] According to the present invention, there is provided a method for receiving and decoding a lost image or video, the method comprising the steps of receiving, at a second computer system, an entropy coded quantized latent transmitted according to the above method, entropy decoding the entropy coded quantized latent using a probability distribution to obtain a quantized latent, and decoding the quantized latent using a second trained neural network to generate an output image, where the output image is an approximation of the input image.

[0138] According to the present invention, there is provided a method for lossy image and video encoding, transmission and decoding, the method comprising: receiving an input image at a first computer system; encoding the input image using a first trained neural network to generate a latent representation; encoding the latent representation using a second trained neural network to generate a hyper latent representation; encoding the hyper latent representation using a third trained neural network to generate a hyper hyper latent representation; transmitting the latent representation, the hyper latent representation and the hyper hyper latent representation to a second computer system; decoding the hyper hyper latent representation using a fourth trained neural network; decoding the hyper hyper latent representation using an output of the fourth trained neural network and a fifth trained neural network; decoding the latent representation using an output of the fifth trained neural network and an output of a sixth trained neural network to generate an output image, the output image being an approximation of the input image.

[0139] The method may further include the step of determining a rate of the input image, and if the determined rate satisfies a predetermined condition, the steps of encoding the hyper-latent representation and decoding the hyper-latent representation are not performed.

[0140] If the determined rate satisfies a predetermined condition, the steps of encoding the hyper-flat representation and decoding the hyper-flat representation are not performed.

[0141] According to the present invention, there is provided a method of training one or more networks, the one or more networks being for use in encoding, transmitting, and decoding lossy images or videos, the method including the steps of: receiving an input image at a first computer system; encoding the input image using a first neural network to generate a latent representation; encoding the latent representation using a second neural network to generate a hyper latent representation; encoding the hyper latent representation using a third neural network to generate a hyper hyper latent representation; decoding the hyper hyper latent representation using a fourth neural network; decoding the hyper latent representation using an output of the fourth neural network and a fifth neural network; decoding the latent representation using an output of the fifth neural network and a sixth neural network to generate an output image, the output image being an approximation of the input image; determining a quantity based on a difference between the output image and the input image; updating parameters of the third and fourth neural networks based on the determined quantity; and repeating the above steps using a plurality of input images to generate third and fourth trained neural networks.

[0142] Parameters of the first, second, fifth and sixth neural networks may not be updated in at least one of the iterations of the steps.

[0143] The method may further include a step of determining a rate of the input image, and if the determined rate satisfies a predetermined condition, the parameters of the first, second, fifth and sixth neural networks are not updated in that iteration of the step.

[0144] The predetermined condition may be that the rate is less than a predetermined value.

[0145] The parameters of the first, second, fifth and sixth neural networks may not be updated after repeating the steps a predetermined number of times.

[0146] Parameters of the first, second, fifth and sixth neural networks may be additionally updated based on the determined quantities to generate first, second, fifth and sixth trained neural networks.

[0147] At least one of the input images may be subjected to at least one of the following operations before other steps are performed: upsampling, a smoothing filter, and a random crop.

[0148] According to the present invention, there is provided a method for image or video encoding and transmission, the method including the steps of receiving an input image at a first computer system, encoding the input image using a first trained neural network to generate a latent representation, encoding the latent representation using a second trained neural network to generate a hyper-latent representation, encoding the hyper-latent representation using a third trained neural network to generate a hyper-hyper-latent representation, and transmitting the latent representation, the hyper-latent representation, and the hyper-hyper-latent representation.

[0149] According to the present invention, there is provided a method for receiving and decoding a lost image or video, the method comprising the steps of receiving at a second computer system the latent representations, the hyper latent representations and the hyper hyper latent representations transmitted according to the above method, decoding the hyper hyper latent representations using a fourth trained neural network, decoding the hyper latent representations using outputs of the fourth trained neural network and the fifth trained neural network, and decoding the latent representations using outputs of the fifth trained neural network and the sixth trained neural network to generate an output image, wherein the output image is an approximation of the input image.

[0150] According to the present invention there is provided a data processing system arranged to carry out any of the above methods.

[0151] According to the present invention there is provided a data processing apparatus arranged to carry out any of the above methods.

[0152] According to the invention there is provided a computer program comprising instructions which, when said program is executed by a computer, cause the computer to carry out any of the methods described above.

[0153] According to the present invention there is provided a computer readable storage medium containing instructions which, when executed by a computer, cause the computer to perform any of the above methods. [Brief description of the drawings]

[0154] Aspects of the present invention will now be described by way of example with reference to the following figures. [Figure 1] 1 shows an example of an image or video compression, transmission and decompression pipeline. [Diagram 2] A further example of an image or video compression, transmission and decompression pipeline including a hypervisor network is shown. [Diagram 3] 1 shows a schematic diagram of an example encoding phase of an Al-based compression algorithm for video and image compression. [Figure 4] 1 shows a schematic diagram of an example decoding phase of an Al-based compression algorithm for video and image compression. [Diagram 5] We present an example distribution of quantization bin size values ​​that can be learned during training of an Al-based compression pipeline. [Figure 6] 13 shows a heatmap illustrating how the learned quantization bin sizes vary across latent channels for a given image. [Figure 7]1 shows a schematic diagram of an example encoding phase of an Al-based compression algorithm utilizing hyperpliers and learned quantization bin sizes. [Figure 8] 1 shows a schematic diagram of an example decoding phase of an Al-based compression algorithm utilizing a hyperplier and learned quantization bin sizes. [Figure 9] Some examples of the inverse gamma distribution are given below. [Figure 10] 1 shows an example overview of a GAN architecture. [Figure 11] We present an example of a standard generative adversarial compression pipeline. [Figure 12] We show the failure modes of the combined GAN and autoencoder architecture. [Figure 13] We present an example compression pipeline using multi-classifier cGAN training with dataset bias. [Figure 14] We show a comparison of reconstructions of the same generative model trained to the same bitrate with and without the multi-discriminative dataset biasing method. [Figure 15] 13 shows an example of the results of bitrate adjustment dataset bias. [Figure 16] An example of an Al-based compression pipeline is shown. [Figure 17] We present a further example of an Al-based compression pipeline. [Figure 18] 1 shows an example of a compression pipeline with quantization using quantization maps. [Figure 19] We present an example of the results of implementing the pipeline with a face detector used to identify regions of interest. [Figure 20] 1 shows different quantization functions Qm for regions identified by the ROI detection network H(x). [Figure 21] We give examples of three typical one-dimensional distributions that can be used in training an Al-based compression pipeline. [Figure 22] A comparison of piecewise linear and piecewise cubic Hermite interpolation is shown. [Figure 23]We present an example in which the autoregressive structure defined by the sparse context matrix L is utilized to parallelize components of a serial decoding path. [Figure 24] 1 illustrates an encoding process using a prediction context matrix Ly in an example of an Al-based compression pipeline. [Diagram 25] 1 illustrates a decoding process using a prediction context matrix Ly in an example of an Al-based compression pipeline. [Figure 26] 1 shows an example of a raster scan order for a single channel image. [Figure 27] It shows a 3 × 3 receptive field in which the next pixel is conditioned on the local variables instead of all the preceding variables. [Figure 28] Here is an example DAG that describes the joint distribution of four variables {y1,y2,y3,y4,}: [Figure 29] We present an example of a two-step AO in which all current variables at each step are conditionally independent and can be evaluated in parallel. [Diagram 30] The directed graph corresponding to AO in FIG. 29 is shown. [Diagram 31] We show an example where a 2 × 2 binary mask kernel subject to constraints (61) and (62) generates four subimages. [Diagram 32] 1 shows an example of an adjacency matrix A that determines the graph connectivity of an AO defined by the binary mask kernel framework. [Diagram 33] This shows the Adam7 interlacing index. [Diagram 34] An example of the visualization of the scale parameter σ where AO is defined as y1, y2, ... y16 is shown. [Diagram 35] An example of a permutation matrix representation is shown below. [Diagram 36] 1 illustrates an example of the concept of ranking tables applied to a binary mask kernel framework. [Figure 37] 1 shows an example of a hierarchical autoregressive ordering based on the wavelet transform. [Figure 38]Diagrammatic representation of various tensors and tensor products. [Figure 39] 3. Diagrammatic representation of an example of tensor decomposition and matrix product state. [Diagram 40] A diagram showing an example of a Locally Purified State. [Diagram 41] 1 is a diagram showing an example of a tensor tree. [Diagram 42] A diagram showing an example of a 2x2 Projected Entangled Pair State. [Diagram 43] 1 is a diagram illustrating an example of a procedure for converting a matrix product state to a canonical form. [Diagram 44] 1 shows an example of an image or video compression pipeline with a tensor network predicted by a hyperencoder / hyperdecoder. [Diagram 45] 1 shows a decompression pipeline for an image or video with a tensor network predicted by a hyperdecoder.

[0155] Compression processes can be applied to any form of information to reduce the amount of data required to store the information, i.e., the file size. Image and video information are examples of information that are compressed. The file size required to store the information is sometimes called the rate, specifically during the compression process referring to the compressed file. Generally, compression can be lossless or lossy. Both forms of compression result in a smaller file size. However, in lossless compression, no information is lost when the information is compressed and then subsequently decompressed. That is, during the decompression process, the original file that stored the information is completely reconstructed. In contrast, in lossy compression, information is lost during the compression and decompression process, and the reconstructed file may differ from the original file. Image and video files that contain image and video data are common targets for compression. JPEG, JPEG2000, AVC, HEVC, and AVI are examples of compression processes for image and video files.

[0156] In compression processes involving images, the input image is sometimes represented as x. Data representing the image may be stored in a tensor of dimensions H×W×C, where H represents the height of the image, W represents the width of the image, and C represents the number of channels of the image. Each H×W data point of the image represents the pixel value of the image at the corresponding location. Each channel C of the image represents a different component of the image at each pixel that is combined when the image file is displayed by a device. For example, an image file may have three channels, one for the red, one for the green, and one for the blue component of the image. In this case, the image information is stored in the RGB color space, which is sometimes referred to as a model or format. Other examples of color spaces and formats are the CMKY and YCbCr color models. However, the channels of an image file are not limited to storing color information, and other information may be represented in the channels. Since video is considered to be a continuous series of images, compression processes applied to images may also be applied to video. Each image that makes up a video may be called a frame of the video.

[0157] Frames of a video may be labeled according to the nature of the frame. For example, frames of a video may be labeled as I-frames and P-frames. An I-frame may be the first frame of a new section of a video. For example, the first frame after a scene transition may be labeled as an I-frame. A P-frame may be a subsequent frame after an I-frame. For example, the background or objects present in a P-frame may not change going from the I-frame to the P-frame. The changes in a P-frame compared to an I-frame going through a P-frame may be described by the movement of objects present in the frame or the perspective movement of the frame.

[0158] The output image may differ from the input image. The difference between the input image and the output image may be referred to as distortion or image quality difference. Distortion may be measured using any distortion function that receives an input image and an output image and provides an output that represents the difference between the input image and the output image in a numerical value. An example of such a method is using the mean squared error (MSE) between the pixels of the input image and the output image, but there are many other methods of measuring distortion, as known to those skilled in the art. The distortion function may include a trained neural network.

[0159] In general, the rate and distortion of a lossy compression process are related: an increase in rate results in a decrease in distortion, and a decrease in rate results in an increase in distortion. A change in distortion can affect the rate in a corresponding way. The relationship between these quantities for a given compression technique can be defined by the rate-distortion equation:

[0160] AI-based compression processes may involve the use of neural networks. A neural network is an operation that can be performed on an input to produce an output. A neural network may be built with multiple layers. The first layer of the network receives the input. The layer performs one or more operations on the input to produce the output of the first layer. The output of the first layer is then passed to the next layer of the network, which performs one or more operations in a similar manner. The output of the final layer is the output of the neural network.

[0161] Each layer of a neural network may be divided into nodes. Each node may receive at least a portion of the inputs from a previous layer and provide outputs to one or more nodes of a subsequent layer. Each node of a layer may perform one or more operations of the layer on at least a portion of the inputs to the layer. For example, a node may receive inputs from one or more nodes of a previous layer. The one or more operations include convolutions, weights, biases, and activation functions. Convolution operations are used in convolutional neural networks. If a convolution operation is present, the convolution may be performed across the inputs to the layer. Alternatively, the convolution may be performed across at least a portion of the inputs to the layer.

[0162] Each of the one or more operations may be defined by one or more parameters associated with each operation. For example, a weight operation may be defined by a weight matrix that defines the weights to be applied to each input from each node in the previous layer to each node in the current layer. In this example, each value in the weight matrix is ​​a parameter of the neural network. A convolution may be defined by a convolution matrix, also known as a kernel. In this example, one or more of the values ​​in the convolution matrix may be a parameter of the neural network. An activation function may also be defined by values ​​that may be parameters of the neural network. The parameters of the network may be varied during training of the network.

[0163] Other characteristics of the neural network may be predetermined and therefore do not change during the training of the network. For example, the number of layers of the network, the number of nodes of the network, one or more operations performed in each layer, and the connections between layers may be predetermined and therefore fixed before the training process occurs. These preset characteristics are sometimes called the hyperparameters of the network. These characteristics are sometimes called the architecture of the network.

[0164] To train a neural network, a training set of inputs for which the expected output (sometimes called ground truth) is known may be used. The initial parameters of the neural network are randomized and the first training inputs are provided to the network. The output of the network is compared to the expected output and, based on the difference between the output and the expected output, the parameters of the network are changed so that the difference between the network's output and the expected output is reduced. This process is repeated for multiple training inputs to train the network. The difference between the network's output and the expected output may be defined by a loss function. The result of the loss function may be calculated using the difference between the network's output and the expected output to determine the gradient of the loss function. Backpropagation of the gradient descent of the loss function may be used to update the parameters of the neural network using the gradient dL / dy of the loss function. Multiple neural networks in the system may be trained simultaneously by backpropagating the gradient of the loss function to each network.

[0165] For AI-based image or video compression, the loss function may be defined by a rate-distortion equation: Loss=D+λ*R, where D is the distortion function, λ is a weighting factor, and R is the rate loss. Lagrange multipliers serve as weights for a particular term in the loss function relative to each other term and can be used to control which terms in the loss function are prioritized when training the network.

[0166] For Al-based image or video compression, a training set of input images can be used. An example of a training set of input images is the KODAK image set (e.g., www.cs.albany.edu / xypan / research / snr / Kodak.html). An example of a training set of input images is the IMAX image set. An example of a training set of input images is the Imagenet dataset (e.g., www.image-net.org / download). An example of a training set of input images is the CLIC Training Dataset P ("professional") and M ("mobile") (e.g., http: / / challenge.compression.cc / tasks / ).

[0167] An example of an Al-based compression process 100 is shown in Figure 1. As a first step in the Al-based compression process, an input image 5 is provided. The input image 5 is encoded by a function f θ The input image is provided to a trained neural network 110 characterized by: . The encoder neural network 110 generates an output based on the input image. This output is called the latent representation of the input image 5. In a second step, the latent representation is quantized in a quantization process 140 characterized by an operation Q, resulting in a quantized latent representation. The quantization process converts the continuous latent representation into a discrete quantized latent representation. An example of a quantization process is a rounding function.

[0168] In a third step, the quantized latent quantities are entropy coded in an entropy coding process 150 to generate a bitstream 130. The entropy coding process may be, for example, range coding or arithmetic coding. In a fourth step, the bitstream 130 may be transmitted over a communication network.

[0169] In a fifth step, the bitstream is entropy decoded in an entropy decoding process 160. The quantized latent quantities are then decoded using a function gθ The quantized latent images are provided to another trained neural network 120, characterized by: The trained neural network 120 generates an output based on the quantized latent images. The output may be the output image of the Al-based compression process 100. The encoder-decoder system may be referred to as an autoencoder.

[0170] The system described above may be distributed across multiple locations and / or devices. For example, the encoder 110 may be located on a device such as a laptop computer, a desktop computer, a smartphone or a server. The decoder 120 may be located on another device, called a recipient device. The system used to encode, transmit and decode an input image 5 to obtain an output image 6 may be referred to as a compression pipeline.

[0171] The Al-based compression process may further include a HyperNetwork 105 for the transmission of meta-information that improves the compression process. The trained neural network 115 acting as JPEG2024528208000003.jpg78 and the hyperdecoder The system is constructed from a trained neural network 125 acting as a latent representation of the JPEG2024528208000004.jpg76. An example of such a system is shown in Figure 2. The building blocks of the system not described further may be assumed to be the same as those described above. A neural network 115 acting as a hyper-decoder receives the latents that are the output of the encoder 110. The hyper-encoder 115 generates an output based on the latent representation, sometimes called the hyper-latent representation. The hyper-latent is then computed as Q h The quantization process 145 is characterized by: h The quantization process 145, characterized by Q, may be the same as the quantization process 140, characterized by Q, described above.

[0172] In a similar manner as described above for the quantized latents, the quantized hyper-latencies are then entropy coded in an entropy coding process 155 to produce a bitstream 135. The bitstream 135 may be entropy decoded in an entropy decoding process 165 to retrieve the quantized hyper-latencies. The quantized hyper-latencies are then used as input to a trained neural network 125, which functions as a hyper-decoder. However, in contrast to the compression pipeline 100, the output of the hyper-decoder may not be an approximation of the input to the hyper-decoder 115. Instead, the output of the hyper-decoder is used to provide parameters for use in the entropy coding process 150 and the entropy decoding process 160 of the main compression process 100. For example, the output of the hyper-decoder 125 may include one or more of the mean, standard deviation, variance, or any other parameters used to describe the probability model of the entropy coding process 150 and the entropy decoding process 160 of the latent 10 representation. 2, for simplicity, only a single entropy decoding process 165 and hyperdecoder 125 are shown, but in practice the decompression process will usually be performed in a separate device, and so there will be duplicates of these processes on the device used for encoding to provide the parameters used by the entropy encoding process 150.

[0173] At any stage of the Al-based compression process 100, further transformations may be applied to the latents and / or hyper-latents. For example, the latents and / or hyper-latents may be converted to residual values ​​before the entropy encoding process 150, 155 is performed. The residual values ​​may be determined by subtracting the mean of the distribution of the latents or hyper-latents from each latent or hyper-latent. The residual values ​​20 may also be normalized.

[0174] To perform training of the Al-based compression process described above, a training set of input images may be used, as described above. During the training process, the parameters of both the encoder 110 and the decoder 120 may be updated simultaneously at each training step. If a hyper-network 105 is also present, the parameters of both the hyper-encoder 115 and the hyper-decoder 125 are additionally updated simultaneously at each training step.

[0175] The training process can further include a generative adversarial network (GAN). When applied to an Al-based compression process, in addition to the compression pipeline described above, the system includes an additional neural network that acts as a discriminator. The discriminator receives an input and outputs a score based on the input to provide an indication of whether the discriminator considers the input to be ground truth or fake. For example, this indication is a score, where a high score is associated with a true input and a low score is associated with a fake input. A loss function that maximizes the difference in the output representation between the input ground truth and the input fake is used to train the discriminator.

[0176] When a GAN is incorporated into the training of the compression process, the output image 6 may be provided to a classifier. The output of the classifier may be used in the loss function of the compression process as a measure of the distortion of the compression process. Alternatively, the classifier may receive both the input image 5 and the output image 6 and use the difference in the output representation in the loss function of the compression process as a measure of the distortion of the compression process. The training of the neural network acting as the classifier and the other neural networks in the compression process may be performed simultaneously. During use of the trained compression pipeline for image or video compression and transmission, the classifier neural network is removed from the system and the output of the compression pipeline is the output image 6.

[0177] Incorporating a GAN into the training process may result in the decoder 120 performing hallucinations. A hallucination process is a process that adds information to the output image 6 that was not present in the input image 5. In one example, fine details may be added to the output image 6 that were not present in the input image 5 or received by the decoder 120. The hallucinations performed may be based on quantized latent information received by the decoder 120.

[0178] As described above, a video is constructed from a sequence of images. The Al-based compression process 100 described above may be applied multiple times to compress, transmit, and decompress the video. For example, each frame of the video may be compressed, transmitted, and decompressed individually. The received frames may then be grouped together to obtain the original video.

[0179] We now describe a number of concepts related to the Al compression process described above. Although each concept is described separately, one or more of the following concepts may be applied in an Al-based compression process as described above.

[0180] Learned quantization Quantization is a critical step in Al-based compression pipelines. Typically, quantization is achieved by rounding data to the nearest integer because some regions of images and videos can tolerate higher information loss, while others require finer-grained detail. In what follows, we describe how the size of the quantization bins can be learned, rather than fixed to rounding to the nearest integer. We detail several architectures to achieve this, including predicting the size of the bins from a hypernetwork, a context module, and an additional neural network. We also document the necessary modifications to the loss function and quantization procedure required to train an Al-based compression pipeline with learned quantization bin sizes, and show how to introduce a Bayesian prior to control the distribution of bin sizes learned during training. We show that the learned quantization bins can be used with or without split quantization. This innovation also allows the distortion gradient to flow through the decoder to the hypernetwork. Finally, we detail a generalized quantization function that improves performance and runtime. Specifically, this innovation allows us to include a context model in the decoder of a compression pipeline, but without the run-time penalty of repeatedly running an arithmetic (or other lossless) decoding algorithm.Our technique for learning the quantization bins is compatible with any method that carries meta-information, including hyperpliers, autoregressive models, and implicit models.

[0181] The following discussion provides an overview of the functionality, scope, and future prospects of learned quantization bins and generalized quantization functions for use in, but not limited to, Al-based image and video compression.

[0182] Compression algorithms can be divided into two phases: encoding and decoding. In the encoding phase, the input data is transformed into latent variables that have a smaller representation (in bits) than the original input variables. In the decoding phase, an inverse transformation is applied to the latent variables to recover the original data (or an approximation of the original data).

[0183] An Al-based compression system must also be trained, which is the procedure of selecting parameters for the Al-based compression system that achieve good compression results (small file size and minimal distortion). During training, parts of the encoding and decoding algorithms are run to determine how to adjust the parameters of the Al-based compression system.

[0184] More precisely, for Al-based compression, the encoding generally takes the form

number

[0185] where x is the data to be compressed (image or video) and f enc is an encoder, typically a neural network with trained parameters θ. The encoder converts the input data x into a latent representation y with lower dimensionality and in a form that is improved for further compression.

[0186] To further compress y and transmit it as a stream of bits, established lossless coding algorithms such as arithmetic coding can be used. Such lossless coding algorithms may require y to be discrete rather than continuous, and may also require knowledge of the probability distribution of the latent representation. To achieve this, a quantification function Q (usually rounded to nearest) is used that transforms the continuous data into a discrete value y (hat).

[0187] The required probability distribution p(y(hat)) is found by fitting a probability distribution to the latent space. The probability distribution can be learned directly, but is often a parametric distribution with parameters determined by a hypernetwork consisting of a hyperencoder and a hyperdecoder. When using a hypernetwork, an additional bitstream z(hat) (also called "side information") is sometimes encoded, transmitted, and decoded:

number

[0188] The encoding process (using a hypernetwork) is shown in Figure 3. Figure 3 shows a schematic diagram of an example of the encoding phase of an Al-based compression algorithm for video and image compression.

[0189] The decryption is done as follows:

number

[0190] Summary: In an arithmetic decoder (or other lossless decoding algorithm), the distribution of latencies p(y(hat)) is used to convert the bitstream to quantized latencies y(hat). Then, the function f dec transforms the quantized latent data into a lossy reconstruction of the input data, denoted by x. In Al-based compression, f dec is usually a neural network depending on the learned parameters θ.

[0191] When using a HyperNetwork, the side information bitstream is first decoded and then used to obtain the parameters needed to construct p(y(hat)) needed to decode the main bitstream. An example of the decoding process (using a HyperNetwork) is shown in Figure 4. Figure 4 shows an example of the decoding phase of an Al-based compression algorithm for video and image compression.

[0192] AI-based compression relies on learning the parameters of the encoding and decoding neural networks using typical optimization techniques with a "loss function". The loss function is chosen to balance the goals of compressing the image or video to a small file size while maximizing the reconstruction quality. Thus, the loss function consists of two terms:

number

[0193] where R determines the cost of encoding the quantified latent according to the distribution p(y(hat)), D measures the quality of the reconstructed image, and λ is a parameter that determines the tradeoff between low file size and reconstruction quality. A typical choice for R is the cross-entropy.

number

[0194] The choice of JPEG2024528208000010.jpg712 is due to quantization, and the latent is rounded to the nearest integer, so that the probability distribution of p(y(hat)) is given by the integral of the (unquantized) latent distribution p(y) from y(hat)-1 / 2 to y(hat)+1 / 2, which is given by the cumulative distribution function p(y(hat)).

[0195] The function D can be chosen as the mean squared error, but can also be a combination of other metrics of perceptual quality such as MS-SSIM, LPIPS, and / or adversarial loss (if using an adversarial neural network to enforce image quality).

[0196] When using a hypernetwork, an additional term can be added to R to represent the cost of transmitting the additional side information:

number

[0197] In summary, note that the loss function depends explicitly on the choice of quantization scheme through the R term, and implicitly, since y(hat) depends on the choice of quantization scheme.

[0198] We now discuss how the learned quantization bins can be used in Al-based image and video compression. The steps described are as follows: The architecture required to learn and predict the size of the quantization bins Modification of standard quantization functions and encoding / decoding processes to incorporate learned quantization bins How to train a neural network using learned quantization bins

[0199] A key step in a typical Al-based image and video compression pipeline is "quantization", where pixels of the latent representation are typically rounded to the nearest integer. This is necessary for algorithms to losslessly encode the bitstream. However, the quantization step introduces its own information loss, which impacts the reconstruction quality.

[0200] It is possible to improve the quantization function by training a neural network to predict the size of the quantization bin that should be used for each latent pixel. Typically, the latent y is rounded to the nearest integer, which corresponds to a "bin size" of 1. That is, all possible values ​​of y in an interval of length 1 are mapped to the same y (hat).

number

[0201] However, this may not be the optimal choice of information loss: for some latent pixels, more information can be ignored without affecting the reconstruction quality too much (equivalently: use bins larger than 1), and for other latent pixels, the optimal bin size is smaller than 1.

[0202] We can solve this problem by predicting the quantization bin size for each pixel in the image. We call this the tensor Run it on JPEG2024528208000013.jpg636 and modify the quantization function as follows:

number

[0203] This is called the "quantized latent residual" Let's call it JPEG2024528208000015.jpg87. Therefore, equation 7 becomes:

number

[0204] Figure 6 is a heatmap showing how the learned quantification bin size varies for a given image across latent channels: Different pixels are predicted to benefit from larger or smaller quantization bins, corresponding to larger or smaller information loss.

[0205] Note that the learned quantization bin size is incorporated into the modification of the quantization function Q, so that any data we want to encode and transmit can take advantage of the learned quantization bin size. For example, instead of encoding the latent y, we can encode the mean-subtracted latent y-μ y If you want to encode

number

[0206] Similarly, hyperlatent, hyperhyperlatent, and other objects that we wish to quantize are all quantized using a modified quantization function Q for a properly trained Δ.Δ can be used.

[0207] Here we discuss several architectures for predicting the size of the quantization bins. A possible architecture is to use a hyper-network to predict the quantization bin size Δ. The bitstream is encoded as follows:

number

[0208] FIG. 7 is an example of a modified encoding process using a hyper-network. It is a schematic diagram showing an example of the encoding phase of an Al-based compression algorithm using hyper-pliers and learned quantization bin sizes for video and image compression.

[0209] At decoding time, the bitstream is losslessly decoded as usual. It is then rescaled by multiplying the two elements with Δ. y The result of this transformation is called y (hat) and is passed to the decoder network as usual.

number

[0210] An example of a modified decoding process using a hyper-network is shown in Figure 8. Figure 8 is a schematic diagram showing an example of the decoding phase of an Al-based compression algorithm utilizing a hyper-plier and a learned quantization bin size for video and image compression.

[0211] Applying the above techniques can improve the rate of the Al-based compression pipeline by 1.5% and the distortion, as measured by MSE, by 1.9%, thus improving the performance of the Al-based compression process.

[0212] We will now elaborate on some variations of the above architecture. The size of the quantization bins of the hyperlatencies can be a learned parameter, and we can include a hyperlatency network that predicts these variables. · The prediction of the quantization bin size can be augmented with so-called “context module” features that use neighboring pixels to improve the prediction of a particular pixel. Hyperdecoder to quantization bin size Δ y After obtaining , this tensor can be further processed with a nonlinear function, which is typically (but is not limited to) a neural network.

[0213] We also emphasize that our method for learning the quantization bins is compatible with all methods that convey meta-information, including hyper-pliers, hyper-hyper-pliers, autoregressive models, and implicit models.

[0214] To train a neural network with the learned quantization bins for Al-based compression, we can modify the loss function. Specifically, the cost of encoding the data described in Equation 5a can be modified as follows:

number

[0215] Instead of integrating over an interval of length 1, From JPEG2024528208000023.jpg616 This means that it is necessary to integrate the probability distribution of the latent material up to JPEG2024528208000024.jpg516.

[0216] Similarly, when using a hyperlatinum network, the term JPEG2024528208000025.jpg537 is modified exactly as the latin coding cost is to incorporate the learned quantization bin size.

[0217] Neural networks are usually trained by a variant of gradient descent that uses backpropagation to update the parameters of the trained network. This requires that the gradients of all layers in the network are calculated, which requires that the layers of the network are built with differentiable functions. However, the quantization function Q and its training bin modification Q Δ is not differentiable due to the presence of the rounding function. In Al-based compression, one of two differentiable approximations to the quantization is used during network training to replace Q(y) (once the network is trained and used for inference, no approximation is used).

number

[0218] When using learned quantization bins, the approximation of the training quantization is:

number

[0219] Instead of choosing one of these differentiable approximations during training, the Al-based compression pipeline uses Use JPEG2024528208000028.jpg613, but use the decoder for training. It can also be trained using "split quantization" by sending JPEG2024528208000029.jpg612. Al-based compression networks can be trained with both split quantization and learned quantization bins.

[0220] First, note that there are two types of split quantization. Hard split quantization: The decoder Receive JPEG2024528208000030.jpg612 Soft split quantization: The decoder I receive JPEG2024528208000031.jpg612, but When computing JPEG2024528208000032.jpg77, the backward pass of backpropagation Use JPEG2024528208000033.jpg613.

[0221] Note that for integer rounding quantization, hard split and soft split quantization are equivalent.

number

[0222] However, when using learned quantization bins, hard split and soft split quantization are not equivalent, since in the backward pass we get:

number

[0223] In all quantization schemes, the rate gradient wrtΔ is negative.

number

[0224] For any quantization scheme, the rate term is Since we receive JPEG2024528208000037.jpg613, JPEG2024528208000038.jpg712. Then, always JPEG2024528208000039.jpg518, so we get

number

[0225] The distortion gradient wrtΔ depends on the quantization scheme. The gradient is:

number

[0226] In soft split quantization, JPEG2024528208000042.jpg812 and This means that the random noise ε in the backward pass is This is because it does not rely on STE rounding, which does JPEG2024528208000044.jpg66. This means that:

number

[0227] Thus, in soft split quantization, the rate gradient makes Δ large, while the distortion gradient is on average zero, so that overall Δ→∞, and the network cannot be trained.

[0228] On the other hand, in hard division quantization JPEG2024528208000046.jpg66 is Since it does not depend on JPEG2024528208000047.jpg612, the result will be as follows.

number

[0229] In summary, when using split quantization with learned quantization bins, use hard split quantization instead of soft split quantization.

[0230] The nontrivial distortion gradients achievable with or without split quantization imply that the distortion gradients flow through the decoder to the hypernetwork, a feature that is typically not possible in models with hypernetworks, but which is introduced by our method of learning the size of the quantization bins.

[0231] In some compression pipelines, it is important (though not always necessary) to control the distribution of values ​​learned for the quantization bin size. If necessary, this is achieved by introducing an additional term in the loss function.

number

[0232] F Δ is the distribution p Δ It is characterized by a choice of (Δ), which, following the terminology used in Bayesian statistics, is called a "prior" with respect to Δ. There are several options for the prior. Any parametric probability distribution over positive numbers. Specifically, the inverse gamma distribution, some examples of which are shown in Figure 9. The distributions are shown for several values ​​of the parameters. The parameters of the distribution can be adapted from other models, chosen a priori, or learned during training. Neural networks that learn priors during training.

[0233] In the previous section, we provided a detailed description of a simplified quantization function Q that utilizes a tensor of bin size Δ.

number

[0234] All these methods can be extended to more generalized quantization functions. In the general case, Q is some invertible function of y and Δ. The encoding is then given by

number

number

[0235] This quantization function is more flexible than Equation 24, resulting in improved performance. The generalized quantization function can also be made context-aware, for example, by incorporating quantization parameters that use an autoregressive context model.

[0236] All methods from the previous sections are compatible with the generalized quantization function framework: · Estimate the necessary parameters from the hypernetwork, the context module and, if necessary, implicit equations. · Modify the loss function appropriately. When using split quantization, use hard split quantization. Optionally, introduce Bayesian priors into the loss function to control its behavior.

[0237] A more flexible generalized quantization function improves performance. In addition to this, the generalized quantization function can depend on autoregressively determined parameters, which means that the quantization depends on the pixels already encoded / decoded.

number

number

[0238] In general, the use of autoregressive context models improves the performance of Al-based compression.

[0239] The autoregressive generalized quantization function is also beneficial from a runtime perspective. Other standard autoregressive models, such as PixelCNN, need to run an arithmetic decoder (or other lossless decoder) every time a pixel is decoded using a context model. This has a severe performance hit in real-world applications of image and video compression. However, the generalized quantization function framework allows us to incorporate autoregressive context models into Al-based compression without the runtime issues that PixelCNN and others have. This is because the Q -1 This is because the arithmetic decoder operates autorecursively on ξ, which has been completely decoded from the bitstream. Therefore, there is no need to run the arithmetic decoder autorecursively, and the execution time problem is solved.

[0240] The generalized quantization function can be any invertible function. For example, · Invertible rational function: Q(·,Δ)=P(·,Δ) / Q(·,Δ), where P, Q are polynomials. Logarithmic, exponential, and trigonometric functions with appropriate domain restrictions so that these functions are invertible. ·Invertible function of matrix Q(·,Δ) · A reversible function containing the context parameters L predicted by the hyperdecoder: Q(·,Δ,L)

[0241] Moreover, in general Q does not need to be closed-form or invertible. For example, Q enc (·,Δ) and Q dec (·,Δ), where these functions do not necessarily have to be inverses of each other, and the entire pipeline can be trained end-to-end. In this case, Q enc and Q dec can be modeled as a neural network, or as a specific process such as a Gaussian process, or a probabilistic graphical model (a simple example: the hidden Markov model).

[0242] To train an Al-based compression pipeline that uses a generalized quantization function, we use many of the same tools described above. · Modify the loss function appropriately (specifically the rate term). When using split quantization, use hard split quantization. Introduce Bayesian priors and / or regularization terms (l1, l2, second moment penalties) into the loss function, optionally controlling the distribution parameter Δ and, if necessary, the distribution of the context parameter L.

[0243] Depending on the choice of generalized quantization function, other tools may be needed to train an Al-based compression pipeline. Techniques from Reinforcement Learning (Q-Learning, Monte Carlo Estimation, DQN, PPO, SAC, DDPG, TD3) General Proximal Gradient Method Continuous relaxation. Here, the discrete quantification residual is approximated by a continuous function. This function can have hyperparameters that control the smoothness of the function, and is modified at different points in the training to improve the training and final performance of the network.

[0244] There are several possibilities for context modeling that are compatible with the generalized quantization function framework. · Q is an autoregressive neural network. The most common example that can be adapted for Q is a PixelCNN style network, but other neural network building blocks such as Resnets, Transformers, Recurrent Neural Networks, and Fully-Connected networks can also be used for autoregressive Q functions. Δ can be predicted as a linear combination of previously decoded pixels.

number

[0245] If we obtain important meta-information, for example from attention mechanisms / focus masks, we can incorporate this into the delta prediction. In this case, the quantization bin size adapts more precisely to the sensitive regions of the image and video, and the knowledge of the sensitive regions is preserved in this meta-information. In this way, less information is lost from perceptually important regions, while information from unimportant regions is ignored, improving performance in a more enhanced way compared to Al-based compression pipelines that do not have adaptive bin sizes.

[0246] We further outline the connection between learned quantization bins and variable-rate models. One form of variable-rate model trains an Al-based compression pipeline with a free hyper-parameter δ that controls the bin size. At inference, δ is transmitted as meta-information to control the transmission rate (cost in bits).

[0247] In the variable-rate framework, δ is a global parameter in the sense that it controls the size of all bins simultaneously. In our innovation, we locally obtain a tensor Δ of predicted bin sizes for each pixel. In addition, variable-rate models that use δ to control the transmission rate are compatible with our framework, since the local prediction can be element-wise scaled by the global prediction as needed to control the rate during interpolation.

number

[0248] Dataset Bias In this section, we detail the training procedure applied in our generative adversarial network framework. This approach allows us to bias our generative compression models for any kind of image data, and control the quality of the resulting images depending on the objects in the images.

[0249] General Adversarial Networks (GANs) have shown good results when applied to a variety of different generative tasks in image, video and audio domains. The approach is inspired by game theory, where two models, a generator and a critic, are pitted against each other, resulting in both becoming stronger. The first model in a GAN is a generator G that takes as input a noise variable z and outputs synthetic data samples x (hat), and the second model is a discriminator D that is trained to distinguish between samples from the real data distribution and data generated by the generator. An example of a high-level architecture of a GAN is shown in Figure 10.

[0250] P x Let P be the data distribution on the real sample x. z Let P be the data distribution on the noise sample z. g Let be the generator distribution on the data x.

[0251] Training a GAN can be depicted as a minimax game in which the following function is optimized:

number

[0252] We apply a generative adversarial approach to the image compression task to Start by considering JPEG2024528208000058.jpg519, where C is the number of channels, and H and W are the height and width in pixels.

[0253] The autoencoder-based compression encoder pipeline is constructed from the following, with the encoder function f θ (x)=y is the latent representation of image x JPEG2024528208000059.jpg721. Q is the quantization function required to transmit y as a bitstream, and the decoder function JPEG2024528208000060.jpg616 is the reconstructed image of the quantized latent y (hat) This is a function that decodes to JPEG2024528208000061.jpg519.

number

[0254] In this case, the encoder f θ , the quantization function Q and the decoder g θ The combination of x can be thought of as a generative network. For notational simplicity, we denote this generative network as G(x). This generative network is complemented by a discriminative network D, which we train together with the generative network in a two-stage approach.

number

[0255] An example of a standard generative adversarial compression pipeline is shown in Figure 11. We then train using a standard compression rate-distortion loss function.

number

[0256] This trained compression network can be complemented with a classifier model to improve the perceptual quality of the output image. In this case, the compression encoder-decoder network can be considered as a generative network, and the two models can be trained using a two-level approach at each iteration. For the classifier architecture, we choose to use a conditional classifier, which produces higher quality reconstructed images. The classifier d(x,y(hat)) in this case is conditioned on the quantized latent y(hat). First, we train the classifier with the classifier loss.

number

[0257] To train the generative network in (32), we augment the rate-distortion loss in (36) by adding an adversarial “de-saturation” loss used to train the generators of GANs.

number

[0258] Adding an adversarial loss to the rate-distortion loss encourages the network to generate natural patterns and textures. Using a combined GAN-autoencoder architecture, we have been able to significantly improve the perceptual quality of the reconstructed images and achieve superior results in image compression. However, despite the impressive overall results of such architectures, there are a number of notable failure modes. It has been observed that these models struggle to compress regions of high visual importance, including but not limited to human faces and text. An example of such a failure mode is shown in Figure 12. The image on the left is the original image, and the image on the right is the reconstructed image synthesized using a generative compression network. Note that the most distortion occurs in the human faces present in the image, i.e., areas with a lot of visual information. To address this issue, we propose a method that can bias the model towards certain types of images, such as faces, to improve the perceptual quality of the reconstructed image.

[0259] In this framework, we train a network on multiple datasets, each with a different classifier. First, we train a network on N additional datasets X1,…X N A good example of such a dataset that would be useful for face modeling would be a dataset consisting of portraits of people. For each dataset X i For the discrimination model D i Each classifier model D i is the dataset X i The encoder-decoder model is trained on the images of all the datasets, while the encoder-decoder model is trained on the images of only the dataset.

number

[0260] Figure 13 shows a diagram of the compression pipeline with multiple classifier cGAN training using dataset bias. i Image of x i is passed through the encoder, digitized, and converted to a bitstream using a range coder. The decoder then decodes the bitstream, yielding x i Then, the original image x and the reconstruction x (hat) are passed to the classifier corresponding to the dataset i.

[0261] For illustrative purposes, we focus on the face failure mode, as shown in Figure 12. Note that all of the techniques described here can be applied to any number of regions of interest in an image (e.g., text, animal faces, eyes, lips, logos, cars, flowers, patterns, etc.).

[0262] As an example, consider biasing a dataset using only one additional dataset. In this case, X1- is a general training dataset and X2- is a dataset containing only portrait images. A comparison of the reconstruction of the same generative model trained to the same bitrate with and without the multi-identifier dataset biasing scheme is shown in Figure 14. The image on the left is a reconstruction of the image synthesized by a standard generative compression network, and the image on the right is a reconstruction of the same image synthesized by the same generative network trained using the bias of the multi-identifier dataset. Using this scheme improves the perceived quality of the human face without compromising the quality of the rest of the image.

[0263] [Table 1]

[0264] [Table 2]

[0265] The above approach can also be used in architectures where a single classifier is used for all datasets. i can be trained more frequently than the generator, amplifying the effect of bias on that dataset.

[0266] Given a generative compression network as above, we now define an architectural modification that allows for higher bit allocation, conditional on the image or frame context. To enhance the effect of dataset bias and vary the perceptual quality of different regions of an image depending on the image subject, we define the Lagrangian coefficients that control the bitrate as a function of the dataset X. i We propose a different training procedure for each model. The generator modifies the loss function in (37) as follows:

number

[0267] This approach trains the model to allocate a higher percentage of the bitstream to face regions of the compressed image. The results of the bias on the bitrate-adjusted dataset can be observed in Figure 15. Figure 15 shows the same results for the face dataset. While keeping JPEG2024528208000071.jpg59, we use a different set of background datasets. JPEG2024528208000072.jpg59 shows three images synthesized by a model trained on the background dataset. The left image was trained at a low bitrate, the middle image at a medium bitrate, and the right image at a high bitrate.

[0268] We extend the method proposed above and propose to use a different distortion function d(x, x(hat)) for each dataset used for bias. This method allows us to focus the model on each specific type of data. For example, we can use a linear combination of the MSE, LPIPS and MS-SSIM metrics as the distortion function.

number

[0269] Coefficients of the different components of the distortion function By modifying JPEG2024528208000074.jpg638, the perceptual quality of the resulting image can be altered, just as the generative compression model can reconstruct different regions of the image differently. Then, Equation 37 is i This can be modified by indexing the distortion function d(x,x(hat)) for

number

[0270] Here we describe the use of salience masks to bias a generative compression model towards particularly important regions of an image. The masks are generated using a separate pre-trained network, whose output can be used to further improve the performance of the compression model.

[0271] Given an image x, we use a binary mask Consider a network H that outputs JPEG2024528208000076.jpg517 Start with:

number

[0272] X i Salient pixels in m are denoted by 1s in m, while 0s indicate regions the network does not need to focus on. This binary mask can be used to further bias the compression network towards those regions. Examples of such important regions include, but are not limited to, human facial features such as eyes and lips. Given m, the input image x can be modified to prioritize these regions. Modified image x H is used as the input of the adversarial compression network. An example of such a compression pipeline is shown in Figure 16.

[0273] Extending the approach proposed above, we propose an architecture that utilizes a pre-trained network that generates a salience mask to bias the compression pipeline. With this approach, the bitrate allocation to different parts of the reconstructed image can be altered by modifying the mask, without retraining the compression network. In this variant, we use the mask m from Equation 41 as an additional input to train the network to allocate more bits to regions marked as salient(1) in m. During the inference phase after the network is trained, the bit allocation can be adjusted by modifying the mask m. An example of such a compression pipeline is shown in Figure 17.

[0274] We further propose a training scheme that ensures that the model is exposed to a wide range of natural image examples. A training dataset is constructed from images from N distinct classes, and each image is labeled accordingly. During training, images are sampled from the dataset according to their class. By sampling images evenly from each class, the model is exposed to under-represented classes and is able to learn the full distribution of natural images.

[0275] Area emphasis State-of-the-art methods of learned image compression, such as architectures based on VAE and GAN, allow for good compression at small bitrates and a significant improvement in the perceptual quality of the reconstructed images. However, despite the impressive overall results of such architectures, there are a number of notable failure modes, as discussed above. It has been observed that these models struggle to compress regions of high visual importance, including but not limited to human faces or text. Examples of such failure modes are shown in Figure 12. The image on the left shows the original image, while the image on the right shows the reconstructed image synthesized using a generative compression network. We note that the most distortion occurs in the human faces present in the image, i.e. the areas of high visual importance.

[0276] We propose an approach to improve the perceptual quality of a region of interest (ROI) by allocating more bits to the bitstream by modifying the quantization bins within the ROI.

[0277] To encode the latent y into a bitstream, we can first quantize it to ensure that it is discrete. We propose to control the bpp allocated to a region with a quantization parameter Δ. Δ is the size of the quantization bin or quantization interval, which represents the coarseness of the quantization in the latent and hyper-latent spaces. The coarser the quantization, the fewer the number of bits allocated to the data.

[0278] Quantization of the latent y is achieved as follows:

number

[0279] We propose to exploit spatially varying deltas to control the coarseness of quantization within an image, which allows us to control the number of bits allocated and therefore the visual quality of different regions of the image. This proceeds as follows:

[0280] We start by considering a function H for detecting regions of interest, usually represented in neural networks. The function H(s) takes as input an image x and a binary mask It outputs JPEG2024528208000079.jpg517. A 1 in m indicates that the corresponding pixel in image x is inside the region of interest, and a zero corresponds to a pixel that is outside it.

[0281] In one example, the network H(x) is trained prior to the training of the compression pipeline, and in another example, it is trained in combination with the encoder-decoder. The map m is used to create a quantization map Δ where each pixel is assigned a quantization parameter. If a pixel has a value in m of 1, the corresponding value in Δ is small. The function Q defined in Equation 42 then quantizes y to y(hat) using the spatial map Δ before encoding it into the bitstream. Such a quantization scheme results in a higher bitrate for the region of interest compared to the rest of the image.

[0282] The proposed pipeline is shown in Fig. 18. Fig. 18 shows an illustration of the proposed compression pipeline with quantization utilizing a quantization map Δ that depends on a binary mask of ROI m. The results of implementing such a pipeline with a face detector, network H(x), used to identify regions of interest, are shown in Fig. 19. The image on the left is a compressed image synthesized with a face assigned as an ROI in the quantization map Δ, and the image on the right is a compressed image with a standard generative compression model.

[0283] In another example, a different quantization function Q m An example of such an arrangement is shown in FIG. 20. FIG. 20 shows the quantization function Q for the whole image and the quantization function Q for the region of interest. m 1 shows a diagram of a proposed compression pipeline using

[0284] Discrete PMFS Encoding and decoding a stream of discrete symbols (such as latent pixels in an Al-based compression pipeline) into a binary bitstream may require access to a discrete probability mass function (PMF). However, it is widely believed that training such a discrete PMF in an Al-based compression pipeline is not possible because training requires access to a continuous probability distribution function (PDF). Therefore, the de facto standard for training an Al-based compression pipeline is to train a continuous PDF and only after training is complete, approximate the continuous PDF with a discrete PMF evaluated at a discrete number of quantization points.

[0285] We describe below how we reverse this procedure, allowing Al-based compression pipelines to be trained directly on discrete PMFs by interpolating them into a continuous real-valued space. Discrete PMFs can be learned or predicted, and can also be parameterized.

[0286] The following discussion outlines the capabilities, scope, and future prospects of discrete probability mass functions and their interpolation for use in, but not limited to, Al-based image and video compression. Below, we provide a high-level description of discrete probability mass functions, inference and training for Al-based compression algorithms, and methods for interpolating functions (such as discrete probability mass functions).

[0287] In the Al-based compression literature, the standard approach to constructing an entropy model is to use the continuous probability density function (PDF) p yThe key is to start with a distribution of y (e.g., Laplace or Gaussian). The Shannon entropy is calculated by dividing the discrete variable y (usually JPEG2024528208000080.jpg69), we can use the PDF as a discrete probability mass function (PMF) JPEG2024528208000081.jpg610 and use it for example in a lossless arithmetic encoder / decoder. This can be done by collecting all (continuous) masses in (unit) bins centered at y (hat).

number

[0288] This approach was first proposed by Johannes Balle, Valero Laparra, and Eero P. Simoncelli. End-to-end optimized image compression. 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings, OpenReview.net, 2017, incorporated herein by reference. The function is an integer It is defined for JPEG2024528208000083.jpg69, but also accepts any real-valued argument. This is very useful during training, when the PDF for the continuous real-valued latent values ​​output by the encoder is needed. Therefore, the "new" function JPEG2024528208000084.jpg523 is defined. This is by definition a PMF defined over integers. This function matches JPEG2024528208000085.jpg55 exactly. JPEG2024528208000086.jpg56 is the PMF on which the end-to-end Al-based compression algorithm is actually trained.

[0289] In summary, a continuous real-valued PDF JPEG2024528208000087.jpg617 is a discrete PMF Converted to JPEG2024528208000088.jpg517 and then trained JPEG2024528208000089.jpg513 to Continuous PDF Figure 21 shows examples of three typical one-dimensional distributions used for training the Al-based compression pipeline. In fact, the original PDF p y is not explicitly used during training or inference. JPEG2024528208000091.jpg56 (used in inference) and There are only two functions: JPEG2024528208000092.jpg56 (used in training).

[0290] This model of thought can also be reversed: instead of starting with a PDF, it is possible to start with discrete PMFs and recover a continuous PDF (which is only used in training) by interpolating the PMFs.

[0291] PMF Given a JPEG2024528208000093.jpg56, this PMF is a function of two vectors of length N, namely, y i (hat) and p i (hat), where i=1...,N represent discrete points. In the old way (where the PMF is defined by a function), JPEG2024528208000094.jpg517. However, in general JPEG2024528208000095.jpg58 can be any non-negative vector whose sum is 1. Vector y i are sorted in ascending order and are not necessarily integer values.

[0292] Here, the query point Given JPEG2024528208000096.jpg619, the query points must be surrounded by the discrete extrema. (Approximate) Training PDF To define JPEG2024528208000097.jpg58, we use the interpolation routine

number

[0293] A non-exhaustive list of possible interpolation routines is: Piecewise constant interpolation (nearest neighbor interpolation) Linear Interpolation Polynomial Interpolation Spline interpolation (such as piecewise cubic interpolation) ·Gaussian Process / Kriging

[0294] In general, functions defined by interpolation are not necessarily exactly PDFs. Depending on the interpolation routine used, the interpolated values ​​may be negative or have no unit mass. However, these problems can be mitigated by choosing an appropriate routine. For example, piecewise linear interpolation preserves mass, preserves positivity, and ensures that the interpolated function is indeed a PDF. Piecewise cubic Hermite interpolation can be constrained to be positive if the interpolated points themselves are positive, as discussed in Randall L Dougherty, Alan S Edelman, and James M Hyman. Cubic and quintic Hermite interpolation that preserve non-negativity, monotonicity, or convexity. Mathematics of Computation, 52(186):471-494, 1989, which is incorporated herein by reference.

[0295] However, piecewise linear interpolation has another problem: its derivatives are piecewise constant, and the interpolation error can be quite severe, as shown, for example, in the left panel of Figure 22. For example, if the PMF is generated by Balle's method, the interpolation error is JPEG2024528208000099.jpg523. Other interpolation schemes, such as piecewise cubic Hermite interpolation, have smaller interpolation errors, as shown in the right diagram of FIG.

[0296] The discrete PMF can be trained directly on AI-based image and compression algorithms. During training, the real-valued latent probability values ​​output by the encoder are interpolated using the discrete values ​​of the PMF model. During training, the PMF is learned by flowing the gradient back from the rate (bitstream size) loss to the parameters of the PMF model.

[0297] PMF models can be trained or predicted. Trained means that the PMF model and its hyperparameters do not depend on the input image. Predicted means that the PMF model conditionally depends on "side information" such as hyperlatencies. In this scenario, the parameters of the PMF may be predicted by the hyperdecoder. Additionally, the PMF can also conditionally depend on neighboring latent pixels (in which case the PMF is said to be a discrete PMF context model). Regardless of how the PMF is represented, the values ​​of the PMF can be interpolated during training to provide estimates of the probability values ​​of real-valued (unquantized) points, which can then be input into the rate loss of the training objective function.

[0298] The PMF model can be parameterized in one of the following ways (although this list is not exhaustive): · The PMF is a categorical distribution, and the probability values ​​of the categorical distribution correspond to a finite number of quantized points on the solid line. The probability values ​​of a categorical distribution correspond to a finite number of quantized points on a real line. A categorical distribution can be parameterized by a vector whose values ​​are then projected onto the probability simplex. This projection can be a softmax-style projection or any other projection onto the probability simplex. ·The PMF can be parameterized with several parameters. For example, if the PMF is defined for N points, n parameters (n < N) can be used to control the values of the PMF. For example, the PMF can be controlled with mean and scale parameters. This can be done, for example, by collecting the mass of a continuous-valued 1D distribution in quantization bins of discrete numbers. ·The PMF can be multivariate, in which case the PMF is defined over a multi-dimensional set of quantization points. ·For any of the previous items, the quantization points can be, for example, several integer values or any interval. The interval of the quantization bins can also be predicted by an auxiliary network such as a hyper-decoder or predicted from the context (adjacent latent pixels). ·For any of the previous items, the parameters that control the PMF can be fixed or predicted by an auxiliary network such as a hyper-decoder or predicted from the context (adjacent latent pixels).

[0299] This framework can be extended in several ways. For example, if the discrete PMF is multivariate (multi-dimensional), a multivariate (multi-dimensional) interpolation scheme can be used to interpolate the values of the PMF to real vector-valued points. For example, multi-linear interpolation (bilinear in 2D, trilinear in 3D, etc.) can be used. Also, multi-cubic interpolation (bicubic interpolation in 2D, tricubic interpolation in 3D, etc.) can be used.

[0300] This interpolation method is not restricted to modeling only discrete-valued PMFs. Any discrete-valued function can be interpolated anywhere in the AI-based compression pipeline, and the techniques described here are not strictly limited to the modeling of probability mass / density functions.

[0301] Context Model In Al-based compression, autoregressive context models have powerful entropy modeling capabilities, but suffer from very short run times because they must be run serially. This paper describes how to overcome this difficulty by predicting the autoregressive model components from a hyperdecoder (and conditioning these components on "side" information). This technique results in an autoregressive system that has impressive modeling capabilities but can run in real time. This real time is achieved by decoupling the autoregressive system from the models required by the lossless decoder. Instead, the autoregressive system solves linear equations at decoding time, which can be solved very quickly using numerical linear algebra techniques. Encoding can also be done quickly by solving simple implicit equations.

[0302] This document outlines the capabilities and scope of current and future uses of autoregressive probability models with linear decoding systems, including but not limited to image and video data compression based on AI and deep learning.

[0303] In AI-based image and video compression, an input image x is mapped to a latent variable y. This latent variable is encoded into a bitstream and sent to the receiver, which decodes the bitstream into the latent variable. The receiver then converts the recovered latent variable into a representation (reconstruction) of the original image x (hat).

[0304] To perform the step of converting the latent to a bitstream, the latent can be quantized to an integer value representation y(hat). This quantized latent y(hat) is converted to a bitstream via a lossless encoder / decoder scheme, such as an arithmetic encoder / decoder or a range encoder / decoder.

[0305] A lossless encoding / decoding scheme may require a model one-dimensional discrete probability mass function (PMF) for each element of the latent quantized variable. Optimal bitstream length (file size) is achieved when this model PMF matches the latent true one-dimensional data distribution.

[0306] Therefore, the file size is closely tied to the power of the model PMF to match the true data distribution. A stronger model PMF leads to smaller file sizes and better compression. In some cases, this leads to better reconstruction error (because for a given file size, more information can be sent to reconstruct the original image). Therefore, much effort has been put into developing stronger model PMFs (often called entropy models).

[0307] The typical approach to model one-dimensional PMF in Al-based compression is to use a parametric one-dimensional distribution, The solution is to use JPEG2024528208000100.jpg519, where θ is the parameter of the one-dimensional PMF. For example, a quantized Laplacian or a quantized Gaussian can be used. In these two cases, θ constitutes the location and scale o-parameters of the distribution. For example, if a quantized Gaussian (Laplacian) is used, the PMF can be expressed as:

number

[0308] A more powerful model can be created by "conditioning" the parameters θ, such as the position μ or scale σ, on other information stored in the bitstream. In other words, rather than statically fixing the parameters of the PMF to be constant across all inputs in an Al-based compression system, the parameters can respond dynamically to the input.

[0309] This is commonly done in two ways. The first is to send extra side information z into the bitstream in addition to y. This variable z is often called a hyper-latent. This variable is decoded in its entirety before decoding y, so that it can be used to encode / decode y. Then μ and σ can be made functions of z, for example, and run through a neural network to return μ and σ. The one-dimensional PMF is then said to be conditional on z, Given by JPEG2024528208000102.jpg532.

[0310] Another approach is to use autoregressive probabilistic models, such as the PixelCNN described in Aaron van den Oord, Nal Kalchbrenner, Lasse Espeholt, Koray Kavukcuoglu, Oriol Vinyals, and Alex Graves. Conditional Image Generation with PixelCNN Decoders. In Daniel D. Lee, Masashi Sugiyama, Ulrike von Luxburg, Isabelle Guyon, and Roman Garnett, editors, Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016, Barcelona, ​​Spain, pages 4790-4798, 2016, which is herein incorporated by refernce, David Minnen, Johannes Balle, and Joint autoregressive and hierarchical priors for learned image compression, these have been widely used in academic papers on AI-based compression. Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicolo Cesa-Bianchi, and Roman Garnett, editors, Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montreal, Canada, pages 10794-10803, 2018. In this framework, the context pixel is used to condition the position μ and scale σ parameters of the PMF at the current pixel.These context pixels are previously decoded pixels close to the current pixel. For example, suppose the previous k pixels have been decoded. Due to inherent spatial correlation in images, these pixels often contain relevant information about the current active pixel. Therefore, these context pixels can be used to improve the prediction of the position and scale of the current pixel. The PMF of the current pixel is: JPEG2024528208000103.jpg569, where μ and σ are functions of the prior k variables (usually a convolutional neural network).

[0311] These two approaches, conditioning on hyperlatent or autoregressive contextual models, both have advantages and disadvantages.

[0312] One of the main advantages of hyper-latent conditioning is that the quantization can be location-shifted. In other words, the quantization bins can be centered around the location parameter μ. With integer-valued bins, the quantization latency becomes

number

[0313] The main advantage of autoregressive context models is that they can utilize context information (decoded neighboring pixels). Since images (and videos) are highly spatially correlated, these neighboring pixels can provide a very accurate and precise prediction of what the current pixel should be. Most of the state-of-the-art academic Al-based compression pipelines use autoregressive context models due to their impressive performance results, when measured in terms of bitstream length and reconstruction error. However, despite their impressive relative performance, they suffer from two problems:

[0314] That is, the PMF of the current pixel depends on all previously decoded pixels. Moreover, the position function μ(·) and the scale function σ(·) are usually large neural networks. These two facts mean that autoregressive context models cannot be run in real-time, taking orders of magnitude more time than the computational requirements needed for real-time performance on edge devices. Thus, in the current situation, autoregressive context models are not commercially viable, despite the fact that they yield impressive compression performance.

[0315] Second, due to the effect of cascading errors, the autoregressive context model is not as straightforward as straight rounding. JPEG2024528208000107.jpg515 must be used. Position shift rounding JPEG2024528208000108.jpg527 is not possible because small floating-point errors introduced early in the decoding path will be amplified during the serial decoding path, leading to wildly different predictions between the encoder and decoder. The absence of position-shift rounding is problematic, and all other building blocks being equal, an autoregressive model with position-shift rounding (if it can be constructed) is likely to be superior to a straight-rounding autoregressive model.

[0316] Therefore, there is a need to develop a PMF modeling framework that combines the advantages of conditioning over hyperlatent (fast execution time, position shift rounding) with the impressive performance of autoregressive models (making strong predictions from pre-decoded context pixels).

[0317] Described below is a technique where we modify the hyperdecoder to use it to additionally predict the parameters of the autoregressive model - in other words, we condition the parameters of the autoregressive model on the hyperlatent z. This is in contrast to the standard setup in autoregressive modeling, where the autoregressive function is static and invariant, and does not change depending on the inputs of the compression pipeline.

[0318] Here we mainly deal with the following quasi-linear setup (the decoding path is linear, while the encoding is not): In addition to the μ and σ predictions, the hyperdecoder can output a sparse matrix L, called the context matrix. This sparse matrix is ​​used for the autoregressive context modeling component of the PMF as follows: Given an order of latent pixels (e.g., raster scan order), we assume that the previous k latent pixels have been encoded / decoded and are available for autoregressive context modeling. Our approach uses a modified position-shift quantization as follows: We quantize by:

number

[0319] The probability model is JPEG2024528208000110.jpg848. In matrix-vector notation, this is:

number

[0320] Note that this is a form of autoregressive context modeling, since the 1D PMF depends on previously decoded latent pixels, whereas only the position parameter depends on previously decoded latent pixels, not the scale parameter.

[0321] Note that the integer values ​​that may actually be encoded by an arithmetic encoder / decoder are the quantized residuals.

number

[0322] Thus, upon decoding, what the arithmetic encoder returns from the bitstream is ξ(hut), not y(hut). y(hut) can then be reconstructed by solving the following simultaneous linear equations:

number

[0323] Solving system (50) is decoupled from the arithmetic decoding process, i.e., the arithmetic decoding process must be performed serially as the bitstream is received, whereas solving (50) is independent of this process and can be performed using any numerical linear algebra algorithm. The decoding passes of the L context modeling step do not have to be a serial procedure and can be performed in parallel.

[0324] Another way to look at this result is that, equivalently, the arithmetic encoder / decoder operates on the ξ (hat) whose position is 0. That is, the arithmetic encoder operates on the residual ξ (hat) rather than on the latent y (hat). In this way, the PMF is JPEG2024528208000115.jpg524. Only after ξ is recovered from the bitstream do we then recover y. However, this procedure can be very fast since the recovery from the bitstream may not be autoregressive (the only dependency is σ, which is not made context / autoregressive dependent). We can then recover y using a highly optimized linear algebra routine to solve (50).

[0325] In both encoding and training the L-context system, (48) can be solved, but the unknown variables y(hut)-y(hut) are not given explicitly and must be determined. In fact, (48) is an implicit system. Here, we outline several possible approaches to finding y(hut) that satisfies (48). The first approach is to solve (48) successively, operating on the pixels according to their order of dependency in the autoregressive model. In this setting, we simply iterate over all pixels in their autoregressive order, and at each iteration apply (47) to extract the quantized latent at the current iteration. Since (48) is an implicit equation, the second approach is to employ an implicit equation solver (implicit coding solver), which is an iterative solver that finds a fixed-point solution to (48). Finally, in certain special cases, the autoregressive structure defined by the sparse context matrix L can be exploited to parallelize components of the serial decoding path. In this approach, we first create a dependency graph (Directed Acyclic Graph) that defines the dependencies between latent pixels. This dependency graph can be constructed based on the sparse structure of the L matrix. We then note that pixels at the same level of the DAG are conditionally independent of each other. Hence, every pixel can be computed in parallel without affecting the computation of other pixels at that level. Thus, in encoding (and training), the graph is iterated by starting from the root node and working through the levels of the DAG. At each level, every node is processed in parallel. This procedure results in a dramatic speedup compared to a naive serial implementation when a parallel computing environment (such as a Graphics Processing Unit or Neural Processing Unit) is available. This procedure is illustrated in Figure 23. The left image in Figure 23 shows the L context parameters associated with the i-th pixel of the example. The neighboring context pixels are the pixel directly above the current pixel and the pixel in the left neighborhood. The right image shows the pixels enumerated in raster scan order. The image below shows an example of constructing a directed acyclic graph (DAG) given the dependencies generated by the L context matrix. Pixels at the same level are conditionally independent of each other and can be encoded / decoded in parallel.

[0326] Many of the techniques described in the previous section can also be applied to the decoding process.

number

[0327] An example of an L context module in an Al-based compression pipeline is described in detail below.

[0328] FIG. 24 shows the prediction context matrix L in an example Al-based compression pipeline. y The figure shows the encoding process using a general implicit solver. In this diagram, an input image x (hat) is fed into an encoder function such as a neural network. The encoder outputs a latent y, which is fed into a hyperencoder, which returns a hyperlatent z.

[0329] The hyperlatency is based on the learned position μ z and the scale σ z The quantized hyper-latency is then sent to the hyper-decoder, which outputs the parameters of the entropy model for y. These include the parameters of the position μ y , scale σ y and the L context matrix L y The residual is computed by solving the coding equation using one of the methods described above. This quantized residual ξ (hat) has zero mean and scale parameter σ yThe MF is sent to the bitstream using the MF with

[0330] Figure 25 shows the prediction context matrix L for an example Al-based compression pipeline. y , a linear equation solver is depicted. In the decoding, the learned position parameter μ z and σ z First, the hyper-latency z is recovered from the bitstream using a 1-D PMF with scale parameter as well as a lossless decoder. Optionally (not shown), the L context module used during encoding can also be employed. The hyper-latency is fed to a hyper-decoder, which then extracts the position μ y , scale σ y and the sparse context matrix L y The residual is the lossless decoder, with zero-mean PMF and scale parameter σ y The quantized latents y are then recovered from the bitstream using a linear decoding system as described above. Finally, the reconstructed image is recovered by feeding the quantized latents y through a decoder function, such as another neural network.

[0331] In the previous sections, we have assumed that L is lower triangular with respect to the pixel decoding order. As a generalization, we relax this assumption and assume a general matrix A that is not necessarily lower triangular. In this case, the encoding equation becomes

number

number

number

[0332] In general, the context function may be nonlinear. For example, the coding problem is

number

number

number

[0333] One interpretation of this latter extension is the implicit PixelCNN. For example, (55) models an autoregressive system when f(·) has a triangular Jacobian (a matrix of first derivatives). However, (55) is more general than this interpretation, and in fact can model not only autoregressive systems, but also stochastic systems with both forward and backward conditional dependence on pixel ordering.

[0334] Learned AR order In Al-based image and video compression, autoregressive modeling is a powerful technique for entropy modeling of the latent space. Context models conditioned on previously decoded pixels are used in modern Al-based compression pipelines. However, the autoregressive ordering in the context model is often predefined and elementary, such as raster-scan ordering, which may impose undesirable biases on the learning. For this reason, we propose a fixed ordering other than raster-scan, a conditional ordering, a learned ordering or a directly optimized ordering as the autoregressive ordering in the context model.

[0335] In mathematical terms, the goal of lossy Al-based compression is to infer a prior probability distribution, an entropy model, that matches as closely as possible to the latent distribution generating the observed data. This can be achieved by training a neural network through an optimization framework such as gradient descent. Entropy modeling underpins the entire Al-based compression pipeline, and better distribution matching corresponds to better compression performance, characterized by lower reconstruction loss and bitrate.

[0336] For image and video data that exhibit large spatial and temporal redundancy, an autoregressive process called context modeling is very useful to exploit this redundancy in entropy modeling. At a high level, the general idea is to condition the explanation of subsequent information using existing available information. The process of conditioning previous variables to realize the next variable implies a specific order of autoregressive information retrieval structure. This concept has proven very powerful in AI-based image and video compression and is commonly used as part of state-of-the-art neural compression architectures.

[0337] However, the ordering of autoregressive structures in Al-based image and video compression, autoregressive ordering (AO for short), may be pre-determined. Such context models often adopt the so-called raster scan order, which naturally follows the data sequence of, for example, image data type (3-dimensional; length x width x channel, such as RGB). Figure 26 shows an example of raster scan ordering for a single channel image. The grey squares are conditionable variables, and the white squares are non-conditionable variables. However, adopting the raster scan order as the basic AO is arbitrary and may be harmful, since it cannot condition on information from below and to the right of the current variable (or pixel). This may cause inefficiency or potential undesirable bias in the learning of neural networks.

[0338] Below we describe a number of AOs, either fixed or learnable, along with a number of different frameworks in which they can be formulated. Context modeling AOs can be generalized across these frameworks, and latent variables can be modeled as follows: (a) We detail the theoretical aspects of Al-based image and video compression and the purpose of autoregressive modeling in context models. (b) Describe and illustrate a range of traditional and non-traditional AOs that can be assigned to context modeling. (c) We describe a number of frameworks in which AO can be trained by network optimization using gradient descent or reinforcement learning methods.

[0339] Al-based image and video compression pipelines usually follow an autoencoder structure, which is constructed by a convolutional neural network (CNN) that constructs the encoding and decoding modules, whose parameters can be optimized by training on natural image and video datasets. The (observed) data is generally denoted as p and is assumed to be distributed according to a data distribution p(x). The feature representation after the encoder module is called latent and denoted as y, which is entropy coded into a bitstream for encoding and vice versa for decoding.

[0340] The true distribution of the latent space p(y|x) is practically not available. This is because the data distribution This is because it is hard to marginalize the joint distribution over y and x to compute JPEG2024528208000126.jpg637. Therefore, we can only find an approximate representation of this distribution, and that is what entropy modeling does.

[0341] The true latent distribution of JPEG2024528208000127.jpg612 can be expressed, without loss of generality, as a joint probability distribution with the conditional dependent variables

number

number

[0342] Applying either of the two concepts imposes constraints that invalidate the equivalence of the joint probability and factorization into conditional components as described in Equation (59), but this is often done to trade off against the complexity of the modeling. The first concept is most often actually done for high-dimensional data, such as modeling the context based on PixelCNN where only local receptive fields are considered. An example of this process is shown in Figure 27, which shows a 3×3 receptive field where the next pixel is conditioned on local variables (where the arrows occur) instead of all the preceding variables. However, in this paper, to generalize the innovation taken up here, the imposition of this constraint is not considered.

[0343] The second concept includes cases where a factored entropy model (conditioned only on deterministic parameters without conditioning on random variables) and a hyperprior entropy model (latent variables are all conditionally independent by conditioning on a set of hyperlatent z) are assumed, both of which have 1-step AO, that is, the inference of the joint distribution is performed in a single step.

[0344] Below, for applications in Al-based image and video compression, three different frameworks for specifying AO for serial execution of any autoregressive process will be described. This includes, but is not limited to, entropy modeling by a context model. Each framework provides (1) a way to define AO and (2) a way to formulate optimization techniques for AO.

[0345] The data can be assumed to be arranged in a two-dimensional format (single channel image or frame, single channel video) with dimensionality M=H×W, where H is the height dimension and W is the width dimension. The concepts presented here are equally applicable to data with multiple channels and multiple frames.

[0346] Graphical models, more specifically Directed Acyclic Graphs (DAGs), are very useful for describing probability distributions and their conditional dependency structure. The graph is constructed with nodes corresponding to the variables of the distribution and directed links (arrows) indicating conditional dependencies (the variable at the arrowhead is conditioned on the variable at the end). As a visual example, the joint distribution that explains the example in Figure 28 looks like this:

number

[0347] The main constraint for a directed graph to correctly describe the joint probabilities is that it does not contain any directed cycles. This means that there should not be any path that starts at any node on the path and ends at the same node, thus resulting in a directed acyclic graph. The raster-scan ordering follows exactly the same structure as shown in Figure 28, which shows an example DAG describing the joint distribution of four variables {y1,y2,y3,y4}, and in equation (60) when the variables are ordered from 1 to N in a raster-scan pattern. Given our assumptions, this is an M-step AO, meaning that M passes or runs of the context model are required to evaluate the complete joint distribution.

[0348] Another AO less than M steps is checkerboard ordering. Figure 29 shows a 2-step AO where all current variables at each step are conditionally independent and can be evaluated in parallel. In Figure 29, at step 1, the distribution of the current pixels is inferred in parallel without conditions. At step 2, the distribution of the current pixels is inferred in parallel and conditioned on all pixels of the previous step. At each step, all current variables are conditionally independent. The corresponding directed graph is shown in Figure 30. As shown in Figure 30, step 1 evaluates the topmost nodes, and step 2 is represented by the bottommost nodes with arrows indicating the conditions.

[0349] The binary mask kernel for autoregressive modeling is a convenient framework for specifying an N-step AO with N << M. Given data y, the binary mask kernel approach requires splitting it into N low-resolution sub-images {y1,…,y N}. Previously, each pixel was defined as variable y i , but here, JPEG2024528208000131.jpg513 is defined as a group of K pixels or variables that are conditionally independent (conditioned by combining for future steps).

[0350] The sub-image y i is, k H ×k W stride of, containing element M i,pq is extracted by convolving the data with a binary mask kernel JPEG2024528208000132.jpg651. For this to define a valid AO, the binary mask kernel must follow the following constraints.

Number

Number

[0351] where 1 kH,kW is size k H ×k W In other words, each mask must be unique and have only a single entry of 1 (with the remaining elements being 0). These are sufficient conditions to ensure that the AO is exhaustive and follows a logical ordering of conditions. To enforce constraints (61) and (62) while establishing the AO, k H k W We can learn logits, one for each position (p,q) in the kernel, and then order the autoregressive process, for example by ranking the logits from high to low. We can also apply the Gumbel softmax trick to eventually force one-hot coding as we reduce the temperature.

[0352] Figure 31 shows an example of a 2 × 2 binary mask kernel subject to constraints (61) and (62) generating four sub-images, conditioned according to the graphical model described in Figure 28 (but with vector variables instead of scalars). This example is very similar to the checkerboard AO described in the previous section, but with some additional intermediate steps. In fact, it is possible to define an AO with a binary mask kernel that exactly reflects the previous checkerboard AO, by not conditioning between y1 and y2, and between y3 and y4. i The choice of whether to condition on can be determined by the adjacency matrix A being strictly lower triangular, and the threshold condition T. In this case, A is more practically of dimension k. H k W ×k H k W Following the previous example, Figure 32 shows how this is possible. Figure 32 shows an example of an adjacency matrix A that determines the graph connectivity of an AO defined in the binary mask kernel framework. If any link has an associated adjacency term smaller than a threshold T, then the link is ignored. When links are ignored, a conditional independence structure emerges, allowing the autoregressive process to be built in fewer steps.

[0353] It is also possible to represent traditional interlacing schemes with binary mask kernels such as Adam7 used in PNG. Figure 33 shows the indices for the Adam7 interlacing scheme. Note that this notation is used as an indication of how the kernels are arranged and grouped, and is not actually used directly in the model. For example, 1 corresponds to one mask in the top left position; 2 corresponds to another single mask in the same position; 3's correspond to two masks grouped together, one each in the same position as the 3's in the index. The same principle applies to the remaining indices. The kernel size is now 8x8, resulting in 64 mask kernels, indexed groupings as shown in Figure 33, resulting in 7 sub-images (hence the 7-step AO). This shows that the binary mask kernel framework is adapted to handle common interlacing algorithms that utilize low-resolution information to generate high-resolution data.

[0354] The raster scan order can also be defined in the framework of binary mask kernels with kernel size H × W. In this case, we have H × W = N steps of AO, where N binary mask kernels of size H × W are organized in raster scan order.

[0355] In summary, the binary mask kernel is well suited for gradient descent-based learning techniques and is related to the further notion of autoregressive ordering in frequency space, as explained below.

[0356] Ranking tables are a third framework for characterizing AOs, and under fixed rankings they are particularly effective for describing M-step AOs without the representational complexity of binary mask kernels. The concept of a ranking table is simple: Given JPEG2024528208000135.jpg513 (flattened and corresponding to the total number of variables), each AO creates a ranking system of elements in q,qi The indexing can be done with the argsort operator, and the ranking can be either descending or ascending, depending on the interpretation of q. The largest index is assigned as y1, the second largest q i so that the index with q is assigned as y2. The indexing is done using the argsort operator, and the ranking is either descending or ascending depending on the interpretation of q.

[0357] q can be an existing quantity that conveys certain information about the source data y, such as the entropy parameter of y (learned or predicted by hyperpreferences), e.g., the scale parameter σ. This is due to the large scale parameter σ ij Regions of high uncertainty associated with variables with y1, y2, …, y 16 Figure 34 shows an example of the visualization of the scale parameters, where after flattening to q, we define a ranking table using the argsort operator to order the variables in descending order by the magnitude of their respective scale parameters.

[0358] q can also be derived from existing quantities, such as the first or second derivative of the location parameter μ. Both of these can be obtained by applying finite differences to obtain the gradient vector (for first derivatives) or Hessian matrix (for second derivatives), obtained before computing q and argsort(q). A ranking is then established by the norm of the gradient vector, the norm of the eigenvalues ​​of the Hessian matrix, or any measure of the curvature of the latent image. Alternatively, in the case of second derivatives, the ranking can be based on the magnitude of the Laplacian, which is equivalent to the trace of the Hessian matrix.

[0359] Finally, q can also be a separate entity altogether: a fixed q can be arbitrarily predefined before training, remaining static or dynamic during training, like a hyperparameter. Alternatively, it can be learned and optimized using gradient descent, or parameterized by a hypernetwork.

[0360] How you access the elements of y depends on whether you want to run the gradient through a ranking operator: · When no gradient is needed : Sort q in ascending / descending order, access elements, and then access elements of y based on that order. · When gradients are needed : The order is the discrete permutation matrix P or its continuous relaxation matrix Represent it as JPEG2024528208000136.jpg55, and multiply it by the sorted y and matrix:y sort =Py

[0361] If the ranking table is optimized based on gradient descent, indexing operators such as argsort or argmax may not be differentiable. Therefore, the permutation matrix We need to use successive relaxations of JPEG2024528208000137.jpg55, which can be implemented with a soft sort operator.

number

[0362] The ranking table concept can be extended to work for binary mask kernels as well: the matrix q will have the same dimensions as the mask kernel itself, and AO will be specified based on the ranking of the elements of q. Figure 36 shows an example visualization of the ranking table concept applied to a binary mask kernel framework.

[0363] Another possible autoregressive model is one defined by a hierarchical transformation of the latent space: in this view, the latent is transformed into a hierarchy of variables, and lower hierarchical levels are conditioned on higher hierarchical levels.

[0364] This concept can be best explained using wavelet decomposition, where a signal is broken down into high and low frequency components. This is done via the wavelet operator W. Let y be a latency of size H × W pixels. 0 The superscript 0 indicates that the latent is at the lowest (or root) level of the hierarchy. After one wavelet transform, the latent is divided into four small images y 1 ll ,y 1 lh ,y 1 hl , and y 1 hh , each of which has size JPEG2024528208000142.jpg521. H and L stand for high-frequency and low-frequency components, respectively. The first letter of the tuple corresponds to the first spatial dimension of the image (e.g., height), the second letter to the second dimension (e.g., width). Thus, for example, y 1 hl is the latent image y, which corresponds to high frequencies in the height dimension and low frequencies in the width dimension. 0 are the wavelet components of

[0365] In matrix notation it looks like this:

number

[0366] Applying this procedure recursively to the low-frequency blocks builds a hierarchical tree of decomposition. Figure 37 shows an example of a procedure with two hierarchies. Figure 37 shows a hierarchical autoregressive order based on the wavelet transform. The top image shows that the forward wavelet transform creates a hierarchy of variables, in this case with two levels. The middle image shows that the transform can be inverted to restore the low-frequency elements of the previous level. The bottom image shows the autoregressive model defined by an example DAG between elements at one level of the hierarchy. Thus, the wavelet transform can be used to create trees with multiple levels.

[0367] The important thing is that if the transformation matrix W is invertible (in fact, in the case of the wavelet transform, W -1 =W T ) The whole procedure is reversible: given the last level of a hierarchy, it is easy to recover the low-frequency components of the previous level by simply applying an inverse transform to the last level. Then, after recovering the low-frequency components of the next level, we apply an inverse transform to the second level, and so on, recovering the original image.

[0368] Now, how can we use this hierarchical structure to construct an autoregressive ordering? At a hierarchical level, an autoregressive ordering is defined between the elements of that level. For example, see the bottom image in Figure 37, where the low frequency components are at the root of the DAG at that level. Note that the autoregressive model can also be applied to the building blocks (pixels) of each variable in the hierarchy. The remaining variables of a level are conditioned on the preceding elements of the level. Then, after describing all the conditional dependencies of the level, the inverse wavelet transform is used to restore the lowest frequency components of the preceding level.

[0369] Another DAG is defined between the elements of the next lowest level and an autoregressive process is applied recursively until the original latent variables are recovered.

[0370] Thus, an autoregressive order is defined on the variables given by the levels of the wavelet transform of the image, using a DAG and an inverse wavelet transform between the elements of the levels of the tree.

[0371] This sequel can be generalized in several ways. It does not have to be a wavelet transform, any reversible transform can be used. This includes: - Wavelets - Permutation matrices, e.g. as defined by a binary mask -Other orthonormal transforms such as the Fast Fourier Transform -Learned invertible matrices - Learned invertible matrices -Invertible matrices predicted by e.g. hyperpreferences. Hierarchical decomposition can also be applied to video, where each level of the tree has eight components corresponding to lll, hll, lhl, hhl, llh, lhh, hhh, hlh, where the first letter denotes the temporal component.

[0372] Augmented Lagrangian Examples of constrained optimization and rate-distortion annealing techniques are described in International Patent Application No. PCT / GB2021 / 052770, which is incorporated herein by reference.

[0373] An Al-based compression pipeline seeks to minimize the rate (R) and distortion (D), with the objective function:

number

[0374] In International Patent Application No. PCT / GB2021 / 052770, this problem is reformulated as a constrained optimization problem. The method for solving this constrained optimization problem is the Augmented Lagrangian method, as described in International Patent Application No. PCT / GB2021 / 052770. The constrained optimization problem is to solve the following: Min D (66) R=c etc. (67) where c is the target compression ratio. Note that D and R are means over the data distribution. We can also use inequality constraints. Furthermore, the roles of R and D can be reversed. Instead, we can minimize the rate subject to a distortion constraint (which can be a system of constraint equations).

[0375] Typically, constrained optimization problems are solved using stochastic linear optimization methods: the objective function is calculated for a small number of training samples (rather than the entire dataset), and then the gradients are calculated. An update step is then performed to change the parameters of the compression algorithm and possibly other parameters related to the constrained optimization, such as the Lagrange multipliers. This process is repeated thousands of times until a suitable convergence criterion is reached. For example, the following steps are performed: · Calculate the loss using SGD, mini-batch SGD, Adam (or any other optimization algorithm used to train neural networks), etc. Run one optimization step on JPEG2024528208000145.jpg747. Update the Lagrange multipliers: JPEG2024528208000146.jpg530, where ε is chosen to be small. · Repeat the above two steps until the Lagrange multipliers converge based on the target rate r0.

[0376] However, there are several issues encountered while training constrained optimization problems in a stochastic small batch first-order optimization setting. First, constraints cannot be computed for the entire dataset at each iteration, but are typically computed only for the small number of training samples used in each iteration. Using such a small number of training samples in each batch can make updates to constrained optimization parameters (such as the Lagrangian Multipliers in Augmented Lagrangian) extremely dependent on the current batch, leading to high variability in training updates, suboptimal solutions, or even instability of the optimization routine.

[0377] An average constraint value (such as the average rate) can be calculated over the N previous iteration steps. This has the beneficial effect of extending the constraint information of many last optimization iteration steps to be applied to the current optimization step, especially with respect to parameter updates related to constrained optimization algorithms (such as updating the Lagrangian multipliers in the augmented Lagrangian). A non-exhaustive list of ways to calculate the average of the last N iterations and apply it to an optimization algorithm is as follows: Maintain a buffer of constraint values ​​for the last N training samples. This buffer can be, for example, JPEG2024528208000147.jpg637 can be used to update the augmented Lagrange multiplier λ at iteration t, where avg is a general average operator. Examples of average operators are: -Arithmetic mean (often simply called "average") -Median -geometric mean -harmonic mean - Exponential moving average -Smoothed moving average -Linear weighted moving average However, any averaging operator may be used.

[0378] Instead of computing the gradients at every training step and immediately applying them to the model weights, we accumulate the gradients for N iterations, and then do one training step after N iterations with an averaging function.

[0379] This averaged constraint value is calculated over the previous N iterations and is used to update the parameters of the training optimization algorithm, such as the Lagrange multipliers.

[0380] A second problem with stochastic linear optimization algorithms is that the dataset may contain images with extremely large or small constraint values ​​(e.g., the rate R is either extremely small or extremely large). The presence of outliers in the dataset may force the function being trained to take the outliers into account, leading to poor fit for more typical samples. For example, updates to the optimization algorithm's parameters (e.g., Lagrange multipliers) may have high variance in the presence of many outliers, causing suboptimal training.

[0381] Some of these outliers can be removed from the above constrained average calculation. Some possible ways to filter (remove) these outliers are: When accumulating training samples, you can replace the mean with a trimmed mean. The difference between the trimmed mean and the regular mean is that in the trimmed version, x% of the values ​​above and below are not taken into account in the mean calculation. For example, you can trim 5% of the top and bottom samples (in sorted order). There is no need to trim throughout the entire training procedure, one can also use regular averaging (without trimming) later in the training procedure. For example, the trimmed averaging can be turned off after 1 million iterations. · Fit an outlier detector every N iterations so that the model can either completely remove or weight those samples from the average.

[0382] Using constrained optimization algorithms such as the Augmented Lagrangian, we can target a particular average constraint (e.g. rate) objective c on the training dataset. However, convergence to that target constraint on the training set does not guarantee that we will end up with the same constraint value on the validation set. This can be caused, for example, by a change between the quantization function used in training and the quantization function used in inference (test / validation). For example, it is common to quantize with uniform noise during training but use rounding in inference (known as "STE"). Ideally, we would like the constraint to be satisfied in inference, but this is difficult to achieve.

[0383] Possible solutions include: At each training step, we compute the constraint target value using the detach operator, where in the forward pass the inference value is used, but in the backward pass (gradient calculation) the training gradient is used. For example, when considering rate constraints, The value JPEG2024528208000148.jpg548 is used, where 'detach' means to detach the value from the automatic differentiation graph. Update the parameters of the constraint algorithm with a “holdout set” of images on which to evaluate the model using the same settings as inference.

[0384] The techniques described above and in International Patent Application PCT / GB2021 / 052770 can also be applied to Al-based video compression. In this case, Lagrangian multipliers can be applied to the rate and distortion associated with each frame of the video used for each training step. One or more of these Lagrangian multipliers can be optimized using the techniques described above. Alternatively, the multipliers may be averaged over multiple frames during the training process.

[0385] The target value of the Lagrangian multiplier may be set to an equal value for each frame of the input video used in the training step. Alternatively, different values ​​may be used. For example, different target values ​​may be used for I-frames and P-frames of a video. A higher target rate may be used for I-frames compared to P-frames. A similar approach may also be applied for B-frames.

[0386] In a manner similar to image compression, the target rate may be initially set to zero for one or more frames of the video used for training. If a target value is set for distortion, the target value may be set to maximize the initial weighting for distortion (e.g., the target rate may be set to 1).

[0387] Tensor Networks Al-based compression relies on modeling discrete probability mass functions (PMFs). These PMFs may seem simple at first glance: our usual mental model starts with one discrete variable X, which has D possible values ​​X1,…,X D Then, the construction of the PMF P(X) is i =P(X i ) Of course, Pi must be non-negative and sum to one, which can be achieved by, for example, using the softmax function For modeling purposes, we can find a set of Ps in this table that fit a particular data distribution. i It doesn't seem that difficult to learn.

[0388] What is the PMF for two variables, X and Y? Entry P ij =P(X i ,Y j ), which is a bit more complicated. Currently, the table contains D 2entries, but still, as long as D is not too large, JPEG2024528208000150.jpg2646 is manageable. Continuing, with three variables, you need a 3d table, with entry P ijk is indexed by a 3-tuple.

[0389] However, this naive table-building "approach" quickly becomes unwieldy when trying to model more than a handful of discrete variables. For example, consider modeling PMF on an RGB 1024x1024 image space, each of which has 256 3 (Each color channel has 256 possible values, and we have 3 color channels). Then the lookup table needed is JPEG2024528208000151.jpg615 entries. In decimal, it is about JPEG2024528208000152.jpg68 There are many approaches to dealing with this problem, but the textbook approach in discrete modeling is to use probabilistic graphical models.

[0390] As an alternative approach, we can model the PMF as a tensor. A tensor is just another word for a big table (but it has algebraic properties that we won't discuss here). A discrete PMF can always be written as a tensor. For example, a 2-tensor (also called a matrix) is an array with two indices, i.e. a two-dimensional table. Thus, the above PMF P for two discrete variables X and Y is ij =P(X i ,Y j ) is a 2-tensor. N-tensor T i1 ,…,T iN is an array with N indices, and if the entries of T are positive and sum to 1, then this is a PMF for N discrete variables. Table 1 compares the standard PMF view with the tensor view for several probabilistic concepts.

[0391] The main attraction of this perspective is that large tensors can be modeled using the framework of tensor networks: tensor networks are used to approximate very high-dimensional tensors as a reduction of several lower-dimensional (and therefore tractable) tensors, i.e., tensor networks are used to provide a low-rank approximation of otherwise intractable tensors. [Table 3]

[0392] For example, if we view matrices as 2-tensors, standard low-rank approximations (such as singular value decomposition (SVD) and principal component analysis (PCA)) are tensor network factorizations. Tensor networks are a generalization of low-rank approximations used in linear algebra to multi-linear maps. An example of the use of tensor networks in probabilistic modeling for machine learning is given in "Ivan Glasser, Ryan Sweke, Nicola Pancotti, Jens Eisert, and J Ignacio Cirac. Expressive power of tensor-network factorizations for probabilistic modeling, with applications from hidden markov models to quantum machine learning. arXiv preprint, arXiv:1907.03741, 2019", which is incorporated herein by reference.

[0393] Tensor networks can be thought of as an alternative to graphical models. There is a correspondence between tensor networks and graphical models, and probabilistic graphical models can be reconstructed as tensor networks, but the reverse is not true. While they cannot be recast as probabilistic graphical models, there exist tensor networks for joint density modeling that have strong performance guarantees and are computable. In many situations, tensor networks are more expressive than traditional probabilistic graphical models such as HMMs. Given a fixed number of parameters, tensor networks experimentally outperform HMMs. Furthermore, for certain low-rank approximations, tensor networks could theoretically outperform HMMs again.

[0394] All other modeling assumptions being equal, tensor networks may be preferred over HMMs.

[0395] An intuitive explanation for this result is that probabilistic graphs factor connections via conditional probabilities, whereas exponential maps By considering only JPEG2024528208000154.jpg642, we are modeling the coupling as a Boltzmann / Gibbs distribution, which may indeed be a restrictive modeling assumption. A completely different approach that tensor networks offer is to model the coupling as an inner product: for some Hermitian positive (semi)definite operator H, JPEG2024528208000155.jpg527 (This modeling approach is inspired by the Born rule for quantum systems). The operator H can be written as a large tensor (or tensor network). The key is that the entries of H are complex. It is not at all clear how (or if) this can be transformed into a graph model. However, it does offer a completely different modeling perspective that is not available elsewhere.

[0396] Let us explain what tensor network decomposition is with a simple example. ijThere is a large D×D matrix T (2 - tensor) having [something], and it is desired to create a low - rank approximation (rank - r approximation with r < D) of T. One way to do this is to find the approximate T (hat).

Number

[0397] In other words, setting T(hat)=AB, where A is a D×r matrix and B is an r×D matrix. We have introduced hidden dimensions shared between A and B, and these are summed. By setting r to be very small compared to dealing with a huge D×D matrix, instead of D 2 parameters, we can get 2Dr parameters, which can significantly save computational time and power. Furthermore, in many modeling scenarios, even if r is made very small, a “sufficient” approximation of T can be obtained.

[0398] Now, let's try to model a 3 - tensor with the same approach. Let T ijk be a D×D×D tensor T given with entries.

Number

[0399] Here A and C are low - rank matrices, and B is a low - rank 3 - tensor. There are two hidden dimensions between A - B and B - C. One is between A and B, and the other is between B and C. In tensor network terms, these hidden dimensions may be called binding dimensions. The summation of dimensions may be called reduction.

[0400] Continuing with this example, a 4 - tensor can be approximated as a product of lower - dimensional tensors, but the index notation becomes cumbersome to write quickly. Instead, use a tensor network diagram, which is a concise way to convey the same calculation diagrammatically.

[0401] In a tensor network diagram, tensors are represented as blocks, with each index dimension represented as an arm, as shown in Figure 38. The dimensionality of a tensor can be found by simply counting the number of free (dangling) arms. The top row of Figure 38 shows, from left to right, a vector, a matrix and an N-tensor. A tensor product (addition / contraction along a particular index dimension) is represented by connecting two tensor arms. Looking at Figure 38 diagrammatically, we can see that the matrix-vector product in the bottom left has one dangling arm, so the resulting product is a tensor, i.e., a vector, as expected. Similarly, the matrix-matrix product in the bottom right has two dangling arms, so the result is a matrix, as expected.

[0402] The tensor decomposition of the three-dimensional tensor T (hat) given by equation (69) can be illustrated as shown in the upper part of Figure 39. Here, a specific element T ijk (hat), we just fix the free index to the desired value and perform the necessary contraction.

[0403] Using this notation, we can explore the potential of tensor network factorizations for use in probabilistic modeling. The key idea is that the true joint distribution of high-dimensional PMFs is infeasible. We must approximate it, and we do so using tensor network factorizations. These tensor network factorizations can be learned to fit the training data. Not all tensor network factorizations are suitable; we may need to constrain the entries of the tensor network to be non-negative and sum to 1.

[0404] An example of such an approach is the use of Matrix Product State (MPS) (also known as tensor training). N ) into a tensor Let's say we want to model it as JPEG2024528208000158.jpg611.

number

[0405] A normalization constant is computed by summing over all possible states to ensure that the entries sum to 1. While it may be impractical to compute this normalization constant for a general N tensor, in the case of MPS, its linear nature allows it to be computed in O(N) time. "Linearity" in this case means that we can operate on the rows of a tensor training and perform tensor products one by one sequentially. (Both tensors and their tensor network approximations are multilinear functions.)

[0406] MPSs are very similar to Hidden Markov Models (HMMs), and in fact there is a correspondence: an MPS with correct entries corresponds exactly to an HMM.

[0407] Further examples of tensor network models are Born Machines and Locally Purified States (LPS), both inspired by quantum systems, which assume the Born rule, which states that the probability of an event X occurring is proportional to the square of its norm under the inner product <·,H·> with some positive (semi-)constant volume Hermitian operator H. This is a powerful probabilistic modelling framework, but has no obvious connection to graphical models.

[0408] The locally purified form (LPS) has the form shown in Figure 40. In LPS, the building blocks A k There is no restriction on the sign of the tensor, it can be positive or negative. In fact, A k can have complex values. In this case, JPEG2024528208000160.jpg54 is the tensor obtained by taking the complex conjugate of the entries of A. α k The dimension is sometimes called the bond dimension, and β k The plane is sometimes called a refined plane.

[0409] The elements of T are guaranteed to be positive by the fact that contraction along the purification dimension yields positive values ​​(for complex z, JPEG2024528208000161.jpg511). {i1,…,i N If we view {\displaystyle \mathbb {I}\) as one giant multi-exponential I, we can see that the LPS is the diagonal of a giant matrix (after collapsing all hidden dimensions), and evaluating the LPS is equivalent to an inner product operating on the state space.

[0410] Like MPS, computing the normalization constants of LPS is fast, taking O(N) time. The Born Machine is a special case of LPS, where the purification dimension is of size 1.

[0411] A tensor tree is another example of a tensor network. At the leaves of the tree, the dangling arms are contracted with data. However, the hidden dimensions are placed in the tree, and the nodes of the tree store tensors. The edges of the tree are the dimensions of the tensors that are contracted. A simple tensor tree is shown in Figure 41. The nodes of the tree store tensors, and the edges represent contractions between tensors. The leaves of the tree have indices that are contracted with data. Tensor trees can be used for multi-resolution and / or multi-scale modeling of probability distributions.

[0412] Each tensor node can have a refinement dimension added to it, such that it is contracted with the complex conjugate of that node, thus defining an inner product according to the Hermitian operator given by the tensor tree and its complex conjugate.

[0413] Another example of a tensor network is Projected Entangled Pair States (PEPS). In this tensor network, tensor nodes are arranged in a regular lattice and are contracted with their immediate neighbors. Each tensor has additional dangling arms (free indices) that are contracted with data (such as latent index values). In a certain sense, PEPS resembles Markov Random Fields and Ising models. A simple example of PEPS on a 2x2 image patch is shown in Figure 42.

[0414] Tensor network computations (such as computing the joint, conditional, or marginal probabilities of a PMF, or computing the entropy of a PMF) can be significantly simplified and sped up by putting the tensors into a canonical form, as explained in more detail below. All of the tensor networks mentioned above can be put into a canonical form.

[0415] Since the basis in which the hidden dimensions are represented is not fixed (so-called gauge freedom), one can simply change the basis in which these tensors are represented: for example, by putting a tensor network in canonical form, one can transform almost any tensor into an orthonormal (unitary) matrix.

[0416] This is possible by sequentially performing a series of decompositions on the tensors of the tensor network. These decompositions include the QR decomposition (and its variants RQ, QL and LQ), SVD decomposition, spectral decomposition (when available), Schur decomposition, QZ decomposition, Takagi decomposition, etc. The procedure for describing a tensor network in canonical form works by decomposing each of the tensors into an orthonormal (unitary) component and other factors. The other factors are contracted with their neighbors, modifying the neighbors. The same procedure is then applied to the neighbors and their neighbors, and so on, until all tensors except one are orthonormal (unitary).

[0417] The remaining tensors that are not orthonormal (unitary) are called core tensors. They are similar to the diagonal matrix of singular values ​​in an SVD decomposition and contain the spectral information of the tensor network. They can be used, for example, to compute the normalization constant of a tensor network, or the entropy of a tensor network.

[0418] Figure 43 shows an example of the procedure for converting an MPS to canonical form, starting from the top. Each core tensor is successively decomposed into a QR decomposition. The R tensor is contracted with the next tensor in the chain. This procedure is repeated until all but the core tensor C are in canonical form.

[0419] The use of tensor networks for probabilistic modeling in Al-based image and video compression is now described in more detail. As mentioned above, in an Al-based compression pipeline, an input image (or video) x is mapped to latent variables y via an encoding function (typically a neural network). The latent variables y are quantized to integer values ​​y (hat) using a quantization function Q. These quantized latent variables are converted to a bitstream using a lossless encoding method such as entropy encoding, as mentioned above. Arithmetic encoding or decoding is one example of such an encoding process and will be used as an example in the further discussion.

[0420] This lossless encoding process requires a probability model: an arithmetic encoder / decoder needs a probability mass function q(y(hat)) to convert integer values ​​into a bitstream. During decoding, a PMF is similarly used to convert the bitstream back into quantized latents, which are then passed through a decoder function (typically a neural network) to return the reconstructed image x(hat).

[0421] The size of the bitstream (compression rate) depends heavily on the quality of the probability (entropy) model: a better, more powerful probability model results in a smaller bitstream for the same quality of the reconstructed image.

[0422] Arithmetic encoders usually operate on one-dimensional PMFs. To accommodate this modeling constraint, the joint PMFs q(y(hat)) are usually assumed to be independent, and the pixel Each of the JPEG2024528208000162.jpg55 is a one-dimensional probability distribution. The joint density is then modeled as

number

[0423] In any case, essentially this modeling approach: JPEG2024528208000165.jpg assumes a 1-D distribution for each of the 55 pixels. This can be limiting. A better approach is to fully model the joint distribution. Then the 1-D distributions needed for the arithmetic encoder / decoder can be calculated as conditional probabilities when encoding or decoding the bitstream.

[0424] A tensor network can be used to model the joint distribution. This can be done as follows: Given the image JPEG2024528208000166.jpg630, each latent pixel is embedded (or lifted) into a high-dimensional space, where integers are represented by vectors that lie on the vertices of a probability simplex. For example, y i D possible integer values Let's say JPEG2024528208000167.jpg677. This embedding is Map JPEG2024528208000168.jpg55 to a D-dimensional one-hot vector, with ones in slots corresponding to integer values ​​and zeros everywhere else.

[0425] For example, each JPEG2024528208000169.jpg55 has values ​​of {-3,-2,-1,0,1,2,3}, Let's say the image is JPEG2024528208000170.jpg612. Then the embedding is The result is JPEG2024528208000171.jpg536.

[0426] Therefore, the embedding is JPEG2024528208000172.jpg649 JPEG2024528208000173.jpg633. In effect, this maps y (a hat) in M-dimensional space to D M Map to dimensional space.

[0427] Now, each of these terms in the embedding can be seen as a dimension that indexes a higher-dimensional tensor. Thus, the approach we take is to model the joint probability density via a tensor network T.

number

[0428] When encoding / decoding, arithmetic encoders / decoders cannot use joint probabilities. Instead, they must use one-dimensional distributions, which can be calculated using conditional probabilities.

[0429] Conveniently, the conditional probabilities can be computed easily by marginalizing the hidden variables, fixing the preconditioning variables, and normalizing, all of which can be easily done using tensor networks.

[0430] For example, suppose we are encoding / decoding in raster scan order. Then, for each pixel, we need the following conditional probabilities: JPEG2024528208000175.jpg565. Each of these conditional probabilities can be easily computed by collapsing the tensor network to the hidden variables, fixing the condition variable indices, and normalizing with an appropriate normalization constant.

[0431] This is a particularly fast procedure if the tensor network is in canonical form, since in this case contraction along the hidden dimension is equivalent to multiplication by unity.

[0432] Tensor networks can be applied to joint probabilistic modeling of PMFs across all latent pixels, or across patches of latent pixels, or to model the joint probabilities across channels of the latent representation, or any combination thereof.

[0433] Joint probability modeling with tensor networks can be easily incorporated into Al-based compression pipelines as follows: The tensor network is learned during end-to-end training and fixed after training. Alternatively, the tensor network or its components can be predicted by a hyperlatent network. The tensor network can additionally or alternatively be used for entropy encoding and decoding of hyperlatents in a hyperlatent network. In this case, the parameters of the tensor network used for entropy encoding and decoding of hyperlatents can be learned during end-to-end training and fixed after training.

[0434] For example, the hyperencoder network can predict the core tensors of the tensor network, patch-by-patch. In this scenario, the core tensors change across pixel patches, while the remaining tensors are learned and fixed across pixel patches. See, for example, Figure 44 showing an Al-based compression encoder with a tensor network predicted by a hyperencoder / hyperdecoder, and Figure 45 showing an Al-based compression decoder with a tensor network predicted by a hyperdecoder for the use of tensor networks in an Al-based compression pipeline. The features corresponding to those shown in Figures 1 and 2 can be assumed to be the same, as discussed above. In these examples, the residual ξ = y-μ is quantized, encoded, and decoded using a tensor network probability model. In this case, the parameters of the tensor network are T y In the examples of Figures 44 and 45, the quantized hyperlatent z is further expressed as T z The signal is encoded and decoded using a Consol network probability model with parameters expressed as:

[0435] Rather than using a hypervisor network to predict tensor network components (or possibly in combination with it), parts of the tensor network may be predicted using a context module that uses previously decoded latent pixels.

[0436] During training of the Al-based compression pipeline with tensor network probabilistic models, the tensor network can be trained for non-integer-valued latents (y instead of y(hat)=Q(y), where Q is the quantization function). To do so, the embedding function e can be defined with non-integer values. For example, the embedding function can be constructed with a tent function that takes values ​​of I at appropriate integer values, is zero at all other integer values, and linearly interpolates between them. This results in multi-linear interpolation. Other real-valued extensions to the embedding scheme can be used as long as they are consistent with the original embedding for integer-valued points.

[0437] The performance of a tensor network entropy model can be improved by some regularization during training. For example, entropy regularization can be used. In this case, we can calculate the entropy of the tensor network, H(q), and add or subtract a multiple of it to the training loss function. Note that the entropy of a normalized form of a tensor network can be easily calculated by calculating the entropy of its core tensors.

[0438] Hyper Hyper Network The current functionality and scope of use of the training technique for auxiliary hyper-hyper priors for use in, but not limited to, image and video data compression based on AI and deep learning is described below.

[0439] A commonly adopted network configuration in AI-based image and video compression is the autoencoder. It is built with an encoder module that converts the input data into a "latent" (y), which is an alternative representation of the input data and is often modeled as a set of pixels, and a decoder module that is intended to receive the set of latents and convert it back to the input data (or as close to it as possible). Due to the high-dimensional nature of the "latent space", where each latent pixel represents one dimension, a so-called "entropy model" p(y(hat)) is used to "fit" a parametric distribution into the latent space. The entropy model is used to convert y(hat) into a bitstream using a lossless arithmetic encoder. The parameters of the entropy model (the "entropy parameters") are learned inside the network. The entropy model can be learned directly or predicted by a hyperplier structure. A diagram of this structure can be seen in Figure 1.

[0440] The entropy parameter is most commonly (but not exclusively) constructed from a location parameter and a scale parameter (often expressed as a positive real value), although of course many more distribution types exist, both parametric and non-parametric, with many different types of parameters.

[0441] The hyperprior structure (Figure 2) introduces an additional set of "latencies" (z) through a set of transformations that predict the entropy parameter. Assuming that y is modeled with a Gaussian distribution, the hyperprior model can be defined as

number

[0442] A hyperplier can be added to a model that already has a trained hyperplier, sometimes called an auxiliary hyperplier. This technique is applied to improve the model's ability to capture low frequency features, particularly but not exclusively, to improve performance on "low rate" images. Low frequency features are present in an image when there are no abrupt color changes through its axis. Thus, an image with many low frequency features will have only one color throughout the image. The amount of low frequency features in an image can be found by extracting the power spectrum of the image.

[0443] The hyper-hyperplier can be trained jointly with the hyper-plier and the entropy parameter. However, using the hyper-hyperplier for all images can be computationally expensive. To give the network the ability to model low frequency features while maintaining performance on non-low rate images, we can employ an auxiliary hyper-plier that is used when the image meets a given characteristic, such as being low rate. An example of a low rate image is one with a bit per pixel (bpp) roughly less than 0.1. An example is shown in Algorithm 3.

[0444] The auxiliary hyper-hyperplier framework allows us to adjust the model only when necessary. Once trained, we can encode a flag into the bitstream indicating that a hyper-hyperplier is needed for this particular image. This approach can be generalized to infinite building blocks of entropy models such as hyper-hyperpliers.

[0445] [Table 4]

[0446] The most straightforward way to train a hyper-hyperplier is to "freeze" an existing trained hyper-hyperplier network, including the encoder and decoder, and optimize only the weights of the hyper-hyperplier modules. In this paper, "freezing" means that the weights of the frozen modules are not trained, and gradients are not accumulated to train the non-frozen modules. By freezing the existing entropy model, the hyper-hyperplier can modify the parameters of the hyper-hyperplier to bias towards low-rate images, such as μ and σ in the normal distribution case.

[0447] Using this training scheme provides several advantages. With fewer parameters, gradients need to be computed and stored, so training time and memory consumption scale better with image size. -Hyperpliers can be trained infinitely, and you can freeze a hyperplier and train another one on top of it.

[0448] One idea is to train the hyperplier network for N iterations first. Once N iterations are reached, we can freeze the entropy model and switch to the hyperplier when the image is low-rate. This will allow the hyperplier model to specialize on the images it is already good at, and the hyperplier will work as intended. Algorithm 4 shows the training scheme. This training can also be done by dividing the image into K blocks of size NxN, and then applying this scheme to the blocks, so that the image is trained only on low-frequency regions, if any.

[0449] Another possibility is to not wait for N iterations to start training the hyperplier, as shown in Algorithm 5.

[0450] [Table 5]

[0451] [Table 6]

[0452] There are various criteria to choose from for classifying an image as low-rate, such as using the rate calculated with a distribution chosen as a prior distribution, using the mean or median of the power spectrum of the image, or using the median or mean of the frequencies obtained by the Fast Fourier Transform.

[0453] Data augmentation can create sufficient data by increasing the number of samples that have the low frequency features associated with low rate images. There are various ways to modify an image. Upsample all images to a fixed size NxN. Use a constant upsampling factor. Randomly sample the upsampling coefficients from a distribution: uniform, Gaussian, Gumbel, Laplacian, Gaussian mixture, geometric, Student's t, nonparametric, chi 2 You can choose any distribution from the following distributions: beta distribution, gamma distribution, Pareto distribution, and Cauchy distribution. Use a smoothing filter on the image. Any filter works: average, weighted average, median, Gaussian, bilateral. Use the actual image size unless it is below a certain threshold N, then upsample to a fixed size and use a constant upsampling factor or sample the upsampling factor from a distribution as described above.

[0454] In addition to upsampling or blurring the image, random cropping may also be performed.

Claims

1. 1. A method of training one or more neural networks, the one or more neural networks being for use in encoding, transmitting, and decoding lossy images or videos, the method comprising: receiving an input image at a first computer system; encoding the input image using a first neural network to generate a latent representation; entropy encoding the latent representation; transmitting the entropy encoded latent representation to a second computer system; entropy decoding the entropy encoded latent representation; decoding the latent representation using a second neural network to generate an output image, the output image being an approximation of the input image; determining a quantity based on a difference between the output image and the input image; updating parameters of the first neural network and the second neural network based on the determined quantity; repeating the steps above with a first set of input images to generate a first trained neural network and a second trained neural network; the entropy decoding of the entropy coded latent representation is performed on a pixel-by-pixel basis; the pixel-by-pixel decoding order is additionally updated based on the determined amount. method.

2. The method of claim 1 , wherein the order of the pixel-by-pixel decoding is based on the latent representation.

3. The method of claim 1 , wherein the entropy decoding of the entropy encoded latent representation comprises operations based on previously decoded pixels.

4. The method of claim 1 , wherein determining the order of the pixel-by-pixel decoding comprises ordering the pixels of the latent representation in a directed acyclic graph.

5. The method of claim 1 , wherein determining the order of the pixel-by-pixel decoding comprises operating on the latent representation using multiple adjacency matrices.

6. The method of claim 1 , wherein determining the order of the pixel-by-pixel decoding comprises dividing the latent representation into a plurality of sub-images.

7. The method of claim 6 , wherein the multiple sub-images are obtained by convolving the latent representation with multiple binary mask kernels.

8. The method of claim 1 , wherein determining the order of the pixel-by-pixel decoding comprises ranking pixels of the latent representation based on a magnitude of a quantity associated with each pixel.

9. The method of claim 8, wherein the quantity associated with each pixel is a position or scale parameter associated with that pixel.

10. The method of claim 8, wherein the quantity associated with each pixel is additionally updated based on the evaluated difference.

11. The method of claim 1 , wherein determining the order of the pixel-by-pixel decoding comprises a wavelet decomposition of a plurality of pixels of the latent representation.

12. The method of claim 11 , wherein the order of the pixel-by-pixel decoding is based on frequency content of the wavelet decomposition associated with the plurality of pixels.

13. encoding the latent representation using a fourth trained neural network to generate a hyper-latent representation; transmitting the hyper-latent representation to the second computer system; decoding the hyper-latent representation using a fifth trained neural network, wherein the order of the pixel-by-pixel decoding is based on an output of the fifth trained neural network; The method of claim 1 .

14. 1. A method for lossy image and video encoding, transmission and decoding, said method comprising: receiving an input image at a first computer system; encoding the input image using a first trained neural network to generate a latent representation; entropy encoding the latent representation; transmitting the entropy encoded latent representation to a second computer system; entropy decoding the entropy encoded latent representation; decoding the latent representation using a second trained neural network to generate an output image, wherein the output image is an approximation of the input image; The first trained neural network and the second trained neural network are trained according to the method of any one of claims 1 to 13. method.

15. 1. A method for lossy image or video encoding and transmission, said method comprising: receiving an input image at a first computer system; encoding the input image using a first trained neural network to generate a latent representation; entropy encoding the latent representation; transmitting the entropy coded latent representation; The first trained neural network is trained according to the method of any one of claims 1 to 13. method.

16. 1. A method for receiving and decoding a lossy image or video, said method comprising: receiving, at a second computer system, the entropy-encoded latent representation transmitted according to the method of claim 15; decoding the latent representation using a second trained neural network to generate an output image, wherein the output image is an approximation of the input image; The second trained neural network is trained according to the method of any one of claims 1 to 13. method.

17. A data processing system configured to perform the method of any one of claims 1 to 13.

18. A data processing apparatus configured to perform the method of claim 15.

19. 16. A computer program comprising instructions that, when executed by a computer, cause the computer to carry out the method of claim 15.

20. A computer-readable storage medium containing instructions that, when executed by a computer, cause the computer to perform the method of claim 15.