Potential coding for end-to-end image / video compression

By considering channel importance, adopting post-conditional entropy coding, channel reordering and channel activity notification in end-to-end neural compression technology, the problem of high redundancy in latent entropy coding is solved, and more efficient image and video compression is achieved.

CN120019664APending Publication Date: 2025-05-16INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072236.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-12
Filing Date
2023-10-03
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing end-to-end neural compression technology has high redundancy in potential entropy coding, which affects compression performance.

Method used

By taking channel importance into account during the encoding process, using post-conditional entropy encoding, improving inter-channel correlation using channel reordering, and signaling channel activity on images/blocks, or performing RDOQ-like processes by optimizing major potentials for specific images.

Benefits of technology

It effectively reduces the redundancy in the quantized potential, improves the performance of latent entropy encoding, and improves the compression efficiency of images and videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019664A_ABST
    Figure CN120019664A_ABST
Patent Text Reader

Abstract

In end-to-end compression, a deep neural network-based encoder may be used to encode an image. An insert output from an encoder is quantized and encoded using a lossless encoder. Advantageously, at least one embodiment allows for improved potential entropy coding by further reducing redundancy in quantized potentials. To this end, at least one embodiment discloses taking into account channel importance by: encoding an indication (or meaning) of channel activity; performing post conditional entropy coding by calculating a conditional probability based on a subsequent context; using channel reordering to improve inter-channel correlation; or a process of the RDOQ class is performed by optimizing the primary potential for a particular image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of European Patent Application No. 22306547.5 filed on October 12, 2022, which is incorporated herein by reference in its entirety. Technical Field

[0002] The present embodiments generally relate to methods and apparatus for end-to-end neural compression of images and videos. Background Art

[0003] End-to-end compression is a compression technique where all components of the process are learned from the given data. End-to-end means: everything is learned from one end (the given data) to the other end (the compressed bitstream). Once the architecture is defined for end-to-end compression, there is no manual engineering work involved in designing the steps. End-to-end neural compression methods have not yet been standardized. Currently, MPEG is exploring these techniques. Summary of the invention

[0004] In end-to-end compression, a deep neural network based encoder may be used to encode an image. The embedding output from the encoder is quantized and encoded using a lossless encoder. Advantageously, at least one embodiment allows for improved latent entropy coding by further reducing redundancy in the quantized latent. To this end, at least one embodiment discloses taking into account channel importance by: encoding an indication (or meaning) of channel activity; performing post-conditional entropy coding by calculating conditional probabilities based on later context; using channel reordering to improve inter-channel correlation; signaling channel activity on an image / block, or performing an RDOQ-like process by optimizing the primary latent for a particular image.

[0005] One or more embodiments also provide a computer program including instructions, which, when executed by one or more processors, cause the one or more processors to perform an encoding method or a decoding method according to any one of the embodiments described herein. One or more of the present embodiments also provide a computer-readable storage medium having instructions stored thereon, the instructions being used for video encoding or decoding according to the methods described herein.

[0006] One or more embodiments also provide a computer-readable storage medium having stored thereon video data generated according to the methods described herein. One or more embodiments also provide a method and apparatus for transmitting or receiving video data generated according to the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 Illustrated is a block diagram of a system within which aspects of the present embodiments may be implemented.

[0008] Figure 2 A block diagram of a general embodiment of an end-to-end neural network-based video compression system is illustrated.

[0009] Figure 3 Illustrated is a training loop of an end-to-end neural network-based video compression system according to an embodiment.

[0010] Figure 4 The encoding process of an end-to-end neural network-based video compression system according to an embodiment is illustrated.

[0011] Figure 5 The decoding process of an end-to-end neural network-based video compression system according to an embodiment is illustrated.

[0012] Figure 6 Picture shows Figure 2 A structural block diagram of a general embodiment of the present invention.

[0013] Figure 7 Illustrated is a representation of a latent tensor in a neural network-based video compression system to which aspects of the present embodiments may be applied.

[0014] Figure 8 The encoding process of an end-to-end neural network-based video compression system according to an embodiment is illustrated.

[0015] Fig. 9 The encoding process of an end-to-end neural network-based video compression system according to an embodiment is illustrated.

[0016] Fig.10 Illustrated is a representation of a latent tensor in a neural network-based video compression system to which aspects of the present embodiments may be applied.

[0017] Fig.11 and 12 Illustrated is a representation of a receptive field of a latent tensor in accordance with at least one embodiment related to RDOQ.

[0018] Fig.13 Two remote devices communicating on a communication network are shown according to an example of the present principles in which various aspects of the embodiments may be implemented.

[0019] Fig.14 The syntax of a signal according to an example of the present principles is shown. DETAILED DESCRIPTION

[0020] Figure 1The block diagram of the example of the system in which various aspects and embodiments can be implemented is illustrated. System 100 can be embodied as the equipment including the various components described below, and is configured to perform one or more of the aspects described in this application. Examples of such equipment include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of system 100 can be embodied in a single integrated circuit, multiple ICs and / or discrete components singly or in combination. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, system 100 is coupled to other systems or to other electronic devices via, for example, a communication bus or by dedicated input and / or output ports. In various embodiments, system 100 is configured to implement one or more of the aspects described in this application.

[0021] The system 100 includes: at least one processor 110 configured to execute instructions loaded therein for implementing various aspects described, for example, in the present application. The processor 110 may include embedded memory, input-output interfaces, and various other circuits as known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). The system 100 includes: a storage device 140, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, a magnetic disk drive, and / or an optical disk drive. The storage device 140 may include an internal storage device, an attached storage device, and / or a network accessible storage device, as non-limiting examples.

[0022] The system 100 includes: an encoder / decoder module 130, which is configured to process data to provide encoded video or decoded video, for example, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents a module (one or more) that can be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of the encoding and decoding modules. In addition, the encoder / decoder module 130 may be implemented as a separate element of the system 100, or may be incorporated into the processor 110 as a combination of hardware and software as known to those skilled in the art.

[0023] Program code to be loaded onto the processor 110 or the encoder / decoder 130 to perform various aspects described in this application may be stored in the storage device 140 and subsequently loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of the various items during the execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0024] In several embodiments, memory internal to the processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be a memory 120 and / or a storage device 140, such as a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations (e.g., for MPEG-2, MPEG-4, HEVC, or VVC).

[0025] Input to the elements of system 100 may be provided through various input devices as indicated in block 105. Such input devices include, but are not limited to: (i) an RF section that receives an RF signal transmitted over the air, such as by a broadcaster; (ii) a composite input terminal; (iii) a USB input terminal; and / or (iv) an HDMI input terminal.

[0026] In various embodiments, the input device of block 105 has associated corresponding input processing elements as known in the art. For example, the RF section may be associated with elements suitable for the following operations: (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band); (ii) down-converting the selected signal; (iii) band-limiting again to a narrower frequency band to select a signal frequency band that may be referred to as a channel in some embodiments; (iv) demodulating the down-converted and band-limited signal; (v) performing error correction; and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments includes one or more elements to perform these functions, such as a frequency selector, a signal selector, a frequency band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include: a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a near-baseband frequency) or to baseband. In a set-top box embodiment, the RF part and its associated input processing element receive the RF signal transmitted on wired (for example, cable) medium, and filter to the desired frequency band again by filtering, down-conversion and perform frequency selection.Various embodiments rearrange the order of (and other) elements described above, remove some in these elements, and / or add other elements that perform similar or different functions.Adding element can include inserting element between existing element, for example inserting amplifier and analog-to-digital converter.In various embodiments, the RF part comprises antenna.

[0027] In addition, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 100 to other electronic devices across the USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, can be implemented, for example, in a separate input processing IC or in the processor 110 when necessary. Similarly, aspects of USB or HDMI interface processing can be implemented in a separate interface IC or in the processor 110 when necessary. The demodulated, error-corrected and demultiplexed stream is provided to various processing elements, including, for example, a processor 110 and an encoder / decoder 130 that operate in conjunction with memory and storage elements to process data streams for presentation on an output device when necessary.

[0028] The various elements of the system 100 may be provided within an integrated housing. Within the integrated housing, the various elements may be interconnected and data transmitted therebetween using a suitable connection arrangement 115, such as an internal bus as known in the art, including an I2C bus, wiring, and printed circuit boards.

[0029] The system 100 includes a communication interface 150 that implements communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data on the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or a network card, and the communication channel 190 may be implemented in, for example, a wired and / or wireless medium.

[0030] In various embodiments, data is streamed to the system 100 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments are received on the communication channel 190 and the communication interface 150 adapted for Wi-Fi communication. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box that delivers data on the HDMI connection of the input box 105 to provide streaming data to the system 100. Still other embodiments use the RF connection of the input box 105 to provide streaming data to the system 100.

[0031] The system 100 can provide output signals to various output devices, including a display 165, a speaker 175, and other peripherals 185. In various examples of embodiments, other peripherals 185 include one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide functions based on the output of the system 100. In various embodiments, signaling such as AV is used to transmit control signals between the system 100 and the display 165, the speaker 175, or other peripherals 185. Links, CEC, or other communication protocols for device-to-device control are implemented with or without user intervention. The output device can be communicatively coupled to the system 100 via a dedicated connection through the corresponding interfaces 160, 170, and 180. Alternatively, the output device can be connected to the system 100 using a communication channel 190 via the communication interface 150. The display 165 and the speaker 175 can be integrated in a single unit with other components of the system 100 in an electronic device (e.g., a television). In various embodiments, the display interface 160 includes a display driver, such as a timing controller (TCon) chip.

[0032] Display 165 and speaker 175 may alternatively be separate from one or more of the other components, such as if the RF portion of input 105 is part of a separate set-top box. In various embodiments where display 165 and speaker 175 are external components, output signals may be provided via dedicated output connections including, for example, an HDMI port, a USB port, or a COMP output.

[0033] Figure 2 A block diagram of an embodiment of an end-to-end neural network-based video compression system is illustrated. Figure 2 A variant embodiment of is a high-level simplified version of an end-to-end NN. The input X to the encoder part of the network can consist of images or frames of a video, parts of images, tensors representing groups of images, tensors representing parts (crops) of groups of images.

[0034] In each case, the input may have one or more components, for example: monochrome, RGB or YCbCr components. In a first step, the input X is fed into the encoder network 210 which applies a function g also known as an analysis transform a (). g a () is typically a sequence of convolutional layers with activation functions. The stride in the convolution or space-to-depth operation can be used to reduce the spatial resolution while increasing the number of channels. The encoder network 210 can be viewed as a learned transformation. In the next step, the output of the transformation is analyzed: Y = g a (X), a 3-way array or 3-dimensional tensor (referred to as a tensor) of latent variables or latent representations Y, is quantized (Q) and entropy encoded (EC) into a binary stream (bitstream) for storage or transmission. For the convenience of notation in this embodiment, it is assumed that EC is embedded in the quantization operation Q. The bitstream is then entropy decoded (ED) to obtain A quantized version of Y. Applying the function g, also known as the synthesis transform s The decoder network 230 of () generates the reconstructed input: According to the quantified potential Approximation to the original X. g s () is typically a sequence of upsampling convolutions (e.g., "deconvolution" or convolution followed by an upsampling filter) or depth-to-spatial operations. The decoder network can be viewed as a learned inverse transform or a denoising and generative transform.

[0035] The encoder network typically consists of a sequence of convolutional layers with strides, allowing the spatial resolution of the input to be reduced while increasing the depth (i.e., the number of channels of the input). Pooling (e.g., average pooling, max pooling, etc.) or squeezing operations (via reshaping and arranging space to depth, where, for example, a tensor of size (N, H, W) is reshaped and arranged into a tensor of size (N*2*2, H / / 2, W / / 2)) can also be used to replace the strided convolutional layers. The encoder network can be viewed as a learned transformation. The output of most analyses, which are in the form of 3-way arrays called 3-D tensors, are called tensors of latent representations or latent variables. From a broader perspective, a set of latent variables constructs a latent space, which is also frequently used in the context of end-to-end compression based on neural networks. The latent representation is quantized and entropy encoded for storage / transmission, in Figure 2 The bitstream is a collection of coded syntax elements and payloads representing bins of quantized symbols that are transmitted to a decoder.

[0036] The decoder first decodes the ED quantized symbols from the bitstream. The decoded latent representation is then transformed into pixels for output through a set of layers typically consisting of (de)convolutional layers (or depth-to-space squeezing operations). The decoder network is thus a learned inverse transform operation on the quantized coefficients. The output of the decoder is a reconstructed image or group of images.

[0037] According to another embodiment, there is a more sophisticated end-to-end neural network-based video compression system. For example, a "super autoencoder" (super prior) can be added to the network to jointly learn a parameterized distribution of latent representations as the output of the encoder. Figure 3-5 Another embodiment of an end-to-end neural network-based video compression system including a super prior is illustrated and will be described below. Therefore, the present principles are not limited to the use of autoencoders. Any end-to-end differentiable codec can be considered.

[0038] Figure 3 The figure shows the training phase of a sophisticated end-to-end compression system. The input image to be compressed x∈R n×n×3 First, the depth encoder (210) uses y=g a (x; φ) process, where φ is the function g to be optimized during the training phase a The output of the encoder y∈R m×m×ois called the principal embedding (or principal latent) of the image. Without loss of generality, we assume here that the image is square (n×n) and the principal embedding is also square (m×m). However, they do not have to be square, but can exist in any shape. The input image can be fed to the encoder at the image level, or can be partitioned into image regions, with the individual image regions as input to the encoder.

[0039] Then, y goes to the quantizer (360) as to get the main code of the image, followed by the dequantization block (365) to obtain the main potential substances after reconstruction Prior art neural models use a hyper-prior entropy model. Specifically, the auxiliary (side) embedding (or auxiliary latent) z∈R k×k×f By another deep neural network (320) through z = h a (y; Φ) to learn, where Φ is the function h to be optimized during the training phase a The trainable parameters of , and the auxiliary embedding is given by Quantization (370) to obtain auxiliary code Followed by dequantization (375) to obtain the reconstructed auxiliary insert These For learning Typically, can be modeled by a Gaussian distribution, where the parameters are obtained by another deep network (340), such as Decompressed image The depth decoder (330) uses get.

[0040] The values ​​of m and k depend on the architecture being defined, and m is typically 1 / 16 of the spatial resolution of the image, and k is typically 1 / 8 of the size of the embedding. The value of o is fixed at 128, for example in the most common architectures. In one example, if the image size is 256×256×3, then typically y will be 16×16×128 and z will be 2×2×128.

[0041] In this model, the auxiliary code The lower bound of the bit length of is calculated by the decomposition entropy model (350). The model accepts As input and learn auxiliary code The probability density function (PDF) of the code is given by . Using this PDF, the model calculates the probability mass function (PMF) of each code under a certain quantization method. The PMF value is sufficient to calculate a lower bound on the bit length of the auxiliary code. On the other hand, the primary code The lower bound of the bit length of is calculated by the hyper-prior entropy model. This model is usually implemented by a Gaussian distribution. Basically, Each of the reconstructed main potential should follow a Gaussian distribution with parameters (μ i ,σ i ) is obtained in the previous step by the auxiliary information. Therefore, the PMF of these Gaussians under the determined quantization method is sufficient to calculate the lower bound of the bit length of the primary code.

[0042] In this setup, the depth encoder (g a (.;φ)), deep decoder (g s (.;θ)), deep hyper-prior encoder (h a (.; Φ)), deep super prior decoder (h s (.; Θ)) and the decomposition entropy model p f (.; ω) consists of multiple neural layers (such as convolutional layers). Each neural layer can be described as a function that first multiplies the input by a tensor, adds a vector called a bias, and then applies a nonlinear function on the resulting value. The shape of the tensor (and other properties) and the type of nonlinear function are called the architecture of the network. We will refer to the values ​​of tensors and biases by the term "weights". The parameters of weights and (if applicable) nonlinear functions are called parameters. The architecture and parameters define a "model". As described above, we refer to their parameters by φ, θ, Φ, Θ, and ω.

[0043] Many end-to-end architectures have been proposed recently. Typically, they are even faster than Figure 3 The ones illustrated in are more complex, but they all retain a deep encoder and decoder. State-of-the-art models can compete with traditional codecs in terms of rate-distortion tradeoff. The model must be trained on a massive database D of images to learn the weights of the encoder, decoder, and entropy model. Typically, the weights are optimized to minimize the training loss: where d(.,.)(315) is a measure of the distortion between the original image and the reconstructed image (e.g., mean square error). The rate term (355, 345) is a lower bound on the bit length of the side information The lower bound on the bit length of the main information The hyperparameter λ controls the tradeoff between the rate term (r) and the distortion term (d). Note here that the rate is based on a lower bound on the bit length of the main information and the side information, but other methods can be used to estimate the rate.

[0044] When the model is trained and the optimal parameters (φ, θ, ω, Φ, Θ) are obtained, Figure 4 Neutralization Figure 5The model is deployed in the encoding and decoding devices shown in the figure. This phase is usually called the testing phase. In the testing phase, both the encoding and decoding devices need the PMF table. In addition to the PMF table, the encoding and decoding devices must also agree on which quantization method to use.

[0045] Figure 4 The encoding process of an end-to-end neural network-based video compression system according to a specific embodiment is illustrated. Figure 4 As illustrated in , in the encoding device, the primary code and the auxiliary code are obtained (410, 420) from a given image in the same way as in the training phase. Later, the encoding device converts the auxiliary code into a bitstream by a lossless encoder, such as an arithmetic encoder AE (490) driven by the learned PMF table provided by the decomposition entropy model (450). In the next step, the auxiliary code obtained from the quantization (470) is dequantized (475) and a reconstructed auxiliary potential is obtained. The super prior entropy decoder (440) decodes the reconstructed auxiliary potential in order to find the distribution of the primary code. Finally, the primary code obtained from the quantization (460) is converted into a bitstream by a lossless encoder, such as an arithmetic encoder AE (480) using these distributions. Finally, the two bitstreams are concatenated using a pointer in between.

[0046] Figure 5 The decoding process of an end-to-end neural network-based video compression system according to a specific embodiment is illustrated. Figure 5 As illustrated in , in the decoding device, the process starts with decoding the auxiliary codes from the bitstream, which can be done by a lossless decoder, such as an arithmetic decoder AD (570), by using the learned PMF table provided by the decomposition entropy model (560). In the next step, the auxiliary codes are dequantized (550) and the reconstructed auxiliary potential is obtained. The super prior entropy decoder (540) accepts the reconstructed auxiliary potential as input in order to find the distribution of the primary codes. These distributions inform the AD (530) how to read the primary codes from the bitstream. Dequantization (520) is then applied on the primary codes to reconstruct the primary potential. The decoding device can obtain the reconstructed image by feeding the reconstructed primary potential to the depth decoder (510).

[0047] Despite the fact that there is no restriction on the input format of the autoencoder as stated earlier, most existing schemes take the entire image or frame as input to transform into a latent representation Y, such as Figure 2-5. In this case, the latent representation represents a 3-dimensional tensor in the latent space by transforming the input image through a (non-)linear transformation using multiple convolutional layers followed by activations. This implies that the spatial redundancy is only factored through the learned transformation operations, which not only limits the coding efficiency but also the application of compression. Moreover, although we mentioned the example of image compression, the present principles are applicable to any compression model that derives and encodes a latent representation from input content (such as motion fields, depth maps, 3D scenes, etc.).

[0048] In the current approach, this artificial neural network (ANN) is constructed using several types of losses (or Figure 3 The loss is trained on the distortion 315 on the image. In a variant, the loss can be based on an "objective" metric (typically, mean squared error (MSE)) or on structural similarity (SSIM). The result may not be as good perceptually as the second type, but the fidelity to the original signal (image) is higher. In another variant, the loss can be based on "subjective" (or subjective by proxy), typically using a generative adversarial network (GAN) or a high-level visual metric via a proxy NN during the training phase.

[0049] In a related problem, such ANN models are trained using several types of training sets. The same network can be first trained on a general training set, allowing satisfactory performance on a wide range of content types, and then it is possible to fine-tune the model using a specific training set for a specific use, thereby improving performance on domain-specific content.

[0050] Figure 6 Picture shows Figure 2 The structural block diagram of the NN autoencoder of the general embodiment of Figure 3 As described by the hyper-prior on , the input image is fed into the encoder, which consists of 3 convolutional layers, each performing 128 3×32D convolutions (with n=128 output channels), downsampling (denoted by / 2), followed by an activation layer (e.g., ReLU or generalized divisive normalization (GDN) etc.). The output of the last layer of the encoder is called the "latent representation" or "latent". The coefficients of the latent are then quantized. The quantized coefficients are (losslessly) entropy encoded to form the payload of the bitstream. At the decoder side, 2D deconvolution is performed to reconstruct the image using either a transposed convolution or a classic upscaling (denoted by ×2) operator followed by a convolution.

[0051] exist Figure 3-5 In an embodiment of the invention, an additional NN called a hyper-prior is used to improve the entropy coding module S by predicting the probability distribution of each parameter of the latent. These methods allow to compare with the above description of Figure 2 and Figure 6 The distribution is better predicted than the general decomposition prior model described in the previous section, in which all potential coefficients for a given channel share the same prior. Another improvement is to use an autoregressive model to predict the prior for a given coefficient using the distribution parameters of previously encoded neighboring potential coefficients, typically by using masked convolutions. Both methods can be combined to further improve the entropy coding of the latents.

[0052] However, in practice, most methods use separate models for each RD (rate-distortion) tradeoff (corresponding to different lambda λ in loss optimization): in: -x is the input image - is the reconstructed image -d(.,.) is the distortion measure, typically MSE -y is the latent (output of the encoder) - It is a quantified potential -Φ is the distribution parameter.

[0053] Therefore, depending on the model architecture and the target lambda λ, the estimated distribution Φ can vary widely from channel to channel. While some channels have high entropy, other channels may not carry any information, i.e., they may consist of all-zero coefficients. Moreover, the main drawbacks of super-prior methods are that they involve additional computations due to the additional NN running in parallel with the main NN, and they incur additional delays in the encoding / decoding process since the decoding of the main latent cannot start before the decoding of the super latent is finished. Considering an embodiment of an autoregressive encoder, it adds even more delay and destroys the parallelism of the scheme due to the causal relationship between the distribution parameters of the latent coefficients, requiring sequential encoding / decoding. Finally, during inference for a specific image, RD optimality is an approximation from the parameters found during the training phase.

[0054] Therefore, there is a need to improve the entropy coding of latent tensors.

[0055] At least some embodiments relate to methods for entropy encoding / decoding latent tensors that improve the compression performance of existing autoencoders without requiring retraining of their parameters. At least some embodiments relate to the retraining process and also present loss improvements. At least some disclosed embodiments allow for further reduction of residual redundancy in the quantized latent.

[0056] Figure 7The representation of a latent tensor for y in a neural network-based video compression system to which aspects of the present embodiment may be applied is illustrated. In most neural network-based compression frameworks, the latent representation y is formed in a 3-dimensional tensor (referred to as a latent tensor or latent). Figure 7 A potential object y∈R of size (m×n×p) is shown m×n×p Example: m×n depends on the input image size, while the number of channels p is a property of the encoder. The encoder has dimension R m×n×p (also denoted as R^(m×n×p)), where m, n, and p are positive integers. Figure 7 The indices i, j, c indicate the dimensions R m×n×p In addition, for each channel c in the range p and for R m×n For each value v at (i, j) in , the encoder obtains the probability distribution D at (i, j, c) and entropy encodes the value v using H = E(v, D).

[0057] At the decoder side, for dimension R m×n For each channel c and for each position (i, j) in channel c, the decoder obtains the current probability distribution D at (i, j, c) and uses the probability distribution D from the bitstreams b and E -1 The value v is entropy decoded by using the symbols resolved from (b, D). For a fully decomposed prior model, the distribution depends only on channel c, so D(i, j, c) = D(c). The distribution can be stored as a parameter function or a complete CDF (cumulative distribution function) and shared between the encoder and decoder. When using a super prior and / or autoregressive model, the parameters of the distribution function are updated for each value of the potential at (i, j, c). The default probability distribution of the channel is either retrieved directly from the training phase or recalculated using a large data set, where for each sample in the large data set, the sample is encoded. The resulting potential is retrieved, and the values ​​in the potential are accumulated in a histogram for each channel of the potential. At the end, the distribution is a normalized histogram for each channel. Alternatively, from the distribution, distribution parameters can be extracted (e.g., assuming a Gaussian distribution).

[0058] At least some embodiments relate to methods for entropy encoding / decoding potentials based on their probability distribution. Advantageously, the potential entropy encoding is improved by further reducing the redundancy in the quantized potentials. To this end, at least one embodiment is described that: - Taking channel importance into account by encoding an indication (or meaning) of channel activity; - performing post-conditional entropy coding by computing conditional probabilities based on later context; -Use channel reordering to improve inter-channel conditional entropy coding; - Use signaling on images / blocks; -Performs an RDOQ-like process by optimizing the primary latents for a specific image.

[0059] In particular, the proposed mechanism and syntax improve the compression performance of existing autoencoders without requiring retraining of their parameters. Additionally, training and loss improvements are presented.

[0060] Figure 8 A block diagram of a generalized embodiment of an entropy coding method in an end-to-end neural network-based video compression scheme is illustrated. In step 810, a latent representation (Y) associated with image data using a neural network is obtained. The latent representation Y is formed of a 3-dimensional tensor (referred to as a latent tensor). According to Figure 7 In the variant illustrated above, the potential is formed by a number p of channels of two-dimensional data m×n. In the following, the spatial position in the 2D space is indicated by the reference index (I, j) and the channel position is indicated by the index c. In step 820, the distribution according to any of the variants described above is also used as input for entropy coding. A probability distribution D(i, j, c) is determined for each value (v) of the potential, and in the case of fully decomposed data, the distribution depends only on the channel D(c). Then, in step 830, the potential is quantized and entropy encoded based on the probability distribution of the potential to generate a as a binary stream (bit stream). According to various embodiments, entropy coding further includes at least one of the following: obtaining an indication of the activity of the channel, obtaining a conditional probability distribution of the potential, reordering the channels based on inter-channel correlation, or performing rate-distortion optimization on the potential.

[0061] Fig. 9 A block diagram of a general embodiment of an entropy decoding method in an end-to-end neural network-based video compression scheme is illustrated. The entropy decoding method mirrors the entropy encoding described above. In step 910. A bitstream is obtained, which contains NN-based encoded data representing a potential (y) associated with image data. The potential representation Y is formed by a 3-dimensional tensor (called a potential tensor), for example, including a number of channels p of two-dimensional data m×n. In step 920, a distribution according to any of the variants described above is also used as input to entropy decoding. Then, in step 930, the potential is dequantized and entropy decoded based on the probability distribution, allowing the image data to be reconstructed using NN-based decoding.

[0062] According to a first embodiment, an indication of the activity of a channel is signaled.

[0063] Since not all channels are necessarily activated for specific content, in a first embodiment, it is proposed to add information in the bitstream to indicate when a specific channel c in R^(m×n×p) has a value that needs to be decoded (i.e., the current value is different from the most likely value of the distribution (e.g., the mean in the case of a Gaussian distribution)).

[0064] Accordingly, in the encoder, a comparison between the current value v at (i, j, c) and the most probable value of the distribution D(i, j, c) or D(c) is performed. In case at least one current value is different from the most probable value of the distribution, the activity for the channel is positive (set to 1) and the channel is effectively encoded. In case all values ​​of channel c are equal to the most probable value of the distribution, the activity A[c] is set to 0. The encoding of the channel is skipped. Furthermore, in order to signal to the decoder that channel c has the most probable value, a channel activity indication is encoded.

[0065] Advantageously, the proposed activity table can be entropy encoded using a priori probabilities on the activity of a particular channel. This probability can be calculated using a large data set to derive statistics on the distribution of potentials. This can be done using a data set for training, where for each sample in the data set, the sample is encoded and a potential is generated. For each channel of the potential, if any value in the channel is different from the most likely value of the channel distribution, the channel is marked as activated. Finally, for the potential data, the channel activity probability is calculated as the average probability over the data set.

[0066] At the decoder, the activity A[c] of the channel is decoded. When the decoded activity A[c] of the channel is positive ("1"), it indicates that the channel contains an encoded value and the channel is entropy decoded. When the channel activity A[c] is zero ("0"), the channel is initialized to the most likely value for the channel. In a variant embodiment, the channels of the potentials are first sorted by activity, as detailed in the section related to channel reordering, from the most active channel to the least active channel. Instead of indicating the activity for each channel, the index of the last active channel is transmitted. According to yet another variant, this index is also entropy encoded by using the probability calculated as before on the training set.

[0067] According to a second embodiment, the conditional probability distribution of the potentials is signaled.

[0068] Fig.10 Illustrated is a representation of a latent tensor in a neural network-based video compression system to which aspects of the present embodiments may be applied.

[0069] For a given model using a simple (non-autoregressive) probability model, the conditional probability for each value of a particular channel is calculated. The context is calculated for each value v at (i, j, c) depending on the causal neighbors. Fig.10 As shown above, in a variant, a context (ctx) of a potential value (v) includes at least one causal spatial neighbor value (T,L) in the same channel. In another variant, the at least one context (k) of the potential value (v) further includes at least one causal inter-channel neighbor value (P). In an embodiment, the context is selected as the number of values ​​in the causal neighborhood that are above a threshold f. Fig.10 In , we call T (top) and L (left) the values ​​of the latent objects in the decoded channel, and the context ctx(f,T,L) is calculated as: If |T-μ|>f and |L-μ|>f: t=2 Otherwise if |T-μ|>f or |L-μ|>f: t=1 Otherwise t=0 where μ is the distribution mean or the most likely value for a channel. In this example, there are 3 contexts. In a variation, a context may be identified by an index k. For each context ctx, a probability distribution is associated, where D(i,j,c)=cdf ctx =p(v|t). In a variant, the conditional probability distribution can be calculated based on offline training on any dataset. Advantageously, the optimal threshold for each of the context modeling for a specific channel c can be calculated offline using an optimization function: f * =arg f min∑H(v,cdf ctx(f,T(v),L(v)) ) where H is the value of the cdf k The entropy of the value v of , and k is the context index. The context index is calculated as k = ctx(f, T(v), L(v)), where f is the threshold, T(v) and L(v) are the values ​​of the top and left values ​​of the current value v (values ​​outside the potential tensor are considered to be 0 or any other arbitrary known value).

[0070] In another variant embodiment, the context also uses causal inter-channel values. For example, the coefficient P at the same spatial position (i, j) in the previously decoded channel is used to calculate the context: ctx(f,T,L,P)=(|T-μ|>f)+(|L-μ|>f)+(|P-μ′|>f) Where μ' is the distribution mean or the most likely value for the channel containing P. In yet another variation, the threshold is different depending on the neighbor position. In another embodiment, the use of the optimal conditional probability model for a given image or region of an image is calculated at the encoder side and signaled in the bitstream. The optimal conditional probability model is determined based on the rate in any of the disclosed models. At the decoder side, the same model is used to entropy decode the value of the potential.

[0071] According to a third embodiment, the channels of the potentials are reordered before encoding.

[0072] In the previously described embodiment, the inter-channel correlation coding was performed using the original order of the channels in the latent tensor. However, since this order is not necessarily optimal for entropy coding, it is proposed in this variant embodiment to reorder the channels using a training dataset. This order is then fixed in the codec and known from the encoder and decoder.

[0073] The new order of channel encoding / decoding is then fixed for a particular encoder / decoder. In another variant, a deep encoder or deep decoder may be adapted to generate / obtain a potential with a new channel order as input, and accordingly, the weights of the last layer of the encoder and the first layer of the decoder may be adapted to take into account the new order for encoding. In one embodiment, the encoding order of the channels is selected to maximize the correlation between consecutive channels. For example, the first channel in the reordered tensor is the channel with the highest average energy averaged over the data set (e.g., which can be calculated as the variance of the values ​​of the coefficients of the channel). Then iteratively, the next channel is selected as the channel with the highest correlation with the previous channel. For example, using the above optimization function that calculates the optimal threshold f for the context of the current value, the correlation between channels n and m is calculated as:

[0074] According to another variant embodiment, the channels may be ordered by decreasing average energy, wherein energy is defined as the sum of the squares of the values ​​in a channel.

[0075] According to yet another variant embodiment, the channels may be sorted by decreasing average activity, wherein activity is defined as the sum of values ​​greater than a threshold f.

[0076] According to yet another variant embodiment, the channels are sorted by iteratively searching for the order that minimizes the coding cost using the context defined in the previous embodiment.

[0077] According to a fourth embodiment, a syntax for implementing signaling channel activity is disclosed.

[0078] According to different variants, channel activity signaling can be done at frame or block / tile level. The following table shows an example of a header for the encoding of a potential object: where the signaled width and height are the original image width and height, P is the number of channels of the latent at the decoder input, and c is the channel index in the latent, and the decimation factor is the spatial decimation performed by the encoder on the original image to obtain a latent space dimension m×n.

[0079] To restore the input latent width and height, apply padding on the original image, for example: m = (height_minus_one + 1 + filling - 1) / / filling Where padding = 2^decimation_factor, and / / indicates integer division. This means that the original image is padded, for example, the new padded image height is: H' = m*filling.

[0080] Typically, padding can be done by centering the image and filling the outside with 0 values. In a variant, the original image is in the upper left. After reconstruction, the output image is cropped according to the input image padding strategy.

[0081] The potential value for the inactive channel is considered to be the most likely potential value in the cdf distribution.

[0082] In a variation, block level signaling is performed. In yet another variation, the active channel for each block uses a conditional probability that depends on the activity of the last coded block, the first block using the same coding as above.

[0083] According to a fifth embodiment, rate-distortion optimization of potentials is performed.

[0084] In conventional codecs, a well-known method for optimizing the coding of coefficients of transformed and quantized residuals is rate-distortion optimized quantization (RDOQ). The correlation between the modification of the quantization of the coefficients and the modification of the distortion is used. It allows slightly changing the quantized coefficients for a specific residual in order to optimize the rate-distortion RD metric: C=R+λD, where R is a measure of the rate of the residual, λ is a Lagrange multiplier, and D is a measure of the distortion.

[0085] To adapt RDOQ to the encoding of the latent, an iterative process is used. In a first step, a portion of the image is encoded using the NN encoder part of the autoencoder. It produces the main latent Y. Then, for each channel c in the latent Y, and for each value v (or coefficient) at the coordinates (i, j) in the channel c, the value is modified by adding an offset value, resulting in a modified latent. Then, the image is reconstructed from the modified latent, and the distortion D' with the original image is calculated, as well as the new cost C' = R + λD'. The new cost is then compared with the previous cost C. If the new cost C' is less than the previous cost C, the modified value is kept in the latent. The process is performed iteratively for each value in the latent. Since only a change in C is required, according to a variant, instead of calculating the full rate of the latent, only the rate induced by the modified value is used. Advantageously, the complexity of the process is reduced.

[0086] Fig.11 1 illustrates a representation of a receptive field of a latent tensor according to at least one embodiment related to RDOQ. Fig.11 In , for a given autodecoder, the coefficient of the latent at coordinate (i,j) in space corresponds to coordinates (S*i+S / 2,S*j+S / 2) in the output, where S is the upscaling factor between the spatial latent dimensions and the reconstructed image (corresponding to the decimation factor mentioned earlier). For the receptive field F, the region of size (2F+1)×(2F+1) centered on M' corresponds to the samples in the output image that are modified when the value at M in the latent is modified.

[0087] According to a particular embodiment, the RDOQ process is optimized by decoding only the portion of the latent comprising the modified values, depending on the receptive field of the decoder. Advantageously, the process is accelerated. That is, instead of running the decoder on the full latent, only the portion of the latent comprising the modified values ​​(or coefficients) is used. Indeed, given the modified coefficients in the latent, only a portion of the reconstructed image is modified, depending on the receptive field of the decoder.

[0088] Fig.12 Another representation of the receptive field of a latent tensor according to at least one embodiment related to RDOQ is illustrated, where only part of the latent is decoded. For a given decoder with a receptive field F and an upscaling factor S (i.e., the final image size is S times larger spatially than the input latent), the size of the latent for cropping inside the original latent is given by: where ceil(x) is a function that computes the smallest integer greater than or equal to x. The clipped latent then has size (2*dL+1)×(2*dL+1)×p, centered spatially around the current value at coordinates (i,j,p), the same as in the original latent.

[0089] When using a cropped latent, only the distortion of the resulting patch A of size (2F+1)×(2F+1) at the center of the reconstructed image is used, as Fig.12 As previously for rate, this approximation does not challenge the RDOQ process since only the difference in distortion between the modified reconstructed image and the previously reconstructed image is involved in the RDOQ comparison.

[0090] In yet another variation, to further speed up the process, the approximation is done by reducing the theoretical receptive field by a certain factor. For example, instead of using F in the above process, F / 2 or F / 3 is used as an approximation.

[0091] According to a specific embodiment, the RDOQ process is accelerated using parallel processing. In a variation, each channel is processed independently and in a parallel process. Whenever the value of a potential is updated, all parallel processes update their potentials. Those skilled in the art will note that this process becomes non-deterministic because the parallel processing order is not controlled. The updated values ​​are still shared between the processes.

[0092] In another variation, parallel processing is done on the same channel by spatially splitting the latents into several parts.

[0093] According to a particular embodiment, the parallel RDOQ is iterated with an approximation of the reduced receptive field. To improve the results of the parallel processing, the entire process is iterated several times to converge towards a better potential value. For example, in multithreading, at each pass, the entire potential is modified. When all processes are completed, a second pass is started on the potential. Further passes can be completed until convergence is reached, for example when the modification of the total cost C is below a given value.

[0094] The following table shows an example of the tradeoffs of the presented approach: method PSNR(dB) Rate(bpp) RDOQ Time original 32.6034 0.315362 0mn 1 pass, accurate receptive field, no threading 32.7609 0.309789 100mn 1 pass, accurate receptive field, multithreading 32.7731 0.3134 7mn 1 pass, approximate receptive field, multithreading 32.7922 0.311289 3mn 3 passes, approximate receptive field, multithreading 32.8223 0.306403 7mn .

[0095] According to certain embodiments, additional heuristics may be used in order to speed up or improve the performance of parallel RDOQ. According to a non-limiting example, if available, the gradient of the output with respect to the input potential in the decoder may be used to guide the potential coefficients to be updated in the process. Thus, the change in distortion with respect to the change in the potential may be calculated, and the optimal potential update with respect to the gradient may be calculated by minimizing the cost C. In another example, to improve the rate first, a step size is selected for each value of the potential, such as it reduces the rate. The cost C is then calculated, and if the cost is less than the original cost, the value is updated.

[0096] Fig.13 Two remote devices communicating on a communication network according to an example of the present principles in which various aspects of the embodiments can be implemented are shown. According to the example of the present principles illustrated in FIG. 17 , in the context of a transmission between two remote devices A and B on a communication network NET, device A includes a processor associated with memory RAM and ROM, which is configured to implement as described with respect to Figure 2 , 3 , 4, 6, 8 described in any of the embodiments of the method for NN encoding, and the device B includes a processor associated with a memory RAM and a ROM, which is configured to implement as described with respect to Figure 2 , 3, 5, 6 or 9. According to an example, the network is a broadcast network adapted to broadcast / transmit encoded images from device A to decoding devices including device B. The A signal intended to be transmitted by device A carries at least one bitstream comprising encoded data representing at least one image, together with metadata allowing application of entropy coding improvement information.

[0097] Fig.14 An example of the syntax of such a signal is shown when the at least one coded image is transmitted over a packet-based transport protocol. Each transmitted packet P comprises a header H and a payload PAYLOAD. The payload PAYLOAD may carry the bitstream described above including metadata with respect to signaling channel activity. In a variant, the payload comprises neural network-based coded data representing image data samples and associated metadata, wherein the associated metadata comprises at least one of an indication of channel activity.

[0098] It should be noted that our approach is not limited to a specific neural network architecture, e.g., the so-called hyper-prior model Figure 3 Alternatively, our approach can be used in other neural network architectures such as fully factorized neural image / video models, implicit neural image / video compression models, recurrent network-based neural image / video compression models, or generative model-based image / video compression methods.

[0099] Various numerical values ​​are used in this application. The specific values ​​are for illustrative purposes, and the described aspects are not limited to these specific values.

[0100] Various methods are described herein, and each of the methods includes one or more steps or actions for implementing the described method. Unless the specific order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions can be modified or combined. In addition, terms such as "first", "second", etc. can be used in various embodiments to modify elements, parts, steps, operations, etc., such as, for example, "first decoding" and "second decoding". The use of such terms does not imply the sequencing of modified operations, unless specifically required. Therefore, in this example, the first decoding need not be performed before the second decoding, but can, for example, occur before, during or in the time period overlapping with the second decoding.

[0101] Various implementations involve decoding. As used in this application, "decoding" may encompass, for example, all or part of a process performed on a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, and inverse transform. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or is intended to generally refer to a broader decoding process will be clear based on the context of the specific description and is considered to be well understood by those skilled in the art.

[0102] Various implementations involve encoding.In a similar manner to the above discussion regarding "decoding", "encoding" as used in this application may encompass all or part of a process performed on an input video sequence, for example, to produce an encoded bitstream.

[0103] The implementation and aspects described herein can be implemented in, for example, a method or process, a device, a software program, a data stream or a signal. Even if only discussed in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the discussed features can also be implemented in other forms (e.g., a device or a program). Apparatus can be implemented, for example, with appropriate hardware, software, and firmware. Methods can be implemented in, for example, a device, such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, such as a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate the communication of information between end users.

[0104] Reference to "one embodiment" or "an embodiment" or "one implementation" or "an implementation" and other variations thereof means that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase "in one embodiment" or "in an embodiment" or "in an implementation" or "in an implementation" and any variations thereof appearing throughout this application do not necessarily all refer to the same embodiment.

[0105] Additionally, the present application may refer to "determining" various pieces of information. Determining information may include, for example, one or more of: estimating information, calculating information, predicting information, or retrieving information from a memory.

[0106] Further, the present application may refer to "accessing" various pieces of information. Accessing information may include, for example, one or more of: receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0107] Additionally, the present application may refer to "receiving" various pieces of information. As with "accessing," receiving is intended to be a broad term. Receiving information may include, for example, one or more of the following: accessing information or retrieving information (e.g., from a memory). Further, "receiving" typically involves, in one way or another, during an operation such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0108] It should be appreciated that use of any of the following “ / ,” “and / or,” and “at least one of…” (e.g., in the case of “A / B,” “A and / or B,” and “at least one of A and B”) is intended to encompass selection of only the first listed option (A), or only the second listed option (B), or both options (A and B). As a further example, in the case of “A, B, and / or C” and “at least one of A, B, and C,” such phrases are intended to encompass selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A and B and C). This can be extended to as many items as listed, as would be clear to one skilled in the art and related fields.

[0109] As will be apparent to one skilled in the art, implementations may generate a variety of signals formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for executing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry a bit stream of the described embodiments. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using a radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is well known. The signal may be stored on a processor readable medium.

Claims

1. A video encoding method, comprising: using a neural network to obtain a latent associated with the image data, the latent comprising a plurality of channels of two-dimensional data; obtaining a probability distribution for each value of the potential; and entropy encoding the latent object based on a probability distribution of the latent object; Wherein the entropy coding further comprises at least one of the following: Signaling an indication of channel activity, Obtain the conditional probability distribution of the potential object, Channel reordering, A rate-distortion optimization is performed on the potentials.

2. The method of claim 1, wherein the method further comprises: obtaining an indication of activity of a channel, wherein for a current channel of the potential, the indication of activity of the current channel indicates that at least one value of the potential is different than a most likely value in a probability distribution for the value of the potential; Obtaining a probability distribution of the channel activity indication; entropy encoding the channel activity indication based on a probability distribution of the channel activity indication; as well as Wherein entropy encoding the latents comprises encoding only channels having positive channel activity indications.

3. The method of claim 2, wherein the method further comprises: sorting the channels of the potential objects according to the values ​​of the probability distribution of the channel activity indications; as well as Wherein the sorted potentials are entropy encoded based on a probability distribution of the potentials.

4. The method of claim 3, wherein entropy encoding the channel activity indication based on the probability distribution of the channel activity indication further comprises: determining the index of the last active channel in the sorted potentials, and An index of a last active channel indication is entropy encoded based on a probability distribution of the channel activity indications.

5. The method of claim 1, wherein the method further comprises: obtaining at least one context of the value of the potential, obtaining a conditional probability distribution for each context for each value of the latent; as well as Wherein the latent is entropy encoded based on a conditional probability distribution of the latent.

6. The method of claim 5, wherein at least one context of the value of the potential comprises at least one causal space neighboring value in the same channel.

7. The method of claim 5, wherein at least one context of the value of the potential further comprises at least one causal inter-channel neighbor value.

8. The method of claim 5, further comprising: determining an optimal conditional probability distribution among the at least one context; as well as Wherein the latents are entropy encoded based on an optimal conditional probability distribution of the latents.

9. The method of claim 1, wherein the method further comprises: sorting the channels of the potential objects according to a channel order; as well as Wherein the sorted potentials are entropy encoded based on a probability distribution of the potentials.

10. The method of claim 9, wherein the channel order is fixed and obtained from offline training.

11. The method of claim 10, wherein the channel order is obtained by maximizing the correlation between consecutive channels of the potentials.

12. The method of claim 1, wherein rate-distortion optimization of the potentials is performed.

13. The method of claim 12, wherein only a portion of a potential corresponding to a receptive field of a modified value in the potential is processed for rate-distortion optimization of the potential.

14. The method of claim 13, wherein a portion of a potential corresponding to a receptive field of a modified value in the potential is further scaled for rate-distortion optimization of the potential.

15. The method of claim 12, wherein the rate-distortion optimization of the potentials is parallelized for each channel.

16. The method of claim 13, wherein rate-distortion optimization of the latent using the scaled portion of the latent is iterated.

17. A method for video decoding, comprising: obtaining encoded data representing a latent object associated with the image data, the latent object comprising a plurality of channels of two-dimensional data; obtaining a probability distribution for each value of the potential; and entropy decoding the encoded data based on a probability distribution of the latent to reconstruct the latent; Wherein the entropy coding further comprises at least one of the following: Signaling an indication of channel activity, Obtain the conditional probability distribution of the potential object, Channel reordering, Rate-distortion optimization is performed on the potentials.

18. An apparatus comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to perform the method of any one of claims 1-17.

19. A signal comprising video data, formed by performing a method as claimed in any one of claims 1 to 16.

20. A computer-readable storage medium having stored thereon instructions for encoding or decoding a video according to the method of any one of claims 1-16.