Method and apparatus for learned video compression
Patent Information
- Application Number
- PCT/EP2026/054416
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-26
- Filing Date
- 2026-02-18
- Publication Date
- 2026-10-01
Smart Images

Figure EP2026054416_01102026_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS FOR LEARNED VIDEO COMPRESSIONTECHNOLOGICAL FIELD
[0001] An example embodiment relates to video and image coder-decoders (codecs) and, more particularly, to codecs for performing learned video compression.BACKGROUND
[0002] A video codec consists of an encoder that transforms the input video into a compressed representation suited for storage and / or transmission and a decoder that can decompress the compressed video representation back into a viewable form. Typically, an encoder discards some information in the original video sequence in order to represent the video in a more compact form, that is, at a lower bitrate.
[0003] Typical hybrid video codecs, for example ITU-T H.263 and H.264 codecs, encode the video information in two phases. Firstly, pixel values in a certain picture area (or “block”) are predicted, for example, by motion compensation in which an area in one of the previously coded video frames that corresponds closely to the block being coded is identified or spatially by using the pixel values around the block to be coded in a specified manner. Secondly, the prediction error, e.g., the difference between the predicted block of pixels and the original block of pixels, is coded. This coding is typically done by transforming the difference in pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT) or a variant of DCT), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate).
[0004] Inter prediction, which may also be referred to as temporal prediction, motion compensation, or motion-compensated prediction, exploits temporal redundancy. In inter prediction the sources of prediction are previously decoded pictures, also known as reference pictures.
[0005] In temporal inter prediction, the sources of prediction are previously decoded pictures in the same scalable layer. In intra block copy (IBC), also known as intra-block-copy prediction, prediction may be applied similarly to temporal inter prediction, but the reference picture is the current picture and only previously decoded samples can be referred in the prediction process.Inter-layer or inter-view prediction may be applied similarly to temporal inter prediction, but the reference picture is a decoded picture from another scalable layer or from another view, respectively. In some cases, inter prediction may refer to temporal inter prediction only, while in other cases inter prediction may refer collectively to temporal inter prediction and any of intra block copy, inter-layer prediction, and inter-view prediction provided that they are performed with the same or similar process than temporal prediction. Inter prediction, temporal inter prediction, or temporal prediction may sometimes be referred to as motion compensation or motion-compensated prediction.
[0006] Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in a spatial or transform domain, that is, either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.
[0007] One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy-coded more efficiently if they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors and only the difference relative to the motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.
[0008] The decoder reconstructs the output video by applying prediction means similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in spatial pixel domain). After applying prediction and prediction error decoding means, the decoder sums up the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) can also apply additional filtering means to improve the quality of the output video before passing the output video for display and / or storing the output video as prediction reference for the forthcoming frames in the video sequence.
[0009] In typical video codecs, the motion information is indicated with motion vectors associated with each motion compensated image block. Each of these motion vectors represents the displacement of the image block in the picture to be coded (in the encoder side) or decoded (in the decoder side) and the prediction source block in one of the previously coded or decoded pictures.In order to represent motion vectors efficiently, motion vectors are typically coded differentially with respect to block specific predicted motion vectors. In typical video codecs, the predicted motion vectors are created in a predefined way, for example calculating the median of the encoded or decoded motion vectors of the adjacent blocks. Another way to create motion vector predictions is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures and signalling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of a previously coded and / o decoded picture can be predicted. The reference index is typically predicted from adjacent blocks and / or or co-located blocks in a temporal reference picture. Moreover, typical high efficiency video codecs employ an additional motion information coding and / or decoding mechanism, often called merging and / o merge mode, where all the motion field information, which includes motion vector and corresponding reference picture index for each available reference picture list, is predicted and used without any modification or correction. Similarly, predicting the motion field information is carried out using the motion field information of adjacent blocks and / or co-located blocks in temporal reference pictures and the motion field information that is used is signalled among a list of motion field candidates filled with motion field information of available adjacent and / or co-located blocks.
[0010] In typical video codecs, the prediction residual after motion compensation is first transformed with a transform kernel (such as DCT) and then coded. The reason for this transformation is that often there still exists some correlation among the residual and the transformation can in many cases reduce this correlation and provide more efficient coding.
[0011] Typical video encoders utilize Lagrangian cost functions to find optimal coding modes, e.g., the desired macroblock mode and associated motion vectors. This kind of cost function uses a weighting factor A. to tie together the (exact or estimated) image distortion due to lossy coding methods and the (exact or estimated) amount of information that is required to represent the pixel values in an image area, such as in accordance with the following equation:C = D + XRwhere C is the Lagrangian cost to be minimized, D is the image distortion (e.g., mean squared error) with the mode and motion vectors considered, and R the number of bits needed to representthe required data to reconstruct the image block in the decoder (including the amount of data to represent the candidate motion vectors).
[0012] Video coding specifications may enable the use of supplemental enhancement information (SEI) messages or the like. Some video coding specifications include SEI network abstraction layer (NAL) units, and some video coding specifications contain both prefix SEI NAL units and suffix SEI NAL units, where the former type can start a picture unit or the like and the latter type can end a picture unit or the like. An SEI NAL unit contains one or more SEI messages, which are not required for the decoding of output pictures but may assist in related processes, such as picture output timing, post-processing of decoded pictures, rendering, error detection, error concealment, and resource reservation. Several SEI messages are specified in H.264 / AVC, H.265 / HEVC, H.266 / VVC, and H.274 / VSEI standards, and the user data SEI messages enable organizations and companies to specify SEI messages for their own use. The standards may contain the syntax and semantics for the specified SEI messages but a process for handling the messages in the recipient might not be defined. Consequently, encoders may be required to follow the standard specifying a SEI message when they create SEI message(s), and decoders might not be required to process SEI messages for output order conformance. One of the reasons to include the syntax and semantics of SEI messages in standards is to allow different system specifications to interpret the supplemental information identically and hence interoperate. It is intended that system specifications can require the use of particular SEI messages both in the encoder and in the decoder, and additionally the process for handling particular SEI messages in the recipient can be specified.BRIEF SUMMARY
[0013] A method, apparatus, and computer program product are disclosed for performing learned video compression in a codec. By performing learned video compression in accordance with an example embodiment, the rate-distortion performance may be improved.
[0014] In an example embodiment, an apparatus is provided that includes a first codec comprising at least a first encoder configured to receive and encode input data to generate a first bitstream and a first decoder configured to generate first reconstructed data based at least on the first bitstream. The apparatus also includes a second codec comprising at least a second encoder configured to receive and encode data to generate a second bitstream and a second decoder configured to generate second reconstructed data based at least on the second bitstream. The data received andencoded by the second encoder comprises or is derived from one or more of the input data, data derived from the input data, the first reconstructed data, data derived from the first reconstructed data, data from which the first reconstructed data is derived, or data representative of one or more intermediate features generated by one or more layers of the first decoder.
[0015] The first encoder may include at least a first neural network based encoder configured to receive and process the input data to generate a first latent tensor, a first quantization operation configured to quantize the first latent tensor to obtain a quantized first latent tensor, and a first lossless or substantially lossless encoder configured to encode the quantized first latent tensor to obtain the first bitstream. In this embodiment, the first decoder may include at least a first lossless or substantially lossless decoder configured to decode the first bitstream to obtain a quantized second latent tensor, a first dequantization operation configured to dequantize the quantized second latent tensor to obtain a second latent tensor, and a first neural network based decoder configured to generate the first reconstructed data based at least on the second latent tensor. The second encoder may also include at least a second neural network based encoder configured to receive and process the data received by the second encoder to generate a third latent tensor, a second quantization operation configured to quantize the third latent tensor to obtain a quantized third latent tensor, and a second lossless or substantially lossless encoder configured to encode the quantized third latent tensor to obtain the second bitstream. And the second decoder may include at least a second lossless or substantially lossless decoder configured to decode the second bitstream to obtain a quantized fourth latent tensor, a second dequantization operation configured to dequantize the quantized fourth latent tensor to obtain a fourth latent tensor, and a second neural network based decoder configured to generate the second reconstructed data based at least on the fourth latent tensor.
[0016] In one embodiment, the first encoder further comprises a first probability model configured to estimate a first probability of respective elements of the quantized first latent tensor, and the first lossless or substantially lossless encoder is configured to encode the quantized first latent tensor to obtain the first bitstream based on the first probability. In this embodiment, the first decoder further comprises a second probability model configured to estimate a second probability of respective elements of a quantized second latent tensor, and the first lossless or substantially lossless decoder is configured to decode the first bitstream to obtain the quantized second latent tensor based at least on the second probability.
[0017] The second encoder may also include a third probability model configured to estimate a third probability of respective elements of the quantized third latent tensor, and the second lossless or substantially lossless encoder is configured to encode the quantized third latent tensor to obtain the second bitstream based at least on the third probability. In this embodiment, the second decoder further comprises a fourth probability model configured to estimate a fourth probability of respective elements of a quantized fourth latent tensor, and the second lossless or substantially lossless decoder is configured to decode the second bitstream to obtain the quantized fourth latent tensor based at least on the fourth probability.
[0018] The apparatus of an example embodiment also includes a neural network configured to process at least one of the data derived from the input data or the data derived from the first reconstructed data to generate a residual. In this embodiment, the second encoder of the second codec is configured to receive and encode the residual to generate the second bitstream.
[0019] The second encoder of the second codec of an example embodiment comprises a neural network that is configured to process at least one of the data derived from the input data or the data derived from the first reconstructed data based on one or more trained parameters of the neural network to generate a residual.
[0020] The second encoder of the second codec of an example embodiment is configured to receive one or more first features derived from the first reconstructed data, extract one or more second features from the input data and process both the one or more first features and the one or more second features to generate the second bitstream.
[0021] The second encoder of the second codec of an example embodiment is configured to receive one or more first features from which the first reconstructed data was derived or one or more first features derived from data from which the first reconstructed data was derived, extract one or more second features from the input data and process both the one or more first features and the one or more second features to generate the second bitstream.
[0022] The one or more first features and the one or more second features may be multi-scale features. In this embodiment, the second encoder of the second codec comprises a second neural network based encoder having one or more attention layers at each scale that are configured to integrate the one or more first features and the one or more second features.
[0023] The first probability model may be the same or substantially same as the second probability model. Additionally, the third probability model may be the same or substantially the same as the fourth probability model.
[0024] In one embodiment, the third probability model and the fourth probability model are also configured to receive auxiliary data comprising one or more of the following:one or more features derived from the first reconstructed data of the first codec, one or more features from which the first reconstructed data of the first codec is derived, an output of a neural network having an input comprising one or more features from which the first reconstructed data of the first codec is derived,the first reconstructed data of the first codec,the first latent tensor or features derived from the first latent tensor,the second latent tensor or features derived from the second latent tensor,the quantized first latent tensor or features derived from the quantized first latent tensor, the quantized second latent tensor or features derived from the quantized second latent tensor,features extracted by the first decoder or by a first neural network based decoder, data derived from other data from which the first reconstructed data was also derived, or one or more scaling values provided by the first probability model or by the second probability model.
[0025] In one embodiment, the second decoder or the neural network based second decoder is also configured to receive auxiliary data comprising one or more of the following:one or more features derived from the first reconstructed data of the first codec, one or more features from which the first reconstructed data of the first codec is derived, an output of a neural network having an input comprising one or more features from which the first reconstructed data of the first codec is derived,the first reconstructed data of the first codec,the quantized first latent tensor or features derived from the quantized first latent tensor, the quantized second latent tensor or features derived from the quantized second latent tensor,the second latent tensor or features derived from the second latent tensor,one or more scaling values provided by the first probability model,features extracted by the first decoder or by a first neural network based decoder, data derived from other data from which the first reconstructed data was also derived, or one or more scaling values provided by the second probability model.
[0026] The second decoder of the second codec of an example embodiment is configured to receive one or more first features from which the first reconstructed data was derived, derive one or more second features from the fourth latent tensor of the second codec and process both the one or more first features and the one or more second features to generate the second reconstructed data.
[0027] In this embodiment, the one or more first features and the one or more second features may be multi-scale features, and the second decoder of the second codec may comprise a second neural network based encoder having one or more attention layers at each scale that are configured to integrate the one or more first features and the one or more second features.
[0028] The apparatus of an example embodiment also includes a post-processing filter configured to provide a filtered output based upon the second reconstructed data of the second codec and one or more auxiliary inputs comprising one or more of the following:one or more features derived from the first reconstructed data of the first codec, one or more features from which the first reconstructed data of the first codec is derived, an output of a neural network based on one or more features from which the first reconstructed data is derived,the first reconstructed data of the first codec,the quantized first latent tensor or one or more features derived from the quantized first latent tensor,the quantized second latent tensor or one or more features derived from the quantized second latent tensor,the quantized third latent tensor or one or more features derived from the quantized third latent tensor,the quantized fourth latent tensor or one or more features derived from the quantized fourth latent tensor,the second latent tensor or one or more features derived from the second latent tensor, the fourth latent tensor or one or more features derived from the fourth latent tensor, one or more intermediate features computed by the second decoder of the second codec,one or more Intermediate features computed by the first decoder of the first codec, data derived from other data from which the first reconstructed data was also derived, data derived from other data from which the second reconstructed data was also derived, or one or more scaling values provided by one or both of the second and fourth probability models.
[0029] In another example embodiment, a method is provided that includes receiving and encoding input data with at least a first encoder of a first codec to generate a first bitstream and generating first reconstructed data based at least on the first bitstream with a first decoder of the first codec. The method also includes receiving and encoding data with at least a second encoder of a second codec to generate a second bitstream and generating second reconstructed data based at least on the second bitstream with a second decoder. The data received and encoded by the second encoder comprises or is derived from one or more of the input data, data derived from the input data, the first reconstructed data, data derived from the first reconstructed data, data from which the first reconstructed data is derived, or data representative of one or more intermediate features generated by one or more layers of the first decoder.
[0030] In a further example embodiment, a non-transitory computer-readable storage medium is provided that includes program instructions stored thereon that are configured to receive and encode input data with at least a first encoder of a first codec to generate a first bitstream and to generate first reconstructed data based at least on the first bitstream with a first decoder of the first codec. The program instructions are also configured to receive and encode data with at least a second encoder of a second codec to generate a second bitstream and to generate second reconstructed data based at least on the second bitstream with a second decoder. The data received and encoded by the second encoder comprises or is derived from one or more of the input data, data derived from the input data, the first reconstructed data, data derived from the first reconstructed data, data from which the first reconstructed data is derived, or data representative of one or more intermediate features generated by one or more layers of the first decoder.
[0031] In yet another example embodiment, an apparatus is provided that includes means for receiving and encoding input data with at least a first encoder of a first codec to generate a first bitstream and means for generating first reconstructed data based at least on the first bitstream with a first decoder of the first codec. The apparatus also includes means for receiving and encoding data with at least a second encoder of a second codec to generate a second bitstream and meansfor generating second reconstructed data based at least on the second bitstream with a second decoder. The data received and encoded by the second encoder comprises or is derived from one or more of the input data, data derived from the input data, the first reconstructed data, data derived from the first reconstructed data, data from which the first reconstructed data is derived, or data representative of one or more intermediate features generated by one or more layers of the first decoder.
[0032] This summary is intended to provide a brief overview of some of the aspects and features according to the subject disclosure. Accordingly, it will be appreciated that the above-described features are merely examples and should not be construed to narrow the scope of the subject disclosure in any way. Other features, aspects, and advantages of the subject disclosure will become apparent from the following detailed description, drawings and claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Having thus described certain example embodiments of the present disclosure in general terms, reference will hereinafter be made to the accompanying drawings, which are not necessarily drawn to scale, and wherein:
[0034] Figure 1 is a block diagram of a video coding pipeline in which at least some components are implemented as a neural network;
[0035] Figure 2 is a block diagram of an end-to-end leaned codec;
[0036] Figure 3 is an illustration of the architecture of a multi-scale progressive probability model;
[0037] Figure 4 is an illustration of the architecture of a prediction model in a multi-scale progressive probability model;
[0038] Figure 5 is an illustration of a neural network-based end-to-end learned video coding system in accordance with previous embodiments;
[0039] Figure 6 is a block diagram of encoder-side operations for overfitting a signal;
[0040] Figure 7 is a block diagram of decoder-side operations for updating a neural network filter using the overfitting signal;
[0041] Figure 8 is a block diagram of a multi-step learned intra-frame codec;
[0042] Figure 9 is a block diagram of a multi-step learned codec model in which both the first-step codec and the second-step codec includes one or more neural networks in accordance with an example embodiment of the present disclosure;
[0043] Figure 10 is a block diagram of a multi-step learned codec model in which both the first-step codec and the second-step codec includes one or more neural networks in accordance with another example embodiment of the present disclosure;
[0044] Figure 11 is a block diagram of an apparatus that may be configured to perform operations in accordance with an example embodiment of the present disclosure;
[0045] Figure 12 is a block diagram of a multi-step learned codec model in accordance with an example embodiment of the present disclosure; and
[0046] Figure 13 is a flowchart demonstrating operations performed, such as by the apparatus of Figure 11, in accordance with an example embodiment of the present disclosure.DETAILED DESCRIPTION
[0047] Some embodiments of the present invention will now be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all, embodiments of the invention are shown. Indeed, various embodiments of the invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Like reference numerals refer to like elements throughout. As used herein, the terms “data,” “content,” “information,” and similar terms may be used interchangeably to refer to data capable of being transmitted, received and / or stored in accordance with embodiments of the present invention. Thus, use of any such terms should not be taken to limit the spirit and scope of embodiments of the present invention.
[0048] Although the terms first, second, third, fourth, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and similarly, a second element could be termed a first element, without departing from the scope of this disclosure. As used herein, the term “and / or,” includes any and all combinations of one or more of the associated listed items.
[0049] When an element is referred to as being “connected,” or “coupled,” to another element, it can be directly connected or coupled to the other element or intervening elements may be present. By contrast, when an element is referred to as being “directly connected,” or “directly coupled,” to another element, there are no intervening elements present. Other words used to describe therelationship between elements should be interpreted in a like fashion (e.g., “between,” versus “directly between,” “adjacent,” versus “directly adjacent,” etc.).
[0050] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the singular forms “a”, “an”, and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises”, “comprising,”, “includes” and / or “including”, when used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0051] It should also be noted that in some alternative implementations, the functions / acts noted may occur out of the order noted in the figures. For example, two figures shown in succession may in fact be executed concurrently or may sometimes be executed in the reverse order, depending upon the functionality / acts involved.
[0052] Specific details are provided in the following description to provide a thorough understanding of example embodiments. However, it will be understood by one of ordinary skill in the art that example embodiments may be practiced without these specific details. For example, systems may be shown in block diagrams so as not to obscure the example embodiments in unnecessary detail. In other instances, well-known processes, structures and techniques may be shown without unnecessary detail in order to avoid obscuring example embodiments.
[0053] Additionally, as used herein, the term ‘circuitry’ may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in analog circuitry and / or digital circuitry); (b) combinations of circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and (c) hardware circuit(s) and / or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when needed for operation. This definition of ‘circuitry’ applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term ‘circuitry’ also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portions of a hardware circuit orprocessor and its (or their) accompanying software and / or firmware. The term ‘circuitry’ also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device or other computing or network device.
[0054] Additionally, as used herein, the terms “model”, “neural network”, “neural net” and “network” are used interchangeably. In addition, the weights of neural networks may be referred to as learnable parameters or parameters.
[0055] Additionally, as used herein, the terms “machine” and “task neural network” are used interchangeably, to mean any process or algorithm (learned or not from data) which analyzes or processes data for a certain task.
[0056] Additionally, as used herein, the terms “encoder-side” and “decoder-side” refer to the physical or abstract entity or device which may contain one or more machines, and may run these one or more machines on some encoded and eventually decoded video representation which is encoded by another physical or abstract entity or device, the “encoder-side device”.
[0057] Additionally, as used herein, the terms “intra frame”, “frame”, and “image” may be used interchangeably. These terms may refer to at least part of the input data and at least part of the output data of a learned intra-frame codec. In one or more embodiments, these terms refer to image as the data type. However, the proposed embodiments can be extended to other types of data such as video, audio, and etc. Additionally, the terms frame, picture and image are used interchangeably herein. For example, the input and output to an end-to-end learned codec may be pictures. The input and output of a NN filter may be pictures. Additionally, the term block, when it means a portion of a picture, may be simply referred to as frame or picture or image. In other words, at least some of the embodiments herein, even when described as applied to a picture, may be applicable also to a block, e.g., to a portion of a picture.
[0058] As defined herein, a “computer-readable storage medium,” which refers to a physical storage medium (e.g., volatile or non-volatile memory device), may be differentiated from a “computer-readable transmission medium,” which refers to an electromagnetic signal.
[0059] In the following description, illustrative embodiments will be described with reference to acts and symbolic representations of operations (e.g., in the form of flow charts, flow diagrams, data flow diagrams, structure diagrams, block diagrams, etc.) that may be implemented as program modules or functional processes include routines, programs, objects, components, data structures,etc., that perform particular tasks or implement particular abstract data types and may be implemented using existing hardware at existing network elements. Such existing hardware may include one or more Central Processing Units (CPUs), digital signal processors (DSPs), application-specific-integrated-circuits, field programmable gate arrays (FPGAs), computers or the like.
[0060] Although a flow chart may describe the operations as a sequential process, many of the operations may be performed in parallel, concurrently or simultaneously. In addition, the order of the operations may be re-arranged. A process may be terminated when its operations are completed but may also have additional steps not included in the figure. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or the main function.
[0061] As disclosed herein, the term “storage medium” or “computer readable storage medium” may represent one or more devices for storing data, including read only memory (ROM), random access memory (RAM), magnetic RAM, core memory, magnetic disk storage mediums, optical storage mediums, flash memory devices and / or other tangible machine readable mediums for storing information. The term “computer-readable medium” may include, but is not limited to, portable or fixed storage devices, optical storage devices, and various other mediums capable of storing, containing or carrying instruction(s) and / or data.
[0062] Furthermore, example embodiments may be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks may be stored in a machine or computer readable medium such as a computer readable storage medium. When implemented in software, a processor or processors will perform the necessary tasks.
[0063] A code segment may represent a procedure, function, subprogram, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via anysuitable means including memory sharing, message passing, token passing, network transmission, etc.
[0064] Recently, neural networks (NNs) have been used in the context of image and video compression in accordance with at least two primary approaches. A neural network (NN) may be described as a computation graph consisting of several layers of computation. Each layer may consist of one or more units, where each unit performs an elementary computation. A unit is connected to one or more other units, and the connection may be associated with a weight. The weight may be used for scaling the signal passing through the associated connection. Weights are learnable parameters, e.g., values which can be learned from training data, or simply parameters. There may be other learnable parameters, such as those of batch-normalization layers.
[0065] In some neural networks, such as convolutional neural networks for image classification, initial layers (those close to the input data) extract semantically low-level features such as edges and textures in images, whereas intermediate layers extract more high-level features. After the feature extraction layers there may be one or more layers performing a certain task, such as classification, semantic segmentation, object detection, denoising, style transfer, super-resolution, etc.
[0066] Neural networks are being utilized in an ever-increasing number of applications for many different types of devices, such as mobile phones. Examples include image and video analysis and processing, social media data analysis, device usage data analysis, etc.
[0067] One property of neural nets (and other machine learning tools) is that they are able to learn properties from input data, e.g., in supervised way or in unsupervised way. Such learning is a result of a training algorithm, or of a meta-level neural network providing the training signal.
[0068] In general, the training algorithm consists of changing some properties of the neural network so that its output is as close as possible to a desired output. For example, in the case of classification of objects in images, the output of the neural network can be used to derive a class or category index which indicates the class or category to which the object in the input image belongs. Training usually happens by minimizing or decreasing the output’s error, also referred to as the loss or loss function. Examples of losses are mean squared error, cross-entropy, etc. In some deep learning techniques, training is an iterative process, where at each iteration the algorithm modifies the weights of the neural net to make a gradual improvement of the network’s output, e.g., to gradually decrease the loss, by means of gradient descent technique. In one example, ateach training iteration, gradients of the loss function with respect to one or more weights or parameters of the NN are computed, for example by a backpropagation technique, and the computed gradients are then used by an optimization routine, such as Adam or Stochastic Gradient Descent (SGD) to obtain an update to the one or more weights or parameters.
[0069] Training a neural network is an optimization process, but the final goal may be different from the typical goal of optimization. In optimization, the only goal is to minimize a function. In machine learning, the goal of the optimization or training process is to make the model learn the properties of the data distribution from a limited training dataset. In other words, the goal is to learn to use a limited training dataset in order to learn to generalize to previously unseen data, that is, data which was not used for training the model. This process is usually referred to as generalization. In practice, data is usually split into at least two sets, the training set and the validation set. The training set is used for training the network, such as to modify its learnable parameters in order to minimize the loss. The validation set is used for checking the performance of the network on data which was not used to minimize the loss, as an indication of the final performance of the model. In particular, the errors on the training set and on the validation set are monitored during the training process to understand if the model is underfitting or overfitting or is being trained properly.
[0070] If the network is learning and, as a result, being properly trained, the training set error should decrease. Otherwise, the model is not learning and is in the regime of underfitting.
[0071] If the network is learning to generalize, the validation set error decreases and is not too much greater than the training set error. If the training set error is low, but the validation set error is much greater than the training set error, or the training set error does not decrease, or the training set error even increases, the model may be in the regime of overfitting. Overfitting means that the model has just memorized the training set’s properties and performs well only on that set, but performs poorly on a set not used for tuning its parameters.
[0072] In a first approach, NNs are used to replace one or more of the components of a traditional codec such as a WC / H.266-compliant codec. As used herein, “traditional” means those codecs whose components and their parameters are typically not learned from data by means of machine learning techniques. In contrast, examples of components of or operations performed by a codec that may be implemented as neural networks are and in-loop filter, intra-frame prediction, interframe prediction, transform and / or inverse transform, probability model for lossless coding, etc.Y1With respect to an in-loop filter, for example, a NN may serve as an additional in-loop filter with respect to the traditional loop filters, or a NN may serve as the only additional in-loop filter, thus replacing any other in-loop filter.
[0073] In a second approach, commonly referred to as “end-to-end learned compression” (or end-to-end learned codec), NNs are used as the main components of the image / video codecs. However, the codec may still comprise components which are not based on machine learning techniques. In this second approach, there may be at least two design options. In a first option, the traditional video coding pipeline is re-used, but most or all the components ae replaced with NNs as shown in Figure 1. In this instance, the forward and inverse transforms are replaced with two neural networks, the loop filter is in the form of a neural network, etc.
[0074] In the second option (also referred to as end-to-end learned coding), the whole coding and decoding pipeline is redesigned as a neural network auto-encoder with a quantization and lossless coding serving as an intermediate portion of the pipeline. In accordance with this option, an encoder NN (also referred to as neural network based encoder, or NN encoder) performs a nonlinear transformation of the input and provides an output that is typically referred to as latent tensor. Thereafter, quantization and lossless encoding of the encoder NN’s output is performed followed by lossless decoding and dequantization. This option also includes a decoder NN (also referred to as neural network based decoder, or NN decoder) that performs a non-linear inverse transformation from dequantized latent tensor to a reconstructed input. Even in end-to-end learned approaches, there may be components which are not learned from data, such as the arithmetic codec.
[0075] Figure 2 illustrates an example neural network-based end-to-end learned coding system, such as an end-to-end learned video coding system or an end-to-end learned image coding system. Even though some examples are provided with respect to coding images or videos, it is to be understood that other types of data may be coded in a similar way, such as audio, speech, text, features, etc. As shown in Figure 2, a typical neural network-based end-to-end learned coding system contains an encoder 10 and a decoder 12.
[0076] The encoder 10 includes an encoder NN 14, a quantizer or quantization operation 16, a probability model 24, and a lossless encoder 20 (for example arithmetic encoder) of a lossless codec 18. The decoder 12 includes a lossless decoder 22 (for example, an arithmetic decoder) of the lossless codec, a probability model 26, a dequantizer or dequantization operation 28, and adecoder NN 29. The probability model present at the encoder side and the probability model present at the decoder side may be same or substantially the same probability models. For example, they may be two copies of the same probability model. As illustrated, the lossless encoder 20 and the lossless decoder 22 form a lossless codec 18. A lossless codec may be an entropy -based lossless codec. An example of a lossless codec is an arithmetic codec, such as a context-adaptive binary arithmetic coding (CAB AC).
[0077] The encoder NN 14 and decoder NN 29 are typically two neural networks, or are mainly comprised of neural network components. The probability models 24, 26 may also be a neural network and / or be comprised of mainly neural network components, and may be referred to as neural network based probability model or learned probability model. In some instances, the term lossless codec may refer to a system that includes the probability model, in addition to, for example, an arithmetic encoder and an arithmetic decoder. The quantizer 16, dequantizer 28 and lossless codec 18 are typically not based on neural network components, but they may also include neural network components in some instances.
[0078] In operation, the encoder NN 14 receives an input x, which may include, for example, an image to be compressed. The encoder NN outputs a latent tensor z. In one example, the latent tensor may be a three dimensional (3D) tensor, where the three dimensions of the tensor represent a channel dimension, a vertical dimension (also sometimes referred to as height dimension) and a horizontal dimension (also sometimes referred to as width dimension). In another example, the latent tensor may be a four dimensional (4D) tensor, where the four dimensions of such tensor represent sample dimension (also sometimes referred to as batch dimension, which is the dimension along which different samples of data can be placed), a channel dimension, a vertical dimension (also sometimes referred to as height dimension) and a horizontal dimension (also sometimes referred to as width dimension). The latent tensor is input to a quantization operation 16, obtaining a quantized latent tensor zq. The quantized latent tensor is lossless-encoded into a bitstream b by the lossless encoder 20, based also on the output of the probability model 24. In particular, the probability model takes as input at least part of the quantized latent tensor and outputs an estimate of a probability or an estimate of a probability distribution or an estimate of one or more parameters of a probability distribution for one or more elements of the quantized latent tensor. The bitstream represents an encoded or compressed version of the input x.
[0079] The bitstream is lossless-decoded by the lossless decoder 22 also based on the output of the probability model 26 present at decoder side, obtaining a quantized latent tensor zq. The quantized latent tensor is dequantized by dequantization operation 28, obtaining a reconstructed latent tensor z. The reconstructed latent tensor is input to a decoder NN 29, obtaining a reconstructed input x, that is, a reconstructed version of the input x. The reconstructed input may also be referred to as reconstructed data, or reconstruction, or decoded data, or decoded input, or decoded output, and the like.
[0080] The neural network components, or a subset of the neural network components, of an end-to-end learned codec may be trained by minimizing a rate-distortion loss function as shown below:L = D + R,
[0081] where D is a distortion loss term, R is a rate loss term, and A is a weight that controls the balance between the two losses. The distortion loss term may be referred to also as reconstruction loss term, or simply reconstruction loss. The rate loss term may be referred to simply as rate loss. The distortion loss term measures the quality of the reconstructed or decoded output, and may include, but is not be limited to, one or more of a mean square error (MSE), a structural similarity (SSIM) or a multi-scale structural similarity (MS-SSIM).
[0082] Other losses include losses derived from the use of a pretrained neural network. For example, error(fl, f2), where fl and f2 are the features extracted by a pretrained neural network for the input data and the decoded data, respectively, and errorQ is an error or distance function, such as LI norm or L2 norm. Other losses may also be derived from the use of a neural network that is trained simultaneously with the end-to-end learned codec. For example, adversarial loss can be used, which is the loss provided by a discriminator neural network that is trained adversarially with respect to the codec, following the settings proposed in the context of Generative Adversarial Networks (GANs) and their variants. Additionally, another loss may be related to a performance of one or more machine analysis tasks or to an estimated performance of one or more machine analysis tasks, where the one or more machine analysis tasks may comprise classification, object detection, image segmentation, instance segmentation, etc. In one example, the estimated performance of one or more machine analysis tasks may include a distortion computed based atleast on a first set of features extracted from an output of the decoder and a second set of features extracted from a respective ground truth data, where the first set of features and the second set of features are output by one or more layers of a pretrained feature-extraction neural network.
[0083] Multiple distortion losses may be used and integrated into the distortion loss term D, such as a weighted sum of MSE and SSIM. The rate loss term may be used to train the encoder NN 14 to output a low-entropy latent tensor, or a latent tensor such that the quantized latent tensor has low entropy, or a latent tensor such that the probability distribution of the quantized latent tensor can be better estimated or predicted by the probability model 24. The rate loss term may be used to train the probability model to better estimate or predict the probability distribution of the quantized latent tensor. One example of the rate loss terms is a rate loss term that is derived from the output of the probability model, and that represents the estimated entropy of the quantized latent representation, which indicates the number of bits necessary to represent the quantized latent tensor. Another example is a sparsification loss, that is, a loss that encourages the quantized latent tensor to comprise many zeros. Examples of a sparsification loss are L0 norm, LI norm, LI norm divided by L2 norm.
[0084] In order to train the neural network components, or a subset of the neural network components, of an end-to-end learned codec, one or more of reconstruction losses may be used, and one or more rate losses may be used. In one example, the one or more reconstruction losses and / or one or more rate losses are combined by means of a weighted sum. Typically, the different loss terms are weighted using different weights, and these weights determine how the final system performs in terms of rate-distortion performance. For example, if more weight is given to the reconstruction losses with respect to the rate losses, the system may learn to compress less but to reconstruct with higher accuracy (as measured by a metric that correlates with the reconstruction losses). These weights are usually considered to be hyper-parameters of the training process, and may be set manually by the person designing the training process, or automatically for example by grid search or by using additional neural networks.
[0085] In one case, the training process may be performed jointly with respect to the distortion loss D and the rate loss R. In another case, the training process may be performed in two alternating phases, where in a first phase only the distortion loss D may be used, and in a second phase only the rate loss R may be used.
[0086] For lossless video / image compression, the system would include only the probability model and lossless encoder and lossless decoder. The loss function would comprise only the rate loss, since the distortion loss is always zero, thereby resulting in no loss of information.
[0087] A probability model may be used in an end-to-end learned codec to estimate the probability distribution of the elements in a latent tensor, which is the output of an encoder. The estimated probability distribution may be used by an arithmetic encoder to encode the latent tensor into a bitstream at the encoding stage, or by an arithmetic decoder to decode the latent tensor from the bitstream at the decoding stage. For lossless image and video compression, the probability model estimates the probability distribution of the elements in the input image or video for the arithmetic encoder to encode the input image or video at the encoding stage; at the decoding stage. The probability model estimates the probability distribution of the elements for the arithmetic decoder to decode the output image or video. As used herein, the term latent tensor or latent representation may also refer to the input image or video in a lossless image or video compression system.
[0088] A multi-scale progressive probability model may partition the elements in a latent tensor into multiple groups. The elements in one group may be processed in parallel and the groups may be processed sequentially. Figure 3 shows the architecture of a multi-scale progressive probability model at the encoding stage. Input latent tensoris first downsampled into a certain number of low-resolution representations, e.g. x^\ x(2). The downsampling operation may use the nearest neighborhood, bilinear, or bicubic algorithm. If the downsampling algorithm is not the nearest neighborhood method, extra information may be transferred from the encoder to the decoder to recover round-off error due to the downsampling operation. The probability distribution of the elements in the representation at the lowest resolution, i.e. x in Figure 3, may be modeled as identically and independently distributed with a Gaussian distribution model, a uniform distribution model, or a mixture of probability distribution models. The probability of elements in the latent tensors in resolution levels other than the lowest one may be modeled by a conditional distribution model (whose parameters are estimated by a prediction model), where the conditioning information (also referred to as context) may comprise the representation at lower resolution levels. In Figure 3, z is an auxiliary output from the prediction model 30 at resolution level i where i = 1,2 , that may be used by the prediction model at the next resolution level (i — 1) as an extra input, z may be a zero tensor, is the estimated parameters of the (conditional) probability distribution model 32 for elements at resolution level i.
[0089] At the decoding stage, the latent tensor at the lowest resolution, such as. x in Figure 3 may be first decoded from the bitstream using the predefined probability distribution model. A multi-scale progressive probability model may use the elements in the latent tensor at resolution level i, e.g., x^las the context to estimate the parameters of the distribution model for the elements in the latent tensor at a higher resolution level, e.g. x^l-1The estimated probability distribution may be used by the arithmetic decoder to decode the elements in the bitstream to obtain the latent tensor at higher resolution level x^l-1The procedure may repeat until all elements in the latent tensor at the highest resolution level, such as x^° are decoded.
[0090] The prediction models at different resolution levels may share the weights or a subset of the weights. In one example, the prediction models at different resolutions are the same or substantially the same.
[0091] To further improve the accuracy of the probability distribution estimation, the elements in the latent tensor at a resolution level may be further partitioned into several groups. The groups may be processed sequentially. The elements in a group are modeled by independent conditional distribution models using the elements that have already been processed as the context. For example, the elements of the latent tensor at a resolution level are processed in steps, where the elements in a group associated to a step are processed in parallel. The architecture of the prediction model is shown in Figure 4.
[0092] In an instance in which N is the number of groups into which the elements in the latent representation at resolution level i are partitioned, Figure 4 shows the prediction model at resolution level i and step j, where j = 1,The distribution predictor 42 predicts the parameters of the probability distribution for the elements in the latent representation at resolution level i and step j.is the auxiliary input for the distribution predictor. For the first step, z*7,1) =z(i+1), where z^+1^ is the auxiliary output from the resolution level i + 1, i.e., z^+1^ =Z(I+I,N(I+1)) ^(i ) -s atensorthatcontains the true values for the elements that have already been processed and the predicted values of the elements that have not been processed at step j. x^ is derived by upsampling x^+1\ m*7,7) is a binary-valued mask tensor with the same shape of xS1^ indicating the positions of the elements in x^ that have the true values. p^ is the estimated parameters of the probability distribution for the elements in group j at resolution level i. At the encoding stage, afteris calculated, the tensor updater component 44 may update the elementsin group j with the corresponding true values to generateand the mask updater component 40 may update the mask tensoraccordingly to generate m^j+1At the last step, i.e., j = N^llet<
[0093] At the decoding stage, the calculated p^1’^ is used to decode the corresponding elements in the bitstream. After a group of elements is decoded, the corresponding values are updated in x1’^ to generateanc[ mask tensoris updated accordingly to generate m^j+1The prediction model repeats this operation insteps until all elements at the resolution level i are processed.
[0094] Reducing the distortion in image and video compression is often intended to increase human perceptual quality, as humans are considered to be the end users, such as by consuming or watching the decoded images or videos. Recently, with the advent of machine learning, such as deep learning, there is a rising number of machines (e.g, autonomous agents) that analyze data independently from humans and that may even take decisions based on the analysis results without human intervention. Examples of such analysis are object detection, scene classification, semantic segmentation, video event detection, anomaly detection, pedestrian tracking, etc. For example, such analysis tasks may be performed by neural networks.
[0095] It is likely that the device where the analysis takes place has multiple “machines” or neural networks (NNs). These multiple machines may be used in a certain combination which is for example determined by an orchestrator sub-system. The multiple machines may be used for example in succession, based on the output of the previously used machine, and / or in parallel. For example, a video may be analyzed by one machine (NN) for detecting pedestrians, by another machine (another NN) for detecting cars, and by another machine (another NN) for estimating the depth of all the pixels in the frames.
[0096] Example use cases and applications are self-driving cars, video surveillance cameras and public safety, smart sensor networks, smart TV and smart advertisement, person re-identification, smart traffic monitoring, drones, etc. In addition to image and video data, automatic analysis and processing are increasingly being performed for other types of data, such as audio, speech, text.
[0097] Compressing (and decompressing) data where the end user is a machine (e.g., a neural network) is commonly referred to as compression or coding for machines. In the case of video data, it is referred to as video compression or coding for machines (VCM). Compressing for machines may differ from compressing for humans, for example, with respect to the algorithmsand technology used in the codec, or the training losses used to train any neural network components of the codec, or the evaluation methodology of codecs. When considering the case of coding for machines, the terms “receiver-side” or “decoder-side” are used to refer to the physical or abstract entity or device which contains one or more machines, and runs these one or more machines on some encoded and eventually decoded video representation which is encoded by another physical or abstract entity or device, the “encoder-side device”.
[0098] Figure 5 is an illustration of a pipeline of Video Coding for Machines. A VCM encoder 50 encodes the input video into a bitstream. A bitrate may be computed from the bitstream, as a measure of the size of the bitstream. A VCM decoder 52 decodes the bitstream that was produced by the VCM encoder. The output of the VCM decoder is referred in the figure as “Decoded data for machines”. This data may be considered as the decoded or reconstructed video. However, in some implementations of this pipeline, this data may not have the same or similar characteristics as the original video which was input to the VCM encoder. For example, this data may not be easily understandable by a human by simply rendering the data onto a screen. The output of VCM decoder is then input to one or more task neural networks. In Figure 5, for the sake of illustrating that there may be any number of task-NNs, there are three example task-NNs, and a non-specified one (Task-NN X). One goal of VCM may be to obtain a low bitrate while guaranteeing that the task-NNs still perform well in terms of the evaluation metric associated to each task.
[0099] In some cases, the VCM decoder 52 may not be present. In one example, the machines are run directly on the bitstream. In some other cases, the VCM decoder may include only a lossless decoding stage, and the lossless decoded data is provided as input to the machines. In yet some other cases, the VCM decoder may include a lossless decoding stage following by a dequantization operation, and the loss-decoded and dequantized data is provided as input to the machines.
[0100] When a conventional video encoder, such as a H.266 / VVC encoder, is used as a VCM encoder, one or more of several approaches may be used to adapt the encoding to be suitable to machine analysis tasks. In a first approach, one or more regions of interest (ROIs) may be detected, such as by an ROI detection method. For example, ROI detection may be performed using a task NN, such as an object detection NN. In some cases, ROI boundaries of a group of pictures or an intra period may be spatially overlaid and rectangular areas may be formed to cover the ROI boundaries. The detected ROIs (or rectangular areas) may be used in one or more ways. For example, the quantization parameter (QP) may be adjusted spatially in a manner that ROIs areencoded using finer quantization step size(s) than other regions. For example, QP may be adjusted CTU-wise. Additionally or alternatively, the video is preprocessed to contain only the ROIs, while the other areas are replaced by one or more constant values or removed. The video may also or alternatively preprocessed so that the areas outside the ROIs are blurred or filtered. In addition or in the alternative, a grid may be formed in a manner that a single grid cell covers a ROI. Grid rows or grid columns that contain no ROIs are downsampled as preprocessing to encoding. The detected ROIs may also be used such that a quantization parameter of the highest temporal sublayer(s) is increased (e.g., coarser quantization is used) when compared to practices for human watchable video. Additionally or alternatively, the original video may be temporally downsampled as preprocessing prior to encoding. In this regard, a frame rate upsampling method may be used as postprocessing subsequent to decoding, if machine analysis at the original frame rate is desired. Also, a filter may be used to preprocess the input to the conventional encoder. The filter may be a machine learning based filter, such as a convolutional neural network.
[0101] In the context of video coding for machines, the terms “machine vision”, “machine vision task”, “machine task”, “machine analysis”, “machine analysis task”, “computer vision”, “computer vision task”, "task network" and “task” may be used interchangeably. Also, in the context of video coding for machines, the terms “machine consumption” and “machine analysis” may be used interchangeably.
[0102] A neural network may be used for filtering or processing input data. Such a neural network may be referred to as a neural network based filter, or simply NN filter. A NN filter may include one or more neural networks, and / or one or more components that may not be categorized as neural networks. The purpose of a NN filter may include, but is not be limited to, visual enhancement, colourization, upsampling, super-resolution, inpainting, temporal extrapolation, generating content, and the like.
[0103] In some video codecs, a neural network may be used as a filter in the encoding and decoding loop (also referred to as the coding loop), and it may be referred to as a neural network loop filter, or a neural network in-loop filter. The NN loop filter may replace all other loop filters of an existing video codec, or may represent an additional loop filter with respect to the already present loop filters in an existing video codec. A neural network filter may be used as a postprocessing filter for a codec, e.g., the neural network filter may be applied to an output of an image or video decoder in order to remove or reduce coding artifacts.
[0104] In one example, a codec is a modified WC / H.266 compliant codec (e.g., a WC / H.266 compliant codec that has been modified and thus it may not be compliant to the WC / H.266) that includes one or more NN loop filters. An input to the one or more NN loop filters may include at least a reconstructed block or frames (referred to as reconstruction) or data derived from a reconstructed block or frame (e.g., the output of a conventional loop filter). The reconstruction may be obtained based on predicting a block or frame (e.g., by means of intra-frame prediction or inter-frame prediction) and performing residual compensation. The one or more NN loop filters may enhance the quality of at least one of their input, so that a rate-distortion loss is decreased. The rate may indicate a bitrate (estimate or real) of the encoded video. The distortion may indicate a pixel fidelity distortion such as the mean-squared error (MSE), the mean absolute error (MAE), the mean average precision (mAP) computed based on the output of a task NN (such as an object detection NN) when the input is the output of the post-processing NN and / or other machine task-related metric, for tasks such as object tracking, video activity classification, video anomaly detection, etc. The enhancement may result into a coding gain, which can be expressed for example in terms of BD-rate or BD-PSNR.
[0105] A neural network filter may be used as a post-processing filter for a codec, e.g., may be applied to an output of an image or video decoder in order to remove or reduce coding artifacts. In one example, the NN filter is used as a post-processing filter where the input includes data that is output by or is derived from an output of a traditional decoder, such as a decoder that is compliant with the WC / H.266 standard. In another example, the NN filter is used as a post-processing filter where the input comprises data that is output by or is derived from an output of a decoder of an end-to-end learned decoder.
[0106] In terms of the input to a NN filter and in the case of filtering images, a filter may take as input at least one or more first images to be filtered and may output at least one or more second images, where the one or more second images are the filtered version of the one or more first images. In one example, the filter takes as input one image and outputs one image. In another example, the filter takes as input more than one image and outputs one image. In another example, the filter takes as input more than one image and outputs more than one image.
[0107] A filter may take as input also other data (also referred to as auxiliary data, or extra data) than the data that is to be filtered, such as data that can aid the filter to perform a better filtering than if no auxiliary data was provided as input. In one example, the auxiliary data includesinformation about prediction data, and / or information about the picture type, and / or information about the slice type, and / or information about a Quantization Parameter (QP) used for encoding, and / or information about boundary strength, etc. In one example, the filter takes as input one image and other data associated to that image, such as information about the quantization parameter (QP) used for quantizing and / or dequantizing that image, and outputs one image.
[0108] A NN filter can be adapted at test time based at least on part of the data to be encoded and / or decoded and / or post-processed. Although, for an NN filter that is considered herein, similar adaptation may be performed for other coding tools and / or post-processing tools that are based on neural network technology. For example, a neural network based intra-frame prediction, or a neural network based inter-frame prediction, etc. Such operation may be referred to, for example, as adaptation, content adaptation, overfitting, finetuning, optimization, specialization, and the like. The NN filter that results from the adaptation process may be referred to, for example, as an adapted filter, content-adapted filter, overfitted filter, finetuned filter, optimized filter, specialized filter, and the like.
[0109] An overfitting process may be performed at encoder side based on a training process. The resulting overfitted filter is then used to derive an overfitting signal, or adaptation signal. The adaptation signal may be compressed and then signaled from encoder to decoder, in or along a bitstream that represents encoded data, such as an encoded image or video. Figure 6 illustrates an example of such encoder-side operations.
[0110] In Figure 6, x represents an input to the NN filter 61 of the overfitting process 60, x represents an output of the NN filter, x represents a ground-truth data associated with x, the “Compute loss” operation 62 is configured to compute a training loss I in order to overfit 63 the NN filter, “Overfit” uses I to overfit the NN filter. As a result of the overfitting process, an overfitted NN filter is obtained, which is used, together with the original NN filter, to derive 64 an overfitting signal, that is, an adaptation signal. The adaptation signal is compressed 66 and signaled 68 to a decoder or receiver.
[0111] At decoder side, the overfitting signal, or a signal derived from the overfitting signal, is used to update the NN filter. The updated NN filter is then used to filter one or more pictures, or one or more blocks. Figure 7 illustrates an example of such decoder or receiver side operations. In this regard, the NN filter that is obtained from the overfitting process at the encoder side may be different from the NN filter that is obtained from the updating process at decoder side. Forexample, one reason may be that the adaptation signal may be compressed in a lossy way. Thus, the former NN filter may be referred to as overfitted filter, an adapted filter or the like, and the latter NN filter may be referred to as updated filter. As shown in Figure 7, the compressed adaptation signal may be decompressed 70 and the original NN filter may be updated 72 based at least in part upon the decompressed adaptation signal to generate the updated NN filter.
[0112] In order to perform overfitting at the encoder side, an initial NN filter begins the adaptation process. In one example, the initial NN filter is a pretrained NN filter, which was pretrained during an offline stage on a sufficiently large dataset. In another example, the initial NN filter is a randomly initialized NN filter. In the adaptation, one or more parameters of the NN filter may be adapted. Examples of such parameters may include, but are not limited to, the following: the bias terms of a convolutional neural network, multiplier parameters that multiply one or more tensors produced by the NN filter, such as one or more feature tensors that are output by respective one or more layers of the NN filter, parameters of the kernels of a convolutional neural network, parameters of an adapter layer and / or one or more arrays or tensors that are used as input to respective one or more layers of the NN filter.
[0113] The adaptation may be performed by means of a training process, e.g., by minimizing a loss function until a stopping criterion is met. The data used for this training process may comprise one or more pictures or blocks of input to the NN filter and associated respective one or more pictures or blocks of ground-truth data. In one example where the filter is an in-loop filter, the input to the NN filter is reconstruction data, after prediction and residual compensation; the ground-truth data is the uncompressed data that is given as input to the encoder. In one example where the filter is a post-processing filter, the input to the NN filter is decoded data (e.g., the output of a video decoder); the ground-truth data is the uncompressed data that is given as input to the encoder.
[0114] The loss function used during the training process may comprise one or more distortion loss functions (also referred to as reconstruction loss functions) and zero or more rate loss functions. A rate loss function may measure, for example, the cost in terms of bitrate of signaling any adaptation signal, such as updates to the parameters of the NN filter. A distortion loss function may comprise one of MSE, MAE, MS-SSIM, video multimethod assessment fusion (VMAF), etc.
[0115] The adaptation signal may be derived based on the adapted NN filter and on the original NN filter, that is, the NN filter before the overfitting process. In one example, the adaptation signalincludes an update to one or more parameters of the NN filter. This update may also be referenced as a weight update, or parameter update. This update may be computed, for example, by subtracting the values of the adapted parameters (that is, the parameters of the adapted NN filter) from the corresponding values of the original parameters (that is, the parameters of the original NN filter). In another example, the adaptation signal includes the parameters (of the NN filter) that were adapted, also referred to as updated parameters, or adapted parameters, or adapted weights, or overfitted parameters, and the like.
[0116] In order to keep the size of the adaptation signal low, the adaptation signal may go through one or more compression steps, such as sparsifi cation, quantization and / or lossless coding. In one example, an encoder that compresses the adaptation signal into a bitstream that is compliant with a neural network compression (NNC) standard, such as movie pictures experts group (MPEG) NNC, may be used.
[0117] The compressed adaptation signal may be signaled from the encoder to the decoder in or along a bitstream that represents encoded image or video data. In one example, the compressed adaptation signal is signaled in an adaptation parameter set (APS) syntax structure of a video coding bitstream. In another example, the compressed adaptation signal is signaled in a supplemental enhancement information (SEI) message of a video coding bitstream. Signaling may include also other information which is associated with the adaptation signal and that may be required for correctly parsing and / or decompressing and / or using the adaptation signal, such as any quantization parameters.
[0118] At decoder side, the signaled compressed adaptation signal is received and decompressed. The decompressed adaptation signal may then be used to update the NN filter. In one example, where the adaptation signal includes a weight update and the weight update includes one or more updates to respective one or more parameters of the NN filter, the one or more updates are added to the one or more parameters. In another example, where the adaptation signal includes one or more updated or adapted parameters, the one or more updated or adapted parameters are used to replace respective one or more parameters of the NN filter.
[0119] Once the NN filter has been updated based on the adaptation signal, the updated NN filter may be used for its purpose. For example, the NN filter may be used for filtering an input picture or an input block.
[0120] A multi-step codec includes two or more steps of coding, or two or more codecs. In one example, where the multi-step codec includes two steps of coding, and it is used to code an image (for example, the multi-step codec is an intra-frame codec, or an image codec), the multi-step codec includes a first codec that is used to initially code the input image and may be referred to as a first-step codec (or simply as a first codec), and a second codec that is used to code data derived at least from the input image and may be referred to as a second-step codec (or simply as a second codec). Figure 8 illustrates an example of multi-step learned intra-frame codec having a first codec 80 and a second codec 81. The input to the first encoder 82 of the first codec may be referred to as one or more of the following terms: input to code, original input, original data, uncompressed data, uncompressed input, and the like. It is to be understood that an input to the multi-step codec, even when referred to as uncompressed (e.g., uncompressed data, or uncompressed input) or as original (e.g., original input, original data), may have been coded and decoded and / or processed, for example may have been processed by a temporal filter such as a motion-compensated temporal filter (MCTF) as part of a pre-processing operation. In one example, the original input may comprise an image which includes both a luma component and a chroma component. The output of the first encoder represents the initial bitstream or first bitstream. The first bitstream is input to the first decoder 83 of the fist codec to obtain an initial reconstruction, such as an initial reconstruction of the luma component.
[0121] An input to the second encoder 85 of the second codec 81 may include data derived at least from the original input or a portion of the original input. The output of the second encoder represents a second bitstream. The second bitstream is input to the second decoder 86 of the second codec, to obtain, as an output of the second decoder 86, a final reconstruction, such as a final reconstruction of the luma component. A multi-step codec may include more than two codecs. For example, there may be a third-step codec, a fourth-step codec, etc.
[0122] Certain embodiments may provide for an improvement of the rate-distortion performance of learned video compression, such as in the context of end-to-end learned intra-frame codecs. However, at least some embodiments are also applicable to end-to-end learned inter-frame codecs, or even end-to-end learned video codecs that do not distinguish, at least explicitly, between intra frame and inter frame. As described herein, an image is referenced as the data type. However, certain embodiments can be extended to other types of data such as video, audio, point cloud data, etc.
[0123] In an example, some of the embodiments described herein may be applied to an end-to-end learned video codec. However, at least some of the embodiments described herein may be applied to codecs, other than to an end-to-end learned intra-frame codec or an end-to-end learned image codec. In one example, at least some of the embodiments described herein may be applied to an end-to-end learned inter-frame codec, e.g., an end-to-end learned codec that encodes and decodes frames based on one or more other frames of a video.
[0124] Additionally, in at least some of the embodiments or examples, when the input or original input comprises an image or a video frame, YUV is the considered format of the image or video frame. However, the various embodiments and examples can also be extended to other formats such as RGB. In YUV format, ‘Y’ represents the ‘luma’ value or component; and ‘UV’ represents the ‘chroma’ values or components (e.g., Cb and Cr). In one example, the input image may be an image in YUV 4:4:4 color format, represented as a 3-dimensional array (or tensor) with size 256x256x3, where the horizontal size is 256 pixels, vertical size is 256 pixels, and 3 channels are for Y, U, V components, respectively. In another example, the input image may be an image in YUV 4:2:0 color format, represented by the combination of a matrix of size 256x256 for the luma component and of a 2-dimensional array (or tensor) of size 128x128x2 for the chroma components.
[0125] In an example embodiment, a multi-step intra-frame codec is provided that includes a first codec (such as a first-step codec) that is used to initially code the input image (which may also be referred as original input or original data in some embodiments or examples), and a second codec (such as a second-step codec) that is used to code the residual or any other components of the original input, or data derived at least from the input image. As noted above, although the multi-step intra-frame codec is described to have first and second-step codecs, the multi-step intra-frame codec may include more than two codecs, for example, there may be a third-step codec, a fourthstep codec, etc.
[0126] It is to be noted that first-step codec and first codec may be used interchangeably. Also, second-step codec and second codec may be used interchangeably.
[0127] In one embodiment, a bitstream (such as a second bitstream) that is output by an encoder of the second-step codec may be derived based at least on one or more of the following:The original input, such as the input image;Data derived from the original input (e.g., from the input image), such as features;An output of the first-step codec, such as an initial reconstruction that is output by a decoder of the first-step codec (e.g., the first decoder);Data derived from the initial reconstruction, such as features;Data from which an output of the first-step codec is derived, such as data from which the initial reconstruction is derived, for example, intermediate features that are generated or output by one or more layers of a decoder of the first-step codec;Data or features generated by the decoder of the first-step codec (e.g., the first decoder), which may not necessarily be used to derive the initial reconstruction; and / orData or features generated by the decoder of the first-step codec (e.g., the first decoder), that are derived from other data or features, where also the initial reconstruction (e.g., output of first decoder) may be derived from the other data or features.
[0128] In one example, the second bitstream is derived based on the input image and the initial reconstruction, by using at least a neural network that is part of the encoder of the second-step codec. In another example, the second bitstream is derived based on the input image and features from which the initial reconstruction is derived, by using at least a neural network that is part of the encoder of the second-step codec. In yet another example, the bitstream is derived based on the input image, the initial reconstruction, and features from which the initial reconstruction is derived, by using at least a neural network that is part of the encoder of the second-step codec. Features may comprise numerical values output by one or more neural network layers, such as tensors of data, depending at least on one or more parameters comprised in the one or more neural network layers. In one example, a feature is a three-dimensional tensor whose values depend on both the one or more parameters comprised in the one or more neural network layers and on one or more inputs to the one or more neural network layers.
[0129] In one embodiment, the input to the encoder of the second-step codec may include at least a residual between the original input (e.g., a block or whole image) and the initial reconstruction from the first-step codec. In one example, the residual may be computed by a neural network prior to the encoder of the second-step codec, where an input to the neural network includes at least the original input and the initial reconstruction, and where the residual may be computed based on one or more parameters of the neural network. As used herein, the term ‘residual’ herein may not necessarily refer to a difference signal, or a signal computed based on a sample-wise difference. Instead, the term ‘residual’ is used herein to describe a signal that may be used at decoder side,e.g., in a decoder of the second-step codec, to obtain a final reconstruction, based for example also on the initial reconstruction or on data from which the initial reconstruction is derived. For example, a residual may be first computed by a neural network prior to the encoder of the second-step codec, then the residual is input to the encoder of the second-step codec to obtain a second bitstream, prior to the second bitstream being decoded by a decoder of the second-step codec, where an input to the decoder of the second-step codec includes also the initial reconstruction or data from which the initial reconstruction is derived. The neural network that computes the residual may be considered either to be part of the first codec, or to be part of the second codec, or not to be part of either of the first codec and second codec. However, the neural network that computes the residual may be considered to be comprised in the multi-step codec that comprises also the first codec and the second codec.
[0130] In one embodiment, the residual may be computed by the encoder of the second-step codec, based on the original input and the initial reconstruction from the first-step codec and based on one or more parameters included in a neural network comprised in the encoder of the second-step codec, where the one or more parameters may have been trained based at least on a training algorithm and on a training dataset. In one example, the original input and the initial reconstruction from the first-step codec, or data derived from the original input and the initial reconstruction from the first-step codec, may be concatenated and processed by at least part of the encoder of the second-step codec. In one example, the residual is computed by using a linear or a non-linear process (e.g., one or more neural network layers) applied to the original input and the initial reconstruction from the first-step codec, which may or may not involve a subtraction or a samplewise difference. In another example, the residual is computed by a neural network based encoder that is comprised in the encoder of the second-step codec. In yet another example, the encoder of the second-step codec may comprise a first set of neural network layers that processes the original input to obtain a processed original input, a second set of neural network layers that processes the initial reconstruction (or data derived from the initial reconstruction, or data from which the initial reconstruction is derived, or data that is output by the decoder of the first-step codec) to obtain a processed initial reconstruction, and a third set of neural network layers that processes the processed original input and the processed initial reconstruction to obtain a latent tensor, where the latent tensor may then be used to derive the second bitstream.
[0131] In one embodiment, an input to the encoder of the second-step codec may include the original input and the initial reconstruction from the first-step codec. In one example, the input to the encoder includes only the original input and the initial reconstruction from the first-step codec.
[0132] In one embodiment, an input to the encoder of the second-step codec may include the original input, and one or more features derived from the initial reconstruction of the first-step codec, which may be referred to as first features. The derived features (or first features) may have the same scale (e.g., spatial resolution) or be multi-scale. The encoder of the second-step codec may extract features (e.g., second features) from the original input and process both these features (second features) and the features input to the encoder (first features). The processing methods may comprise, but may not be limited to, one or more of the following: computing the residual, summation, concatenation, applying neural network layers, utilizing attention mechanisms, and other techniques. In one example, a neural network may be applied to extract multi-scale features from the initial reconstruction of the first-step codec, and the extracted features are then fed into the encoder of the second-step codec. The encoder of the second-step codec of this example extracts multi-scale features from the original input, and integrates them with the features from the initial reconstruction of the first-step codec using attention layers at each scale.
[0133] In one embodiment, an input to the encoder of the second-step codec may include the original input, and one or more features (e.g., third features), where the initial reconstruction of the first-step codec was derived from the third features, or the third features are derived from same or substantially same data or features (e.g., fourth features) from which the initial reconstruction was derived. The third features may have the same scale (e.g., one scale or resolution) or be multiscale (e.g., multiple scales or resolutions). In another embodiment, the third features may be processed by additional neural networks before being fed into the encoder of second-step codec. For example, prediction neural networks may be applied before feeding. The encoder of the second-step codec may extract features from the original input and process both these features and the third features. The processing methods may comprise, but may not be limited to, one or more of the following: computing a residual (e.g., by sample-wise difference), summation, concatenation, applying neural network layers, utilizing attention mechanisms, and other techniques. In one example, multi-scale features (e.g., the third features) from which the initial reconstruction of the first-step codec was derived are fed into the encoder of the second-step codec.The encoder of the second-step codec extracts multi-scale features from the original input, and integrates them with the third features using attention layers at each scale.
[0134] Figure 9 depicts an example of this embodiment in which the multi-step codec includes a first-step codec 90 and a second-step codec 91. For simplicity, in the example illustrated in Figure 9, any quantization, dequantization, lossless encoding and lossless decoding is not considered, and the output of the encoder of the first-step codec and the input to the decoder of the first-step codec is a first latent tensor, and the output of the encoder of the second-step codec and the input to the decoder of the second-step codec is a second latent tensor. The first-step encoder neural network 92 of the first-step codec receives the input data xyuv (e.g., an input image comprising YUV format) and generates the first latent tensor zyl, such as by applying one or more convolutional (conv) operations. The first latent tensor (or data derived from the first latent tensor, such as a quantized first latent tensor) is provided to a first-step probability model 94 to estimate the probability of each element of the first latent tensor zyl, where this probability may be used for lossless encoding of the first latent tensor (not shown in Figure 9). The first latent tensor (or data derived from the first latent tensor, such as a dequantized first latent tensor) is provided to a first-step decoder neural network 93 for generating the initial reconstruction xylfrom the first-step codec 90. As shown, the first-step decoder neural network generates the initial reconstruction xylby applying one or more deconvolutional (deconv) operations. It is to be noted that the number and type of neural network layers described here is only an example. The second-step encoder neural network 95 of the second-step codec receives input data xyuv and features from the deconvolutional (deconv) operations of the first-step decoder neural network and generates the second latent tensor zr. In the illustrated embodiment, as a way of example, the second-step encoder neural network includes convolutional (conv) operations and attention layers. The second latent tensor zr (or data derived by it, such as a quantized second latent tensor) is provided to the second-step probability model 97 to estimate the probability of each element of the second latent tensor zr, where this probability may be used for lossless encoding of the second latent tensor. The second latent tensor, or data derived from it, is provided to the second-step decoder neural network 96 for generating the final reconstruction xyfrom the second-step codec. In another embodiment, an input to the encoder of the second-step codec may consist of any combination of inputs as previously described.
[0135] In one embodiment, an input to a probability model 97 in the second-step codec 91 may comprise one or more auxiliary inputs, for example in addition to any previously decoded data such as previously decoded latent tensor elements. The one or more auxiliary inputs may include, but are not limited to, one or more of the following: one or more features which may be samescale or multi-scale derived from the initial reconstruction of the first-step codec; one or more features, which may be same-scale or multi-scale, that are used to derive the initial reconstruction of the first-step codec; the output of any neural network whose input comprises one or more features (which may be same-scale or multi-scale) used to derive the initial reconstruction of the first-step codec; the initial reconstruction of the first-step codec; a latent tensor that is used to derive the initial reconstruction, such as a dequantized latent tensor that is obtained by lossless-decoding a first bitstream that is input to a decoder of the first-step codec, or data / features derived from such latent tensor; scaling values that may be output by the first step probability model and / or any combination of the above.
[0136] In one embodiment, a decoder 96 of the second-step codec 91 may also receive as input one or more second auxiliary inputs, for example in addition to the second bitstream or to a dequantized latent tensor. The one or more second auxiliary inputs to the decoder of the second-step codec may include, but are not limited to, one or more of the following:one or more features which may be same-scale or multi-scale derived from the initial reconstruction of the first-step codec;one or more features, that may be same-scale or multi-scale, and that are used to derive the initial reconstruction of the first-step codec;the output of any neural network whose input includes one or more features (that may be same-scale or multi-scale) used to derive initial reconstruction of the first-step codec;the initial reconstruction of the first-step codec;the first latent tensor that is used to derive the initial reconstruction, such as a dequantized latent tensor that is obtained by lossless-decoding a first bitstream that is input to a decoder of the first-step codec, or data and / or features derived from such latent tensor;the second latent tensor or features derived from the first latent tensor;the quantized first latent tensor or features derived from the quantized first latent tensor; the quantized second latent tensor or features derived from the quantized second latent tensor;one or more scaling values that may be output by the first-step probability model; and / or any combination of the foregoing.
[0137] In another embodiment depicted in Figure 10, a first-step codec 100 includes a first-step encoder neural network 102, a first-step probability model 104 and a first-step decoder neural network 103 and a second-step codec 101 includes a second-step encoder neural network 105, a second-step probability model 107 and a second-step decoder neural network 106. For simplicity, in the example illustrated in Figure 10, any quantization, dequantization, lossless encoding and lossless decoding is not considered, and the output of the encoder of the first-step codec and the input to the decoder of the first-step codec is a first latent tensor, and the output of the encoder of the second-step codec and the input to the decoder of the second-step codec is a second latent tensor. In this embodiment, the decoder of the second-step codec may be configured to process the dequantized latent tensor of the second-step codec and one or more of the auxiliary inputs, as described above. The processing methods may comprise extracting features, computing the residual, summation, concatenation, applying neural network layers, utilizing attention mechanisms, and other techniques. In one example, multi-scale features from which the initial reconstruction of the first-step codec was derived are fed into the decoder of the second-step codec. The decoder of the second-step codec is configured to decode the multi-scale features from the dequantized latent tensor of the second-step codec, and integrates the dequantized latent tensor and the multi-scale features using attention layers at each scale.
[0138] In one embodiment, a post-processing filter, or post-filter for short, may be used as part of multi-step codec that includes a first-step codec and a second-step codec, or as part of the second-step codec, or outside of and separate from the first-step codec and the second-step codec, or outside of and separate from the multi-step codec. In one example, an input to the post-filter is an output of the decoder of the second-step codec, or data derived from the output of the decoder. In another example, an input to the post-filter is an output of the first-step codec and / or the second-step codec. The inputs to the post filter may include, but may not be limited to, one or more of the following:the reconstruction of the second-step codec;one or more features, which may be same-scale or multi-scale, derived from the initial reconstruction of the first-step codec;one or more features, that may be same-scale or multi-scale, and that are used to derive the initial reconstruction of the first-step codec;the output of any neural network whose input includes one or more features that may be same-scale or multi-scale and that are used to derive the initial reconstruction of the first-step codec;the initial reconstruction of the first-step codec;a quantized latent tensor that is used to derive the initial reconstruction, such as a quantized latent tensor that is obtained by lossless-decoding a first bitstream that is input to a decoder of the first-step codec, or data / features derived from such latent tensor;a quantized latent tensor that is used to derive a final reconstruction, such as a quantized latent tensor that is obtained by lossless-decoding the second bitstream that is input to the decoder of the second-step codec, or data / features derived from the latent tensor;intermediate features computed by the decoder of the second-step codec; intermediate features computed by the decoder of the first-step codec;scaling values that may be output by the 1st step and / or 2nd step probability model; and / or any combination of lists above.
[0139] In some embodiments, the initial reconstruction of the first-step codec may not be utilized at inference time such that the first-step codec is used only during the training of the first-step codec. At inference time, one or more features (that may be same-scale or multi-scale), which were used while training the first-step codec to derive the initial reconstruction of the first-step codec based at least on one or more neural network layers, are used as input to one or more components of the second-step codec. In an additional embodiment, the one or more neural network layers may be dropped or not used or not be included in the first-step codec at the inference time.
[0140] In any of the above embodiments, the first-step codec may receive as input the original input and may output one or more of the following:an initial reconstruction that comprises a lower resolution than the original input;an initial reconstruction that comprises a lower quality than the original input;an initial reconstruction that targets or is optimized for a first quality metric, such as a first objective quality metric (e.g., PSNR). In one example, the final reconstruction that is output by the second-step codec may target or may be optimized for a second quality metric, such as a second objective quality metric (e.g., VMAF), or a subjective quality metric; and / oran initial reconstruction that comprises a first color space, such as YUV. In one example the final reconstruction that is output by the second-step codec may include a second color space, such as RGB.
[0141] In one embodiment, the first-step codec and the second-step codec may target different aspects of an image or video, or different objectives, or different qualities. In one example, the first-step codec may target a first spatial resolution and the second-step codec may target a second spatial resolution. In another example, the first-step codec may target a first quality level and the second-step codec may target a second quality level. In yet another example, the first-step codec may code a first frequency band and the second-step codec may code a second frequency band. In yet another example, the first-step codec may target a first quality metric and the second-step codec may target a second quality metric. In yet another example, the first-step codec may reconstruct data in a first color space, and the second-step codec may reconstruct data in a second color space. In yet another example, the first-step codec may code or reconstruct a first portion of an image or video, and the second-step codec may code or reconstruct a second portion of the image or video. In yet another example, the first-step codec may code or reconstruct a first portion of bits, or a first precision, of an image or video, and the second-step codec may code or reconstruct a second portion of bits, or a second precision, of the image or video.
[0142] In one embodiment, the first step codec may be the chroma codec, and the second step codec may be the luma codec. In another embodiment, the first step codec may be the chroma codec, and the second step codec may be the residual chroma codec. In another embodiment, the first step codec may be the luma codec and the second step codec may be the chroma codec. In yet another embodiment, the first step codec may be the luma codec and the second step codec may be the residual luma codec.
[0143] In one embodiment, when some of the embodiments are applied to an end-to-end learned inter-frame codec or a codec that encodes and decodes frames based on one or more other frames of a video, one of more steps may target coding the temporal information, such as motion vectors between one or more frames or features, while the other steps may target spatial information. In one example, the first-step codec may target motion, the second-step codec may target luma and the third-step codec may target chroma.
[0144] Figure 11 illustrates an example block diagram of an apparatus 110 that may be configured to operate as or be embodied as a multi-step codec in accordance with one or more embodimentsof the present disclosure. The apparatus includes, for example, at least one processor 112 and at least one memory 114 storing instructions 115 that, when executed by the at least one processor, cause the apparatus at least to perform the method or methods as disclosed herein, and one or more embodiments thereof. In an example, the at least one memory and the instructions (e.g., a computer program code, software), are configured, with the at least one processor, to cause the apparatus to perform the method or methods as disclosed herein, and one or more embodiments thereof.
[0145] A processor 112 may include circuitry, or be constituted as circuitry or circuitries, the circuitry or circuitries being configured to perform phases of methods in accordance with one or more example embodiments described herein. As used in this application, the term “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations, such as implementations in only analog and / or digital circuitry, and (b) combinations of hardware circuits and software, such as, as applicable: (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a user equipment, to perform various functions) and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation. This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
[0146] The memory 114 may be implemented using any suitable data storage technology. The memory may include a database for storing data. The memory may be at least in part external to apparatus 110 but accessible to apparatus.
[0147] The instructions 115 may be included in a computer readable medium or a non-transitory computer readable medium. A term non-transitory, as used herein, is a limitation of the medium itself (e.g., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., random access memory, RAM, vs. read only memory, ROM).
[0148] In some examples, the apparatus 110 may include a radio interface 116. The radio interface may provide the apparatus with communication capabilities. The radio interface may include a receiver configured to receive information in accordance with at least one cellular or non-cellular standard. The radio interface may include a transmitter configured to transmit information in accordance with at least one cellular or non-cellular standard. The receiver may include more than one receiver. The transmitter may include more than one transmitter. The radio interface may include a transceiver configured to receive and transmit information in accordance with at least one cellular or non-cellular standard.
[0149] In some examples, the apparatus 110 may include a user interface 118 including, for example, at least one of a keypad, a microphone, a touch display, a display, a speaker, etc. The user interface may be used to control the apparatus by the user. The user interface may be external to the apparatus. For example, the apparatus may be connected to another device, such as a computer, either via wireless or wired connection, and the apparatus is controlled by the user via the computer.
[0150] In at least one embodiment, at least some of the processes described herein may be carried out by an apparatus including means for carrying out at least some of the described processes. Means for performing methods as disclosed herein may include software and / or hardware components of the apparatus 110. For example, the at least one processor 112, the memory 114, and the computer program code form means for carrying out the method or methods as disclosed herein, and one or more embodiments thereof. The term “means” as used in the description and in the claims may refer to one or more individual elements configured to perform the corresponding recited functionality or functionalities, or it may refer to several elements that perform such functionality or functionalities. Furthermore, several functionalities recited in the claims may be performed by the same individual means or the same combination of means. For example, performing such functionality or functionalities may be caused in an apparatus by a processor that executes instructions stored in a memory of the apparatus.
[0151] The apparatus 110 may be embodied by a multi-step codec as shown in Figure 12. In this embodiment, the apparatus, such as the at least one processor 112, is configured to function as a first codec 1200 including at least a first encoder 1204 configured to receive and encode input data (e.g., an original input) to generate a first bitstream and a first decoder 1206 configured to generate first reconstructed data based at least on the first bitstream. The apparatus, such as the at least oneprocessor 112, is also configured to function as a second codec 1202 including at least a second encoder 1224 configured to receive and encode data to generate a second bitstream and a second decoder 1226 configured to generate second reconstructed data based at least on the second bitstream. The data received and encoded by the second encoder comprises or is derived from one or more of the input data, data derived from the input data, the first reconstructed data, data derived from the first reconstructed data, data from which the first reconstructed data is derived, data derived from other data from which the first reconstructed data is also derived, or data representative of one or more intermediate features generated by one or more neural network layers comprised in the first decoder.
[0152] As shown in Figure 12, the first encoder 1204 includes at least a first neural network based encoder 1208 configured to receive and process the input data (e.g., original input, which may be for example the same as an input to the multi-step codec or to the first codec or to an encoder of the multi-step codec) to generate a first latent tensor, a first quantization operation 1210 configured to quantize the first latent tensor to obtain a quantized first latent tensor, and a first lossless or substantially lossless encoder 1212 configured to encode the quantized first latent tensor to obtain the first bitstream. The first decoder 1206 of this embodiment also includes at least a first lossless or substantially lossless decoder 1216 configured to decode the first bitstream to obtain a quantized second latent tensor, a first dequantization operation 1220 configured to dequantize the quantized second latent tensor to obtain a second latent tensor (which may be referred to also as dequantized second latent tensor), and a first neural network based decoder 1222 configured to generate the first reconstructed data based at least on the second latent tensor. The first reconstructed data may be also referred to as initial reconstruction. It is to be noted that the quantized first latent tensor and the quantized second latent tensor may be the same or substantially the same, meaning that values comprised in the quantized first latent tensor may be the same or substantially the same as values comprised in the quantized second latent tensor. Additionally, the second encoder 1224 includes at least a second neural network based encoder 1228 configured to receive and process the data received by the second encoder to generate a third latent tensor, a second quantization operation 1230 configured to quantize the third latent tensor to obtain a quantized third latent tensor, and a second lossless or substantially lossless encoder 1232 configured to encode the quantized third latent tensor to obtain the second bitstream. And, the second decoder 1226 includes at least a second lossless or substantially lossless decoder 1236 configured to decode the secondbitstream to obtain a quantized fourth latent tensor, a second dequantization operation 1240 configured to dequantize the quantized fourth latent tensor to obtain a fourth latent tensor (which may also be referred to as dequantized fourth latent tensor), and a second neural network based decoder 1242 configured to generate the second reconstructed data (which may be referred to also as final reconstruction) based at least on the fourth latent tensor. It is to be noted that the quantized third latent tensor and the quantized fourth latent tensor may be the same or substantially the same, meaning that values comprised in the quantized third latent tensor may be the same or substantially the same as values comprised in the quantized fourth latent tensor.
[0153] The first encoder 1204 of the illustrated embodiment may further include a first probability model 1214 configured to estimate a first probability of respective elements of the quantized first latent tensor. The first lossless or substantially lossless encoder 1212 is configured to encode the quantized first latent tensor to obtain the first bitstream based on the first probability. Additionally, the first decoder 1206 may further include a second probability model 1218 configured to estimate a second probability of respective elements of a quantized second latent tensor. The first lossless or substantially lossless decoder 1216 is configured to decode the first bitstream to obtain the quantized second latent tensor based at least on the second probability. In at least some embodiments, the first probability model is the same as or substantially the same as the second probability model. In one example, the first probability model may be a copy of or otherwise identical to the second probability model.
[0154] As also shown in Figure 12, the second encoder 1224 may further include a third probability model 1234 configured to estimate a third probability of respective elements of the quantized third latent tensor. The second lossless or substantially lossless encoder 1232 is configured to encode the quantized third latent tensor to obtain the second bitstream based at least on the third probability. And, the second decoder 1226 may further include a fourth probability model 1238 configured to estimate a fourth probability of respective elements of a quantized fourth latent tensor. In addition, the second lossless or substantially lossless decoder 1236 is configured to decode the second bitstream to obtain the quantized fourth latent tensor based at least on the fourth probability. In at least some embodiments, the third probability model is the same as or substantially the same as the fourth probability model.
[0155] It is to be noted that a neural network based encoder or decoder, such as the first or second neural network based encoder or the first or second neural network based decoder, may be a neuralnetwork, such as a convolutional neural network, or a Transformer neural network, or a neural network that comprises, but may not be limited to, one or more of the following: convolutional layers, Transformer layers, attention layers, normalization layers, fully-connected layers, nonlinear activation functions, and the like. A neural network based encoder may map or convert or transform an input (for example, an image) to a latent tensor (e.g., features). A neural network based decoder may map or convert or transform a latent tensor (e.g., features) to data in a target domain, such as in a same domain as the input to a neural network based encoder that obtained the latent tensor, for example an image.
[0156] It is to be noted that, in some cases, the first codec and / or the second codec may not comprise a probability model (e.g., the first probability model, the second probability model, the third probability model, or the fourth probability model), for example when lossless coding is performed based on factorized probabilities or based on a factorized prior.
[0157] In one embodiment, the apparatus 110 also includes a neural network configured to process at least one of the data derived from the input data or the data derived from the first reconstructed data to generate a residual. In this embodiment, the second encoder 1224 of the second codec 1202 is configured to receive and encode the residual to generate the second bitstream.
[0158] The second encoder 1224 of the second codec 1202 of an example embodiment includes a neural network that is configured to process at least one of the data derived from the input data or the data derived from the first reconstructed data based on one or more trained parameters of the neural network to generate a residual.
[0159] In another embodiment, the second encoder 1224 of the second codec 1202 is configured to receive one or more first features derived from the first reconstructed data, extract one or more second features from the input data and process both the one or more first features and the one or more second features to generate the second bitstream.
[0160] The second encoder 1224 of the second codec 1202 of an example embodiment is configured to receive one or more first features from which the first reconstructed data was derived or one or more first features derived from data from which the first reconstructed data was derived, extract one or more second features from the input data and process both the one or more first features and the one or more second features to generate the second bitstream.
[0161] The one or more first features and the one or more second features may be multi-scale features. In this embodiment, the second neural network based encoder 1228 of the second encoder1224 includes one or more atention layers at each scale that are configured to integrate the one or more first features and the one or more second features, such as shown in Figure 9.
[0162] In an example embodiment, the third probability model 1234 and the fourth probability model 1238 are also configured to receive auxiliary data. The auxiliary data includes one or more of the following: one or more features derived from the first reconstructed data of the first codec 1200, one or more features from which the first reconstructed data of the first codec is derived, an output of a neural network having an input comprising one or more features from which the first reconstructed data of the first codec is derived, the first reconstructed data of the first codec, the first latent tensor or features derived from the first latent tensor, the second latent tensor or features derived from the second latent tensor, the quantized first latent tensor or features derived from the quantized first latent tensor, the quantized second latent tensor or features derived from the quantized second latent tensor, features extracted by the first decoder or by the first neural network based decoder, data derived from other data from which the first reconstructed data was also derived, or one or more scaling values provided by the first probability model or by the second probability model.
[0163] The second decoder 1226 or the neural network based second decoder 1242 is also configured to receive auxiliary data. The auxiliary data includes one or more of the following: one or more features derived from the first reconstructed data of the first codec 1200, one or more features from which the first reconstructed data of the first codec is derived, an output of a neural network having an input comprising one or more features from which the first reconstructed data of the first codec is derived, the first reconstructed data of the first codec, the quantized first latent tensor or features derived from the quantized first latent tensor, the quantized second latent tensor or features derived from the quantized second latent tensor, the second latent tensor or features derived from the second latent tensor, features extracted by the first decoder or by the first neural network based decoder, data derived from other data from which the first reconstructed data was also derived, one or more scaling values provided by the first probability model 1214, or one or more scaling values provided by the second probability model 1218.
[0164] The second decoder 1226 of the second codec 1202 of an example embodiment is configured to receive one or more first features from which the first reconstructed data was derived, derive one or more second features from the fourth latent tensor of the second codec and processboth the one or more first features and the one or more second features to generate the second reconstructed data.
[0165] In an example embodiment, the one or more first features and the one or more second features are multi-scale features. In this embodiment, the second neural network-based decoder 1242 of the second decoder 1226 includes one or more attention layers at each scale that are configured to integrate the one or more first features and the one or more second features, such as shown in Figure 10.
[0166] The apparatus of an example embodiment may also be configured to serve as a postprocessing filter configured to provide a filtered output based upon the second reconstructed data of the second codec 1202 and one or more auxiliary inputs comprising one or more of the following: one or more features derived from the first reconstructed data of the first codec 1200, one or more features from which the first reconstructed data of the first codec is derived, an output of a neural network based on one or more features from which the first reconstructed data is derived, the first reconstructed data of the first codec, the quantized first latent tensor or one or more features derived from the quantized first latent tensor, the quantized second latent tensor or one or more features derived from the quantized second latent tensor, the quantized third latent tensor or one or more features derived from the quantized third latent tensor, the quantized fourth latent tensor or one or more features derived from the quantized fourth latent tensor, the second latent tensor or one or more features derived from the second latent tensor, the fourth latent tensor or one or more features derived from the fourth latent tensor, one or more intermediate features computed by the second decoder 1226 of the second codec, one or more intermediate features computed by the first decoder 1206 of the first codec 1200, data derived from other data from which the first reconstructed data was also derived, data derived from yet other data from which the second reconstructed data was also derived, and / or one or more scaling values provided by one or both of the second probability model 1218 and the fourth probability model 1238.
[0167] In some embodiments, the term “substantially lossless” is used, for example substantially lossless coding, or substantially lossless compression, or substantially lossless encoder, or substantially lossless decoder, or substantially lossless codec. “Substantially lossless” compression may refer to a compression method in which the original information or signal that is compressed is preserved with such high fidelity that any discrepancies between the original information and the decoded information are minimal and do not materially impair its function or utility. Whileminor variations or losses may occur after compression, these may be considered to be confined within predetermined acceptable thresholds such that the essential characteristics and operational effectiveness of the original information remain essentially unchanged.
[0168] In some embodiments, the “substantially lossless encoder” and “substantially lossless decoder” are the respective methodologies used for substantially lossless compression and decompression.
[0169] In some embodiments, two models, neural networks, or tensors are considered “substantially the same” if their structural configurations, internal representations, e.g. parameters and weights, and functional outputs, whenever applicable, are preserved with such high fidelity that any discrepancies between them are minimal and do not materially impair their intended performance or utility. While minor variations in parameters, architecture, or numerical values may be present, these differences are confined within predetermined acceptable thresholds such that the essential characteristics and operational effectiveness of the models or tensors remain essentially unchanged.
[0170] Figure 13 illustrates an example flowchart of a method to which one or more examples disclosed herein may be applied. The method may be computer-implemented. The method may be performed by an apparatus 110, such as shown in Figure 11, that may be embodied as or may be part of a multi-step codec, such as illustrated in Figure 12. In this example embodiment, the apparatus includes means, such as the processor 112, the first encoder 1204 or the like, for receive and encode input data to generate a first bitstream. See block 1302. The apparatus of this example embodiment also includes means, such as the processor 112, the first decoder 1206 or the like, for generating first reconstructed data based at least on the first bitstream. See block 1304.
[0171] The apparatus 110 also includes means, such as the processor 112, the second encoder 1224 or the like, for receiving and encoding data to generate a second bitstream. In this regard, the received data includes or is derived from one or more of the input data, data derived from the input data, the first reconstructed data, data derived from the first reconstructed data, data from which the first reconstructed data is derived, or data representative of one or more intermediate features generated by one or more layers of the first decoder 1206. See block 1306. The apparatus of this example embodiment further includes means, such as the processor 112, the second decoder 1226 or the like, for generating second reconstructed data based at least on the second bitstream. See block 1308.
[0172] Figure 13 is a flowchart depicting a method according to an example embodiment of the present disclosure. It will be understood that each block of the flowchart and combination of blocks in the flowchart may be implemented by various means, such as hardware, firmware, processor, circuitry, and / or other communication devices associated with execution of software including one or more computer program instructions. For example, one or more of the procedures described above may be embodied by computer program instructions. In this regard, the computer program instructions which embody the procedures described above may be stored by a memory device 114 of an apparatus 110 employing an embodiment and executed by a processor 112. As will be appreciated, any such computer program instructions may be loaded into a computer or other programmable apparatus (for example, hardware) to produce a machine, such that the resulting computer or other programmable apparatus implements the functions specified in the flowchart blocks. These computer program instructions may also be stored in a computer-readable memory that may direct a computer or other programmable apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture the execution of which implements the function specified in the flowchart blocks. The computer program instructions may also be loaded into a computer or other programmable apparatus to cause a series of operations to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide operations for implementing the functions specified in the flowchart blocks.
[0173] Accordingly, blocks of the flowchart support combinations of means for performing the specified functions and combinations of operations for performing the specified functions for performing the specified functions. It will also be understood that one or more blocks of the flowchart, and combinations of blocks in the flowchart, can be implemented by special purpose hardware-based computer systems which perform the specified functions, or combinations of special purpose hardware and computer instructions.
[0174] Many modifications and other embodiments set forth herein will come to mind to one skilled in the art to which this disclosure pertains having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the disclosure is not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims
[0175] Moreover, although the foregoing descriptions and the associated drawings describe example embodiments in the context of certain example combinations of elements and / or functions, it should be appreciated that different combinations of elements and / or functions may be provided by alternative embodiments without departing from the scope of the appended claims. In this regard, for example, different combinations of elements and / or functions than those explicitly described above are also contemplated as may be set forth in some of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
THAT WHICH IS CLAIMED:
1. An apparatus comprising:a first codec comprising at least a first encoder configured to receive and encode input data to generate a first bitstream and a first decoder configured to generate first reconstructed data based at least on the first bitstream; anda second codec comprising at least a second encoder configured to receive and encode data to generate a second bitstream and a second decoder configured to generate second reconstructed data based at least on the second bitstream,wherein the data received and encoded by the second encoder comprises or is derived from one or more of the input data, data derived from the input data, the first reconstructed data, data derived from the first reconstructed data, data from which the first reconstructed data is derived, or data representative of one or more intermediate features generated by one or more layers of the first decoder.
2. An apparatus according to Claim 1, wherein the first encoder comprises at least a first neural network based encoder configured to receive and process the input data to generate a first latent tensor, a first quantization operation configured to quantize the first latent tensor to obtain a quantized first latent tensor, and a first lossless or substantially lossless encoder configured to encode the quantized first latent tensor to obtain the first bitstream and wherein the first decoder comprises at least a first lossless or substantially lossless decoder configured to decode the first bitstream to obtain a quantized second latent tensor, a first dequantization operation configured to dequantize the quantized second latent tensor to obtain a second latent tensor, and a first neural network based decoder configured to generate the first reconstructed data based at least on the second latent tensor; andthe second encoder comprises at least a second neural network based encoder configured to receive and process the data received by the second encoder to generate a third latent tensor, a second quantization operation configured to quantize the third latent tensor to obtain a quantized third latent tensor, and a second lossless or substantially lossless encoder configured to encode the quantized third latent tensor to obtain the second bitstream and wherein the second decoder comprises at least a second lossless or substantially lossless decoder configured to decode thesecond bitstream to obtain a quantized fourth latent tensor, a second dequantization operation configured to dequantize the quantized fourth latent tensor to obtain a fourth latent tensor, and a second neural network based decoder configured to generate the second reconstructed data based at least on the fourth latent tensor.
3. An apparatus according to Claim 2, wherein the first encoder further comprises a first probability model configured to estimate a first probability of respective elements of the quantized first latent tensor, and wherein the first lossless or substantially lossless encoder is configured to encode the quantized first latent tensor to obtain the first bitstream based on the first probability, and wherein the first decoder further comprises a second probability model configured to estimate a second probability of respective elements of a quantized second latent tensor, and wherein the first lossless or substantially lossless decoder is configured to decode the first bitstream to obtain the quantized second latent tensor based at least on the second probability.
4. An apparatus according to Claims 2 or 3, wherein the second encoder further comprises a third probability model configured to estimate a third probability of respective elements of the quantized third latent tensor, and wherein the second lossless or substantially lossless encoder is configured to encode the quantized third latent tensor to obtain the second bitstream based at least on the third probability, and wherein the second decoder further comprises a fourth probability model configured to estimate a fourth probability of respective elements of a quantized fourth latent tensor, and wherein the second lossless or substantially lossless decoder is configured to decode the second bitstream to obtain the quantized fourth latent tensor based at least on the fourth probability.
5. An apparatus according to Claim 1 further comprising a neural network configured to process at least one of the data derived from the input data or the data derived from the first reconstructed data to generate a residual, wherein the second encoder of the second codec is configured to receive and encode the residual to generate the second bitstream.
6. An apparatus according to Claim 1 wherein the second encoder of the second codec comprises a neural network that is configured to process at least one of the data derived from the input data or the data derived from the first reconstructed data based on one or more trained parameters of the neural network to generate a residual.
7. An apparatus according to any one of Claims 1 to 6, wherein the second encoder of the second codec is configured to receive one or more first features derived from the first reconstructed data, extract one or more second features from the input data and process both the one or more first features and the one or more second features to generate the second bitstream.
8. An apparatus according to any one of Claims 1 to 6, wherein the second encoder of the second codec is configured to receive one or more first features from which the first reconstructed data was derived or one or more first features derived from data from which the first reconstructed data was derived, extract one or more second features from the input data and process both the one or more first features and the one or more second features to generate the second bitstream.
9. An apparatus according to Claims 7 or 8, wherein the one or more first features and the one or more second features are multi-scale features, and wherein the second encoder of the second codec comprises a second neural network based encoder having one or more attention layers at each scale that are configured to integrate the one or more first features and the one or more second features.
10. An apparatus according to claim 3, wherein the first probability model is same or substantially same as the second probability model.
11. An apparatus according to claim 4, wherein the third probability model is same or substantially the same as the fourth probability model.
12. An apparatus according to any one of Claims 4 or 11, wherein the third probability model and the fourth probability model are also configured to receive auxiliary data comprising one or more of the following:one or more features derived from the first reconstructed data of the first codec, one or more features from which the first reconstructed data of the first codec is derived, an output of a neural network having an input comprising one or more features from which the first reconstructed data of the first codec is derived,the first reconstructed data of the first codec,the first latent tensor or features derived from the first latent tensor,the second latent tensor or features derived from the second latent tensor,the quantized first latent tensor or features derived from the quantized first latent tensor, the quantized second latent tensor or features derived from the quantized second latent tensor,features extracted by the first decoder or by a first neural network based decoder, data derived from other data from which the first reconstructed data was also derived, or one or more scaling values provided by the first probability model or by the second probability model.
13. An apparatus according to any one of Claims 1 to 12, wherein the second decoder or the neural network based second decoder is also configured to receive auxiliary data comprising one or more of the following:one or more features derived from the first reconstructed data of the first codec, one or more features from which the first reconstructed data of the first codec is derived, an output of a neural network having an input comprising one or more features from which the first reconstructed data of the first codec is derived,the first reconstructed data of the first codec,the quantized first latent tensor or features derived from the quantized first latent tensor, the quantized second latent tensor or features derived from the quantized second latent tensor,the second latent tensor or features derived from the second latent tensor,one or more scaling values provided by the first probability model,features extracted by the first decoder or by a first neural network based decoder, data derived from other data from which the first reconstructed data was also derived, orone or more scaling values provided by the second probability model.
14. An apparatus according to Claim 1, wherein the second decoder of the second codec is configured to receive one or more first features from which the first reconstructed data was derived, derive one or more second features from the fourth latent tensor of the second codec and process both the one or more first features and the one or more second features to generate the second reconstructed data.
15. An apparatus according to Claim 14, wherein the one or more first features and the one or more second features are multi-scale features, and wherein the second decoder of the second codec comprises a second neural network based encoder having one or more attention layers at each scale that are configured to integrate the one or more first features and the one or more second features.
16. An apparatus according to any one of Claims 1 to 15, further comprising a postprocessing filter configured to provide a filtered output based upon the second reconstructed data of the second codec and one or more auxiliary inputs comprising one or more of the following:one or more features derived from the first reconstructed data of the first codec, one or more features from which the first reconstructed data of the first codec is derived, an output of a neural network based on one or more features from which the first reconstructed data is derived,the first reconstructed data of the first codec,the quantized first latent tensor or one or more features derived from the quantized first latent tensor,the quantized second latent tensor or one or more features derived from the quantized second latent tensor,the quantized third latent tensor or one or more features derived from the quantized third latent tensor,the quantized fourth latent tensor or one or more features derived from the quantized fourth latent tensor,the second latent tensor or one or more features derived from the second latent tensor, the fourth latent tensor or one or more features derived from the fourth latent tensor, one or more intermediate features computed by the second decoder of the second codec, one or more Intermediate features computed by the first decoder of the first codec, data derived from other data from which the first reconstructed data was also derived, data derived from other data from which the second reconstructed data was also derived, orone or more scaling values provided by one or both of the second and fourth probability models.
17. A method comprising:receiving and encoding input data with at least a first encoder of a first codec to generate a first bitstream;generating first reconstructed data based at least on the first bitstream with a first decoder of the first codec;receiving and encoding data with at least a second encoder of a second codec to generate a second bitstream; andgenerating second reconstructed data based at least on the second bitstream with a second decoder,wherein the data received and encoded by the second encoder comprises or is derived from one or more of the input data, data derived from the input data, the first reconstructed data, data derived from the first reconstructed data, data from which the first reconstructed data is derived, or data representative of one or more intermediate features generated by one or more layers of the first decoder.
18. A method according to Claim 17, wherein receiving and processing the input data with the at least the first encoder comprises receiving and processing the input data with at least a first neural network based encoder to generate a first latent tensor, wherein the method further comprises quantizing the first latent tensor with a first quantization operation to obtain aquantized first latent tensor, wherein receiving and processing the input data with the at least the first encoder comprises encoding the quantized first latent tensor with a first lossless or substantially lossless encoder to obtain the first bitstream, wherein generating the first reconstructed data comprises decoding the first bitstream to obtain a quantized second latent tensor with at least a first lossless or substantially lossless decoder, wherein the method further comprises dequantizing the quantized second latent tensor with a first dequantization operation to obtain a second latent tensor, and wherein generating the first reconstructed data further comprises generating the first reconstructed data with a first neural network based decoder based at least on the second latent tensor; andwherein receiving and processing data with at least the second encoder comprises receiving and processing the data with at least a second neural network based encoder to generate a third latent tensor, wherein the method further comprises quantizing the third latent tensor with a second quantization operation to obtain a quantized third latent tensor, wherein receiving and processing data with at least the second encode further comprises encoding the quantized third latent tensor with a second lossless or substantially lossless encoder to obtain the second bitstream, wherein generating the second reconstructed data comprises decoding the second bitstream with at least a second lossless or substantially lossless decoder to obtain a quantized fourth latent tensor, wherein the method further comprises dequantizing the quantized fourth latent tensor with a second dequantization operation to obtain a fourth latent tensor, and wherein generating the second reconstructed data further comprises generating the second reconstructed data with a second neural network based decoder based at least on the fourth latent tensor.
19. A method according to Claim 18, wherein the method further comprises estimating a first probability of respective elements of the quantized first latent tensor with a first probability model of the first encoder, wherein encoding the quantized first latent tensor with the first lossless or substantially lossless encoder comprises encoding the quantized first latent tensor to obtain the first bitstream based on the first probability, wherein the method further comprises estimating a second probability of respective elements of a quantized second latent tensor with a second probability model of the first decoder, and wherein decoding the first bitstream with thefirst lossless or substantially lossless decoder comprises decoding the first bitstream to obtain the quantized second latent tensor based at least on the second probability.
20. A method to Claims 18 or 19, wherein the method further comprises estimating a third probability of respective elements of the quantized third latent tensor with a third probability model of the second encoder, wherein encoding the quantized third latent tensor with the second lossless or substantially lossless encoder comprises encoding the quantized third latent tensor to obtain the second bitstream based at least on the third probability, wherein the method further comprises estimating a fourth probability of respective elements of a quantized fourth latent tensor with a fourth probability model of the second decoder, and wherein decoding the second bitstream with the second lossless or substantially lossless decoder comprises decoding the second bitstream to obtain the quantized fourth latent tensor based at least on the fourth probability.
21. A method according to Claim 17 further comprising processing at least one of the data derived from the input data or the data derived from the first reconstructed data with a neural network to generate a residual; and receiving and encoding the residual with the second encoder of the second codec to generate the second bitstream.
22. A method according to Claim 17 further comprising processing, with a neural network of the second encoder of the second coded, at least one of the data derived from the input data or the data derived from the first reconstructed data based on one or more trained parameters of the neural network to generate a residual.
23. A method according to any one of Claims 17 to 22, further comprising receiving, with the second encoder of the second codec, one or more first features derived from the first reconstructed data, extracting one or more second features from the input data and processing both the one or more first features and the one or more second features to generate the second bitstream.
24. A method according to any one of Claims 17 to 22, further comprising receiving, with the second encoder of the second coded, one or more first features from which the first reconstructed data was derived or one or more first features derived from data from which the first reconstructed data was derived, extracting one or more second features from the input data and processing both the one or more first features and the one or more second features to generate the second bitstream.
25. A method according to Claims 23 or 24, wherein the second encoder of the second codec comprises a second neural network based encoder having one or more attention layers, wherein the one or more first features and the one or more second features are multi-scale features, and wherein the method further comprises integrating the one or more first features and the one or more second features with the one or more attention layers at each scale.
26. A method according to claim 19, wherein the first probability model is same or substantially same as the second probability model.
27. A method according to claim 10, wherein the third probability model is same or substantially the same as the fourth probability model.
28. A method according to any one of Claims 20 or 27, wherein the third probability model and the fourth probability model are also configured to receive auxiliary data comprising one or more of the following:one or more features derived from the first reconstructed data of the first codec, one or more features from which the first reconstructed data of the first codec is derived, an output of a neural network having an input comprising one or more features from which the first reconstructed data of the first codec is derived,the first reconstructed data of the first codec,the first latent tensor or features derived from the first latent tensor,the second latent tensor or features derived from the second latent tensor,the quantized first latent tensor or features derived from the quantized first latent tensor,the quantized second latent tensor or features derived from the quantized second latent tensor,features extracted by the first decoder or by a first neural network based decoder, data derived from other data from which the first reconstructed data was also derived, or one or more scaling values provided by the first probability model or by the second probability model.
29. A method according to any one of Claims 17 to 28, further comprising receiving auxiliary data with the second decoder or the neural network based second decoder, and wherein the auxiliary data comprises one or more of the following:one or more features derived from the first reconstructed data of the first codec, one or more features from which the first reconstructed data of the first codec is derived, an output of a neural network having an input comprising one or more features from which the first reconstructed data of the first codec is derived,the first reconstructed data of the first codec,the quantized first latent tensor or features derived from the quantized first latent tensor, the quantized second latent tensor or features derived from the quantized second latent tensor,the second latent tensor or features derived from the second latent tensor,one or more scaling values provided by the first probability model,features extracted by the first decoder or by a first neural network based decoder, data derived from other data from which the first reconstructed data was also derived, or one or more scaling values provided by the second probability model.
30. A method according to Claim 17, further comprising receiving, with the second decoder of the second codec, one or more first features from which the first reconstructed data was derived, deriving one or more second features from the fourth latent tensor of the second codec and processing both the one or more first features and the one or more second features to generate the second reconstructed data.6031. A method according to Claim 30, wherein the one or more first features and the one or more second features are multi-scale features, wherein the second decoder of the second codec comprises a second neural network based encoder having one or more attention layers, and wherein the method further comprises integrating the one or more first features and the one or more second features with the one or more attention layers at each scale.
32. A method according to any one of Claims 17 to 31, further comprising providing a filtered output with a post-processing filter based upon the second reconstructed data of the second codec and one or more auxiliary inputs comprising one or more of the following:one or more features derived from the first reconstructed data of the first codec, one or more features from which the first reconstructed data of the first codec is derived, an output of a neural network based on one or more features from which the first reconstructed data is derived,the first reconstructed data of the first codec,the quantized first latent tensor or one or more features derived from the quantized first latent tensor,the quantized second latent tensor or one or more features derived from the quantized second latent tensor,the quantized third latent tensor or one or more features derived from the quantized third latent tensor,the quantized fourth latent tensor or one or more features derived from the quantized fourth latent tensor,the second latent tensor or one or more features derived from the second latent tensor, the fourth latent tensor or one or more features derived from the fourth latent tensor, one or more intermediate features computed by the second decoder of the second codec, one or more Intermediate features computed by the first decoder of the first codec, data derived from other data from which the first reconstructed data was also derived, data derived from other data from which the second reconstructed data was also derived, or61one or more scaling values provided by one or both of the second and fourth probability models.