Signaling frame encoding by neural network
Patent Information
- Application Number
- PCT/EP2026/053135
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-20
- Filing Date
- 2026-02-06
- Publication Date
- 2026-08-27
Smart Images

Figure EP2026053135_27082026_PF_FP_ABST
Abstract
Description
[0001] SIGNALING FRAME ENCODING BY NEURAL NETWORK
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the priority to EP Patent Application No. 25305229.4, filed February 20, 2025, the entire disclosure of which is incorporated herein by reference.
[0003] BACKGROUND
[0004] The present application is related to generally relate to image, video and / or 3D scene compression combining data coded with traditional standardized codec and neural network based end-to-end codec. The present embodiments relate to a method and an apparatus for encoding, decoding, transmitting metadata used by a decoder for reconstructing at least one part of a picture end-to-end coded using neural networks.
[0005] BRIEF SUMMARY
[0006] In various implementations, methods and devices are disclosed that indicate to a decoder that at least a part of a picture was coded using a neural network-based auto-encoder while other frames or other parts of the picture are coded with a traditional standardized codec.
[0007] Briefly stated, in one embodiment a method of video encoding is disclosed that comprises encoding an picture of a video, wherein at least one part of the picture is end-to-end coded using a neural network; encoding one or more syntax elements for reconstructing the at least one part of the picture end-to-end coded using a neural network; and generating a bitstream comprising the coded picture, the one or more syntax elements for reconstructing the at least one part of the picture end-to-end coded using the neural network.
[0008] In one embodiment, a method of video decoding or rendering is disclosed that comprises decoding one or more syntax elements for reconstructing at least one part of an picture end-to-end coded using a neural network; reconstructing the picture of a video using the one or more syntax elements, wherein at least one part of the picture is reconstructed using a neural network.
[0009] One or more embodiments also provide an apparatus for video encoding, decoding or rendering comprising one or more processors and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to perform the encoding, decoding or rendering method according to any of the embodiments described herein.
[0010] One or more embodiments also provide a computer program comprising instructions which when executed by one or more processors cause the one or more processors to perform theencoding, decoding or rendering method according to any of the embodiments described herein. One or more of the present embodiments also provide a computer readable storage medium having stored thereon instructions for encoding, decoding or rendering a video according to the methods described herein.
[0011] One or more embodiments also provide a computer readable storage medium having stored thereon video data generated according to the methods described above. One or more embodiments also provide a method and apparatus for transmitting or receiving the video data generated according to the methods described herein.
[0012] BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The following detailed description will be better understood when read in conjunction with the appended drawings, in which there are shown examples of one or more of the multiple embodiments of the present disclosure. It should be understood, however, that the embodiments described herein are not limited to the precise arrangements and instrumentalities shown in the drawings. In the drawings:
[0014] FIG. 1 is a block diagram illustrating an example system according to one or more embodiments of the present disclosure;
[0015] FIG. 2 is a block diagram illustrating an example video encoder according to one or more embodiments of the present disclosure;
[0016] FIG. 3 is a block diagram illustrating an example video decoder according to one or more embodiments of the present disclosure;
[0017] FIG. 4 is a block diagram illustrating an example end-to-end compression system according to one or more embodiments of the present disclosure;
[0018] FIG. 5 a and FIG. 5b are block diagrams illustrating an encoding method according to one or more embodiments of the present disclosure; and
[0019] FIG. 6a and FIG. 6b are block diagrams illustrating a decoding method according to one or more embodiments of the present disclosure.
[0020] DETAILED DESCRIPTION
[0021] In describing the various embodiments of the present disclosure, certain terminology is used herein for convenience only and should not be considered as limiting such embodiments. In the drawings, the same reference numerals are employed for designating the same elements throughout the several figures and the present description.Referring to the drawings, there is shown in FIG. 1 a block diagram illustrating an example system 100 in which embodiments of the present disclosure can be implemented. The system 100 may be an electronic device including, for example, a personal computer, laptop computer, mobile phone, tablet computer, multimedia set-top box, digital television receiver, personal video recording system, connected home appliance, vehicle control and / or entertainment system, and server. One or more elements of the system 100, singly or in combination, may be implemented as an integrated circuit (IC), multiple ICs, and / or discrete components. For example, in one embodiment, the processing, encoding and / or decoding elements of system 100 are distributed across multiple ICs and / or discrete components. In some embodiments, the system 100 is communicatively coupled to and / or in communication with other systems or devices, via, for example, a communications bus or dedicated input / output ports.
[0022] One or more of the elements of system 100 may be provided within an integrated housing, with such elements being interconnected and able to transmit data therebetween using any suitable connection arrangement 115 generally known in the art, including, for example, an internal bus (e.g., I2C bus), wiring, and printed circuit boards.
[0023] The system 100 includes at least one processor 110 configured to execute instructions for implementing the embodiments described herein, including signal / data coding and processing. The processor 110 may be a general-purpose processor or microprocessor, digital signal processor (DSP), one or more microprocessors in association with a DSP core, a controller, a microcontroller, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), a state machine, and the like. The processor 110 may include at least one central processing unit (CPU), embedded memory, input and output interfaces, and other circuitries.
[0024] The system 100 includes at least one memory 120, for example, a volatile memory device and / or anon-volatile memory device. The system 100 includes a storage device 140, that may be or include non-volatile memory and / or dynamic volatile memory, including EEPROM, ROM, PROM, RAM, DRAM, SRAM, DDR, flash, magnetic disk drives, solid state drives (SSD) and / or optical disk drives. The storage device 140 may be or include, for example, an internal storage device, an attached storage device, and / or a network accessible storage device. Although shown separately, the memory 120 and the storage device 140 may be collocated, integrated together, or otherwise combined.
[0025] The system 100 includes an encoder / decoder module 130 configured to process video data and to provide encoded video data or decoded video data. The encoder / decoder module 130 may include one or more processors and / or memory (not shown). Although FIG. 1 depicts the encoder / decoder module 130 as a separate element of system 100, it will be understood that theprocessor 110 and the encoder / decoder module 130 may be collocated and / or integrated together as a combination of hardware and / or software, e.g., in an electronic package or chip. The encoder / decoder module 130 may be or include one or more modules that may be included in one or more separate devices that perform encoding and / or decoding functions.
[0026] Instructions for execution by the processor 110 and / or the encoder / decoder module 130 may be stored in the storage device 140 and subsequently loaded into memory 120 for execution by the processor 110. In some embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more items when performing the processes disclosed herein. Such items may include input video, decoded video or portions thereof, bitstreams, matrices, variables, operational logic, and intermediate and / or final results from processing of equations, formulas, or operations.
[0027] In some embodiments, the memory of the processor 110 and / or the encoder / decoder module 130 is used to store instructions and / or provide working memory for video encoding and decoding functions. In some embodiments, memory external to the processor 110 and / or the encoder / decoder module 130 (e.g., the memory 120 and / or the storage device 140) is used for one or more of these functions and / or, for example, to store the operating system of a television.
[0028] The system 100 may obtain or receive information via one or more input devices, interfaces, and / or ports as indicated in input block 105. Examples of the input devices include a radio frequency (RF) device for transmitting and / or receiving RF signals over various media, for example, RF signals received over the air from a broadcaster; component video (COMP) inputs; a Universal Serial Bus (USB) input; and / or a High-Definition Multimedia Interface (HDMI) input. Other examples include composite video input (not shown). In some embodiments, the input devices are associated with respective input processing elements, e.g., those generally known in the art. For example, the RF device may be associated with elements suitable for selecting a desired frequency (e.g., selecting or band-limiting a signal) or performing error correction on the signal. The USB and / or HDMI inputs may include respective interface processors and transceivers (or transmitters and receivers) for coupling the system 100 to other devices via USB and / or HDMI ports or connections. Various forms of input processing may be implemented, for example, by and / or within a separate input processing device or the processor 110.
[0029] The system 100 includes a communication interface 150 that enables wired and / or wireless communication with other devices, e.g., via a communication channel 190. The communication interface 150 may include one or more transceivers, modems, network cards and the like. The communication channel 190 may be or include wired and / or wireless mediums.In some embodiments, data may be streamed to the system 100 via wired and / or wireless networks. Examples of such wireless networks include cellular, Bluetooth or Wi-Fi (e.g., IEEE 802.11) networks. The wired and / or wireless networks may include one or more base stations (e.g., cellular base stations, access points, etc.), and / or user equipment (e.g. cellular user equipment, stations, etc.), and / or other network elements that communicate with the system 100 via the communication interface 150 and communication channel 190, whereby the system 100 may obtain data streamed from streaming applications (e.g., OTT services) via various networks, including the Internet. In some embodiments, data is streamed to the system 100 via the input block 105 (e.g., using a set-top box that delivers data via the HDMI connection or the RF connection). In some embodiments, data is received by the system 100 in a non-streaming manner.
[0030] The system 100 may provide one or more output signals to one or more output devices. The output devices may include a display device 165 (e.g., touchscreen display, monitor, etc.), an audio device 175 (e.g., speakers), and other peripheral devices 185, including, for example, a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide a function based on the output of the system 100. The display device 165 can be for a television, tablet, laptop, mobile phone, head-mounted display, or other device. In some embodiments, control signals are communicated between the system 100 and the display device 165, the audio device 175, and / or the peripheral devices 185, enabling device-to-device control with or without user intervention. The output devices may couple to and / or communicate with the system 100 via dedicated connections via respective display, audio, and peripheral interfaces 160, 170, 180. Alternatively, the output devices may couple to and / or communicate with the system 100 via the communication channel 190 and the communication interface 150.
[0031] The display device 165 and the audio device 175 may be collocated, integrated, or otherwise combined with the other components of system 100 in a single unit (e.g., a television). Alternatively, the display device 165 and the audio device 175 may be separate from one or more of the other components of the system 100. In embodiments in which the display device 165 and the audio device 175 are external components, the output signals may be provided via dedicated outputs and / or connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0032] FIG. 2 is a block diagram illustrating an example video encoder 200 that may be employed by the system 100 (e.g., via the encoder / decoder module 130) described with respect to FIG. 1. The video encoder 200 may be an encoder that employs video compression technologies, standards, specification, or protocols, including Advanced Video Coding (AVC, H.264 / MPEG-4), High Efficiency Video Coding (HEVC, H.265), Versatile Video Coding (VVC, H.266), Essential Video Coding (EVC, MPEG-5), AOMedia Video 1 (AVI), VP9, or the Enhanced CompressionModel (ECM), and variations or improvements thereof. Those skilled in the art will understand that the various embodiments described herein are not limited to a specific standard and can be applied to other standards and recommendations, as well as extensions thereof.
[0033] Some embodiments disclosed herein are described with reference to a coding unit (CU) or block of a video frame (or a video image or picture) to which coding tools may be applied by the video encoder 200 and / or by the video decoder 300 (described below with reference to FIG. 3). Generally, embodiments described herein may be applied to a video region formed by a video partition of any shape or size. The video region may be a video slice, a coding tree unit (CTU), or a CU (to which inter prediction or intra prediction can be applied), or a partition thereof, each of which can include samples of a luma component, F, and chroma components, U and V (also denoted herein by C).
[0034] Referring generally to FIG. 2 and the video encoder 200, video data (e.g., one or more video frames) is encoded generally as described below. Prior to encoding, video data may be pre-processed by a precoding processor (not shown). The pre-processing may include, for example, applying a color model transform to the input color components of the input video data (e.g., conversion from RGB 4:4:4 to YUV 4:2:0) or mapping the color components of the input video data to obtain a signal distribution that is more resilient to compression (for instance, applying a histogram equalizer and / or a denoising filter to one or more of the video data’s color components).
[0035] The pre-processing may include associating metadata (for example, a supplemental enhancement information (SEI) message) with the video data that can be attached to a coded video bitstream. After pre-processing, if any, an image (frame) to be encoded is partitioned into CUs (blocks) by an image partitioner 202.
[0036] In general, a CU includes a luma block and associated chroma blocks. As such, functions of the video encoder 200 described herein as applied to a CU refer generally to the luma block and the respective chroma blocks. The CUs may be encoded using an intra prediction mode performed by an intra predictor 260. In intra prediction mode, the content of a CU in a frame is predicted based on content from one or more other CUs of the same frame (or region), using reconstructed blocks of other CUs output from an adder 255. The CUs may also or alternatively be encoded using an inter prediction mode, in which motion estimation and motion compensation are performed by a motion estimator 275 and a motion compensator 270, respectively. In inter prediction mode, the content of a CU in a frame is predicted based on content from one or more reconstructed areas of reference frames, available from a reference picture buffer 280.
[0037] The video encoder 200 selects or otherwise determines at 205 which prediction mode (intra prediction mode and / or inter prediction mode) to use for encoding a CU. The selected predictionmode may be enhanced (e.g., filtered) by a prediction enhancer 285. Based on the selected mode, a prediction for the CU is generated. A residual block is determined based on the prediction (i.e., prediction block, predicted CU) and the input CU. In some embodiments, such determination is made by a subtractor 210.
[0038] The residual block or a partition thereof (e.g., a transform block) is transformed into transform coefficients by a transformer 220. The transform coefficients are quantized by a quantizer 230. An entropy encoder 245 performs entropy encoding of the quantized transform coefficients and coding parameters (e.g., syntax elements including motion vectors and other control data) to form a bitstream of coded video data.
[0039] In addition to coding the original video blocks as described herein, the video encoder 200 reconstructs the coded blocks to provide references for future predictions. Thus, quantized transform coefficients (from the quantizer 230) are de-quantized by an inverse quantizer 240, and inverse transformed by an inverse transformer 250, to reconstruct (decode) the residual blocks. The reconstructed residual blocks and prediction blocks are combined (e.g., by the adder 255) to form reconstructed blocks. Thus, the video encoder 200 performs decoding operations through which the encoded images (frames) are reconstructed.
[0040] In-loop filters 265 may be applied to the reconstructed image (formed by the reconstructed blocks). The filtered reconstructed image(s) are stored in the reference picture buffer 280 and used by the motion estimator 275 and motion compensator 270, as explained above. The in-loop filters 265 can be applied to the reconstructed samples of an image to reduce distortions introduced by the encoding process. For example, a deblocking filter (DBF), bilateral filter (BIF), sample adaptive offset (SAG), and / or adaptive loop filter (ALF) can be applied to reduce encoding artifacts.
[0041] FIG. 3 is a block diagram illustrating an example of video decoder 300 that may be employed by the system 100 (e.g., via the encoder / decoder module 130) described with respect to FIG. 1. Generally, operational features of the video decoder 300 are reciprocal to operational features of the video encoder 200. In the video decoder 300, a coded video bitstream (e.g., generated by the video encoder 200 or another video encoding device or process) is entropy-decoded by an entropy decoder 330 to obtain transform coefficients, motion vectors, and other coding parameters. Based on the coding parameters, an image partitioner 335 divides the picture accordingly. The quantized transform coefficients are de-quantized by an inverse quantizer 340 and inverse transformed by an inverse transformer 350 to decode (reconstruct) respective residual blocks. Depending on the selected prediction mode, a predicted block can be obtained at 370 from an intra predictor 360 (i.e., intra prediction) or from a motion compensator 375 (i.e., interprediction) and may be enhanced (e.g., filtered) by a prediction enhancer 390, generating a prediction block. The reconstructed residual blocks are combined with prediction blocks (e.g. by an adder 355), resulting in reconstructed blocks.
[0042] In-loop filters 365 (e.g., DBF, BIF, SAG, and / or ALF) can be applied to the reconstructed image (formed by the reconstructed blocks), to output reconstructed (decoded) video. The filtered reconstructed image is also stored in a reference picture buffer 380 for reference by the motion compensator 375.
[0043] A post-decoding processor (not shown) can process the reconstructed video data. For example, post-decoding processing can include an inverse color model transform (e.g., conversion from YUV 4:2:0 to RGB 4:4:4) or an inverse mapping to reverse the mapping process performed by the pre-encoding processor described with respect to FIG. 2. The post-decoding processor can use metadata derived by the pre-encoding processor and / or signaled in the video bitstream.
[0044] Recent additions to video compression technology include various industry standards, versions of reference software and / or documentation such as Enhanced Compression Model (ECM) being developed by the JVET (Joint Video Exploration Team) group. The aim is to make further improvements to the existing VVC (Versatile Video Coding) standard and VSEI specification.
[0045] Recent improvements to video compression technology further includes neural network end-to-end compression. End-to-end compression is a compression technique where all components of the processes are learned from given data. End-to-end means that everything is learned from one end (given data) to another end (compressed bitstream). Once the architecture is defined for end-to-end compression, there is no manual engineering work on designing the steps. End-to-end neural compression methods are not standardized yet, although MPEG is currently exploring these technologies.
[0046] Feature(s) associated with auto-encoder are provided herein.
[0047] FIG. 4 is a block diagram illustrating an example end-to-end compression system according to one or more embodiments of the present disclosure. An input image to be compressed, x, is first processed by a deep encoder 410. The output of the encoder, y, is called the embedding of the image. This embedding is converted into a bitstream 450 by going through a quantizer, Q, and then through an arithmetic encoder, AE. The resulting bitstream 450 is decoded by going through an arithmetic decoder, AD, to reconstruct the quantized embedding, y. The reconstructed quantized embedding can be processed by a deep decoder 460 to obtain the decompressed image, x.The deep encoder 410 and decoder 460 are composed of multiple neural layers, such as convolutional layers. Each neural layer can be described as a function that first multiply the input by a tensor, add a vector called the bias and then apply a nonlinear function on the resulting values. The shape (and other characteristics) of the tensor and the type of non-linear functions are called the architecture of the network. We will denote the values of the tensor and the bias by the term “weights”. The weights and, if applicable, the parameters of the non-linear functions, are called the parameters. The architecture and the parameters define a “model”. Typically, the encoder and decoder are fixed and supposed to be known at encoding and decoding times. We will denote the layers of the decoder by llt... lit... , lnand its parameters by 6.
[0048] Many end-to-end architectures have been proposed recently. Typically, they are more complex than FIG. 4, but they all retain the deep encoder and decoder. State of the art models can compete with traditional codecs such as VVC in terms of rate distortion tradeoffs.
[0049] A model M must be trained on massive databases D of images to learn the weights of the encoder and decoder. Typically, the weights are optimized to minimize a training loss LTM, D) = Ez~D[-log(pM(y)) + Ad(x,x)], (1) where pMdenotes the probability of the quantized embedding according to M (thus this term is the theoretical lower bound on bitstream size for the encoded quantized embeddings), d(. , . ) denotes a measure of the distortion between the original and the reconstructed image (for example the mean square error) and A denotes a parameter controlling the trade-off between the rate (r) and distortion (d) terms.
[0050] Typically, an architecture is trained several times, using different values for A, to yield a set of models {MJ with different r / d trade-offs. Usually, different architectures yield models with different r / d points. To compare these architectures, the r / d points of each architecture are interpolated, resulting in a function d(r) for each architecture that provides a distortion estimate for any rate value.
[0051] The deep decoder can decode any type of image. In other words, it performs well on average for all images, but it is likely to be suboptimal for any single image. Recently, it has been shown that it is possible to improve the rate distortion trade-off for a single video by retraining the decoder specifically for this video and by transmitting weight updates 6 for the decoder in addition to the quantized embeddings for intra frames of the video. Before decoding the quantized embedding, 6 is added to 0. This approach did not achieve rate distortion improvements for single images because of the increased code size due to the inclusion of the weights updates. The loss used for fine tuning can for example be:
[0052] 8,X) = - log(pA(<5)) + / ? d(x, x(5)), (2)where pA(. ) denotes a probability density overweight update, x(5) the image reconstructed by the decoder whose weights have been updated by 6 and a trade-off between the two losses. In “Overfitting for Fun and Profit: Instance-Adaptive Data Compression” (by T. van Rozendaal, I. A. M. Huijben, and T. S. Cohen, in International Conference on Learning Representations, 2021) an additional term is added to the loss to enforce a global sparsity constraint on 8. so that a lot of weight update are the same value (0), to make encoding more efficient.
[0053] In this case, decoding the image is a two-step process. The decoder is first updated with the weight updates. Then y is decoded by the updated decoder.
[0054] Feature(s) associated with INR are provided herein.
[0055] Another neural network -based compression method has recently emerged, called implicit neural representation (INR). At the core of INR approaches is a neural network, the INR network. The INR network is trained to predict a pixel value of an image, x(i '), based on the pixel’s coordinates (i,j) - that is,
[0056]
[0057] = (r>g>b) (or0(i,j) = (y, u, v)). During a training phase of an INR network, the parameters 9 (or a subset of them) are determined. This is done via an optimization process through which the parameters 9 that minimize a cost function can be determined. For example, the following cost function can be used:
[0058] Cost = D (x, f0) + AR (9) (3)
[0059] where, D is a distortion measure, measuring the fidelity of the estimated pixel values, provided by f0. relative to the ground truth, that is, the corresponding pixel values from the original image, denoted by x. And, where R is the resulting bitrate of the encoded parameters 9. A trade-off parameter A can be set to determine the balance between D and R. Note that the distortion measure D can be any metric that measure the distance (or similarity) between the original image x and its estimated version provided by f0, such as a mean squared error metric or a learned perceptual image patch similarity (LPIPS) metric. For example, a mean squared error metric can be expressed as:
[0060]
[0061] > where, W and H are the width and height of the image x that the INR network is trained to predict. The optimization of the network parameters 0, according to equation (3), is typically performed by a machine learning optimization technique, applying, for example, a batch gradient descent algorithm. Following the training of the INR network and using the optimal parameters 9 (obtained via the optimization process), the INR network can be applied to predict a pixel value based on its corresponding coordinate values.In INR approaches to compression, the weights of the network are typically quantized, entropy coded and transmitted to the decoder. This is reflected in the rate term of the loss. In some INR approaches, some latent values may be associated to pixel coordinates and used as input of the network. Thus, these latent values may be trained, quantized, entropy coded and transmitted as well.
[0062] Recently, it has been proposed to combine neural network-based auto-encoder with traditional codec by allowing the encoder to choose between a traditional codec such as VTM (a codec developed for VVC) or a neural network-based auto encoder to encode each I-frame, also called end-to-end coder or end-to-end neural network in the following. In traditional block-based hybrid video coding standards such as VVC, a standardized syntax allows for signaling the coding tools active in the coded video chunks. Currently, there is no syntax elements allowing to signal the use of a neural network auto-encoder to encode part of the video. Thus, the decoder currently has no way to decode such a bitstream as it has no way to know which decoding method to use and for which part of the video. The skilled in the art will recognize that the concept of intra coded frame is meaningful for traditional codec as the ones designed by JVET. Indeed, the notion of inter frame does not really make sense for neural network auto-encoder as an auto-encoder does not typically start from a previous encoded frame as in meant by inter frame prediction. Therefore, the present principles may be generalized to frames, independently of their type inter, intra, as long as at least a partition of the picture (ie a I-picture, an I-slice, or I-CTU...) may be independently input to a neural network auto-encoder.
[0063] Accordingly, the present principles propose a) a bitstream that contains additional information used while encoding all or a part of a frame of a video using an end-to-end codec, including SEI syntax, and b) a decoding method that can cause a decoder using this information to reconstruct the video.
[0064] While we focus the description on neural network auto-encoder for clarity, it should be understood that it can be used for other neural network-based approaches such as INR-based methods.
[0065] Feature(s) associated with encoding are provided herein.
[0066] FIG. 5 a and FIG. 5b are block diagrams illustrating an encoding method according to one or more embodiments of the present disclosure. FIG. 5a and FIG. 5b may represents a same encoding method 500 from a different perspective, the encoding method 500 is compliant with any of the variants of the syntax proposed hereafter.For instance, in the embodiment of FIG 5a, frames may be encoded one by one while all or at least a part of a frame of a video may be coded using an end-to-end codec as described above. The encoding 500 combines neural network-based auto-encoder with traditional codec, such as a standardized codec (including encoder and / or decoder), for instance HEVC, or VVC. The method 500 thus takes a current frame or picture to encode as input.
[0067] In a step 501, at least one part of the picture is end-to-end coded using a neural network. As shown in FIG. 5b, determining the at least one part of the picture is an encoder choice. Then, in a step 502, one or more syntax elements that allows a decoder to reconstruct the at least one part of the picture end-to-end coded using a neural network are also encoded. Various embodiments of such syntax elements are detailed hereafter. Then, in a step 503, a bitstream is generated that comprises the coded picture, the one or more syntax elements for reconstructing the at least one part of the picture end-to-end coded using the neural network. For instance, an intra coded picture may combine part of the I-picture that is coded by a neural network-based auto-encoder and the remaining part of the I-picture that is coded by a standardized codec. For instance, an inter predicted picture may combine part (I-slice) of the picture that is coded by a neural network-based auto-encoder and the remaining part of the picture that is coded by a standardized codec.
[0068] In the embodiment of FIG. 5b, frames may also be encoded one by one. For each input frame, a bitstream is produced as follows. In a first step 510, the frame type (or partition type such as slice type) is checked. For instance, if the frame is not I-frame, the encoder encodes 530 the frame with the traditional codec and may use one or more previously reconstructed frames as reference frame. For instance, a reference frame or at least one part of a reference frame may have been coded using end-to-end neural network. Otherwise, if the current frame to encode is an I-frame, in a step 520, the encoder chooses between traditional codec and at least one end-to-end (e2e) neural network-based codec. As an example among well-known routines in the domain of video compression, the selection step 520 may involve encoding this frame (or part of the frame) using different codecs (e2e codec or traditional codec), evaluating the performance of different codecs on this frame based on different metrics including bitrate, distortion and / or decoding complexity, comparing these performances, and choosing the codec that maximizes one metric or a combination of metrics. In a variant, choosing a method for encoding may be directly based on the characteristics of this frame, for example through an expert designed algorithm or a machine learning model. As an example, one would evaluate all possible codecs in terms of bitrate and distortion and select the value that minimizes the cost associated to the encoding and reconstruction of that part of the signal as a function of the codec for that frame. Upon determining which one or more parts of the I frame are encoded with e2e codec, the encoding method mayencode 560 the one or more parts of the I frame with the e2e codec and encode 550 the remaining one or more parts of the I frame using a traditional codec. Besides, in a step 540, information about the choice made by the encoder is encoded into the bitstream using an appropriate syntax. The following section provides a description of examples of such syntax elements. The syntax elements may include indication whether the frame is encoded by the traditional codec or an e2e one. In the latter case, additional information describing the e2e codec may be included in the bitstream as well. Finally, the bitstream is formed by combining at least the syntax elements indicating the codec selected by the encoder for I-frame, the coded video data generated by the codec selected by the encoder in step 530, 550 or 560.
[0069] Feature(s) associated with new syntax description are provided herein.
[0070] At least three pieces of information are being made available to the decoder for reconstructing a picture of a video comprising at least one part of an intra picture being end-to-end coded using a neural network. A first element indicates whether an end-to-end codec is used either for a current picture as a whole, or for each of the one or more parts (e.g. slice, tile) of the current picture. A second element indicates which codec is used, and thus how to reconstruct the frame data. Finally, a third element comprises the encoded current frame.
[0071] Advantageously, the proposed syntax allows a decoder to process a hybrid bitstream combining traditional and neural network-based coded data, wherein the bitstream is a single-layer bitstream. Besides, some syntax elements allow for enhanced processing at the decoder since some properties (for instance the partitioning of the picture data) of the data issued from neural networkbased codec are distinct from the properties of the standard codec data and may raise issue at decoding (for instance on IDR or GDR mechanism).
[0072] In a variant embodiment, the elements may be encoded in the picture header for the use of the end-to-end codec at picture level, or in a slice header for the use of the end-to-end codec at slice level. According to another variant, the second element describing e2e used codec may be encoded in a Network Abstraction Layer (NAL) payload. According to yet another variant, the third element comprising encoded frame elements may be embedded in a slice payload.
[0073] In some other variants, the use of the codec at the picture level may be further specified. For example, the encoder may choose that the e2e codec is allowed and enabled frame by frame, or used for all I frames.
[0074] According to a first embodiment of the syntax elements, the use of the end-to-end codec is defined at the picture level. For instance, the one or more syntax elements for reconstructing at least one part of an Intra picture end-to-end coded using a neural network comprises an indication(ph_nn_allowed_flag) enabling end-to-end neural network coding for at least one part of the picture. According to a variant, the one or more syntax elements for reconstructing at least one part of an Intra picture end-to-end coded using a neural network comprises an indication specifying (ph_only_nn_slice) that all parts of the picture are end-to-end neural network coded. In that case, some optimization in the signaling may be further implemented. In the first embodiment, the one or more syntax elements may be signaled in a picture header. For instance, the picture header may be modified so that the decoder knows that some parts of the picture may be encoded using a neural codec. A first flag may enable the use of such a codec for part of the picture. A second flag may specify that all parts of the pictures are encoded using a neural codec. This may reduce signaling. An example of the modified picture header syntax structure may be as follows (added syntax is highlighted by italic & underlines in the table):
[0075] &&
[0076] <
[0077] &&
[0078] <
[0079]
[0080] &&
[0081] &&
[0082] &&
[0083] < <
[0084] &&
[0085] && &&
[0086]
[0087] >
[0088] >
[0089] &&
[0090] >
[0091] &&
[0092] >
[0093] &&
[0094] >
[0095]
[0096] >
[0097] >
[0098] &&
[0099] &&
[0100]
[0101] <
[0102]
[0103] Table 1: an example of picture header syntax
[0104] The semantic of the modified picture header may be as follows:
[0105] ph nn allowed flag indicates whether some parts (e.g. some slices, tiles...) of the current picture may be encoded by a neural network based codec. ph_nn_allowed_flag equal to 1 specifies that part of the current picture may be encoded using a neural network based method. ph_nn_allowed_flag equal to 0 specifies that the current picture is not encoded using a neural network based method.
[0106] ph_only_nn_slice indicates whether all slices of the current picture are encoded by a neural network based codec. ph_only_nn_slice equal to 1 specifies all slices of the current picture are encoded using a neural network based method. ph_only_nn_slice equal to 0 specifies that some slices of the current picture may not be encoded using a neural network based method.
[0107] In a variant of the first embodiment, ph only nn slice may not be present.
[0108] In another variant of the first embodiment, additional parts of the picture header may be disabled when ph only nn slice is equal to 1 thus indicating that the whole picture is encoded with end-to-end neural network. As non-limiting examples, these elements may include the alf syntax elements, the virtual boundaries syntax elements, the deblocking syntax elements. Examples of conditional disabling of at least one part of the picture header when the flag (ph only nn slice) specifies that all parts of the picture are end-to-end neural network coded are illustrated in table 2 (highlighted by underlined and italic font):
[0109] && &&
[0110] && &&
[0111] &&
[0112]
[0113] Table 2: an example of E2E conditional picture header syntax In another variant of the first embodiment, the syntax elements related to the use of a neural based codec may be signaled in another NAL unit, such as in a slice layer unit, or at the sequence parameter set (SPS) level.
[0114] According to a second embodiment of the syntax elements, the use of the end-to-end codec is defined at the slice level. For instance, the one or more syntax elements may comprise anindication enabling end-to-end neural network coding at a given level of the picture, the given level comprises a slice level, a coding tree level, a coding unit or block level (notwithstanding the above-described picture level of the first embodiment). In this embodiment, the use of an end-to-end codec is further signaled at another level, for example in a slice header. This syntax element would allow encoding an I-frame using a traditional codec for some parts of the frame and an end-to-end codec for other parts of the frame. An example of a modified slice header syntax is provided below (added syntax is in italic & underlined):
[0115] && >
[0116] && >
[0117] <
[0118] && >
[0119] && &&
[0120] &&
[0121] <
[0122]
[0123] && && &&
[0124] &&
[0125] && &&
[0126] && >
[0127] && >
[0128] <
[0129] >
[0130] &&
[0131] &&
[0132] && >
[0133] && >
[0134] &&
[0135] &&
[0136] &&
[0137] &&
[0138]
[0139] &&
[0140] &&
[0141] && && && &&
[0142] &&
[0143] &&
[0144] && &&
[0145] <
[0146] >
[0147] <
[0148]
[0149] Table 3: an example of slice header syntaxAn example of the semantic is provided hereafter:
[0150] sh_nn_flag indicates the codec used to encode the slice. sh_nn_flag equal to 1 specifies that the current slice is encoded using a neural network-based method and that the payload contains a bitstream generated using an e2e codec. sh_nn_ flag equal to 0 specifies that the current slice is not encoded using a neural network-based method and thus the payload contains data created by a traditional codec. When ph_only_nn_slice equals 1 and sh slice type = = I, sh_nn_flag should be considered to be set to 1.
[0151] In a variant, the syntax element indicating that an end-to-end codec is used may also be signalled at different level of a partition of a frame, for example in a tile or a coding tree unit header.
[0152] In another variant, the syntax element indicating that an end-to-end codec is used may also be signalled in further division of a coding tree unit, for example in a coding unit also referred to as block level.
[0153] In another variant, an indication that a slice is coded using an end-to-end codec may be encoded by specifying a new slice type sh_slice_type value. For instance, an indication (sh_slice_type = N) is signaled that a slice is coded using a new slice type corresponding to an end-to-end neural network coded intra slice. Thus, there may be two possible values for I frames: one for I-frame not encoded using end-to-end based codec and one for I-frame encoded using end-to-end based codec. In this variant, sh_slice_type specifies the coding type of the slice according to the following table. In that case, some tests on sh slice type == I would be replaced by sh_slice_type == 11| sh_slice_type == N.
[0154] Table - Name association to sh_slice_type
[0155]
[0156] Table 4: an example of name association to sh slice type To facilitate reading in the above table 3, we define a variable nn_slice that is true if the slice is encoded using an e2e codec. In a final specification, it may be replaced by sh_nn_flag or sh slice type == N.
[0157] In yet other variants of the second embodiment, parts of the slice header may be disabled when nn_slice is 1. These elements include the alf syntax elements. As an example, this would lead to the following modification in slice header syntax table: && &&
[0158]
[0159] Table 5: an example of E2E conditional slice header syntax In some variants, an indication that a slice is coded using an end-to-end codec may be encoded by specifying a new nal_unit_type. For instance, a new type (nal_unit_type = 4) could be used to signal that a slice is coded using an end-to-end neural network coded intra slice. This could be advantageously useful to allow high level system modules to perform specific action on that slice (for example using a neural network accelerator hardware to decode it).
[0160] For instance, an example of anewNAL type, corresponding to table entry 4 as non-limiting value, may be described as in the modified NAL unit table below:
[0161]
[0162]
[0163] Table 6: an example of NAL unit type table
[0164] According to a third embodiment, the bitstream may be modified at the level of the slice layer to include encoded data, namely video data corresponding to the current frame either encoded by traditional codec or encoded by auto-encoder. This element corresponds to the third element that is being made available to the decoder for reconstructing a picture of a video comprising at least one part of a picture being end-to-end coded using a neural network. For instance, we propose modifying the bitstream at the level of the slice layer, based on the nn slice value to provide the encoded bitstream of a picture slice to the decoder.
[0165]
[0166] Table 7: an example of E2E conditional slice header syntax for slice datann_slice_data() is a payload containing the data allowing the e2e codec to decode the frame. The exact method to use these bytes depends on the codec used. The payload will typically contain the bytes generated by using an entropy coder to code the intermediate features of an autoencoder.
[0167] In some variants, for example when a syntax element indicating that an end-to-end codec is used may also be defined and signalled in another division of a frame, e.g. in a tile or a coding tree unit header or a coding unit (a coding block), the data payload may be divided between traditional codec data and e2e codec data at the level of said division.
[0168] According to a fourth embodiment of the syntax elements, one or more parameters allowing to determine the end-to-end codec at the decoder are signaled. This element corresponds to the second element that is being made available to the decoder for reconstructing a picture of a video comprising at least one part of a picture being end-to-end coded using a neural network Indeed, the decoder needs to know which e2e codec was used to encode the frame, the slice or the other partition of the signal. Several embodiments are possible. This piece of information may have been transmitted by another part of the system. It may have been transmitted at the SPS level. It may also be signalled as a tag URI with syntax and semantics as specified in IETF RFC 4151. In that case, the tag URI may for example be included in the picture header or the slice header. This piece of information may also be signalled in a new NAL syntax type, which requires a new NAL type as described below. Part of the information signalled about the e2e codec includes a description of the neural network used. An example of syntax for this element below as well.
[0169] According to a variant of the fourth embodiment, a new NAL Type NN_NUT describing the end-to-end neural network may be defined. The decoder is thus informed on which end-to-end network was used to encode the frame. For instance, an example of a new NAL type, corresponding to table entry 28 as non-limiting value, may be described as in the modified NAL unit table below:
[0170]
[0171]
[0172]
[0173] Table 8: another example of NAL unit type table
[0174] When nal unit type is equal to “NN NUT”, a neural network is provided in the bitstream.
[0175] According to a variant, the description of the codec may be included in another nal unit rather than creating a new one. For example, it may be included in a picture header or in a slice header.
[0176] According to yet another variant of the fourth embodiment related to encoding the neural network itself, an element may indicate how the neural network is signaled or encoded in the bitstream. It may for example be a value associated to a specific format such as NNC or Open Neural Network Exchange format and / or a particular version of a format.
[0177] According to a variant, a data structure related to the configuration of the neural network inference engine is defined. The data structure may for example include indications on the precision to use for the operations, the memory necessary to store the network or the inference engine to use.
[0178] Table 9 provides an example of syntax describing the network to be used to decode the frame:
[0179]
[0180]
[0181] Table 9: an example of data structure describing the end-to-end neural network nn_mode_idc indicates the format used to encode the Neural network. A value equals to 0 indicates that this SEI message contains an ISO / IEC 15938-17 bitstream. A value of 1 indicates that this SEI message contains a neural network encoded by a format identified by the tag URI nn tag uri.
[0182] nn_tag_uri contains a tag URI with syntax and semantics as specified in IETF RFC 4151 identifying the format and associated information about the network or an update encoded here. nn_uri contains a URI with syntax and semantics as specified in IETF Internet Standard 66 identifying the neural network used as a network or an update relative to the network.
[0183] nn_complexity_info_present_flag equal to 1 specifies that one or more syntax elements that indicate the complexity of the network are present. nn_complexity_info_present_flag equal to 0 specifies that no syntax elements that indicates the complexity of the network are present.
[0184] nn_parameter_type_idc equal to 0 indicates that the neural network uses only integer parameters. nn_parameter_type_flag equal to 1 indicates that the neural network may use floating point or integer parameters. nn_parameter_type_idc equal to 2 indicates that the neural network uses only binary parameters. nn_parameter_type_idc equal to 3 is reserved for future use.
[0185] nn_log2_parameter_bit_length_minus3 equal to 0, 1, 2, and 3 indicates that the neural network does not use parameters of bit length greater than 8, 16, 32, and 64, respectively. When nn_parameter_type_idc is present and nn_log2_parameter_bit_length_minus3 is not present the neural network does not use parameters of bit length greater than 1.
[0186] nn_num_parameters_idc indicates the maximum number of neural network parameters for the network in units of a power of 2 (or another power of 2). nn_num_parameters_idc equal to 0 indicates that the maximum number of neural network parameters is unknown. The value nn_num_parameters_idc shall be in the range of 0 to 63, inclusive.
[0187] If the value of nn_num_parameters_idc is greater than zero, the variable maxNumParameters is derived as follows:
[0188] maxNumParameters = ( 2 « nn_num_parameters_idc ) - 1
[0189] The number of neural network parameters of the post-processing filter shall be less than or equal to maxNumParameters.
[0190] nn_num_mac_operations_idc greater than 0 indicates that the maximum number of multiply-accumulate operations per inference of network is less than or equal tonn num mac operations idc. nn num mac operations idc equal to 0 indicates that the maximum number of multiply-accumulate operations of the network is unknown. The value of nn_num_mac_operations_idc shall be in the range of 0 to 232- 1, inclusive.
[0191] nn_total_kilobyte_size greater than 0 indicates a total size in kilobytes required to store the uncompressed parameters for the neural network. The total size in bits is a number equal to or greater than the sum of bits used to store each parameter. nn_total_kilobyte_size is the total size in bits divided by 8000, rounded up. nn total kilobyte size equal to 0 indicates that the total size required to store the parameters for the neural network is unknown. The value of nn_total_kilobyte_size shall be in the range of 0 to 232- 1, inclusive.
[0192] nn_reserved_zero_bit_b shall be equal to 0 in bitstreams conforming to this edition of this document. Decoders shall ignore SEI messages in which nn reserved zero bit b is not equal to 0.
[0193] nn_payload_byte[i] contains the i-th byte of a bitstream conforming to ISO / IEC 15938-17. The byte sequence nn_payload_byte[i] for all present values of i shall be a complete bitstream that conforms to ISO / IEC 15938-17.
[0194] According to an embodiment, the data structure describing the end-to-end neural network may comprises at least one of a model type of the neural network, a type of computation of the neural network, a format of an encoding of the neural network, an indication of a tag URI identifying a format of the neural network, an indication of a URI identifying the neural network, an indicator indicating whether complexity information relating to the neural network is present or not, an indication of a type of parameters used by the neural network, an indication relating to a bit length of parameters used by the neural network, an indicator indicating a maximum number of parameters for the neural network, an indication of a maximum number of multiply-accumulate operations per inference of the neural network, or an indication of a size for storing uncompressed parameters of the neural network.
[0195] Feature(s) associated with decoding are provided herein.
[0196] FIG. 6a and FIG. 6b are block diagrams illustrating a decoding method according to one or more embodiments of the present disclosure. FIG. 6a and FIG. 6b may represent a same decoding method 600 from a different perspective, the decoding method 600 is compliant with any of the variants of the syntax proposed previously.
[0197] For instance, in the embodiment of FIG 6a, a bitstream is received that comprises the three pieces of information previously described for reconstructing a picture of a video comprising at least one part of an intra picture being end-to-end coded using a neural network. In a step 601, theone or more syntax elements are decoded for reconstructing at least one part of an intra picture end-to-end coded using a neural network. In a step 602, the intra picture of a video is decoded using the one or more syntax elements, wherein at least one part of the intra picture is reconstructed using a neural network. In a step, not shown on FIG. 6a, the neural network may be reconstructed in the decoder according to information for instance signaled in a data structure.
[0198] FIG. 6b illustrates a more detailed embodiment of the decoding method. For instance, frames may be processed one by one. For each frame (or picture) to decode, the bitstream is processed as follows. In a step 610, the frame or slice type is checked. If the frame or slice is not an I-frame or I-slice (no), the decoder decodes the frame or slice with the traditional codec. If the frame or slice is an I-frame and end-to-end neural network coding for at least a part of the picture is enabled (ph_nn_allowed_flag = 1), then in a step 620, the one or more syntax elements related to reconstructing at least one part of an Intra picture end-to-end coded using a neural network are retrieved form the bitstream and decoded. As previously described, the syntax elements may be signaled as bitstream syntax at various level including picture header, slice header, coding tree header. The step 620 may comprise parsing the indication (sh_nn_flag) enabling end-to-end neural network coding at a given level of the picture, for instance slice level. According to different variant, syntax element may include whether the frame or slice is encoded by the traditional codec or an e2e one (nn_slice, sh_nn_flag and / or ph_only_nn_slice as described above). In case a current slice is end-to-end neural network coded (sh nn flag = 1), additional information describing the e2e codec may be decoded from the bitstream as well, such as the particular codec used. In case a current slice is coded using the traditional codec (e.g. nn_slice=0, or sh_nn_flag=0, or ph_nn_allowed_flag=0), the frame bitstream is processed 650 by the decoder of a traditional codec to reconstruct the frame. In case a current slice is coded using an e2e codec (e.g. nn_slice=l, sh_nn_flag=l and / or ph_only_nn_slice=l), the corresponding neural network may be prepared for decoding in step 630. This may involve decoding the parameters or weights of the network from the bitstream or loading them from memory. It may also involve initializing the network in an appropriate computing resource which may be a neural network accelerator chip. Then in a step 640, the slice data in the bitstream is processed by the e2e codec decoder to reconstruct the current slice or more generally the part of the picture corresponding to the given level using the neural network. Multiple decoded slices issued from traditional codec 650 or for e2e codec may be combined to form a decoded picture. As such, the frame is decoded from the bitstream by the codec selected by the encoder.Feature(s) associated with generalization to other neural network-based approaches are provided herein.
[0199] To illustrate the application of the present principles to other neural network-based approaches, we discuss how the above-described approach can be modified to apply to apply to or include INR-based approaches. The syntax elements introduced to indicate that a frame, a slice or another part of a picture is encoded using a neural based approach can be used as they are. These elements are for example ph only nn slice, ph nn allowed flag, ph nn allowed flag, sh_nn_flag, nn_slice_ The slice header, picture header or nn network rbsp may be modified to include a syntax element describing the type of codec used (e.g. auto-encoder or INR) if more than one type is possible.
[0200] Alternatively, some or all the syntax elements introduced to indicate that a frame, a slice or another part of a picture is encoded using a neural based approach may be duplicated so that one set of syntax elements that indicate the presence or absence of an auto-encoder based codec and the presence or absence of an INR-based codec.
[0201] Alternatively, the syntax elements introduced to indicate that a frame, a slice or another part of a picture is encoded using a neural based approach may be modified from a boolean to an unsigned integer, with 0 meaning no neural network-based codec is or may be used, 1 indicating the element applies to auto-encoder based codec, 2 indicating the element applies to INR-based codecs, 3 indicating it applies to both.
[0202] Alternatively, additional syntax elements may be introduced to explicitly list whether each codec is concerned by the flag.
[0203] If an INR-based codec is used for a slice, nn slice dataQ will contain data describing the INR model. As an example, this may include all or some elements of the syntax proposed in the European patent application 24306011.8 filed on June, 25, 2024 by the same applicant or in the European Patent Application 23305937.7 filed on June, 13, 2023 by the same applicant. Some of these elements that describe properties of the model rather than quantized weight or latent value payloads could also be included in the slice header, picture header or new NAL type.
[0204] One or more embodiments provide a computer program comprising instructions which when executed by one or more processors cause such processors to perform the encoding and / or decoding methods according to any of the embodiments described above. One or more embodiments also provide a computer readable storage medium having stored thereon instructions for encoding or decoding video data according to the methods described above.One or more embodiments provide a computer readable storage medium having stored thereon video data generated according to the methods described above. One or more embodiments also provide a method and apparatus for transmitting or receiving video data generated according to the methods described above.
[0205] The embodiments described herein may be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (e.g., as a method), the implementation of such features may also be implemented in other forms. An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. Corresponding methods may be implemented in, for example, a processor.
[0206] Various methods and aspects described herein can be used to modify one or more modules. For example, the intra predictors and inter predictors described with respect to FIGs. 2 and 3 may be implemented as one or more modules and modified according to the various embodiments of the present disclosure.
[0207] The various embodiments described herein provide at least the following features, devices or aspects, alone or on any combination, across various claim categories and types:
[0208] i. Encoding, into coded video data, syntax elements that can enable the decoder to decode the coded video data, according to any of the embodiments described herein. ii. A bitstream that includes one or more of the described syntax elements, or variations thereof, whether transmitted, stored, or otherwise made available.
[0209] iii. Creating, transmitting, receiving, and / or decoding of the bitstream.
[0210] iv. An electronic device (e.g., TV, set-top box, mobile phone, tablet, etc.) that tunes a channel to receive a bitstream or that receives such bitstream over the air. The electronic device decodes the syntax elements from the bitstream, and, optionally, displays (e.g., via a monitor or other type of display) a resulting image. Various numeric values are used in the present application. Such specific values are for example purposes and the embodiments described are not limited to these specific values.
[0211] Various methods are described herein, and such methods comprise one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., for example, a “firstdecoding” and a “second decoding”. Use of such terms does not imply an order to the operations unless specifically required.
[0212] The present disclosure may refer to “determining” various pieces of information.
[0213] Determining information may include one or more of, for example, estimating, calculating, predicting, or retrieving (e.g., from memory) the information.
[0214] The present disclosure may refer to “accessing” various pieces of information. Accessing information may include one or more of, for example, receiving, retrieving (e.g., from memory), storing, moving, copying, calculating, determining, predicting, or estimating the information. Similarly, the present disclosure may refer to “receiving” various pieces of information. Receiving information may include one or more of, for example, accessing or retrieving (e.g., from memory) the information.
[0215] “Decoding,” as used herein, encompasses all or part of the processes performed, for example, on an encoded sequence to produce an output suitable for display. In some embodiments, such processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, etc. Whether the phrase “decoding process” is intended to refer to a subset of operations or generally to the broader decoding process will be clear based on the context of the specific description and will be well understood by those skilled in the art.
[0216] “Encoding,” as used herein, encompasses all or part of the processes performed, for example, on input video data an order to produce an encoded bitstream. Additionally, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “encoded” or “coded” may be used interchangeably, the terms “image,” “picture,” “sub-picture,” “slice,” and “frame” may be used interchangeably, and the terms “pixel” and “sample” may be used interchangeably.
[0217] The present disclosure refers to information, for example, syntax elements, that can be transmitted or stored. Such information can be packaged or arranged in a variety of manners, including for example manners common in video standards such as putting the information into a sequence parameter set (SPS), a picture parameter set (PPS), a network abstraction layer (NAL) unit, a header (for example, a NAL unit header, or a slice header), or an SEI message. Other manners are also available, including, for example, manners that are common for system level or application-level standards such as signaling the information into one or more of the following:
[0218] i. session description protocol (SDP), for example as described in RFCs and / or used in conjunction with real-time transport protocol (RTP) transmission.
[0219] ii. hypertext transfer protocol (HTTP) live Streaming (HLS) manifest transmitted over HTTP.iii. dynamic adaptive streaming over HTTP (DASH) media presentation description (MPD) descriptors, for example as used in DASH and transmitted over HTTP. iv. RTP header extensions, for example as used during RTP streaming.
[0220] v. International Organization for Standardization (ISO) base media file format, for example, as used in Omnidirectional MediA Format (OMAF).
[0221] As used herein, “signal” and “signaling” refer to, among other things, indicating information to a decoder. For example, in some embodiments the encoder signals a quantization matrix for de-quantization, whereby the same parameter is used for both encoding and decoding. In some embodiments, the signaling may be explicit, such that information (e.g., a particular parameter) is transmitted to the decoder enabling the decoder to use the same particular parameter. In some embodiments, the signaling may be implicit, in that the information (e.g., a particular parameter) is indicated based on other information at or transmitted to the decoder or derived or selected by the decoder based on information available at the decoder. By not transmitting the information (e.g., the particular parameter), a bit savings is thus realized in some embodiments. In some embodiments, one or more syntax elements or flags are used to signal information to a decoder. While the preceding relates to the verb form of the word “signal”, the word “signal” can also be used herein as a noun.
[0222] In some embodiments, signals may be produced that are formatted to carry information that may be stored or transmitted. Such information may include, for example, instructions for performing a method, or data produced by one of the described implementations (e.g., a bitstream of a described embodiment). Such a signal may be formatted, for example, as an electromagnetic wave or as a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links and may be stored on a processor-readable medium.
[0223] It is to be understood that use of any of the following “ / ”, “and / or”, and “at least one of’ is intended to encompass all possible selections of listed items, taken either individually or in any combination thereof.
[0224] While specific embodiments have been described in the foregoing description in connection with the accompanying drawings, it should be understood that embodiments described herein are examples only and should not be taken as limiting the scope of the present disclosure or the following claims. Although features and elements are described herein in particular combinations, those of ordinary skill in the art will appreciate that such features or elements maybe used alone or in any combination with the other features and elements. It is understood, therefore, that the overall teachings of the present disclosure are not limited to the particular embodiments, implementations, and examples disclosed herein, but are intended to cover variations, modifications, and alternatives as defined by the appended claims and any and all equivalents thereof.
Claims
CLAIMS1. A method, comprisingencoding a picture of a video, wherein at least one part of the picture is end-to-end coded using a neural network;encoding one or more syntax elements for reconstructing the at least one part of the picture end-to-end coded using a neural network; andgenerating a bitstream comprising the coded picture, the one or more syntax elements for reconstructing the at least one part of the picture end-to-end coded using the neural network.
2. An apparatus, comprising one or more processors, wherein said one or more processors is operable to:encode a picture of a video, wherein at least one part of the picture is end-to-end coded using a neural network;encode one or more syntax elements for reconstructing the at least one part of the picture end-to-end coded using a neural network; andgenerate a bitstream comprising the coded picture, the one or more syntax elements for reconstructing the at least one part of the picture end-to-end coded using the neural network.
3. A method, comprisingdecoding one or more syntax elements for reconstructing at least one part of a picture end-to-end coded using a neural network; andreconstructing the picture of a video using the one or more syntax elements, wherein at least one part of the picture is reconstructed using a neural network.
4. An apparatus, comprising one or more processors, wherein said one or more processors is operable to decode one or more syntax elements:decode one or more syntax elements for reconstructing at least one part of a picture end-to-end coded using a neural network; andreconstruct the picture of a video using the one or more syntax elements, wherein at least one part of the picture is reconstructed using a neural network.
5. The method of any of claims 1 or 3, or the apparatus of any of claims 2 or 4, wherein the one or more syntax elements comprise at least one of:36an indication(ph_nn_allowed_flag) enabling end-to-end neural network coding for at least a part of the picture;an indication specifying (ph only nn slice) that all parts of the picture are end-to-end neural network coded.
6. The method of any one of claims 1, 3 or 5, or the apparatus of any one of claims 2, 4-5, wherein the one or more syntax elements are signaled in one of a picture header, a slice header, a coding tree unit header.
7. The method of any one of claims 1, 3 or 5-6, or the apparatus of any one of claims 2, 4- 6, wherein at least one part of the picture header is conditionally disabled when the indication specifies (ph only nn slice) that all parts of the picture are end-to-end neural network coded.
8. The method of any one of claims 1, 3 or 5-7, or the apparatus of any one of claims 2, 4- 7, wherein the one or more syntax elements further comprise an indication (sh_nn_flag) enabling end-to-end neural network coding at a given level of the picture, the given level comprises a slice level, a coding tree level, a block level.
9. The method of any one of claims 1, 3 or 5-7, or the apparatus of any one of claims 2, 4-7, an indication (sh_slice_type = N) that a slice is coded using a new slice type corresponding to an end-to-end neural network coded Intra slice.
10. The method of any one of claims 1, 3 or 5-7, or the apparatus of any one of claims 2, 4-7, further comprising a new NAL Type related to a data structure describing the end-to-end neural network.
11. The method or the apparatus of claim 10, wherein the data structure describing the end-to-end neural network comprises at least one of:a model type of the neural network,a type of computation of the neural network,a format of an encoding of the neural network,an indication of a tag URI identifying a format of the neural network,an indication of a URI identifying the neural network,37an indicator indicating whether complexity information relating to the neural network is present or not,an indication of a type of parameters used by the neural network,an indication relating to a bit length of parameters used by the neural network,an indicator indicating a maximum number of parameters for the neural network, an indication of a maximum number of multiply-accumulate operations per inference of the neural network, oran indication of a size for storing uncompressed parameters of the neural network.
12. The method of any of claims 3, 5 and 7 or the apparatus of any of claims 4, 5 and 7, wherein reconstructing the picture of the video using the one or more syntax elements and the neural network further comprises:when the indication (ph nn allowed flag) enables end-to-end neural network coding for at least a part of the picture, parsing the indication (sh nn flag) enabling end-to-end neural network coding at the given level of the picture; andwhen the indication (sh nn flag) enables end-to-end neural network coding at the given level of the picture, reconstructing the at least a part of the picture corresponding to the given level using the neural network; andwhen the indication (sh nn flag) disables end-to-end neural network coding at the given level of the picture, reconstructing the at least a part of the picture corresponding to the given level using a standard decoder.
13. A bitstream comprising data representative of a coded picture of a video, wherein at least one part of the picture is end-to-end coded using a neural network and data representative of one or more syntax elements for reconstructing the at least one part of the picture end-to-end coded using a neural network14. The bitstream of claim 13, wherein the one or more syntax elements comprise at least one of:an indication (ph nn allowed flag) enabling end-to-end neural network coding for at least a part of the picture;an indication specifying (ph only nn slice) that all parts of the picture are end-to-end neural network coded.
15. The bitstream of claim 13, further comprising data representative of a data structure describing the end-to-end neural network, wherein the data structure describing the end-to-end neural network comprises at least one of:a model type of the neural network,a type of computation of the neural network,a format of an encoding of the neural network,an indication of a tag URI identifying a format of the neural network,an indication of a URI identifying the neural network,an indicator indicating whether complexity information relating to the neural network is present or not,an indication of a type of parameters used by the neural network,an indication relating to a bit length of parameters used by the neural network,an indicator indicating a maximum number of parameters for the neural network, an indication of a maximum number of multiply-accumulate operations per inference of the neural network, oran indication of a size for storing uncompressed parameters of the neural network.