Inter-frame coding
By employing end-to-end learned encoders for intra-coded frames and conventional encoders for inter-coded frames with filtering, the encoding and decoding of video sequences are optimized, addressing inefficiencies in existing technologies and improving compression and decoding quality.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NOKIA TECHNOLOGIES OY
- Filing Date
- 2026-01-08
- Publication Date
- 2026-07-23
Smart Images

Figure IB2026050140_23072026_PF_FP_ABST
Abstract
Description
INTER-FRAME CODINGTECHNICAL FIELD
[0001] The example and non-limiting embodiments relate generally to inter-frame coding.BACKGROUND
[0002] It is known to perform encoding and decoding functions as part of a codec.SUMMARY
[0003] The following summary is merely intended to be illustrative. The summary is not intended to limit the scope of the claims.
[0004] Example 1 : An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: using an end-to-end learned encoder for encoding: a picture or frame of a video sequence as an intra-coded frame; or a block of the picture or frame of the video sequence as an intra-coded block; and using a conventional encoder for encoding: another picture or frame of the video sequence as inter-coded frame; or a block of the another picture or frame of the video sequence as an inter-coded block. In an example, the end-to-end learned encoder includes an end-to-end learned encoder module or circuit and the conventional encoder includes a conventional encoder module or circuit.
[0005] Example 2: The apparatus of example 1 , wherein the frame comprises an input intra frame.
[0006] Example 3: The apparatus of example 2, wherein the apparatus is further caused to perform: performing filtering operation on the input intra frame or a portion of the input intra frame by using a filter. In an example, the filter includes a filter module or circuit.
[0007] Example 4: The apparatus of example 3, wherein the filter comprises a pre-processing filter.
[0008] Example 5: The apparatus of example 3, wherein the filtering operation comprises one or more of: smoothing the input intra frame: removing high-frequency signals from the input intra frame: making the intra frame similar or substantially similar to one or more first other frames of the video sequence that comprises the input intra frame: receiving the one or more first other frames as an input; or receiving data derived from the one or more first other frames.
[0009] Example 6: The apparatus of example 5, wherein the one or more first other frames comprise: one or more previous frames with respect to the input intra frame in output order, display order, or in coding order; and / or one or more next frames with respect to the input intra frame, in output order, display order, or in coding order.
[0010] Example 7: The apparatus of example 2, wherein the apparatus is further caused to perform: filtering, with a motion-compensated temporal filtering, the input intra frame.
[0011] Example 8: The apparatus of example 7, wherein the apparatus is further caused to perform: training the end-to-end learned encoder by using a dataset that is based on a motion-compensated temporal filtering process.
[0012] Example 9: The apparatus of example 8, wherein the dataset comprises: an input data and a ground truth filtered by the motion-compensated temporal filtering; the input data filtered the motion-compensated temporal filtering; or the ground truth filtered by the motion-compensated temporal filtering.
[0013] Example 10: The apparatus of any of the examples 7 to 9, wherein the apparatus is further caused to perform: generating a first indicator for indicating whether the input intra frame is filtered by the motion-compensated temporal filtering; and signaling the first indicator to a decoder comprising an end-to-end learned decoder, wherein the first indicator is intended to be used as an auxiliary input to the end-to-end learned decoder. In example, the decoder includes a decoder module or circuit and the end-to-end learned decoder includes an end-to-end learned decoder module or circuit.
[0014] Example 11: The apparatus of any of the examples 2 to 10, wherein the apparatus is further caused to perform: providing the input intra frame and the input intra frame filtered by the motion-compensated temporal filtering as input to the end-to-end learned encoder.
[0015] Example 12: The apparatus of any of the examples 2 to 6, wherein the apparatus is further caused to perform: generating a second indicator for indicating to a decoder whether a pre-processing filter has been used to filter the input intra frame; and signaling the second indicator to the decoder, wherein the second indicator indicates to the decoder whether an loop filter is to be applied to an output of an end-to-end learned decoder comprised in the decoder.
[0016] Example 13: The apparatus of example 12, wherein the apparatus is further caused to perform: training the loop filter by using output of the end-to-end learned decoder as an input and using the input intra frame filtered by a motion-compensated temporal filtering as a ground truth.
[0017] Example 14: The apparatus of example 13, wherein: the end-to-end learned encoder and the loop filter are trained separately: the end-to-end learned encoder and the loop filter are trained jointly: the end-to-end learned encoder is trained first followed by training of the loop filter with the end-to-end learned encoder is kept frozen; or the end-to-end learned encoder is trained first followed by training of the loop filter, and wherein the end-to-end learned encoder is finetuned during the training of the end-to-end learned loop.
[0018] Example 15: The apparatus of any of example 13 or 14, wherein the apparatus is further caused to perform: encoding subsequent inter-frame by using an output of the loop filter as a reference frame, wherein the subsequent inter-frame comprises the another picture or frame.
[0019] Example 16: The apparatus of any of examples 12 to 14, wherein, a first inter frame which is closer to the intra-coded frame is coded by using a final output frame of the end-to-end learned decoder as a reference frame, and wherein a second inter-frame which is farther from the intra-coded frame is coded by using a reference frame generated by the end-to-end learned decoder as the reference frame.
[0020] Example 17: An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving a bitstream comprising an intra-coded frame and an inter-coded frame, wherein the intra-coded frame is encoded by an end-to-end learned encoder, and wherein the inter-coded frame is encoded by a conventional encoder: decoding the intra-coded frame by using an end-to-end learned decoder for obtaining a decoded intra frame; anddecoding the inter-coded frame, based at least on the decoded intra frame, by using a conventional decoder. In an example, the end-to-end learned encoder includes an end-to-end learned encoder module or circuit, the conventional encoder includes a conventional encoder module or circuit, the end-to-end learned decoder includes an end-to-end learned decoder module or circuit, and the conventional decoder includes a conventional decoder module or circuit.
[0021] Example 18: The apparatus of example 17, wherein the intra-coded frame comprises an encoded input intra frame and the inter-coded frame comprises an encoded input inter frame.
[0022] Example 19: The apparatus of example 18, wherein a filtering operation is performed on the input intra frame or a portion of the input intra frame by a filter.
[0023] Example 20: The apparatus of example 19, wherein the filter comprises a pre-processing filter.
[0024] Example 21 : The apparatus of example 19, wherein the filtering operation comprises one or more of: smoothing the input intra frame: removing high-frequency signals from the input intra frame: making the input intra frame similar or substantially similar to one or more first other frames of the video sequence that comprises the input intra frame: receiving the one or more first other frames as an input; or receiving data derived from the one or more first other frames.
[0025] Example 22: The apparatus of example 21 , wherein the one or more first other frames comprises: one or more previous frames with respect to the input intra frame in output order, display order, or in coding order; and / or one or more next frames with respect to the input intra frame, in output order, display order, or in coding order.
[0026] Example 23: The apparatus of example 18, wherein the input intra frame is filtered with a motion-compensated temporal filtering.
[0027] Example 24: The apparatus of example 23, wherein the apparatus is further caused to perform: training the end-to-end learned decoder by using a dataset that is based on a motion-compensated temporal filtering process.
[0028] Example 25: The apparatus of example 24, wherein the dataset comprises: an input data and a ground truth filtered by the motion-compensated temporal filtering; the input data filtered the motion-compensated temporal filtering; or the ground truth filtered by the motion-compensated temporal filtering.
[0029] Example 26: The apparatus of any of the examples 23 to 25, wherein the apparatus is further caused to perform: receiving a first indicator for indicating whether the input intra frame is filtered by the motion-compensated temporal filtering; and using the first indicator as an auxiliary input to the end-to-end learned decoder.
[0030] Example 27: The apparatus of any of the examples 18 to 22, wherein the apparatus is further caused to perform: receiving a second indicator for indicating whether a pre-processing filter has been used to filter the input intra frame; and applying an loop filter to an output of the end-to-end learned decoder based on the second indicator.
[0031] Example 28: The apparatus of example 27, wherein the apparatus is further caused to perform: training the loop filter by using an output of the end-to-end learned decoder as an input and using the intra frame filtered by a motion-compensated temporal filtering as a ground truth.
[0032] Example 29: The apparatus of example 28, wherein: the end-to-end learned decoder and the loop filter are trained separately: the end-to-end learned decoder and the loop filter are trained jointly: the end-to-end learned decoder is trained first followed by training of the loop filter with the end-to-end learned decoder is kept frozen; or the end-to-end learned decoder is trained first followed by training of the loop filter, and wherein the end-to-end learned decoder is finetuned during the training of the end-to-end learned loop.
[0033] Example 30: The apparatus of any of example 28 or 29, wherein the apparatus is further caused to perform: decoding subsequent inter-frame by using the output of the loop filter as a reference frame, wherein the subsequent inter-frame comprises the inter-coded frame.
[0034] Example 31 : The apparatus of any of examples 27 to 29, wherein, a first inter frame which is closer to the intra-coded frame is decoded by using the decoded intra frame as a reference frame, and wherein a second inter-frame which is farther from the intra-coded frame is decoded by using a reference frame generated by the end-to-end learned decoder as the reference frame.
[0035] Example 32: The apparatus any of the examples 17 to 31, wherein the apparatus is further caused to perform: decoding a subsequent inter-frame by using the decoded intra frame of the end-to-end learned decoder as a reference frame.
[0036] Example 33: The apparatus any of the examples 17 to 31, wherein the apparatus is further caused to perform: generating a reference frame.
[0037] Example 34: The apparatus of example 33, wherein the apparatus is further caused to perform: combining the decoded intra frame and the reference frame for generating a new reference frame.
[0038] Example 35: A method comprising: using an end-to-end learned encoder for encoding: a picture or frame of a video sequence as an intra-coded frame; or a block of the picture or frame of the video sequence as an intra-coded block; and using a conventional encoder for encoding: another picture or frame of the video sequence as inter-coded frame; or a block of the another picture or frame of the video sequence as an inter-coded block.
[0039] Example 36: The method of example 35, wherein the frame comprises an input intra frame.
[0040] Example 37 : The method of example 36 further comprising: performing filtering operation on the input intra frame or a portion of the input intra frame by using a filter.
[0041] Example 38: The method of example 37, wherein the filter comprises a pre-processing filter.
[0042] Example 39: The method of example 37, wherein the filtering operation comprises one or more of: smoothing the input intra frame: removing high-frequency signals from the input intra frame: making the intra frame similar or substantially similar to one or more first other frames of the video sequence that comprises the input intra frame: receiving the one or more first other frames as an input; or receiving data derived from the one or more first other frames.
[0043] Example 40: The method of example 39, wherein the one or more first other frames comprise: one or more previous frames with respect to the input intra frame in output order, display order, or in coding order; and / or one or more next frames with respect to the input intra frame, in output order, display order, or in coding order.
[0044] Example 41 : The method of example 36 further comprising: filtering, with a motion-compensated temporal filtering, the input intra frame.
[0045] Example 42: The method of example 41 further comprising: training the end-to-end learned encoder by using a dataset that is based on a motion-compensated temporal filtering process.
[0046] Example 43: The method of example 42, wherein the dataset comprises: an input data and a ground truth filtered by the motion-compensated temporal filtering; the input data filtered the motion-compensated temporal filtering; or the ground truth filtered by the motion-compensated temporal filtering.
[0047] Example 44: The method of any of the examples 41 to 43 further comprising: generating a first indicator for indicating whether the input intra frame is filtered by the motion-compensated temporal filtering; and signaling the first indicator to a decoder comprising an end-to-end learned decoder, wherein the first indicator is intended to be used as an auxiliary input to the end-to-end learned decoder.
[0048] Example 45: The method of any of the examples 36 to 44 further comprising: providing the input intra frame and the input intra frame filtered by the motion-compensated temporal filtering as input to the end-to-end learned encoder.
[0049] Example 46: The method of any of the examples 36 to 40 further comprising: generating a second indicator for indicating to a decoder whether a pre-processing filter has been used to filter the input intra frame; and signaling the second indicator to the decoder, wherein the second indicator indicates to the decoder whether an loop filter is to be applied to an output of an end-to-end learned decoder comprised in the decoder.
[0050] Example 47 : The method of example 46 further comprising: training the loop filter by using output of the end-to-end learned decoder as an input and using the input intra frame filtered by a motion-compensated temporal filtering as a ground truth.
[0051] Example 48: The method of example 47, wherein: the end-to-end learned encoder and the loop filter are trained separately: the end-to-end learned encoder and the loop filter are trained jointly: the end-to-end learned encoder is trained first followed by training of the loop filter with the end-to-end learned encoder is kept frozen; or the end-to-end learned encoder is trained first followed by training of the loop filter, and wherein the end-to-end learned encoder is finetuned during the training of the end-to-end learned loop.
[0052] Example 49: The method of any of example 47 or 48 further comprising: encoding subsequent inter-frame by using an output of the loop filter as a reference frame, wherein the subsequent inter-frame comprises the another picture or frame.
[0053] Example 50: The method of any of examples 46 to 48, wherein, a first inter frame which is closer to the intra-coded frame is coded by using a final output frame of the end-to-end learned decoder as a reference frame, and wherein a second inter-frame which is farther from the intra-coded frame is coded by using a reference frame generated by the end-to-end learned decoder as the reference frame.
[0054] Example 51 : A method comprising: receiving a bitstream comprising an intra-coded frame and an inter-coded frame, wherein the intra-coded frame is encoded by an end-to-end learned encoder, and wherein the inter-coded frame is encoded by a conventional encoder: decoding the intra-coded frame by using an end-to-endlearned decoder for obtaining a decoded intra frame; and decoding the inter-coded frame, based at least on the decoded intra frame, by using a conventional decoder.
[0055] Example 52: The method of example 51, wherein the intra-coded frame comprises an encoded input intra frame and the inter-coded frame comprises an encoded input inter frame.
[0056] Example 53: The method of example 52, wherein a filtering operation is performed on the input intra frame or a portion of the input intra frame by a filter.
[0057] Example 54: The method of example 53, wherein the filter comprises a pre-processing filter.
[0058] Example 55: The method of example 53, wherein the filtering operation comprises one or more of: smoothing the input intra frame: removing high-frequency signals from the input intra frame: making the input intra frame similar or substantially similar to one or more first other frames of the video sequence that comprises the input intra frame: receiving the one or more first other frames as an input; or receiving data derived from the one or more first other frames.
[0059] Example 56: The method of example 55, wherein the one or more first other frames comprises: one or more previous frames with respect to the input intra frame in output order, display order, or in coding order; and / or one or more next frames with respect to the input intra frame, in output order, display order, or in coding order.
[0060] Example 57: The method of example 52, wherein the input intra frame is filtered with a motion-compensated temporal filtering.
[0061] Example 58: The method of example 57 further comprising: training the end-to-end learned decoder by using a dataset that is based on a motion-compensated temporal filtering process.
[0062] Example 59: The method of example 58, wherein the dataset comprises: an input data and a ground truth filtered by the motion-compensated temporal filtering; the input data filtered the motion-compensated temporal filtering; or the ground truth filtered by the motion-compensated temporal filtering.
[0063] Example 60: The method of any of the examples 57 to 59 further comprising: receiving a first indicator for indicating whether the input intra frame is filtered by the motion-compensated temporal filtering; and using the first indicator as an auxiliary input to the end-to-end learned decoder.
[0064] Example 61 : The method of any of the examples 52 to 56 further comprising: receiving a second indicator for indicating whether a pre-processing filter has been used to filter the input intra frame; and applying an loop filter to an output of the end-to-end learned decoder based on the second indicator.
[0065] Example 62: The method of example 61 further comprising: training the loop filter by using an output of the end-to-end learned decoder as an input and using the intra frame filtered by a motion-compensated temporal filtering as a ground truth.
[0066] Example 63: The method of example 62, wherein: the end-to-end learned decoder and the loop filter are trained separately: the end-to-end learned decoder and the loop filter are trained jointly: the end-to-end learned decoder is trained first followed by training of the loop filter with the end-to-end learned decoder is kept frozen; or the end-to-end learned decoder is trained first followed by training of the loop filter, and wherein the end-to-end learned decoder is finetuned during the training of the end-to-end learned loop.
[0067] Example 64: The method of any of example 62 or 63 further comprising: decoding subsequent inter-frame by using the output of the loop filter as a reference frame, wherein the subsequent inter-frame comprises the inter-coded frame.
[0068] Example 65: The method of any of examples 61 to 63, wherein, a first inter frame which is closer to the intra-coded frame is decoded by using the decoded intra frame as a reference frame, and wherein a second inter-frame which is farther from the intra-coded frame is decoded by using a reference frame generated by the end-to-end learned decoder as the reference frame.
[0069] Example 66: The method any of the examples 51 to 65 further comprising: decoding a subsequent inter-frame by using the decoded intra frame of the end-to-end learned decoder as a reference frame.
[0070] Example 67 : The method any of the examples 51 to 65 further comprising: generating a reference frame.
[0071] Example 68: The method of example 67 further comprising: combining the decoded intra frame and the reference frame for generating a new reference frame.
[0072] Example 69: An apparatus comprising means for performing methods as described in any of the examples 35 to 50.
[0073] Example 70: An apparatus comprising means for performing methods as described in any of the examples 51 to 68.
[0074] Example 71: A computer readable medium comprising program instructions which when executed by an apparatus cause the apparatus to perform methods as described in any of the examples 35 to 50.
[0075] Example 72: The computer readable medium of example 71, wherein the computer readable medium comprises a non-transitory computer readable medium.
[0076] Example 73: A computer readable medium comprising program instructions which when executed by an apparatus cause the apparatus to perform methods as described in any of the examples 51 to 68.
[0077] Example 74: The computer readable medium of example 73, wherein the computer readable medium comprises a non-transitory computer readable medium.
[0078] Example 75: An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: generating reference data by using a neural network, where the neural network was trained by using processed images as ground truth.
[0079] Example 76: The apparatus of example 75, wherein the neural network is comprised in one of a loop filter, a reference frame generator, an intra-frame codec, or a motion compensation.
[0080] Example 77 : The apparatus of any of the examples 75 or 76, wherein the processed images used as ground truth are encoded and decoded by using a first quality level that is higher than a second quality level that is used for obtaining an input frame to the neural network.
[0081] Example 78: The apparatus of any of the examples 75 or 76, wherein the processed images have been filtered by a motion-compensated temporal filter.
[0082] Example 79: A method comprising: generating reference data by using a neural network, where the neural network was trained by using processed images as ground truth.
[0083] Example 80: The method of example 79, wherein the neural network is comprised in one of a loop filter, a reference frame generator, an intra-frame codec, or a motion compensation.
[0084] Example 81 : The method of any of the examples 79 or 80, wherein the processed images used as ground truth are encoded and decoded by using a first quality level that is higher than a second quality level that is used for obtaining an input frame to the neural network.
[0085] Example 82: The method of any of the examples 79 or 80, wherein the processed images have been filtered by a motion-compensated temporal filter.BRIEF DESCRIPTION OF THE DRAWINGS
[0086] The foregoing examples and other features are explained in the following description, taken in connection with the accompanying drawings, wherein:
[0087] FIG. 1 is a block diagram of one possible and non-limiting example system in which the example embodiments may be practiced;
[0088] FIG. 2 is a diagram illustrating features as described herein;
[0089] FIG. 3 is a diagram illustrating features as described herein;
[0090] FIG. 4 is a diagram illustrating features as described herein;
[0091] FIG. 5 is a diagram illustrating features as described herein;
[0092] FIG. 6 is a diagram illustrating features as described herein;
[0093] FIG. 7 is a diagram illustrating features as described herein;
[0094] FIG. 8 is a diagram illustrating features as described herein;
[0095] FIG. 9 is a diagram illustrating an example apparatus, which may be implemented in hardware, configured to implement the examples described herein;
[0096] FIG. 10 is a diagram illustrating an example of non-volatile memory media used to store instructions that implement the examples described herein;
[0097] FIG. 11 is a flowchart illustrating an example method as described herein; and
[0098] FIG. 12 is a flowchart illustrating another example method as described herein.
[0099] FIG. 13 is a flowchart illustrating another example method as described herein.DETAILED DESCRIPTION OF EMBODIMENTS
[0100] The following abbreviations that may be found in the specification and / or the drawing figures are defined as follows:3GPP third generation partnership project4G fourth generation5G fifth generation5GC 5G core networkAPS adaptation parameter setAR augmented realityCABAC context-adaptive binary arithmetic codingCDMA code division multiple accessCPU central processing unitcRAN cloud radio access networkDCT discrete cosine transformE2E end-to-endeNB (ore Node B) evolved Node B (e.g., an LTE base station)EN-DC E-UTRA-NR dual connectivityen-gNB or En-gNB node providing NR user plane and control plane protocol terminations towards the UE, and acting as secondary node in EN- DCE-UTRA evolved universal terrestrial radio access, i.e., the LTE radio access technologyFDMA frequency division multiple accessGAN generative adversarial networkgNB (or g Node B) base station for 5G / NR, i.e., a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GCGPU graphical processing unitGSM global systems for mobile communicationsHMD head-mounted displayIBC intra block copyIEEE Institute of Electrical and Electronics EngineersIMD integrated messaging deviceIMS instant messaging serviceloT Internet of ThingsJVET Joint Video Expert TeamLTE long term evolutionMAE mean absolute errormAP mean average precisionMMS multimedia messaging serviceMPEG-I Moving Picture Experts Group immersive codec familyMR mixed realityMSE mean squared errorMS-SSIM multiscale structure similarity index measureNAL network abstraction layerng or NG new generationng-eNB or NG-eNB new generation eNBNN neural networkNNC neural network codingNR new radioN / Wor NW networkO-RAN open radio access networkPC personal computerPDA personal digital assistantPSNR peak signal-to-noise ratioQP quantization parameterROI region of interestSEI supplemental enhancement informationSGD stochastic gradient descentSMS short messaging serviceSSIM structure similarity index measureTCP-IP transmission control protocol-internet protocolTDMA time division multiple accessUE user equipment (e.g., a wireless, typically mobile device) UMTS universal mobile telecommunications systemUSB universal serial busVCM video coding for machinesVMAF Video Multimethod Assessment FusionVNR virtualized network functionVR virtual realityWC volumetric video codingWLAN wireless local area network
[0101] The following describes suitable apparatus and possible mechanisms for practicing example embodiments of the present disclosure. Accordingly, reference is first made to FIG. 1, which shows an example block diagram of an apparatus 50 (for example, an electronic device or a user equipment). The apparatus may be configured to perform various functions such as, for example, gathering information by one or more sensors, encoding and / or decoding information, receiving and / or transmitting information, analyzing information gathered or received by the apparatus, or the like. A device configured to encode a video scene may (optionally) comprise one or more microphones for capturing the scene and / or one or more sensors, such as cameras, for capturing information about the physical environment in which the scene is captured. Alternatively, a device configured to encode a video scene may be configured to receive information about an environment in which a scene is captured and / or a simulated environment. A device configured to decode and / or render the video scene may be configuredto receive a Moving Picture Experts Group immersive codec family (MPEG-I) bitstream comprising the encoded video scene. A device configured to decode and / or render the video scene may comprise one or more speakers / audio transducers and / or displays, and / or may be configured to transmit a decoded scene or signals to a device comprising one or more speakers / audio transducers and / or displays. A device configured to decode and / or render the video scene may comprise a user equipment, a head / mounted display, or another device capable of rendering to a user an AR, VR and / or MR experience.
[0102] The apparatus 50 may for example be a mobile terminal or user equipment of a wireless communication system. Alternatively, the electronic device may be a computer or part of a computer that is not mobile. It should be appreciated that example embodiments of the present disclosure may be implemented within any electronic device or apparatus which may process data. The apparatus 50 may comprise a device that can access a network and / or cloud through a wired or wireless connection. The apparatus 50 may comprise one or more processors / controllers 56, one or more memories 58, and one or more radio interface circuitry 52 interconnected through one or more buses. The one or more processors / controllers 56 may comprise a central processing unit (CPU) and / or a graphical processing unit (GPU). Each of the one or more radio interface circuitry 52 includes a receiver and a transmitter. The one or more buses may be address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. A “circuit” may include dedicated hardware or hardware in association with software executable thereon. The one or more transceivers may be connected to one or more antennas 44. The one or more memories 58 may include computer program code. The one or more memories 58 and the computer program code may be configured to, with the one or more processors / controllers 56, cause the apparatus 50 to perform one or more of the operations as described herein.
[0103] The apparatus 50 may connect to a node of a network. The network node may comprise one or more processors, one or more memories, and one or more transceivers interconnected through one or more buses. Each of the one or more transceivers includes a receiver and a transmitter. The one or more buses may be address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. The one or more transceivers may be connected to one or more antennas. The one or more memories may include computer program code. The one or more memories and the computer program code may be configured to, with the one or more processors, cause the network node to perform one or more of the operations as described herein.
[0104] The apparatus 50 may comprise a microphone 36 or any suitable audio input which may be a digital or analogue signal input. The apparatus 50 may further comprise an audio output device 38 which in example embodiments of the present disclosure may be any one of: an earpiece, speaker, or an analogue audio or digital audio output connection. The apparatus 50 may also comprise a battery (or in other example embodiments of the present disclosure the device may be powered by any suitable mobile energy device such as solar cell, fuel cell, or clockwork generator). The apparatus 50 may further comprise a camera 42 or other sensor capable of recording or capturing images and / or video. Additionally or alternatively, the apparatus 50 may further comprise a depth sensor. The apparatus 50 may further comprise a display 32. The apparatus 50 may further comprise an infraredport for short range line of sight communication to other devices. In other example embodiments of the present disclosure the apparatus or the apparatus 50 may further comprise any suitable short-range communication solution such as for example a BLUETOOTH™ wireless connection or a USB / firewire wired connection.
[0105] It should be understood that an apparatus 50 configured to perform example embodiments of the present disclosure may have fewer and / or additional components, which may correspond to what processes the apparatus 50 is configured to perform. For example, an apparatus configured to encode a video might not comprise a speaker or audio transducer and may comprise a microphone, while an apparatus configured to render the decoded video might not comprise a microphone and may comprise a speaker or audio transducer.
[0106] Referring now to FIG. 1, the apparatus 50 may comprise a processors / controllers 56, processor or processor circuitry for controlling the apparatus 50. The processors / controllers 56 may be connected to memory 58 which in example embodiments of the present disclosure may store both data in the form of image and audio data and / or may also store instructions for implementation on the processors / controllers 56. The processors / controllers 56 may further be connected to codec circuitry 54 suitable for carrying out coding and / or decoding of audio and / or video data or assisting in coding and / or decoding carried out by the controller.
[0107] The apparatus 50 may further comprise a card reader 48 and a smart card 46, for example a UICC and UICC reader, for providing user information and being suitable for providing authentication information for authentication and authorization of the apparatus 50 at a network. The apparatus 50 may further comprise an input device 34, such as a keypad, one or more input buttons, or a touch screen input device, for providing information to the processors / controllers 56.
[0108] The apparatus 50 may comprise a radio interface circuitry 52 (for example, transceivers) connected to the controller and suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system, or a wireless local area network. The apparatus 50 may further comprise an antenna 44 connected to the radio interface circuitry 52 for transmitting radio frequency signals generated at the radio interface circuitry 52 to other apparatus(es) and / or for receiving radio frequency signals from other apparatus(es).
[0109] The apparatus 50 may comprise a microphone 36, camera 42, and / or other sensors capable of recording or detecting audio signals, image / video signals, and / or other information about the local / virtual environment, which are then passed to the codec circuitry 54 or the processors / controllers 56 for processing. The apparatus 50 may receive the audio / image / video signals and / or information about the local / virtual environment for processing from another device prior to transmission and / or storage. The apparatus 50 may also receive either wirelessly or by a wired connection the audio / image / video signals and / or information about the local / virtual environment for encoding / decoding. The structural elements of apparatus 50 described above represent examples of means for performing a corresponding function.
[0110] The memory 58 may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The memory 58 may be a non-transitory memory. The memory 58 may be means forperforming storage functions. The processors / controllers 56 may be or comprise one or more processors, which may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on a multi-core processor architecture, as non-limiting examples. The processors / controllers 56 may be means for performing functions.
[0111] The apparatus 50 may be configured to perform capture of a volumetric scene according to example embodiments of the present disclosure. For example, the apparatus 50 may comprise a camera 42 or other sensor capable of recording or capturing images and / or video. The apparatus 50 may also comprise one or more radio interface circuitry 52 to enable transmission of captured content for processing at another device. Such an apparatus 50 may or may not include all the modules illustrated in FIG. 1.
[0112] The apparatus 50 may be configured to perform processing of volumetric video content according to example embodiments of the present disclosure. For example, the apparatus 50 may comprise a processors / controllers 56 for processing images to produce volumetric video content, a processors / controllers 56 for processing volumetric video content to project 3D information into 2D information, patches, and auxiliary information, and / or a codec circuitry 54 for encoding 2D information, patches, and auxiliary information into a bitstream for transmission to another device with radio interface circuitry 52. Such an apparatus 50 may or may not include all the modules illustrated in FIG. 1.
[0113] The apparatus 50 may be configured to perform encoding or decoding of 2D information representative of volumetric video content according to example embodiments of the present disclosure. For example, the apparatus 50 may comprise a codec circuitry 54 for encoding or decoding 2D information representative of volumetric video content. Such an apparatus 50 may or may not include all the modules illustrated in FIG. 1.
[0114] The apparatus 50 may be configured to perform rendering of decoded 3D volumetric video according to example embodiments of the present disclosure. For example, the apparatus 50 may comprise a controller for projecting 2D information to reconstruct 3D volumetric video, and / or a display 32 for rendering decoded 3D volumetric video. Such an apparatus 50 may or may not include all the modules illustrated in FIG. 1.
[0115] With respect to FIG. 2, an example of a system within which example embodiments of the present disclosure can be utilized is shown. The system 10 comprises multiple communication devices which can communicate through one or more networks. The system 10 may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network (such as a GSM, UMTS, E-UTRA, LTE, CDMA, 4G, 5G, 6G network etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a BLUETOOTH™ personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and / or the Internet. A wireless network may implement network virtualization, which is the process of combining hardware and software network resources and network functionality into a single, software-based administrative entity, a virtual network. Network virtualization involves platform virtualization, often combined with resource virtualization. Network virtualization is categorized as either external, combining many networks, or parts of networks, into a virtual unit, or internal, providing network-like functionality to softwarecontainers on a single system. For example, a network may be deployed in a tele cloud, with virtualized network functions (VNF) running on, for example, data center servers. For example, network core functions and / or radio access network(s) (e.g. CloudRAN, O-RAN, edge cloud) may be virtualized. Note that the virtualized entities that result from the network virtualization are still implemented, at some level, using hardware such as processors and memories, and also such virtualized entities create technical effects.
[0116] It may also be noted that operations of example embodiments of the present disclosure may be carried out by a plurality of cooperating devices (e.g. cRAN).
[0117] The system 10 may include both wired and wireless communication devices and / or electronic devices suitable for implementing example embodiments of the present disclosure.
[0118] For example, the system shown in FIG. 2 shows a mobile telephone network 11 and a representation of the internet 28. Connectivity to the internet 28 may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.
[0119] The example communication devices shown in the system 10 may include, but are not limited to, an apparatus 15, a combination of a personal digital assistant (PDA) and a mobile telephone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, a notebook computer 22, and a head-mounted display (HMD) 17. The apparatus 50 may comprise any of those example communication devices. In an example embodiment of the present disclosure, more than one of these devices, or a plurality of one or more of these devices, may perform the disclosed process(es). These devices may connect to the internet 28 through a wireless connection 2.
[0120] The example embodiments of the present disclosure may also be implemented in a set-top box; i.e. a digital TV receiver, which may / may not have a display or wireless capabilities, in tablets or (laptop) personal computers (PC), which have hardware and / or software to process neural network data, in various operating systems, and in chipsets, processors, DSPs and / or embedded systems offering hardware / software based coding. The example embodiments of the present disclosure may also be implemented in cellular telephones such as smart phones, tablets, personal digital assistants (PDAs) having wireless communication capabilities, portable computers having wireless communication capabilities, image capture devices such as digital cameras having wireless communication capabilities, gaming devices having wireless communication capabilities, music storage and playback appliances having wireless communication capabilities, Internet appliances permitting wireless Internet access and browsing, tablets with wireless communication capabilities, as well as portable units or terminals that incorporate combinations of such functions.
[0121] Some or further apparatus may send and receive calls and messages and communicate with service providers through a wireless connection 25 to a base station 24, which may be, for example, an eNB, gNB, access point, access node, other node, etc. The base station 24 may be connected to a network server 26 that allows communication between the mobile telephone network 11 and the internet 28. The system may include additional communication devices and communication devices of various types.
[0122] The communication devices may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocol-internet protocol (T CP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), BLUETOOTH™, IEEE 802.11, 3GPP Narrowband loT and any similar wireless communication technology. A communications device involved in implementing various example embodiments of the present disclosure may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.
[0123] In telecommunications and data networks, a channel may refer either to a physical channel or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire, whereas a logical channel may refer to a logical connection over a multiplexed medium, capable of conveying several logical channels. A channel may be used for conveying an information signal, for example a bitstream, which may be a MPEG-I bitstream, from one or several senders (or transmitters) to one or several receivers.
[0124] Having thus introduced one suitable but non-limiting technical context for the practice of the example embodiments of the present disclosure, example embodiments will now be described with greater specificity.
[0125] Fundamentals of neural networks
[0126] Features as described herein may generally relate to neural networks. A neural network (NN) may be described as a computation graph consisting of several layers of computation. In an example of a NN, each layer may consist of one or more units, where each unit may perform an elementary computation. A unit may be connected to one or more other units, and the connection may be associated with a weight. The weight may be used for scaling the signal passing through the associated connection. Weights are learnable parameters, i.e., values which can be learned from training data. There may be other learnable parameters, such as those of batchnormalization layers. Example embodiments of the present disclosure may or may not relate to, or involve, NN comprising multiple layers of computation.
[0127] In some neural networks, such as convolutional neural networks for image classification, initial layers (those close to the input data) may extract semantically low-level features such as edges and textures in images, whereas intermediate layers may extract higher-level features. After the feature extraction layers, there may be one or more layers performing a certain task, such as classification, semantic segmentation, object detection, denoising, style transfer, super-resolution, etc. Example embodiments of the present disclosure may or may not relate to, or involve, convolutional neural networks.
[0128] Neural networks are being utilized in an ever-increasing number of applications for many different types of devices, such as mobile phones. Examples include image and video analysis and processing, social media data analysis, device usage data analysis, etc.
[0129] One property of neural nets / networks (and other machine learning tools) is that they are able to learn properties from input data, e.g., in a supervised way or in an unsupervised way. Such learning may be a result of a training algorithm, or may be achieved by means of another neural network providing the training signal (sometimes, this latter approach may be referred to as “meta learning”).
[0130] In general, the training algorithm may consist of changing some properties of the neural network so that its output is as close as possible to a desired output. For example, in the case of classification of objects in images, the output of the neural network may be used to derive a class or category index which may indicate the class or category to which the object in the input image belongs. Training may comprise minimizing or decreasing the output’s error, also referred to as the loss or loss function. Examples of losses are mean squared error, crossentropy, etc. Example embodiments of the present disclosure may or may not relate to, or involve, neural networks trained according to a training algorithm.
[0131] In recent deep learning techniques, training may be an iterative process, where at each iteration the algorithm may modify the weights of the neural net to make a gradual improvement of the network’s output, i.e., to gradually decrease the loss, for example by means of a gradient descent technique. In an example, at each training iteration, gradients of the loss function with respect to one or more weights or parameters of the NN may be computed, for example by a backpropagation technique; the computed gradients may then be used by an optimization routine, such as Adam or Stochastic Gradient Descent (SGD) to obtain an update to the one or more weights or parameters.
[0132] In the present disclosure, the terms “model”, “neural network”, “neural net’ and “network” are used interchangeably. In the present disclosure, the weights of neural networks may sometimes be referred to as learnable parameters or simply as parameters.
[0133] Training a neural network may be regarded as an optimization process, but the final goal may be different from the typical goal of optimization. In optimization, the main goal is to minimize a function. In machine learning, the goal of the optimization or training process is to make the model learn the properties of the data distribution from a limited training dataset. In other words, the goal is to learn to use a limited training dataset in order to learn to generalize to previously unseen data, i.e., data which was not used for training the model. This is usually referred to as generalization. In practice, data is usually split into at least two sets, the training set and the validation set. The training set is used for training the network, i.e., to modify its learnable parameters in order to minimize the loss. The validation set is at least partially different from the training set. The validation set is used for checking the performance of the network on data which was not used to minimize the loss, as an indication of the final performance of the model. In particular, the errors on the training set and on the validation set may be monitored during the training process to understand the following:
[0134] - If the network is learning at all - in this case, the training set error should decrease, otherwise the model is in the regime of underfitting.
[0135] - If the network is learning to generalize - in this case, also the validation set error needs to decrease and to be not too much higher than the training set error. If the training set error is low, but the validation set error is much higher than the training set error, or it does not decrease, or it even increases, the model may be in the regime of overfitting. This means that the model has just memorized the training set’s properties and performs well only on that set, but performs poorly on a set not used for tuning its parameters.
[0136] Fundamentals of video / image coding
[0137] Features as described herein may generally relate to video or image coding. A video codec consists of an encoder that transforms the input video into a compressed representation suited for storage / transmission, and a decoder that can decompress the compressed video representation back into a viewable form. Typically, the encoder discards some information in the original video sequence in order to represent the video in a more compact form (that is, at a lower bitrate).
[0138] Typical hybrid video codecs, for example ITU-T H.263 and H.264, encode the video information in two phases. Firstly, pixel values in a certain picture area (or “block”) are predicted, for example by motion compensation means (i.e. finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded), or by spatial means (i.e. using the pixel values around the block to be coded in a specified manner). Secondly, the prediction error, i.e. the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the difference in pixel values using a specified transform (e.g. Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients, and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (i.e. picture quality) and the size of the resulting coded video representation (i.e. file size or transmission bitrate).
[0139] Inter prediction, which may also be referred to as temporal prediction, motion compensation, or motion-compensated prediction, exploits temporal redundancy. In inter prediction, the sources of prediction are previously decoded pictures (a.k.a. reference pictures).
[0140] In temporal inter prediction, the sources of prediction are previously decoded pictures in the same scalable layer. In intra block copy (IBC; a.k.a. intra-block-copy prediction), prediction may be applied similarly to temporal inter prediction, but the reference picture is the current picture, and only previously decoded samples can be referred in the prediction process. Inter-layer or inter-view prediction may be applied similarly to temporal inter prediction, but the reference picture is a decoded picture from another scalable layer or from another view, respectively. In some cases, inter prediction may refer to temporal inter prediction only, while in other cases inter prediction may refer collectively to temporal inter prediction and any of intra block copy, inter-layer prediction, and inter-view prediction, provided that they are performed with the same or similar process as temporal prediction. Inter prediction, temporal inter prediction, or temporal prediction may sometimes be referred to as motion compensation or motion-compensated prediction.
[0141] Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in the spatial or transform domain, i.e., either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.
[0142] One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy-coded more efficiently if they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors, and only the difference relative to the motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.
[0143] The decoder reconstructs the output video by applying prediction means similar to the encoder to form a predicted representation of the pixel blocks (e.g. using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (e.g. inverse operation of the prediction error coding recovering the quantized prediction error signal in the spatial pixel domain). After applying prediction and prediction error decoding means, the decoder sums up the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) can also apply additional filtering means to improve the quality of the output video before passing it for display and / or storing it as prediction reference for the forthcoming frames in the video sequence.
[0144] In typical video codecs, the motion information is indicated with motion vectors associated with each motion compensated image block. Each of these motion vectors represents the displacement of the image block in the picture to be coded (in the encoder side) or decoded (in the decoder side) and the prediction source block in one of the previously coded or decoded pictures. In order to represent motion vectors efficiently, those are typically coded differentially with respect to block specific predicted motion vectors. In typical video codecs, the predicted motion vectors are created in a predefined way, for example calculating the median of the encoded or decoded motion vectors of the adjacent blocks. Another way to create motion vector predictions is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in the temporal reference pictures, and signaling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index is typically predicted from adjacent blocks and / or or co-located blocks in the temporal reference picture. Moreover, typical high efficiency video codecs employ an additional motion information coding / decoding mechanism, often called merging / merge mode, where all the motion field information, which includes the motion vector and corresponding reference picture index for each available reference picture list, is predicted and used without any modification / correction. Similarly, predicting the motion field information is carried out using the motion field information of adjacent blocks and / or co-located blocks in the temporal reference pictures, and the used motion field information is signaled among a list of motion field candidates filled with motion field information of available adjacent / co-located blocks.
[0145] In typical video codecs, the prediction residual after motion compensation is first transformed with a transform kernel (like DCT) and then coded. The reason for this is that, often, there still exists some correlation among the residual and transform can in many cases help reduce this correlation and provide more efficient coding.
[0146] Typical video encoders utilize Lagrangian cost functions to find optimal coding modes, e.g. the desired Macroblock mode and associated motion vectors. This kind of cost function uses a weighting factor A to tie together the (exact or estimated) image distortion due to lossy coding methods and the (exact or estimated) amount of information that is required to represent the pixel values in an image area:C = D + AR
[0147] where C is the Lagrangian cost to be minimized, D is the image distortion (e.g. Mean Squared Error) with the mode and motion vectors considered, and R the number of bits needed to represent the required data toreconstruct the image block in the decoder (including the amount of data to represent the candidate motion vectors).
[0148] Video coding specifications may enable the use of supplemental enhancement information (SEI) messages or alike. Some video coding specifications include SEI network abstraction layer (NAL) units, and some video coding specifications contain both prefix SEI NAL units and suffix SEI NAL units, where the former type can start a picture unit or alike, and the latter type can end a picture unit or alike. An SEI NAL unit contains one or more SEI messages which are not required for the decoding of output pictures but may assist in related processes, such as picture output timing, post-processing of decoded pictures, rendering, error detection, error concealment, and resource reservation. Several SEI messages are specified in H.264 / AVC, H.265 / HEVC, H.266AA / C, and H.274A / SEI standards, and the user data SEI messages enable organizations and companies to specify SEI messages for their own use. The standards may contain the syntax and semantics for the specified SEI messages, but a process for handling the messages in the recipient might not be defined. Consequently, encoders may be required to follow the standard specifying a SEI message when they create SEI message(s), and decoders might not be required to process SEI messages for output order conformance. One of the reasons to include the syntax and semantics of SEI messages in standards is to allow different system specifications to interpret the supplemental information identically, and hence interoperate. System specifications may require the use of particular SEI messages both in the encoding end and in the decoding end, and additionally the process for handling particular SEI messages in the recipient may be specified.
[0149] Information on neural network based imaqe / video coding
[0150] Features as described herein may generally relate to use of NN to code images and / or videos. Recently, neural networks (NNs) have been used in the context of image and video compression, by following mainly two approaches.
[0151] In a first approach, NNs are used to replace one or more of the components of a traditional codec, such as a WC / H.266-compliant codec. Here, “traditional” or “legacy” means those codecs whose components and their parameters are typically not learned from data by means of machine learning techniques. Examples of components that may be implemented as neural networks are: an in-loop filter, for example a NN that works as an additional in-loop filter with respect to the traditional loop filters, or a NN that works as the only additional in-loop filter, thus replacing any other in-loop filter; Intra-frame prediction; inter-frame prediction; transform and / or inverse transform; probability model for lossless coding; etc.
[0152] In a second approach, commonly referred to as “end-to-end learned compression” (or end-to-end learned codec), NNs are used as the main components of the image / video codecs. However, the codec may still comprise components which are not based on machine learning techniques. In this second approach, two design options are as follows:
[0153] - Option 1 : re-use the traditional video coding pipeline, but replace most or all the components with NNs. Referring now to FIG. 3, illustrated is an example of an end-to-end learned codec that includes NNs replacing some components of the traditional video coding pipeline. Input signal (x) (302) may be combined (303) with other information and provided to a neural transform (304), which may also receive input from an encoder parametercontrol (306). Output of the neural transform (304) may be provided for quantization (308), and then for inverse quantization / neural transform (310) as well as entropy coding (312) to a bitstream (314). Entropy coding (312) may be performed based on input from the encoder parameter control (306).
[0154] The output of the inverse quantization / neural transform (310) may be combined with other information, and provided to a neural intra codec (316) and to a deep loop filter (324). The neural intra codec (316) may also receive input from the encoder parameter control (306), and may comprise an encoder (318), intra coding (320), and a decoder (322).
[0155] The deep loop filter (324) may also receive input from the encoder parameter control (306), and may provide output to a decode picture buffer (326), which may produce an enhanced reference frame (328) based, at least partially, on one or more reconstructed frames (330). The decode picture buffer (326) may provide output for inter prediction (332), which may provide output based, at least partially, on input from the encoder parameter control (306) and ME / MC (336), Gnet(Cnet( )) (334).
[0156] In the example of FIG. 3, the forward and inverse transforms were replaced with two neural networks (304, 310), the neural intra codec (316) comprises a neural network, and the deep loop filter (324) is a neural network.
[0157] - Option 2: re-design the whole pipeline as a neural network auto-encoder with a quantization and lossless coding in the middle part. This option may also be referred to as end-to-end learned coding. The codec may comprise the following:
[0158] - Encoder NN (also referred to as a neural network based encoder, or NN encoder): performs a nonlinear transformation of the input. The output is typically referred to as a latent tensor.
[0159] - Quantization and lossless encoding of the encoder NN’s output.
[0160] - Lossless decoding and dequantization.
[0161] - Decoder NN (also referred to as a neural network based decoder, or NN decoder): performs a nonlinear inverse transformation from dequantized latent tensor to a reconstructed input.
[0162] It is to be understood that even in end-to-end learned approaches, there may be components which are not learned / trained from data, such as the arithmetic codec.
[0163] Further information on neural network-based end-to-end learned video coding
[0164] Features as described herein may generally relate to NN-based end-to-end (E2E) learned video codecs. Referring now to FIG. 4, illustrated is an example of neural network-based end-to-end learned coding, such as an end-to-end learned video coding system or an end-to-end learned image coding system.
[0165] Even though some examples are provided with respect to coding images or videos, it is to be understood that other types of data may be coded in a similar way, such as audio, speech, text, features, etc. As shown in FIG. 4, a typical neural network-based end-to-end learned coding system comprises an encoder (405) and a decoder (460).
[0166] The encoder (405) comprises an encoder NN (415), a quantizer or quantization operation (425), a probability model (435), a lossless encoder (445) (for example arithmetic encoder). The decoder (460) comprisesa lossless decoder (455) (for example, an arithmetic decoder), a probability model (465), a dequantizer or dequantization operation (475), and a decoder NN (485).
[0167] It is to be noted that the probability model (435) present at encoder side and the probability model (465) present at decoder side may be the same or substantially the same. For example, they may be two copies of the same probability model. The probability model (435, 465) may also be a neural network and / or may mainly comprise neural network components, and may be referred to as a neural network based probability model or learned probability model.
[0168] The lossless encoder (445) and the lossless decoder (455) form a lossless codec (440). A lossless codec may be an entropy-based lossless codec. An example of a lossless codec is an arithmetic codec, such as a context-adaptive binary arithmetic coding (CABAC). Sometimes, the term lossless codec may refer to a system that comprises also the probability model, in addition to, for example, an arithmetic encoder and an arithmetic decoder.
[0169] The encoder NN (415) and the decoder NN (485) may typically be two neural networks, or may mainly comprise neural network components.
[0170] The quantization operation (425), dequantization operation (475) and lossless codec (440) are typically not based on neural network components, but may potentially comprise neural network components.
[0171] In the example of FIG. 4, the encoder NN (415) may take an inputx (410), which may comprise, for example, an image to be compressed. The encoder N 475N (415) may output a latent tensor z (420). In an example, the latent tensor may be a 3D tensor, where the three dimensions of such tensor may represent a channel dimension, a vertical dimension (also sometimes referred to as height dimension) and a horizontal dimension (also sometimes referred to as width dimension). In another example, the latent tensor may be a 4D tensor, where the four dimensions of such tensor may represent sample dimension (also sometimes referred to as batch dimension, which is the dimension along which different samples of data can be placed), a channel dimension, a vertical dimension (also sometimes referred to as height dimension) and a horizontal dimension (also sometimes referred to as width dimension). In yet another example, in the case of compressing a signal with a temporal dimension such as a video, the latent tensor may be a 4D tensor, where the four dimensions of such tensor may represent a channel dimension, a vertical dimension (also sometimes referred to as height dimension), a horizontal dimension (also sometimes referred to as width dimension), and a temporal dimension. The latent tensor (420) may be input to a quantization operation (425), obtaining a quantized latent tensor zq(430). The quantized latent tensor zq(430) may be lossless-encoded into a bitstream b (450) by the lossless encoder (445), based also on an output of the probability model (435). In particular, the probability model may take as input at least part of the quantized latent tensor zq(430) and may output an estimate of a probability, or an estimate of a probability distribution, or an estimate of one or more parameters of a probability distribution, for one or more elements of the quantized latent tensor. The bitstream (450) may represent an encoded or compressed version of the input x (410).
[0172] The bitstream (450) may be lossless-decoded by the lossless decoder (455) also based on an output of the probability model (465) present at decoder side, obtaining a quantized latent tensor zq(470). The quantized latent tensor may be dequantized by a dequantization operation(475), obtaining a reconstructed latent tensor z(480). The reconstructed latent tensor (480) may be input to a decoder NN (485), obtaining a reconstructed input x (490), i.e., a reconstructed version of the input x (410). The reconstructed input x (490) may also be referred to as reconstructed data, or reconstruction, or decoded data, or decoded input, or decoded output, and the like.
[0173] FIG. 4 presents a simplified description of an end-to-end learned codec; more sophisticated designs, or variations of this design, are possible.
[0174] The neural network components, or a subset of the neural network components, of an end-to-end learned codec may be trained by minimizing a rate-distortion loss function:L = D + AR
[0175] where D is a distortion loss term, R is a rate loss term, and A is a weight that controls the balance between the two losses. The distortion loss term may be referred to also as reconstruction loss term, or simply reconstruction loss. The rate loss term may be referred to simply as rate loss.
[0176] The distortion loss term measures the quality of the reconstructed or decoded output, and may comprise (but may not be limited to) one or more of the following:
[0177] - Mean square error (MSE)
[0178] - Structure similarity index measure(SSIM)
[0179] - Multiscale structure similarity index measure (MS-SSIM)
[0180] - Losses derived from the use of a pretrained neural network. For example, error(f1, f2), where f1 and f2 are the features extracted by a pretrained neural network for the input data and the decoded data, respectively, and error() is an error or distance function, such as L1 norm or L2 norm.
[0181] - Losses derived from the use of a neural network that is trained (substantially) simultaneously with the end-to-end learned codec. For example, adversarial loss can be used, which is the loss provided by a discriminator neural network that is trained adversarially with respect to the codec, following the settings proposed in the context of Generative Adversarial Networks (GANs) and their variants.
[0182] - Loss that is related to a performance of one or more machine analysis tasks or to an estimated performance of one or more machine analysis tasks, where the one or more machine analysis tasks may comprise classification, object detection, image segmentation, instance segmentation, etc. In an example, the estimated performance of one or more machine analysis tasks may comprise a distortion computed based at least on a first set of features extracted from an output of the decoder, and a second set of features extracted from a respective ground truth data, where the first set of features and the second set of features are output by one or more layers of a pretrained feature-extraction neural network.
[0183] Multiple distortion losses may be used and integrated into D, such as a weighted sum of MSE and SSIM.
[0184] The rate loss term may be used to train the encoder NN to output a low-entropy latent tensor, or a latent tensor such that the quantized latent tensor has low entropy, or a latent tensor such that the probability distribution of the quantized latent tensor may be better estimated or predicted by the probability model.
[0185] The rate loss term may be used to train the probability model to better estimate or predict the probability distribution of the quantized latent tensor.
[0186] Examples of the rate loss terms include the following:
[0187] - In an example, the rate loss term may be derived from the output of the probability model, and it may represent the estimated entropy of the quantized latent representation, which may indicate the number of bits necessary to represent the quantized latent tensor.
[0188] - A sparsification loss, i.e., a loss that encourages the quantized latent tensor to comprise many zeros. Examples are L0 norm, L1 norm, L1 norm divided by L2 norm.
[0189] In order to train the neural network components, or a subset of the neural network components, of an end-to-end learned codec, one or more of reconstruction losses may be used, and one or more rate losses may be used. In an example, the one or more reconstruction losses and / or one or more rate losses may be combined by means of a weighted sum. Typically, the different loss terms are weighted using different weights, and these weights determine how the final system performs in terms of rate-distortion performance. For example, if more weight is given to the reconstruction losses with respect to the rate losses, the system may learn to compress less, but to reconstruct with higher accuracy (e.g. as measured by a metric that correlates with the reconstruction losses). These weights are usually considered to be hyper-parameters of the training process, and may be set manually by the person designing the training process, or automatically, for example by grid search or by using additional neural networks.
[0190] In one case, the training process may be performed jointly with respect to the distortion loss D and the rate loss R. In another case, the training process may be performed in two alternating phases, where in a first phase only the distortion loss D may be used, and in a second phase only the rate loss R may be used.
[0191] For lossless video / image compression, the system may only comprise the probability model and lossless encoder and lossless decoder. The loss function would comprise only the rate loss, since the distortion loss is always zero (i.e., no loss of information).
[0192] In the present disclosure, inference phase, or inference stage, or inference time, or test time, are referred to the phase when a neural network or a codec is used for its purpose, such as encoding and decoding an input image.
[0193] Information on Video Coding for Machines (VCM)
[0194] Features as described herein may generally relate to video coding for machines (VCM). Reducing the distortion in image and video compression is often intended to increase human perceptual quality, as humans are considered to be the end users, i.e. consuming / watching the decoded images or videos. Recently, with the advent of machine learning, especially deep learning, there is a rising number of machines (i.e., autonomous agents) that analyze data independently from humans, and may even make decisions based on the analysis results without human intervention. Examples of such analysis are object detection, scene classification, semantic segmentation, video event detection, anomaly detection, pedestrian tracking, etc. For example, such analysis tasks may be performed by neural networks.
[0195] It is likely that the device where the analysis takes place has multiple “machines” or neural networks (NNs). These multiple machines may be used in a certain combination which is, for example, determined by an orchestrator sub-system. The multiple machines may be used, for example, in succession, based on the output ofthe previously used machine, and / or in parallel. For example, a video may be analyzed by one machine (NN) for detecting pedestrians, by another machine (another NN) for detecting cars, and by another machine (another NN) for estimating the depth of all the pixels in the frames.
[0196] Example use cases and applications are self-driving cars, video surveillance cameras and public safety, smart sensor networks, smart TV and smart advertisement, person re-identification, smart traffic monitoring, drones, etc. In addition to image and video data, automatic analysis and processing is increasingly being performed for other types of data, such as audio, speech, text.
[0197] Compressing (and decompressing) data where the end user comprises machines (e.g., neural networks) is commonly referred to as compression or coding for machines. In the case of video data, it is referred to as video compression or coding for machines (VCM). Compressing for machines may differ from compressing for humans, for example, with respect to the algorithms and technology used in the codec, or the training losses used to train any neural network components of the codec, or the evaluation methodology of codecs.
[0198] It is to be understood that, when considering the case of coding for machines, the term “receiverside” or “decoder-side” refer to the physical or abstract entity or device which comprises one or more machines, and runs these one or more machines on some encoded and eventually decoded video representation which is encoded by another physical or abstract entity or device, the “encoder-side device”.
[0199] Referring now to FIG. 5, illustrated is an example of a pipeline of video coding formachines. A VCM encoder (510) may encode the input video (505) into a bitstream (515). A bitrate (525) may be computed (520) from the bitstream (515), as a measure of the size of the bitstream. A VCM decoder (530) may decode the bitstream (515) that was produced by the VCM encoder (510).
[0200] The output of the VCM decoder (530) may be referred to as “Decoded data for machines” (535). This data may be considered as the decoded or reconstructed video. However, in some implementations of this pipeline, this data may not have same or similar characteristics as the original video which was input to the VCM encoder. For example, this data may not be easily understandable by a human by simply rendering the data onto a screen, if such rendering is possible.
[0201] The output or decoded data for machines (535) of the VCM decoder (530) may then be input to one or more task neural networks (540, 545, 550, 555). In FIG. 5, for the sake of illustrating that there may be any number of task-NNs, there are three example task-NNs, and a non-specified one (Task-NN X, 555). One goal of VCM may be to obtain a low bitrate while guaranteeing that the task-NNs still perform well (580, 585, 590, 595) in terms of the evaluation metric associated to each task (560, 565, 570, 575).
[0202] It is to be understood that, in some cases, the VCM decoder may not be present. In an example, the machines may be run directly on the bitstream. In some other cases, the VCM decoder may comprise only a lossless decoding stage, and the lossless decoded data may be provided as input to the machines. In yet some other cases, the VCM decoder may comprise a lossless decoding stage following by a dequantization operation, and the loss-decoded and dequantized data may be provided as input to the machines.
[0203] When a conventional video encoder, such as a H.266 / WC encoder, is used as a VCM encoder, one or more of the following approaches may be used to adapt the encoding to be suitable to machine analysis tasks:
[0204] - One or more regions of interest (ROIs) may be detected. An ROI detection method may be used. For example, ROI detection may be performed using a task NN, such as an object detection NN. In some cases, ROI boundaries of a group of pictures or an intra period may be spatially overlaid and rectangular areas may be formed to cover the ROI boundaries. The detected ROIs (or rectangular areas, likewise) may be used in one or more of the following ways: the quantization parameter (QP) may be adjusted spatially in a manner that ROIs are encoded using finer quantization step size(s) than other regions. For example, QP may be adjusted CTU-wise; the video may be preprocessed to contain only the ROIs, while the other areas may be replaced by one or more constant values or removed; the video may be preprocessed so that the areas outside the ROIs are blurred or filtered; or, a grid may be formed in a manner that a single grid cell covers a ROI. Grid rows or grid columns that contain no ROIs may be down-sampled as preprocessing to encoding.
[0205] - Quantization parameter of the highest temporal sublayer(s) may be increased (i.e. coarser quantization is used) when compared to practices for human watchable video.
[0206] - The original video may be temporally down-sampled as preprocessing prior to encoding. A frame rate up-sampling method may be used as postprocessing subsequent to decoding, if machine analysis at the original frame rate is desired.
[0207] - A filter may be used to preprocess the input to the conventional encoder. The filter may be a machine learning based filter, such as a convolutional neural network.
[0208] It is to be understood that, in the context of video coding for machines, the terms “machine vision”, “machine vision task”, “machine task”, “machine analysis”, “machine analysis task”, “computer vision”, “computer vision task”, "task network" and “task” may be used interchangeably. Also, it is to be understood that, in the context of video coding for machines, the terms “machine consumption” and “machine analysis” may be used interchangeably.
[0209] Neural network based filtering
[0210] A neural network may be used for filtering or processing input data. Such a neural network may be referred to as a neural network based filter, or simply as a NN filter. A NN filter may comprise one or more neural networks, and / or one or more components that may not be categorized as neural networks (i.e. may be categorized as traditional or legacy components that are not trained based on data using machine learning techniques). The purpose of a NN filter may comprise (but may not be limited to) visual enhancement, colorization, up-sampling, super-resolution, inpainting, temporal extrapolation, generating content, or the like.
[0211] In some video codecs, a neural network may be used as filter in the encoding and decoding loop (also referred to simply as coding loop), and it may be referred to as a neural network loop filter, or a neural network in-loop filter. The NN loop filter may replace all other loop filters of an existing video codec, or may represent an additional loop filter with respect to the already present loop filters in an existing video codec.
[0212] A neural network filter may be used as a post-processing filter for a codec, e.g., may be applied to an output of an image or video decoder in order to remove or reduce coding artifacts.
[0213] In an example, a codec is a modified WC / H.266 compliant codec (e.g., a WC / H.266 compliant codec that has been modified and thus it may not be compliant to the WC / H.266) that comprises one or more NNloop filters. An input to the one or more NN loop filters may comprise at least a reconstructed block or frames (simply referred to as reconstruction) or data derived from a reconstructed block or frame (e.g., the output of a conventional loop filter). The reconstruction may be obtained based on predicting a block or frame (e.g., by means of intra-frame prediction or inter-frame prediction) and performing residual compensation. The one or more NN loop filters may enhance the quality of at least one of their input, so that a rate-distortion loss is decreased. The rate may indicate a bitrate (estimate or real) of the encoded video. The distortion may indicate a pixel fidelity distortion such as the following:
[0214] - Mean-squared error (MSE).
[0215] - Mean absolute error (MAE).
[0216] - Mean Average Precision (mAP) computed based on the output of a task NN (such as an object detection NN) when the input is the output of the post-processing NN.
[0217] - Other machine task-related metric, for tasks such as object tracking, video activity classification, video anomaly detection, etc.
[0218] The enhancement may result into a coding gain, which may be expressed for example in terms of BD-rate or BD-PSNR (peak signal-to-noise ratio).
[0219] A neural network filter may be used as a post-processing filter for a codec, e.g., may be applied to an output of an image or video decoder in order to remove or reduce coding artifacts. In an example, the NN filter may be used as a post-processing filter where the input comprises data that is output by or is derived from an output of a traditional decoder, such as a decoder that is compliant with the WC / H.266 standard. In another example, the NN filter may be used as a post-processing filter where the input comprises data that is output by or is derived from an output of a decoder of an end-to-end learned decoder.
[0220] an examplean examplean examplean examplean examplean examplean examplean examplean examplean example
[0221] Information on the hybrid video coding system
[0222] Recent research on end-to-end (E2E) learned image codecs has demonstrated superior performance compared to the state-of-the-art conventional codecs like WC. However, in the domain of video coding, E2E learned video codecs have yet to reach the performance levels achieved by non-learned codecs such as WC, especially when evaluated on the Random-Access configuration. Moreover, conventional video codecs provide broader application support and are characterised by lower decoding complexity.
[0223] A hybrid video coding framework leverages the high performance of the E2E learned image codec (LI C) while retaining the advantages of the conventional video codec for efficient video coding. I n such a framework, the reconstructed LIC-coded intra frames are used as reference pictures to code inter frames by a conventional codec.
[0224] Thus, in the present disclosure, the term hybrid video codec or hybrid video coding framework refers to a video codec that may use an end-to-end learned image codec (LIC) to encode and decode one or more pictures or frames of a video sequence, or one or more blocks of one or more pictures or frames of a videosequence. A picture coded by an LIC may be an intra-coded frame, may also be referred to as intra frame. A block of a picture, where the block is coded by an LIC, may be an intra-coded block.
[0225] In an example, one or more intra frames of a video sequence are coded by an LIC, and one or more inter frames of a video sequence are coded by a conventional (e.g., non-learned) codec, such as a WC / H.266-compliant codec.
[0226] In an example, an encoder may determine whether a frame is to be coded by an LIC or by a conventional codec, for example based on a rate-distortion cost (e.g., the encoder may choose to code a frame by using a codec that yields the lower rate-distortion cost).
[0227] In an example, a frame, such as an intra frame, may be coded by using both an LIC codec and a conventional codec, as follows. First, the intra frame is coded by using an LIC codec, to obtain a first bitstream and a first reconstructed intra frame. Then, a residual is computed based at least on the intra frame and the first reconstructed intra frame, for example by using a pixel-wise difference operation. The residual is coded by using a conventional codec, to obtain a second bitstream and a reconstructed residual. A second reconstructed intra frame is obtained based on the first reconstructed intra frame and the reconstructed residual, for example by using a pixel-wise summation operation. The first and the second bitstreams represent an encoded intra frame. The second reconstructed intra frame represents a reconstructed or decoded intra frame.
[0228] In an example, a frame, such as an intra frame, may be coded by using both an LIC codec and a conventional codec, as follows. First, the intra frame is coded by using a conventional codec, to obtain a first bitstream and a first reconstructed intra frame. Then, a residual is computed based at least on the intra frame and the first reconstructed intra frame, for example by using a pixel-wise difference operation. The residual is coded by using an LIC codec, to obtain a second bitstream and a reconstructed residual. A second reconstructed intra frame is obtained based on the first reconstructed intra frame and the reconstructed residual, for example by using a pixel-wise summation operation. The first and the second bitstreams represent an encoded intra frame. The second reconstructed intra frame represents a reconstructed or decoded intra frame.
[0229] In an example the first and second bitstreams of the previous two examples may be carried or signaled in a base layer and an enhancement layer, respectively, by using a scalable coding approach or profile of an image or video codec.
[0230] It is to be understood that, even if at least some embodiments described herein refer to a hybrid video coding framework that codes inter frames by using a conventional codec, the at least some embodiments may be realized or applied to also when inter frames are coded by a learned inter-frame codec, such as an end-to-end learned inter-frame codec, or an inter-frame codec that combines conventional coding tools with learned coding tools (e.g., a WC / H.266-compliant decoder that has been extended or modified to include a neural network based loop filter).
[0231] Information on motion compensated temporal filtering (MCTF)
[0232] Motion compensated temporal filtering (MCTF) is a pre-processing approach applied on uncompressed video prior to encoding, to improve compression efficiency. By accounting for motion acrossconsecutive video frames, MCTF offers a sophisticated approach to reduce temporal redundancies while preserving motion details, making it invaluable in both compression standards and high-quality video applications.
[0233] Videos consist of sequences of frames that often contain repetitive or similar information over time. Exploiting this temporal redundancy is essential for effective video compression. However, traditional temporal filtering methods apply operations frame-by-frame or across frames without considering motion, which would introduce artifacts such as blurring or ghosting in dynamic scenes with significant motion. Motion compensation addresses these challenges by aligning frame content along estimated motion trajectories. This ensures that temporal filtering operates on corresponding features of moving objects, preserving sharpness and avoiding artifacts.
[0234] In practice, MCTF aligns blocks in a frame with matching blocks in neighboring frames. These aligned reconstructed blocks can better predict neighboring blocks during inter-frame prediction coding, which can reduce the inter-prediction residual energy and require fewer bits for encoding after quantization.
[0235] Example embodiments of the present disclosure may consider image and video as the data types. However, this is not limiting; the example embodiments may be extended to other types of data, such as audio.
[0236] In the present disclosure, the terms frame, picture and image may be used interchangeably.
[0237] At least some of the embodiments described herein, the term signal, data, frame, picture, and image may be used interchangeably to indicate an input or output.
[0238] In the present disclosure, an end-to-end learned intra codec may be referred to also as LIC. And the conventional video codec may be referred to also as CVC.
[0239] In the present disclosure, neural network layers may be simply referred to as layers, or as a set of layers.
[0240] It is to be noted that, while at least some embodiments herein refer to an intra frame or an input intra frame, the at least some embodiments may be applied to other types of frames than intra frames.
[0241] If not otherwise stated, “an encoder” or “the encoder” may refer to an encoder of the hybrid video coding framework, and “a decoder” or “the decoder” may refer to a decoder of the hybrid video coding framework.
[0242] Embodiments
[0243] In an embodiment, a hybrid video coding framework may comprise one or more intra-frame codecs and one or more inter-frame codecs, where zero, one or more of the one or more intra-frame codecs may comprise an end-to-end learned intra-frame codec (e.g., an LIC codec), and zero, one or more of the one or more intra-frame codecs may comprise a conventional intra-frame codec, and zero, one or more of the one or more interframe codecs may comprise an end-to-end learned inter-frame codec, and zero, one or more of the one or more inter-frame codecs may comprise a conventional inter-frame codec, and where an input to the end-to-end learned intra-frame codec or to the end-to-end learned inter-frame codec, or an output of the end-to-end learned intra-frame codec or of the end-to-end learned inter-frame codec, may be filtered by a filter, such as a smoothing filter, or an MCTF filter, and the like, or may be coded (e.g., encoded and decoded) by a codec, such as a conventional codec, by using a high quality level.
[0244] In an embodiment, in the hybrid video coding framework, an input intra frame to the LIC encoder may be filtered with a filter. In an example, the filter may be a pre-processing filter. In another example, the filter smooths the input intra frame. In yet another example, the filter removes high-frequency signals from the input intra frame. In yet another example, the filter makes the input intra frame more similar to one or more other frames in a video sequence that comprises the input intra frame. In one additional embodiment, the filter may take as input one or more other frames, such as one or more previous frames with respect to the input intra frame, in output / display order, or in coding order, and / or one or more next frames with respect to the input intra frame, in output / display order, or in coding order. In an additional embodiment, the filter may take as input data derived from one or more other frames, such as features extracted based on one or more previous frames with respect to the input intra frame, in output / display order, or in coding order, and / or features based on one or more next frames with respect to the input intra frame, in output / display order, or in coding order.
[0245] FIG. 6 illustrates an example of a hybrid video coding framework 600, in accordance with an embodiment. In this embodiment of the hybrid video coding framework 600, an input intra frame or an original frame 602 to an LIC encoder 604 is filtered with a Motion-Compensated Temporal Filtering (MCTF) 606 to generate an MCTF-filtered frame 608. The MCTF-filtered frame 608 is provided as an input to the LIC encoder 604, which generates a bitstream 610. The bitstream 610 is provided to an LIC decoder 612. An output frame 614, which is an output of the LIC decoder 612, would serve as both a final output intra frame 616 and a reference frame 618 for subsequent inter-frame coding. Final output intra frame refers to a reconstructed or decoded intra frame that is output by a decoder of the hybrid video coding framework and may be used for display.
[0246] In an example, the original frame may be an intra frame in the input video sequence. The MCTF may be included in a conventional video codec that is part of the hybrid video coding framework, such as in the VTM encoder software (VTM is the reference software for the WC / H.266 standard, and it includes both an encoder and a decoder). Since MCTF may be QP-dependent, the choice of QP may correspond to the coding quality of the video sequence, or to other quality levels. The LIC model may be trained using a dataset that takes into account the MCTF process, such as in one of the following ways: (i) both the input data and the ground truth are filtered by MCTF; (ii) the input data is filtered by MCTF; (iii) the ground truth is filtered by MCTF. In another example, the LIC model may initially be trained on an unfiltered dataset, and then finetuned with MCTF-filtered images.
[0247] In an embodiment, an indicator may be input to the LIC encoder to indicate whether the input frame to the LIC encoder is MCTF-filtered. Additionally or alternatively, the indicator may be signaled to the decoder and used as an auxiliary input to the LIC decoder. The presence or the value of the indicator may be determined by the video encoder based on rate-distortion optimization.
[0248] In an additional embodiment, when an encoder indicates to a decoder whether it has used a preprocessing filter, such as MCTF, on an input intra frame, the decoder may determine, based on the indication, whether to apply a filter (that may be referred to as LIC loop filter) on an output of the LIC decoder or not. Alternatively, an encoder may indicate to a decoder whether an LIC loop filter is to be applied to an output of the LIC decoder. In an example, a decoder comprises an LIC loop filter that may be applied to a reconstructed intra frame that is output by an LIC decoder; when the decoder receives an indication that an intra frame was not filteredby MCTF, the decoder may apply the LIC loop filter to a corresponding reconstructed intra frame that is output by the LIC decoder.
[0249] In an embodiment, in the hybrid video coding framework, the input to the LIC encoder includes both the original unfiltered intra frame and the corresponding MCTF-filtered intra frame.
[0250] In an embodiment, the output of the LIC decoder consists of the final output intra frame (e.g., the reconstructed intra frame that is output for display) and the reference frame which is used for subsequent interframe coding.
[0251] FIG. 7 illustrates an example of a hybrid video coding framework 700, in accordance with another embodiment. In this embodiment of the hybrid video coding framework 700, an input to an LIC encoder 702 includes both an original unfiltered intra frame or original frame 704 and a corresponding MCTF-filtered intra frame or MCTF-filtered frame 706. The LIC encoder 702 generates a bitstream 708, based at least on the original unfiltered intra frame 704 and the corresponding MCTF-filtered frame 706. The MCTF-filtered frame may include an MCTF filtered intra frame. The bitstream 708 is provided to an LIC decoder 710. The output of the LIC decoder includes a final output intra frame 712 (e.g., the reconstructed intra frame that is output for display) and a reference frame 714 which is used for subsequent inter-frame coding.
[0252] In an example, when training such an LIC model, the input and output of the LIC model may consist of two frames, respectively. The input to the LIC model may include the original unfiltered frame and the MCTF-filtered frame, while the output of the LIC model may include the final output frame and the reference frame. Two training losses may be computed, by using respective two ground truths. A first ground truth for computing a first loss based on the final output frame may comprise the original unfiltered frame, whereas a second ground truth for computing a second loss based on the reference frame may comprise the MCTF-filtered frame. In another example, the ground truth for the reference frame may be a frame coded by a conventional video codec (CVC) at a higher quality level than a quality level used to code the original frame or the MCTF-filtered frame or the inter frames (e.g., a frame coded by VTM with a QP equal to 22, when the sequence QP for coding the video sequence by using the hybrid video coding framework is equal to 37). During training, the distortion loss may comprise a combination, such as a weighted sum or a weighted average, of the first loss and the second loss. In an example, when the distortion loss comprises a weighted sum or a weighted average of the first loss and the second loss, where a first weight is used to multiply the first loss and a second weight is used to multiply the second loss, the second weight (e.g., the weight that multiplies the loss computed based on the reference frame) may be higher than the first weight (e.g., the weight that multiplies the loss computed based on the final output frame). Here, the distortion loss and / or the first loss and / or the second loss may utilize metrics such as MSE, MS-SSIM, or other suitable loss functions.
[0253] In an example, a training of the LIC codec may include encouraging the LIC codec to derive the output frame primarily from the original unfiltered input frame and the reference frame primarily from the MCTF-filtered input frame.
[0254] In an example, two LIC codecs may be used as part of the encoder of the hybrid video coding framework, where a first LIC codec gets as input an original (uncompressed and unfiltered) intra frame and outputsa final output frame, and a second LIC codec gets as input an MCTF-filtered intra frame and outputs a reference frame.
[0255] In another example, a motion magnitude map of the input frame may be provided as an auxiliary data to the LIC model. The motion magnitude map may represent the magnitude of the motion in different regions of the input frame. During training, the loss for the reference frame can be weighted based on this motion magnitude map. In regions with minimal motion, the pixels may be reconstructed with high quality to preserve fine details. Conversely, in areas with significant motion, higher weights may be assigned to the loss, prioritizing smoother reconstruction to be served as better reference.
[0256] In yet another example, motion estimation information may be provided as an auxiliary data to the LIC model. The motion estimation information may, for example, be derived with the MCTF process. The motion estimation information may, for example, comprise loss or similarity values, such as sum of absolute differences, between a current block and a motion-compensated reference block. During training, the training loss for the reference frame can be weighted based on this motion estimation information. In regions with small loss or high similarity between a current block and a motion-compensated reference block, the pixels may be reconstructed with high quality to preserve fine details. Conversely, in areas with significant loss or low similarity, higher weights may be assigned to the training loss, prioritizing smoother reconstruction to be served as better reference.
[0257] FIG. 8 illustrates an example of a hybrid video coding framework 800, in accordance with still another embodiment. In this embodiment of the hybrid video coding framework 800, an input to an LIC encoder 802 is an original, unfiltered intra frame 804. The LIC encoder 802 generates a bitstream 806, based at least on the original frame 804. The bitstream 806 is provided to an LIC decoder 808. The output of the LIC decoder, an output frame 810, would serve as a final output intra frame 812. Additionally, a loop filter 814 gets as input at least the output frame 810, where an output of the loop filter 814 is used as a reference frame 816 for subsequent inter-frame coding.
[0258] In an example, the LIC model may be trained with a dataset of unfiltered images. During the training of the loop filter, the input to the loop filter may be the LIC-decoded data, while the ground truth may be the frames filtered by MCTF. In another example, the ground truth may consist of CVC-decoded frames at a higher quality level. Here, the CVC may refer to WC.
[0259] In an example, the LIC model and the loop filter may be separately trained from scratch. In another example, the LIC model and the loop filter may be trained jointly. Alternatively, the LIC model may be trained first, followed by training the loop filter with the LIC model kept frozen. In yet another example, the LIC model may be trained initially and then finetuned during the training of the loop filter.
[0260] In one additional embodiment, the loop filter may take, as input, one or more other reconstructed frames other than the output frame from the LIC decoder.
[0261] In an embodiment, an output of the LIC decoder is used as a final output frame and as input to a combination operation, where the combination operation combines the final output frame and a reconstructed residual to obtain a reference frame. In an example, the reconstructed residual may be output by a conventional decoder that is part of a conventional codec, where an encoder of the conventional decoder gets as input a residual,and where the residual is computed based on an original frame and the final output frame, such as by using a pixel-wise difference operation. The combination operation may be a pixel-wise summation operation.
[0262] In an embodiment, the reference frame and the final output frame may be utilized differently when coding or predicting different inter-frames. Inter frames which are closer to an LIC-coded intra-frame, e.g., in output / display order or in coding order, may use the final output frame as a reference for inter-frame coding or prediction, while inter frames which are farther from the LIC-coded intra-frame may use the reference intra frame as a reference for inter-frame coding or prediction. In an example, the decision is based on the temporal layer of the inter frame, e.g., inter frames that are associated with a temporal layer ID higher than a temporal layer threshold value use the LIC-decoded reference frame as reference frame, while other inter frames use the LIC decoded final output frame as reference frame. The temporal layer threshold value may be predetermined and available at the encoder and decoder (e.g., specified in a standard specification), or determined by the encoder and signaled to the decoder.
[0263] In another embodiment, the LIC generated reference frame may be marked as non-output and added to the decoded picture buffer to be used as a reference frame of other inter frames.
[0264] In another embodiment, the reference frame and the final output frame may be combined linearly to generate a new reference frame. For instance, weighted bi-prediction in WC could be used for combining the reference frame and the final output frame and obtaining a new reference frame. The weights assigned to the reference frame and the final output frame may be determined and signalled by the encoder. When the new reference frame is intended solely for reference and not for output, it can be marked accordingly to ensure it is not output.
[0265] Some embodiments have been described in relation to one reference frame reconstructed by the LIC encoder or the LIC decoder per an output time instance. It is to be understood that embodiments may be similarly realized when more than one reference frame is reconstructed. For example, different reference frames may have a different strength of the motion magnitude map. Some embodiments have been described in relation to one final output frame reconstructed by the LIC encoder or the LIC decoder per an output time instance. It is to be understood that embodiments may be similarly realized when more than one final output frame is reconstructed. For example, reconstruction of one final output frame may be trained for machine task(s) and reconstruction of another final output frame may be trained for human viewing.
[0266] In an example, an LIC decoder may output one or more reference frames to be used as reference for inter-frame coding or prediction of respective one or more inter frames.
[0267] In an example, an LIC decoder may output two or more reference frames, where one or more of the two or more reference frames may be used as reference for inter-frame coding or prediction of one or more first inter frames and other one or more of the two or more reference frames may be used as reference for inter-frame coding or prediction of other one or more second inter frames. In another example, an LIC decoder outputs two reference frames, where a first reference frame comprised in the two reference frames is used as reference for inter-frame coding of a first set of inter frames and a second reference frame comprised in the two reference frames is used for inter-frame coding of a second set of inter frames.
[0268] In an embodiment, an LIC encoder may take as input one or more other frames, or data derived from one or more other frames (e.g., features), in addition to an input intra frame to be coded.
[0269] In an embodiment, an LIC decoder may take as input one or more reconstructed frames, or data derived from one or more reconstructed frames (e.g., features), in addition to a latent tensor.
[0270] In an embodiment, a pre-processing filter, such as an MCTF, may be applied to a portion of an input frame to an LIC encoder. In an example, an MCTF is applied to the luma component of an input intra frame to an LIC encoder, whereas the chroma component of the input intra frame to the LIC encoder is not filtered.
[0271] In an embodiment, two or more pre-processing filters, ora pre-processing filter with two or more sets of parameters (such as an MCTF with two or more sets of parameters like QPs or which input frames are used), are applied to respective two or more portions of an input frame. In an example, an MCTF with a first set of parameters is applied to a luma component of an input intra frame, and the MCTF with a second set of parameters is applied to a chroma component of the input intra frame.
[0272] In an embodiment, a filter, such as a loop filter of a video codec, is used to filter at least a portion of an input frame to an LIC encoder. In an example, a loop filter of a WC / H.266-compliant codec is used to filter an input intra frame to an LIC encoder. In another example, a neural network based loop filter that is part of a video codec used to code at least inter frames may be used to filter an input intra frame to an LIC encoder, where the inter frames and the input intra frame may be comprised in a video sequence.
[0273] In an embodiment, where a residual is computed based on an input frame to an LIC encoder and an output frame of an LIC decoder, the residual may be filtered by a filter, such as one of the filtered described in other embodiments or examples (e.g., an MCTF), to obtain a filtered residual, where the filtered residual is input to another encoder, such as a conventional encoder.
[0274] In an embodiment, a loop filter, such as a neural network based loop filter, is trained by using, as ground truth, images that are filtered by a filter, such as by an MCTF.
[0275] In an embodiment, a loop filter, such as a neural network based loop filter, is trained by using, as ground truth, images that are encoded and decoded by using a higher first quality level than a second quality level that is used for obtaining an input picture to the loop filter. In an example, when an input picture to a loop filter is coded with a WC-compliant codec by using a QP equal to 37, a ground truth is a corresponding picture (e.g., a picture that was derived from the same picture that the input picture was derived from) that is coded with a WC-compliant codec by using a QP equal to 22.
[0276] In an embodiment, a process that outputs data that may be used as a reference frame, such as a loop filter, or a reference frame generation, or an end-to-end learned codec, or a motion compensation, may be optimized or trained by using, as ground truth or target, data that has been filtered by a filter, such as a smoothening filter, an MCTF, and the like.
[0277] In an embodiment, a process that outputs data that may be used as a reference frame, such as a loop filter, or a reference frame generation, or an end-to-end learned codec, or a motion compensation, may be optimized or trained by using, as ground truth or target, data that has been coded by using a first quality level that is higher than a second quality level that is used to code an input frame to the process.
[0278] FIG. 9 is a diagram illustrating an example apparatus 900, which may be implemented in hardware, configured to implement the examples described herein. The apparatus 900 comprises at least one processor 902 (e.g., an FPGA and / or CPU), at least one memory 904 including computer program code 905, the computer program code 905 having instructions to carry out the methods described herein, wherein the at least one memory 904 and the computer program code 905 are configured to, with the at least one processor 902, cause the apparatus 900 to implement circuitry, a process, component, module, or function (implemented with control module 906) to implement the examples described herein, including implementing hybrid video coding. Optionally included encoder 908 of the control module 906 implements encoding based on the examples described herein, and optionally included decoder 910 implements decoding based on the examples described herein. The at least one memory 904 may be a non-transitory memory, a transitory memory, a volatile memory (e.g. RAM), or a non-volatile memory (e.g., ROM).
[0279] The apparatus 900 includes a display and / or I / O interface 912, which includes user interface (Ul) circuitry and elements, that may be used to display features or a status of the methods described herein (e.g., as one of the methods is being performed or at a subsequent time), or to receive input from a user such as with using a keypad, camera, touchscreen, touch area, microphone, biometric recognition, one or more sensors, etc. The apparatus 900 includes one or more communication e.g. network (N / W) interfaces (l / F(s)) 914. The communication l / F(s) 914 may be wired and / or wireless and communicate over the Internet / other network(s) via any communication technique including via one or more links 916. The communication l / F(s) 914 may comprise one or more transmitters or one or more receivers.
[0280] The transceiver 918 comprises one or more transmitters 920 and one or more receivers 922. The transceiver 918 and / or communication l / F(s) 914 may comprise standard well-known components such as an amplifier, filter, frequency-converter, (de)modulator, and encoder / decoder circuitries and one or more antennas, such as antennas 924 used for communication over wireless link 926.
[0281] The control module 906 of the apparatus 900 comprises one of or both parts 906-1 and / or 906-2, which may be implemented in a number of ways. The control module 906 may be implemented in hardware as control module 906-1, such as being implemented as part of the at least one processor 902. The control module 906-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the control module 906 may be implemented as control module 906-2, which is implemented as computer program code (having corresponding instructions) 905 and is executed by the at least one processor 902. For instance, the at least one memory 904 store instructions that, when executed by the at least one processor 902, cause the apparatus 900 to perform one or more of the operations as described herein. Furthermore, the at least one processor 902, the at least one memory 904, and example algorithms (e.g., as flowcharts and / or signaling diagrams), encoded as instructions, programs, or code, are means for causing performance of the operations described herein.
[0282] The apparatus 900 to implement the functionality of control module 906 may correspond to any of the apparatuses depicted herein. Alternatively, apparatus 900 and its elements may not correspond to any of theother apparatuses depicted herein, as apparatus 900 may be part of a self-organizing / optimizing network (SON) node or other node, such as a node in a cloud.
[0283] The apparatus 900 may also be distributed throughout the network including within and between apparatus 900 and any network element (such as a base station and / or terminal device and / or user equipment).
[0284] Interface 928 enables data communication and signaling between the various items of apparatus 900, as shown in FIG. 9. For example, the interface 928 may be one or more buses such as address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. Computer program code (e.g. instructions) 905, including control module 906 may comprise object-oriented software configured to pass data or messages between objects within computer program code 905. The apparatus 900 need not comprise each of the features mentioned, or may comprise other features as well. The various components of apparatus 900 may at least partially reside in a housing 930, or a subset of the various components of apparatus 900 may at least partially be located in different housings, which different housings may include housing 930.
[0285] FIG. 10 is a diagram illustrating representation of non-volatile memory media 1000a (e.g. computer / compact disc (CD) or digital versatile disc (DVD)) and 1000b (e.g. universal serial bus (USB) memory stick) and 1000c (e.g. cloud storage for downloading instructions and / or parameters 1002 or receiving emailed instructions and / or parameters 1002) storing instructions and / or parameters 1002 which when executed by a processor allows the processor to perform one or more of the operations of the methods described herein. Instructions and / or parameters 1002 may represent or correspond to a non-transitory computer readable medium.
[0286] FIG. 11 is a flowchart illustrating an example method 1100 as described herein. At 1102, using an end-to-end learned encoder for encoding: a picture or frame of a video sequence as an intra-coded frame; or a block of the picture or frame of the video sequence as an intra-coded block. At 1104, the method 1100 includes using a conventional encoder for encoding: another picture or frame of the video sequence as inter-coded frame; or a block of the another picture or frame of the video sequence as an inter-coded block.
[0287] In an embodiment, the frame comprises an input intra frame.
[0288] In an embodiment, the method 1100 may further include performing filtering operation on the input intra frame or a portion of the input intra frame by using a filter. Some examples of the filter include, but are not limited to, a pre-processing filter and a motion compensated temporal filter.
[0289] The method 1100 may be performed with an encoding apparatus, such as the apparatus 900, or any encoding apparatus described herein.
[0290] FIG. 12 is a flowchart illustrating an example method 1200 as described herein. At 1202, the method 1200 includes receiving a bitstream comprising an intra-coded frame and an inter-coded frame. The intra-coded frame is encoded by an end-to-end learned encoder and the inter-coded frame is encoded by a conventional encoder. At 1204, the method 1200 includes decoding the intra-coded frame by using an end-to-end learned decoder for obtaining a decoded intra frame. At 1206, the method 1200 includes decoding the inter-coded frame, based at least on the decoded intra frame, by using a conventional decoder.
[0291] In an embodiment, the intra-coded frame comprises an encoded input intra frame and the intercoded frame comprises an encoded input inter frame.
[0292] In an embodiment, a filtering operation is performed on the input intra frame or a portion of the input intra frame by a filter. Some examples of the filter include, but are not limited to, a pre-processing filter and a motion compensated temporal filter.
[0293] The method 1200 may be performed with a decoding apparatus, such as the apparatus 900, or any decoding apparatus described herein.
[0294] FIG. 13 is a flowchart illustrating another example method as described herein. At 1302, the method 1300 includes generating reference data by using a neural network, where the neural network was trained by using processed images as ground truth.
[0295] The method 1300 may be performed with the apparatus 900, or any decoding apparatus described herein.
[0296] The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e. tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM).
[0297] It should be understood that the foregoing description is only illustrative. Various alternatives and modifications can be devised by those skilled in the art. For example, features recited in the various dependent claims could be combined with each other in any suitable combination(s). In addition, features from different embodiments described above could be selectively combined into a new embodiment. Accordingly, the description is intended to embrace all such alternatives, modification and variances which fall within the scope of the appended claims.
Claims
CLAIMSWhat is claimed is:
1. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: using an end-to-end learned encoder for encoding:a picture or frame of a video sequence as an intra-coded frame; ora block of the picture or frame of the video sequence as an intra-coded block; and using a conventional encoder for encoding:another picture or frame of the video sequence as inter-coded frame; or a block of the another picture or frame of the video sequence as an inter-coded block.
2. The apparatus of claim 1 , wherein the frame comprises an input intra frame.
3. The apparatus of claim 2, wherein the apparatus is further caused to perform: performing filtering operation on the input intra frame or a portion of the input intra frame by using a filter.
4. The apparatus of claim 3, wherein the filter comprises a pre-processing filter.
5. The apparatus of claim 3, wherein the filtering operation comprises one or more of:smoothing the input intra frame;removing high-frequency signals from the input intra frame;making the intra frame similar or substantially similar to one or more first other frames of the video sequence that comprises the input intra frame;receiving the one or more first other frames as an input; orreceiving data derived from the one or more first other frames.
6. The apparatus of claim 5, wherein the one or more first other frames comprise:one or more previous frames with respect to the input intra frame in output order, display order, or in coding order; and / orone or more next frames with respect to the input intra frame, in output order, display order, or in coding order.
7. The apparatus of claim 2, wherein the apparatus is further caused to perform: filtering, with a motion- compensated temporal filtering, the input intra frame.
8. The apparatus of claim 7, wherein the apparatus is further caused to perform: training the end-to- end learned encoder by using a dataset that is based on a motion-compensated temporal filtering process.
9. The apparatus of claim 8, wherein the dataset comprises:an input data and a ground truth filtered by the motion-compensated temporal filtering; the input data filtered the motion-compensated temporal filtering; orthe ground truth filtered by the motion-compensated temporal filtering.
10. The apparatus of any of the claims 7 to 9, wherein the apparatus is further caused to perform:generating a first indicator for indicating whether the input intra frame is filtered by the motion- compensated temporal filtering; andsignaling the first indicator to a decoder comprising an end-to-end learned decoder, wherein the first indicator is intended to be used as an auxiliary input to the end-to-end learned decoder.
11. The apparatus of any of the claims 2 to 10, wherein the apparatus is further caused to perform:providing the input intra frame and the input intra frame filtered by the motion-compensated temporal filtering as input to the end-to-end learned encoder.
12. The apparatus of any of the claims 2 to 6, wherein the apparatus is further caused to perform: generating a second indicator for indicating to a decoder whether a pre-processing filter has been used to filter the input intra frame; andsignaling the second indicator to the decoder, wherein the second indicator indicates to the decoder whether a loop filter is to be applied to an output of an end-to-end learned decoder comprised in the decoder.
13. The apparatus of claim 12, wherein the apparatus is further caused to perform: training the loop filter by using output of the end-to-end learned decoder as an input and using the input intra frame filtered by a motion-compensated temporal filtering as a ground truth.
14. The apparatus of claim 13, wherein:the end-to-end learned encoder and the loop filter are trained separately;the end-to-end learned encoder and the loop filter are trained jointly;the end-to-end learned encoder is trained first followed by training of the loop filter with the end-to- end learned encoder is kept frozen; orthe end-to-end learned encoder is trained first followed by training of the loop filter, and wherein the end-to-end learned encoder is finetuned during the training of the end-to-end learned loop.
15. The apparatus of any of claim 13 or 14, wherein the apparatus is further caused to perform: encoding subsequent inter-frame by using an output of the loop filter as a reference frame, wherein the subsequent inter-frame comprises the another picture or frame.
16. The apparatus of any of claims 12 to 14, wherein, a first inter frame which is closer to the intra-coded frame is coded by using a final output frame of the end-to-end learned decoder as a reference frame, and wherein a second inter-frame which is farther from the intra-coded frame is coded by using a reference frame generated by the end-to-end learned decoder as the reference frame.
17. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving a bitstream comprising an intra-coded frame and an inter-coded frame, wherein the intra- coded frame is encoded by an end-to-end learned encoder, and wherein the inter-coded frame is encoded by a conventional encoder;decoding the intra-coded frame by using an end-to-end learned decoder for obtaining a decoded intra frame; anddecoding the inter-coded frame, based at least on the decoded intra frame, by using a conventional decoder.
18. The apparatus of claim 17, wherein the intra-coded frame comprises an encoded input intra frame and the inter-coded frame comprises an encoded input inter frame.
19. The apparatus of claim 18, wherein a filtering operation is performed on the input intra frame or a portion of the input intra frame by a filter.
20. The apparatus of claim 19, wherein the filter comprises a pre-processing filter.
21. The apparatus of claim 19, wherein the filtering operation comprises one or more of:smoothing the input intra frame;removing high-frequency signals from the input intra frame;making the input intra frame similar or substantially similar to one or more first other frames of the video sequence that comprises the input intra frame;receiving the one or more first other frames as an input; orreceiving data derived from the one or more first other frames.
22. The apparatus of claim 21 , wherein the one or more first other frames comprises:one or more previous frames with respect to the input intra frame in output order, display order, or in coding order; and / orone or more next frames with respect to the input intra frame, in output order, display order, or in coding order.
23. The apparatus of claim 18, wherein the input intra frame is filtered with a motion-compensated temporal filtering.
24. The apparatus of claim 23, wherein the apparatus is further caused to perform: training the end-to- end learned decoder by using a dataset that is based on a motion-compensated temporal filtering process.
25. The apparatus of claim 24, wherein the dataset comprises:an input data and a ground truth filtered by the motion-compensated temporal filtering;the input data filtered the motion-compensated temporal filtering; orthe ground truth filtered by the motion-compensated temporal filtering.
26. The apparatus of any of the claims 23 to 25, wherein the apparatus is further caused to perform:receiving a first indicator for indicating whether the input intra frame is filtered by the motion- compensated temporal filtering; andusing the first indicator as an auxiliary input to the end-to-end learned decoder.
27. The apparatus of any of the claims 18 to 22, wherein the apparatus is further caused to perform:receiving a second indicator for indicating whether a pre-processing filter has been used to filter the input intra frame; andapplying an loop filter to an output of the end-to-end learned decoder based on the second indicator.
28. The apparatus of claim 27, wherein the apparatus is further caused to perform: training the loop filter by using an output of the end-to-end learned decoder as an input and using the intra frame filtered by a motion-compensated temporal filtering as a ground truth.
29. The apparatus of claim 28, wherein:the end-to-end learned decoder and the loop filter are trained separately;the end-to-end learned decoder and the loop filter are trained jointly;the end-to-end learned decoder is trained first followed by training of the loop filter with the end-to- end learned decoder is kept frozen; orthe end-to-end learned decoder is trained first followed by training of the loop filter, and wherein the end-to-end learned decoder is finetuned during the training of the end-to-end learned loop.
30. The apparatus of any of claim 28 or 29, wherein the apparatus is further caused to perform: decoding subsequent inter-frame by using the output of the loop filter as a reference frame, wherein the subsequent inter-frame comprises the inter-coded frame.
31. The apparatus of any of claims 27 to 29, wherein, a first inter frame which is closer to the intra-coded frame is decoded by using the decoded intra frame as a reference frame, and wherein a second inter-frame which is farther from the intra-coded frame is decoded by using a reference frame generated by the end-to-end learned decoder as the reference frame.
32. The apparatus any of the claims 17 to 31, wherein the apparatus is further caused to perform:decoding a subsequent inter-frame by using the decoded intra frame of the end-to-end learned decoder as a reference frame.
33. The apparatus any of the claims 17 to 31, wherein the apparatus is further caused to perform:generating a reference frame.
34. The apparatus of claim 33, wherein the apparatus is further caused to perform:combining the decoded intra frame and the reference frame for generating a new reference frame.
35. A method comprising:using an end-to-end learned encoder for encoding:a picture or frame of a video sequence as an intra-coded frame; ora block of the picture or frame of the video sequence as an intra-coded block; and using a conventional encoder for encoding:another picture or frame of the video sequence as inter-coded frame; or a block of the another picture or frame of the video sequence as an inter-coded block.
36. The method of claim 35, wherein the frame comprises an input intra frame.
37. The method of claim 36 further comprising: performing filtering operation on the input intra frame or a portion of the input intra frame by using a filter.
38. The method of claim 37, wherein the filter comprises a pre-processing filter.
39. The method of claim 37, wherein the filtering operation comprises one or more of:smoothing the input intra frame;removing high-frequency signals from the input intra frame;making the intra frame similar or substantially similar to one or more first other frames of the video sequence that comprises the input intra frame;receiving the one or more first other frames as an input; orreceiving data derived from the one or more first other frames.
40. The method of claim 39, wherein the one or more first other frames comprise:one or more previous frames with respect to the input intra frame in output order, display order, or in coding order; and / orone or more next frames with respect to the input intra frame, in output order, display order, or in coding order.
41. The method of claim 36 further comprising: filtering, with a motion-compensated temporal filtering, the input intra frame.
42. The method of claim 41 further comprising: training the end-to-end learned encoder by using a dataset that is based on a motion-compensated temporal filtering process.
43. The method of claim 42, wherein the dataset comprises:an input data and a ground truth filtered by the motion-compensated temporal filtering;the input data filtered the motion-compensated temporal filtering; orthe ground truth filtered by the motion-compensated temporal filtering.
44. The method of any of the claims 41 to 43 further comprising:generating a first indicator for indicating whether the input intra frame is filtered by the motion- compensated temporal filtering; andsignaling the first indicator to a decoder comprising an end-to-end learned decoder, wherein the first indicator is intended to be used as an auxiliary input to the end-to-end learned decoder.
45. The method of any of the claims 36 to 44 further comprising: providing the input intra frame and the input intra frame filtered by the motion-compensated temporal filtering as input to the end-to-end learned encoder.
46. The method of any of the claims 36 to 40 further comprising:generating a second indicator for indicating to a decoder whether a pre-processing filter has been used to filter the input intra frame; andsignaling the second indicator to the decoder, wherein the second indicator indicates to the decoder whether an loop filter is to be applied to an output of an end-to-end learned decoder comprised in the decoder.
47. The method of claim 46 further comprising: training the loop filter by using output of the end-to-end learned decoder as an input and using the input intra frame filtered by a motion-compensated temporal filtering as a ground truth.
48. The method of claim 47, wherein:the end-to-end learned encoder and the loop filter are trained separately;the end-to-end learned encoder and the loop filter are trained jointly;the end-to-end learned encoder is trained first followed by training of the loop filter with the end-to- end learned encoder is kept frozen; orthe end-to-end learned encoder is trained first followed by training of the loop filter, and wherein the end-to-end learned encoder is finetuned during the training of the end-to-end learned loop.
49. The method of any of claim 47 or 48 further comprising: encoding subsequent inter-frame by using an output of the loop filter as a reference frame, wherein the subsequent inter-frame comprises the another picture or frame.
50. The method of any of claims 46 to 48, wherein, a first inter frame which is closer to the intra-coded frame is coded by using a final output frame of the end-to-end learned decoder as a reference frame, and wherein a second inter-frame which is farther from the intra-coded frame is coded by using a reference frame generated by the end-to-end learned decoder as the reference frame.
51. A method comprising:receiving a bitstream comprising an intra-coded frame and an inter-coded frame, wherein the intra- coded frame is encoded by an end-to-end learned encoder, and wherein the inter-coded frame is encoded by a conventional encoder;decoding the intra-coded frame by using an end-to-end learned decoder for obtaining a decoded intra frame; anddecoding the inter-coded frame, based at least on the decoded intra frame, by using a conventional decoder.
52. The method of claim 51 , wherein the intra-coded frame comprises an encoded input intra frame and the inter-coded frame comprises an encoded input inter frame.
53. The method of claim 52, wherein a filtering operation is performed on the input intra frame or a portion of the input intra frame by a filter.
54. The method of claim 53, wherein the filter comprises a pre-processing filter.
55. The method of claim 53, wherein the filtering operation comprises one or more of:smoothing the input intra frame;removing high-frequency signals from the input intra frame;making the input intra frame similar or substantially similar to one or more first other frames of the video sequence that comprises the input intra frame;receiving the one or more first other frames as an input; orreceiving data derived from the one or more first other frames.
56. The method of claim 55, wherein the one or more first other frames comprises:one or more previous frames with respect to the input intra frame in output order, display order, or in coding order; and / orone or more next frames with respect to the input intra frame, in output order, display order, or in coding order.
57. The method of claim 52, wherein the input intra frame is filtered with a motion-compensated temporal filtering.
58. The method of claim 57 further comprising: training the end-to-end learned decoder by using a dataset that is based on a motion-compensated temporal filtering process.
59. The method of claim 58, wherein the dataset comprises:an input data and a ground truth filtered by the motion-compensated temporal filtering;the input data filtered the motion-compensated temporal filtering; orthe ground truth filtered by the motion-compensated temporal filtering.
60. The method of any of the claims 57 to 59 further comprising:receiving a first indicator for indicating whether the input intra frame is filtered by the motion- compensated temporal filtering; andusing the first indicator as an auxiliary input to the end-to-end learned decoder.
61. The method of any of the claims 52 to 56 further comprising:receiving a second indicator for indicating whether a pre-processing filter has been used to filter the input intra frame; andapplying an loop filter to an output of the end-to-end learned decoder based on the second indicator.
62. The method of claim 61 further comprising: training the loop filter by using an output of the end-to- end learned decoder as an input and using the intra frame filtered by a motion-compensated temporal filtering as a ground truth.
63. The method of claim 62, wherein:the end-to-end learned decoder and the loop filter are trained separately;the end-to-end learned decoder and the loop filter are trained jointly;the end-to-end learned decoder is trained first followed by training of the loop filter with the end-to- end learned decoder is kept frozen; orthe end-to-end learned decoder is trained first followed by training of the loop filter, and wherein the end-to-end learned decoder is finetuned during the training of the end-to-end learned loop.
64. The method of any of claim 62 or 63 further comprising: decoding subsequent inter-frame by using the output of the loop filter as a reference frame, wherein the subsequent inter-frame comprises the inter-coded frame.
65. The method of any of claims 61 to 63, wherein, a first inter frame which is closer to the intra-coded frame is decoded by using the decoded intra frame as a reference frame, and wherein a second inter-frame which is farther from the intra-coded frame is decoded by using a reference frame generated by the end-to-end learned decoder as the reference frame.
66. The method any of the claims 51 to 65 further comprising: decoding a subsequent inter-frame by using the decoded intra frame of the end-to-end learned decoder as a reference frame.
67. The method any of the claims 51 to 65 further comprising: generating a reference frame.
68. The method of claim 67 further comprising:combining the decoded intra frame and the reference frame for generating a new reference frame.
69. An apparatus comprising means for performing methods as claimed in any of the claims 35 to 50.
70. An apparatus comprising means for performing methods as claimed in any of the claims 51 to 68.
71. A computer readable medium comprising program instructions which when executed by an apparatus cause the apparatus to perform methods as claimed in any of the claims 35 to 50.
72. The computer readable medium of claim 71, wherein the computer readable medium comprises a non-transitory computer readable medium.
73. A computer readable medium comprising program instructions which when executed by an apparatus cause the apparatus to perform methods as claimed in any of the claims 51 to 68.
74. The computer readable medium of claim 73, wherein the computer readable medium comprises a non-transitory computer readable medium.
75. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: generating reference data by using a neural network, where the neural network was trained by using processed images as ground truth.
76. The apparatus of claim 75, wherein the neural network is comprised in one of a loop filter, a reference frame generator, an intra-frame codec, ora motion compensation.
77. The apparatus of any of the claims 75 or 76, wherein the processed images used as ground truth are encoded and decoded by using a first quality level that is higher than a second quality level that is used for obtaining an input frame to the neural network.
78. The apparatus of any of the claims 75 or 76, wherein the processed images have been filtered by a motion-compensated temporal filter.
79. A method comprising:generating reference data by using a neural network, where the neural network was trained by using processed images as ground truth.
80. The method of claim 79, wherein the neural network is comprised in one of a loop filter, a reference frame generator, an intra-frame codec, ora motion compensation.
81. The method of any of the claims 79 or 80, wherein the processed images used as ground truth are encoded and decoded by using a first quality level that is higher than a second quality level that is used for obtaining an input frame to the neural network.
82. The method of any of the claims 79 or 80, wherein the processed images have been filtered by a motion-compensated temporal filter.