Signaling information for temporal extrapolation

The implementation of neural network-based signaling information for temporal extrapolation in video codecs addresses the lack of efficient post-processing in existing technologies, enhancing visual quality and reducing latency in applications like cloud gaming and autonomous driving.

WO2025149229A1PCT designated stage expired Publication Date: 2025-07-17NOKIA TECHNOLOGIES OY

Patent Information

Application Number
PCT/EP2024/084246
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-10
Filing Date
2024-12-02
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing neural network post-filtering technologies lack efficient mechanisms for temporal extrapolation and post-processing operations, particularly in video coding, which can enhance visual quality and reduce latency in applications like cloud gaming and autonomous driving.

Method used

Implementing signaling information for temporal extrapolation using neural networks, including NNPFC and NNPFA SEI messages, to enhance post-processing operations in video codecs, allowing for frame rate upsampling and temporal interpolation/extrapolation.

Benefits of technology

Improves visual quality and reduces latency by enabling efficient temporal extrapolation and post-processing, facilitating applications such as cloud gaming and autonomous driving with enhanced video rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024084246_17072025_PF_FP_ABST
    Figure EP2024084246_17072025_PF_FP_ABST
Patent Text Reader

Abstract

An apparatus configured to: receive signaling information for performing a post-processing operation, wherein the post-processing operation comprises use of at least one neural network; receive at least one input; and provide the at least one input to the at least one neural network based, at least partially, on the signaling information. An apparatus configured to: generate signaling information for performing a post-processing operation, wherein the post-processing operation comprises use of at least one neural network; and provide the signaling information for performing the post-processing operation.
Need to check novelty before this filing date? Find Prior Art

Description

SIGNALING INFORMATION FOR TEMPORAL EXTRAPOLATIONTECHNICAL FIELD

[0001] The example and non-limiting embodiments relate generally to neural network post-filtering and, more particularly, to signaling information for performing temporal extrapolation.BACKGROUND

[0002] It is known, in neural network post-filtering, to use neural-network post-filter characteristics (NNPFC) and neural-network post-filter activation (NNPFA) supplemental enhancement (SEI) messages.SUMMARY

[0003] The following summary is merely intended to be illustrative. The summary is not intended to limit the scope of the claims.

[0004] In accordance with one aspect, an apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: generate signaling information for performing a post-processing operation, wherein the post-processing operation comprises use of at least one neural network; and provide the signaling information for performing the post-processing operation.

[0005] In accordance with one aspect, a method comprising: generating, with an encoder, signaling information for performing a post-processing operation, wherein the post-processing operation comprises use of at least one neural network; and providing the signaling information for performing the post-processing operation.

[0006] In accordance with one aspect, an apparatus comprising means for: generating signaling information for performing a post-processing operation, wherein the post-processing operation comprises use of at least one neural network; and providing the signaling information for performing the post-processing operation.

[0007] In accordance with one aspect, a non-transitory computer-readable medium comprising program instructions stored thereon for performing at least the following: generating signaling information for performing a post-processing operation, wherein the post-processing operation comprises use of at least one neural network; and causing providing of the signaling information for performing the post-processing operation.

[0008] In accordance with one aspect, an apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: receive signaling information for performing a post-processing operation, wherein the post-processing operation comprises use of at least one neural network; receive at least one input; and provide the at least one input to the at least one neural network based, at least partially, on the signaling information.

[0009] In accordance with one aspect, a method comprising: receiving, with a decoder, signaling information for performing a post-processing operation, wherein the post-processing operation comprises use of at least one neural network; receiving at least one input; and providing the at least one input to the at least one neural network based, at least partially, on the signaling information.

[0010] In accordance with one aspect, an apparatus comprising means for: receiving signaling information for performing a post-processing operation, wherein the post-processing operation comprises use of at least one neural network; receiving at least one input; and providing the at least one input to the at least one neural network based, at least partially, on the signaling information.

[0011] In accordance with one aspect, a non-transitory computer-readable medium comprising program instructions stored thereon for performing at least the following: causingreceiving of signaling information for performing a post-processing operation, wherein the postprocessing operation comprises use of at least one neural network; causing receiving of at least one input; and causing providing of the at least one input to the at least one neural network based, at least partially, on the signaling information.

[0012] According to some aspects, there is provided the subject matter of the independent claims. Some further aspects are defined in the dependent claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The foregoing aspects and other features are explained in the following description, taken in connection with the accompanying drawings, wherein:

[0014] FIG. 1 is a block diagram of one possible and non-limiting example system in which the example embodiments may be practiced;

[0015] FIG. 2 is a block diagram of one possible and non-limiting exemplary system in which the example embodiments may be practiced;

[0016] FIG. 3 is a diagram illustrating features as described herein;

[0017] FIG. 4 is a flowchart illustrating steps as described herein;

[0018] FIG. 5 is a flowchart illustrating steps as described herein; and

[0019] FIG. 6 illustrates a generic block diagram of a generative face video codec.DETAILED DESCRIPTION OF EMBODIMENTS

[0020] The following abbreviations that may be found in the specification and / or the drawing figures are defined as follows:3 GPP third generation partnership project4G fourth generation5G fifth generation5GC 5G core networkAl artificial intelligenceAOM Alliance for Open MediaAPS adaptation parameter setAR augmented realityAVC Advanced Video CodingCDMA code division multiple accessCLVS coded layer video sequenceCPU central processing unit cRAN cloud radio access networkCVS coded video sequenceDCT Discrete Cosine TransformDPB Decoded Picture Buffer eNB (or eNodeB) evolved Node B (e.g., an LTE base station)EN-DC E-UTRA-NR dual connectivity en-gNB or En-gNB node providing NR user plane and control plane protocol terminations towards the UE, and acting as secondary node in EN- DCE-UTRA evolved universal terrestrial radio access, i.e., the LTE radio access technologyFDMA frequency division multiple accessFW frame-wiseGDR Gradual Decoding RefreshGFV generative face videoGFVC generative face video characteristics gNB (or gNodeB) base station for 5G / NR, i.e., a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GCGPU graphical processing unitGRA Gradual random accessGSM global systems for mobile communicationsHEVC High Efficiency Video CodingHMD head-mounted displayIBC intra block copyIEC International Electrotechnical CommissionIEEE Institute of Electrical and Electronics EngineersIMD integrated messaging deviceIMS instant messaging service loT Internet of ThingsIRAP intra random access pointISO International Organisation for StandardizationITU-T International Telecommunication UnionJCT-VC Joint Collaborative Team - Video CodingJVET Joint Video Experts TeamLIE long term evolutionMMS multimedia messaging serviceMPEG-I Moving Picture Experts Group immersive codec familyMR mixed realityMSE mean squared errorMVC Multiview Video CodingNAL network abstraction layer ng or NG new generation ng-eNB or NG-eNB new generation eNBNN neural networkNNPF neural-network post-filterNNPFA neural network post-filter activationNNPFC neural network post-filter characteristicsNNPFI neural network post-filter inputNNR MPEG Neural Network compression and Representation standard(ISO / IEC 15938-17)NR new radioN / W or NW networkOBU open bitstream unitO-RAN open radio access networkPC personal computerPDA personal digital assistantPIR Progressive Intra RefreshPOC picture order countPPS picture parameter setPSNR peak signal to noise ratioRAP random access pointRBSP raw byte sequence payloadSAD sum of absolute differencesSATD sum of absolute transformed differencesSEI supplemental enhancement informationSMS short messaging serviceSPS sequence parameter setSSIM structural similaritySVC Scalable Video CodingTCP-IP transmission control protocol-internet protocolTDMA time division multiple accessTID temporal identifierUE user equipment (e.g., a wireless, typically mobile device)UMTS universal mobile telecommunications systemURI uniform resource identifierUSB universal serial busVCEG Video Coding Experts GroupVCL Video Coding LayerVMAF Video Multi-Method Assessment FusionVNR virtualized network functionVPS video parameter setVR virtual realityVSEI versatile supplemental enhancement informationVTM WC Test ModelVTNN visual temporal extrapolation neural networkVUI video usability informationWC versatile video codingWLAN wireless local area network

[0021] The following describes suitable apparatus and possible mechanisms for practicing example embodiments of the present disclosure. Accordingly, reference is first made to FIG. 1, which shows an example block diagram of an apparatus 50. The apparatus may be configured to perform various functions such as, for example, gathering information by one or more sensors, encoding and / or decoding information, receiving and / or transmitting information, analyzing information gathered or received by the apparatus, or the like. A device configured to encode a video scene may (optionally) comprise one or more microphones for capturing the scene and / or one or more sensors, such as cameras, for capturing information about the physical environment in which the scene is captured. Alternatively, a device configured to encode a video scene may be configured to receive information about an environment in which a scene is captured and / or a simulated environment. A device configured to decode and / or render the video scene may be configured to receive a Moving Picture Experts Group immersive codec family (MPEG-I) bitstream comprising the encoded video scene. A device configured to decode and / or render the video scene may comprise one or more speakers / audio transducers and / or displays, and / or may be configured to transmit a decoded scene or signals to a device comprising one or more speakers / audio transducers and / or displays. A device configured to decode and / or render the video scene may comprise a user equipment, a head / mounted display, or another device capable of rendering to a user an AR, VR and / or MR experience.

[0022] The electronic device 50 may for example be a mobile terminal or user equipment of a wireless communication system. Alternatively, the electronic device may be a computer or part of a computer that is not mobile. It should be appreciated that example embodiments of the present disclosure may be implemented within any electronic device or apparatus which may process data. The electronic device 50 may comprise a device that can access a network and / or cloud through a wired or wireless connection. The electronic device 50 may comprise one or more processors 56, one or more memories 58, and one or more transceivers 52 interconnected through one or more buses. The one or more processors 56 may comprise a central processing unit (CPU) and / or a graphical processing unit (GPU). Each of the one or more transceivers 52 includes a receiver and a transmitter. The one or more buses may be address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. A “circuit” may include dedicated hardware or hardware in association with software executable thereon. The one or more transceivers may be connected to one or more antennas 44. The one or more memories 58 may include computer program code. The one or more memories 58 and the computer program code may be configured to, with the one or more processors 56, cause the electronic device 50 to perform one or more of the operations as described herein.

[0023] The electronic device 50 may connect to a node of a network. The network node may comprise one or more processors, one or more memories, and one or more transceivers interconnected through one or more buses. Each of the one or more transceivers includes a receiver and a transmitter. The one or more buses may be address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. The one or more transceivers may be connected to one or more antennas. The one or more memories may include computer program code. The one or more memories and the computer program code may be configured to, with the one or more processors, cause the network node to perform one or more of the operations as described herein.

[0024] The electronic device 50 may comprise a microphone 36 or any suitable audio input which may be a digital or analogue signal input. The electronic device 50 may further comprisean audio output device 38 which in example embodiments of the present disclosure may be any one of: an earpiece, speaker, or an analogue audio or digital audio output connection. The electronic device 50 may also comprise a battery (or in other example embodiments of the present disclosure the device may be powered by any suitable mobile energy device such as solar cell, fuel cell, or clockwork generator). The electronic device 50 may further comprise a camera 42 or other sensor capable of recording or capturing images and / or video. Additionally or alternatively, the electronic device 50 may further comprise a depth sensor. The electronic device 50 may further comprise a display 32. The electronic device 50 may further comprise an infrared port for short range line of sight communication to other devices. In other example embodiments of the present disclosure the apparatus 50 may further comprise any suitable short-range communication solution such as for example a BLUETOOTH™ wireless connection or a USB / firewire wired connection.

[0025] It should be understood that an electronic device 50 configured to perform example embodiments of the present disclosure may have fewer and / or additional components, which may correspond to what processes the electronic device 50 is configured to perform. For example, an apparatus configured to encode a video might not comprise a speaker or audio transducer and may comprise a microphone, while an apparatus configured to render the decoded video might not comprise a microphone and may comprise a speaker or audio transducer.

[0026] Referring now to FIG. 1, the electronic device 50 may comprise a controller 56, processor or processor circuitry for controlling the apparatus 50. The controller 56 may be connected to memory 58 which in example embodiments of the present disclosure may store both data in the form of image and audio data and / or may also store instructions for implementation on the controller 56. The controller 56 may further be connected to codec circuitry 54 suitable for carrying out coding and / or decoding of audio and / or video data or assisting in coding and / or decoding carried out by the controller.

[0027] The electronic device 50 may further comprise a card reader 48 and a smart card 46, for example a UICC and UICC reader, for providing user information and being suitable for providing authentication information for authentication and authorization of the user / electronic device 50 at a network. The electronic device 50 may further comprise an input device 34, suchas a keypad, one or more input buttons, or a touch screen input device, for providing information to the controller 56.

[0028] The electronic device 50 may comprise radio interface circuitry 52 connected to the controller and suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system, or a wireless local area network. The apparatus 50 may further comprise an antenna 44 connected to the radio interface circuitry 52 for transmitting radio frequency signals generated at the radio interface circuitry 52 to other apparatus(es) and / or for receiving radio frequency signals from other apparatus(es).

[0029] The electronic device 50 may comprise a microphone 38, camera 42, and / or other sensors capable of recording or detecting audio signals, image / video signals, and / or other information about the local / virtual environment, which are then passed to the codec 54 or the controller 56 for processing. The electronic device 50 may receive the audio / image / video signals and / or information about the local / virtual environment for processing from another device prior to transmission and / or storage. The electronic device 50 may also receive either wirelessly or by a wired connection the audio / image / video signals and / or information about the local / virtual environment for encoding / decoding. The structural elements of electronic device 50 described above represent examples of means for performing a corresponding function.

[0030] The memory 58 may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The memory 58 may be a non-transitory memory. The memory 58 may be means for performing storage functions. The controller 56 may be or comprise one or more processors, which may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on a multicore processor architecture, as non-limiting examples. The controller 56 may be means for performing functions.

[0031] The electronic device 50 may be configured to perform capture of a volumetric scene according to example embodiments of the present disclosure. For example, the electronic device 50 may comprise a camera 42 or other sensor capable of recording or capturing images and / or video. The electronic device 50 may also comprise one or more transceivers 52 to enable transmission of captured content for processing at another device. Such an electronic device 50 may or may not include all the modules illustrated in FIG. 1.

[0032] The electronic device 50 may be configured to perform processing of volumetric video content according to example embodiments of the present disclosure. For example, the electronic device 50 may comprise a controller 56 for processing images to produce volumetric video content, a controller 56 for processing volumetric video content to project 3D information into 2D information, patches, and auxiliary information, and / or a codec 54 for encoding 2D information, patches, and auxiliary information into a bitstream for transmission to another device with radio interface 52. Such an electronic device 50 may or may not include all the modules illustrated in FIG. 1.

[0033] The electronic device 50 may be configured to perform encoding or decoding of 2D information representative of volumetric video content according to example embodiments of the present disclosure. For example, the electronic device 50 may comprise a codec 54 for encoding or decoding 2D information representative of volumetric video content. Such an electronic device 50 may or may not include all the modules illustrated in FIG. 1.

[0034] The electronic device 50 may be configured to perform rendering of decoded 3D volumetric video according to example embodiments of the present disclosure. For example, the electronic device 50 may comprise a controller for projecting 2D information to reconstruct 3D volumetric video, and / or a display 32 for rendering decoded 3D volumetric video. Such an electronic device 50 may or may not include all the modules illustrated in FIG. 1.

[0035] With respect to FIG. 2, an example of a system within which example embodiments of the present disclosure can be utilized is shown. The system 10 comprises multiple communication devices which can communicate through one or more networks. The system 10may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network (such as a GSM, UMTS, E-UTRA, LIE, CDMA, 4G, 5G network etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a BLUETOOTH™ personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and / or the Internet. A wireless network may implement network virtualization, which is the process of combining hardware and software network resources and network functionality into a single, software-based administrative entity, a virtual network. Network virtualization involves platform virtualization, often combined with resource virtualization. Network virtualization is categorized as either external, combining many networks, or parts of networks, into a virtual unit, or internal, providing network-like functionality to software containers on a single system. For example, a network may be deployed in a tele cloud, with virtualized network functions (VNF) running on, for example, data center servers. For example, network core functions and / or radio access network(s) (e.g. CloudRAN, O-RAN, edge cloud) may be virtualized. Note that the virtualized entities that result from the network virtualization are still implemented, at some level, using hardware such as processors and memories, and also such virtualized entities create technical effects.

[0036] It may also be noted that operations of example embodiments of the present disclosure may be carried out by a plurality of cooperating devices (e.g. cRAN).

[0037] The system 10 may include both wired and wireless communication devices and / or electronic devices suitable for implementing example embodiments of the present disclosure.

[0038] For example, the system shown in FIG. 2 shows a mobile telephone network 11 and a representation of the internet 28. Connectivity to the internet 28 may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.

[0039] The example communication devices shown in the system 10 may include, but are not limited to, an apparatus 15, a combination of a personal digital assistant (PDA) and a mobiletelephone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, a notebook computer 22, and a head-mounted display (HMD) 17. The electronic device 50 may comprise any of those example communication devices. In an example embodiment of the present disclosure, more than one of these devices, or a plurality of one or more of these devices, may perform the disclosed process(es). These devices may connect to the internet 28 through a wireless connection 2.

[0040] The example embodiments of the present disclosure may also be implemented in a set-top box; i.e. a digital TV receiver, which may / may not have a display or wireless capabilities, in tablets or (laptop) personal computers (PC), which have hardware and / or software to process neural network data, in various operating systems, and in chipsets, processors, DSPs and / or embedded systems offering hardware / software based coding. The example embodiments of the present disclosure may also be implemented in cellular telephones such as smart phones, tablets, personal digital assistants (PDAs) having wireless communication capabilities, portable computers having wireless communication capabilities, image capture devices such as digital cameras having wireless communication capabilities, gaming devices having wireless communication capabilities, music storage and playback appliances having wireless communication capabilities, Internet appliances permitting wireless Internet access and browsing, tablets with wireless communication capabilities, as well as portable units or terminals that incorporate combinations of such functions.

[0041] Some or further apparatus may send and receive calls and messages and communicate with service providers through a wireless connection 25 to a base station 24, which may be, for example, an eNB, gNB, access point, access node, other node, etc. The base station 24 may be connected to a network server 26 that allows communication between the mobile telephone network 11 and the internet 28. The system may include additional communication devices and communication devices of various types.

[0042] The communication devices may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmissioncontrol protocol-internet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), BLUETOOTH™, IEEE 802.11, 3GPP Narrowband loT and any similar wireless communication technology. A communications device involved in implementing various example embodiments of the present disclosure may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.

[0043] In telecommunications and data networks, a channel may refer either to a physical channel or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire, whereas a logical channel may refer to a logical connection over a multiplexed medium, capable of conveying several logical channels. A channel may be used for conveying an information signal, for example a bitstream, which may be a MPEG-I bitstream, from one or several senders (or transmitters) to one or several receivers.

[0044] Features as described herein may generally relate to processes occurring at a decoder. FIG. 3 illustrates diagrams of an example encoder (302) and an example decoder (340). In the encoder (302), input pictures (304) may be divided into a CU or CTU (306), a prediction block may be subtracted (308) to form a residual (310), which may be transformed (312) and quantized (314) before coding (316) as compressed bits (318) into a bitstream. The quantized transformed coefficients may also be dequantized / inverse quantized (320) and inverse transformed (322), and combined with the output of a prediction block (324). The result may then be used in intra prediction (326) and may be (e.g. in parallel) in-loop filtered (328), included in a decoded picture buffer (330), and used in inter prediction (332). In the decoder (340), compressed bits (342) may be decoded (344), dequantized (346), inverse transformed (348), and combined with the output of a prediction block (350). The result may then be used in intra prediction (352) and may be (e.g. in parallel) in-looped filtered (354), included in a decoded picture buffer (356), and used in inter prediction (358). The contents of the decoded picture buffer (356) may be output (360).

[0045] As shown in FIG. 3, the decoder (340) may actually be part of the coding loop of the encoder (302) in a reverse way (e.g. 320-332). The quantized transformed coefficients in the encoder (e.g. after quantization block (314), or the output of CAB AC (344) in the decoder) maybe dequantized (320, 346) and inverse transformed (322, 348), generating the coded residual block (e.g. 324, 350). The (intra or inter) prediction block (326, 332, 352, 358) may then be added to the coded residual block (324, 350), generating the reconstructed block. In-loop filtering may be performed over the reconstructed block (328, 354), forming the final reconstructed block. The final reconstructed blocks may be stored in a decoded picture buffer (330, 356) for output (360), as well as for possible use of future coding.

[0046] Features as described herein may generally relate to a video codec. A video codec comprises an encoder that transforms the input video into a compressed representation suited for storage / transmission and a decoder that can decompress the compressed video representation back into a viewable form. An encoder may discard some information in the original video sequence in order to represent the video in a more compact form (that is, at lower bitrate).

[0047] The H.264 / AVC standard was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunications Standardization Sector of International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of International Organisation for Standardization (ISO) / International Electrotechnical Commission (IEC). The H.264 / AVC standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, also known as MPEG-4 Part 10 Advanced Video Coding (AVC). Extensions of the H.264 / AVC include Scalable Video Coding (SVC) and Multiview Video Coding (MVC).

[0048] The High Efficiency Video Coding (H.265 / HEVC a.k.a. HEVC) standard was developed by the Joint Collaborative Team - Video Coding (JCT-VC) of VCEG and MPEG. The standard was published by both parent standardization organizations, and it is referred to as ITU- T Recommendation H.265 and ISO / IEC International Standard 23008-2, also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC). Later versions of H.265 / HEVC included scalable, multiview, fidelity range, three-dimensional, and screen content coding extensions which may be abbreviated SHVC, MV-HEVC, REXT, 3D-HEVC, and SCC, respectively.

[0049] Versatile Video Coding (H.266 a.k.a. WC), defined in ITU-T Recommendation H.266 and equivalently in ISO / IEC 23090-3, (also referred to as MPEG-I Part 3) is a video compression standard developed as the successor to HEVC. A reference software for WC is the WC Test Model (VTM).

[0050] A specification of the AVI bitstream format and decoding process were developed by the Alliance for Open Media (AOM). The AVI specification was published in 2018. AOM is reportedly working on the AV2 specification.

[0051] ITU-T Recommendation H.274, which is equivalent to ISO / IEC 23002-7, may be called "versatile supplemental enhancement information messages for coded video bitstreams" and be referred to as "versatile supplemental enhancement information" or VSEI. The VSEI standard specifies the syntax and semantics of video usability information (VUI) parameters and supplemental enhancement information (SEI) messages. The VUI parameters and SEI messages defined in the VSEI standard are designed to be conveyed within coded video bitstreams in a manner specified in a video coding specification or to be conveyed by other means determined by the specifications for systems that make use of such coded video bitstreams. The VSEI standard is intended for use with WC coded video bitstreams, although it is drafted in a manner intended to be sufficiently generic that it may also be used with other types of coded video bitstreams. VUI parameters and SEI messages may, for example, assist in processes related to decoding, display or other purposes.

[0052] An elementary unit for the input to a video encoder and the output of a video decoder, respectively, in most cases is a picture. A picture given as an input to an encoder may also be referred to as a source picture, and a picture decoded by a decoder may be referred to as a decoded picture or a reconstructed picture.

[0053] The source and decoded pictures are each comprised of one or more sample arrays, such as one of the following sets of sample arrays:Luma (Y) only (monochrome),Luma and two chroma (YCbCr or YcgCo),Green, Blue and Red (GBR, also known as RGB),Arrays representing other unspecified monochrome or tri-stimulus color samplings (for example, YZX, also known as XYZ).

[0054] A component may be defined as an array or single sample from one of the three sample arrays (luma and two chroma) that compose a picture, or the array or a single sample of the array that compose a picture in monochrome format.

[0055] Hybrid video codecs, for example ITU-T H.263 and H.264, may encode the video information in two phases. Firstly, pixel values in a certain picture area (or “block”) are predicted for example by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Secondly the prediction error, i.e., the difference between the predicted block of pixels and the original block of pixels, is coded. This may be done by transforming the difference in pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate).

[0056] Inter prediction, which may also be referred to as temporal prediction, motion compensation, or motion-compensated prediction, exploits temporal redundancy. In inter prediction the sources of prediction are previously decoded pictures.

[0057] In temporal inter prediction, the sources of prediction are previously decoded pictures in the same scalable layer. In intra block copy (IBC; a.k.a. intra-block-copy prediction), prediction may be applied similarly to temporal inter prediction but the reference picture is the current picture and only previously decoded samples can be referred in the prediction process. Inter-layer or interview prediction may be applied similarly to temporal inter prediction, but the reference picture isa decoded picture from another scalable layer or from another view, respectively. In some cases, inter prediction may refer to temporal inter prediction only, while in other cases inter prediction may refer collectively to temporal inter prediction and any of intra block copy, inter-layer prediction, and inter-view prediction, provided that they are performed with the same or similar process as temporal prediction. Inter prediction, temporal inter prediction, or temporal prediction may sometimes be referred to as motion compensation or motion-compensated prediction.

[0058] Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in the spatial or transform domain, i.e., either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.

[0059] One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy-coded more efficiently if they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors, and only the difference relative to the motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.

[0060] The decoder reconstructs the output video by applying prediction means similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in spatial pixel domain). After applying prediction and prediction error decoding means, the decoder sums up the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) can also apply additional filtering means to improve the quality of the output video before passing it for display and / or storage as a prediction reference for the forthcoming frames in the video sequence.

[0061] Image and video codecs may use a set of filters, which may enhance the visual quality of the predicted visual content. Filters may be applied either in-loop or out-of-loop, or both. In-loop filters (which may be also called loop filters) are used in reconstructing prediction reference that may be used for predicting forthcoming video signal. In other words, in the case of in-loop filters, the filter applied on one block in the currently encoded frame may affect the encoding of another block in the same frame and / or in another frame which is predicted from the current frame. An in-loop filter may affect the bitrate and / or the visual quality. In fact, an enhanced block may cause a smaller residual (difference between original block and predicted-and-filtered block), thus requiring less bits to be encoded. An out-of-the loop filter (which may also be called a postprocessing filter or a post-filter) may be applied on a frame or part of a frame after it has been reconstructed, the filtered visual content may not be used as a source for prediction, and thus it may only impact the visual quality of the frames that are output by the decoder.

[0062] In video codecs, the motion information may be indicated with motion vectors associated with each motion compensated image block. Each of these motion vectors represents the displacement of the image block in the picture to be coded (in the encoder side) or decoded (in the decoder side) and the prediction source block in one of the previously coded or decoded pictures. In order to represent motion vectors efficiently, those may be coded differentially with respect to block specific predicted motion vectors. In video codecs, the predicted motion vectors may be created in a predefined way, for example calculating the median of the encoded or decoded motion vectors of the adjacent blocks. Another way to create motion vector predictions is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures and signaling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index is typically predicted from adjacent blocks and / or or colocated blocks in a temporal reference picture. Moreover, high efficiency video codecs can employ an additional motion information coding / decoding mechanism, often called merging / merge mode, where all the motion field information, which includes motion vector and corresponding reference picture index for each available reference picture list, is predicted and used without any modification / correction. Similarly, predicting the motion field information may be carried out using the motion field information of adjacent blocks and / or co-located blocks in temporalreference pictures, and the used motion field information is signaled among a list of motion field candidate list filled with motion field information of available adjacent / co-located blocks.

[0063] In video codecs, the prediction residual after motion compensation may be first transformed with a transform kernel (like DCT) and then coded. The reason for this is that, often, there still exists some correlation among the residual and transform can in many cases help reduce this correlation and provide more efficient coding.

[0064] Video encoders may utilize Lagrangian cost functions to find optimal coding modes, e.g., the desired coding mode for a block, block partitioning, and associated motion vectors. This kind of cost function uses a weighting factor A. to tie together the (exact or estimated) image distortion due to lossy coding methods and the (exact or estimated) amount of information that is required to represent the pixel values in an image area:C = D + XR where C is the Lagrangian cost to be minimized, D is the image distortion (e.g., Mean Squared Error) with the mode and motion vectors considered, and R the number of bits needed to represent the required data to reconstruct the image block in the decoder (including the amount of data to represent the candidate motion vectors). The rate R may be the actual bitrate or bit count resulting from encoding. Alternatively, the rate R may be an estimated bitrate or bit count. One possible way of the estimating the rate R is to omit the final entropy encoding step and use, for example, a simpler entropy encoding or an entropy encoder where some of the context states have not been updated according to previously encoding mode selections.

[0065] Conventionally used distortion metrics may comprise, but are not limited to, peak signal-to-noise ratio (PSNR), mean squared error (MSE), sum of absolute differences (SAD), sum of absolute transformed differences (SATD), and structural similarity (SSIM), typically measured between the reconstructed video / image signal (that is or would be identical to the decoded video / image signal) and the “original” video / image signal provided as input for encoding.

[0066] Features as described herein may generally relate to transmission of data, for example in a bitstream or along a bitstream. The phrase along the bitstream (e.g., indicating along the bitstream) or along a coded unit of a bitstream (e.g., indicating along a coded tile) may be used in claims and described embodiments to refer to transmission, signaling, or storage in a manner that the ‘out-of-band’ data is associated with, but not included within, the bitstream or the coded unit, respectively. The phrase decoding along the bitstream or along a coded unit of a bitstream or alike may refer to decoding the referred out-of-band data (which may be obtained from out-of-band transmission, signaling, or storage) that is associated with the bitstream or the coded unit, respectively. For example, the phrase along the bitstream may be used when the bitstream is included in a container file, such as a file conforming to the ISO Base Media File Format, and certain file metadata is stored in the file in a manner that associates the metadata to the bitstream, such as boxes in the sample entry for a track including the bitstream, a sample group for the track including the bitstream, or a timed metadata track associated with the track including the bitstream. In another example, the phrase along the bitstream may be used when the bitstream is made available as a stream over a communication protocol and a media description, such as a streaming manifest, is provided to describe the stream.

[0067] A bitstream may be defined as a sequence of bits, which may in some coding formats or standards be in the form of a network abstraction layer (NAL) unit stream or a byte stream, that forms the representation of coded pictures and associated data forming one or more coded video sequences.

[0068] A bitstream format may comprise a sequence of syntax structures. A bitstream format may constrain the order of syntax structures in the bitstream.

[0069] A syntax element may be defined as an element of data represented in the bitstream. A syntax structure may be defined as zero or more syntax elements present together in the bitstream in a specified order.

[0070] Syntax structures may be specified, for example, using arithmetic, logical, relational, bit-wise, and assignment operators similar to those available in many programming languages.For example, & may indicate a bit-wise ‘AND’ operation. Furthermore, syntax structures may be specified with reference to mathematical functions.

[0071] Syntax structures and semantics may use the values of variables derived from the values of syntax elements. Naming conventions may be defined for variables. For example, variables may be named by a mixture of lower case and upper case letter and without any underscore characters. Variables starting with an upper case letter may be derived for the decoding of the current syntax structure and all depending syntax structures. Variables starting with an upper case letter may, in some cases, be used in the decoding process for later syntax structures without mentioning the originating syntax structure of the variable. Variables starting with a lower case letter may only be used in relation to the syntax structure or function they have been defined for.

[0072] An elementary unit for the output of a video encoder and the input of a video decoder, respectively, may be a network abstraction layer (NAL) unit. For transport over packet-oriented networks or storage into structured files, NAL units may be encapsulated into packets or similar structures. A bytestream format encapsulating NAL units may be used for transmission or storage environments that do not provide framing structures. The bytestream format may separate NAL units from each other by attaching a start code in front of each NAL unit. To avoid false detection of NAL unit boundaries, encoders may run a byte-oriented start code emulation prevention algorithm, which may add an emulation prevention byte to the NAL unit payload, when a start code would have occurred otherwise. In order to enable straightforward gateway operation between packet and stream-oriented systems, start code emulation prevention may be performed regardless of whether the bytestream format is in use or not. A NAL unit may be defined as a syntax structure including an indication of the type of data to follow and bytes including that data in the form of a raw byte sequence payload interspersed as necessary with emulation prevention bytes. A raw byte sequence payload (RBSP) may be defined as a syntax structure including an integer number of bytes that is encapsulated in a NAL unit. An RBSP is either empty or has the form of a string of data bits including syntax elements followed by an RBSP stop bit and followed by zero or more subsequent bits equal to 0.

[0073] A bitstream may be defined to logically include a syntax structure, such as a NAL unit, when the syntax structure is transmitted along the bitstream but may be included in the bitstream according to the bitstream format. A bitstream may be defined to natively comprise a syntax structure, when the bitstream includes the syntax structure.

[0074] In some formats or standards, a first bitstream may be followed by a second bitstream in the same logical channel, such as in the same file or in the same connection of a communication protocol. An elementary stream (in the context of video coding) may be defined as a sequence of one or more bitstreams.

[0075] In some coding formats or standards, a bitstream may be in the form of a network abstraction layer (NAL) unit stream or a byte stream, that forms the representation of coded pictures and associated data forming one or more coded video sequences.

[0076] In some coding formats, such as AVI, a bitstream may comprise a sequence of open bitstream units (OBUs). An OBU comprises a header and a payload, wherein the header identifies a type of the OBU. Furthermore, an OBU may include the OBU size, which may follow the OBU header and precede the OBU payload within the OBU, wherein the OBU size may be indicative of a size of the payload in bytes.

[0077] A bitstream of OBUs may be formatted as a so-called low-overhead bitstream format, which comprises a sequence of OBUs, wherein the OBU header indicates the OBU size (without the size field) in bytes. Alternatively, a bitstream of OBUs may be formatted as a so-called length delimited bitstream format, in which each temporal unit starts with a size field (temporal unit size) indicating the size of the temporal unit payload in bytes, and each frame unit starts with a size field (frame unit size) indicating the size of the frame unit payload in bytes, and each OBU is preceded by a size field indicating the size of the OBU in bytes.

[0078] The OBU types may include the following.Sequence Header includes information that applies to the entire sequence and whether to enable certain coding tools.Temporal Delimiter indicates the frame presentation time stamp. All displayable frames following a temporal delimiter OBU will use this time stamp, until the next temporal delimiter OBU arrives. A temporal delimiter and its subsequent OBUs of the same time stamp are referred to as a temporal unit. In the context of scalable coding, the compression data associated with all representations of a frame at various spatial and fidelity resolutions will be in the same temporal unit.- Frame Header sets up the coding information for a given frame, including signalling inter or intraframe type, indicating the reference frames and signalling probability model update method.Tile Group includes the tile data associated with a frame. Each tile can be independently decoded. The collective reconstructions form the reconstructed frame after potential loop filtering.- Frame includes the frame header and tile data. The frame OBU is largely equivalent to a frame header OBU and a tile group OBU but allows less overhead cost.- Metadata carries information, such as high dynamic range, scalability, and timecode.Tile List includes tile data similar to a tile group OBU. However, each tile here has an additional header that indicates its reference frame index and position in the current frame. This allows the decoder to process a subset of tiles and display the corresponding part of the frame, without the need to fully decode all the tiles in the frame.

[0079] In some coding standards, NAL units include a header and payload. The NAL unit header indicates the type of the NAL unit. In some coding standards, the NAL unit header indicates a scalability layer identifier (e.g., called nuh_layer_id in H.265 / HEVC and H.266 / VVC), which may be used, e.g., for indicating spatial or quality layers, views of a multiview video, orauxiliary layers (such as depth maps or alpha planes). In some coding standards, the NAL unit header includes a temporal sublayer identifier, which may be used for indicating temporal subsets of the bitstream, such as a 30-frames-per-second subset of a 60-frames-per-second bitstream.

[0080] Bitstreams or coded video sequences may be encoded to be temporally scalable as follows. Each picture may be assigned to a particular temporal sub-layer. A temporal sub-layer may be equivalently called a sub-layer, temporal sublayer, sublayer, or temporal level. Temporal sub-layers may be enumerated, e.g., from 0 upwards. The lowest temporal sub-layer, sub-layer 0, may be decoded independently. Pictures at temporal sub-layer 1 may be predicted from reconstructed pictures at temporal sub-layers 0 and 1. Pictures at temporal sub-layer 2 may be predicted from reconstructed pictures at temporal sub-layers 0, 1, and 2, and so on. In other words, a picture at temporal sub-layer N does not use any picture at temporal sub-layer greater than N as a reference for inter prediction. The bitstream created by excluding all pictures greater than or equal to a selected sub-layer value and including pictures remains conforming.

[0081] Each picture of a temporally scalable bitstream may be assigned with a temporal identifier (also known as TID, temporal layer identifier, temporal sublayer identifier, or temporal layer ID), which may be, for example, assigned to a variable Temporalld. The temporal identifier may, for example, be indicated in a NAL unit header or in an OBU extension header. Temporalld equal to 0 corresponds to the lowest temporal level. The bitstream created by excluding all coded pictures having a Temporalld greater than or equal to a selected value and including all other coded pictures remains conforming. Consequently, a picture having Temporalld equal to tid value does not use any picture having a Temporalld greater than tid value as a prediction reference.

[0082] NAL units may be categorized into Video Coding Layer (VCL) NAL units and non- VCL NAL units. VCL NAL units are typically coded slice NAL units.

[0083] A non-VCL NAL unit may be, for example, one of the following types: a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a supplemental enhancement information (SEI) NAL unit, an access unit delimiter, an end of sequence NAL unit, an end of bitstream NAL unit, or a filler data NAL unit.Parameter sets may be needed for the reconstruction of decoded pictures, whereas many of the other non-VCL NAL units are not necessary for the reconstruction of decoded sample values.

[0084] Video coding specifications may enable the use of supplemental enhancement information (SEI) messages or alike. Some video coding specifications include SEI network abstraction layer (NAL) units, and some video coding specifications contain both prefix SEI NAL units and suffix SEI NAL units, where the former type can start a picture unit or alike and the latter type can end a picture unit or alike. An SEI NAL unit contains one or more SEI messages, which are not required for the decoding of output pictures but may assist in related processes, such as picture output timing, post-processing of decoded pictures, rendering, error detection, error concealment, and resource reservation. Several SEI messages are specified in H.264 / AVC, H.265 / HEVC, H.266 / VVC, and H.274 / VSEI standards, and the user data SEI messages enable organizations and companies to specify SEI messages for their own use. The standards may contain the syntax and semantics for the specified SEI messages, but a process for handling the messages in the recipient might not be defined. Consequently, encoders may be required to follow the standard specifying a SEI message when they create SEI message(s), and decoders might not be required to process SEI messages for output order conformance. One of the reasons to include the syntax and semantics of SEI messages in standards is to allow different system specifications to interpret the supplemental information identically and hence interoperate. It is intended that system specifications can require the use of particular SEI messages both in the encoding end and in the decoding end, and additionally the process for handling particular SEI messages in the recipient can be specified.

[0085] Video usability information (VUI) may be defined as a syntax structure that identifies properties of interpretation of decoded pictures for display purposes, particularly including color representation information. VUI parameters may include, but may not be limited to, one or more of the following: progressive source indication flag, interlaced source indication flag, frame packing indication flag, projected content indication flag, sample aspect ratio information, overscan information, color primaries, transfer characteristics, matrix coefficients for color conversion, sample value range indication (e.g., indicative of full range or a studio range), chromasample location information. VUI parameters may be carried as part of a parameter set, such as a sequence parameter set or a video parameter set.

[0086] ITU-T Recommendation H.274, which is equivalent to ISO / IEC 23002-7, may be called "versatile supplemental enhancement information messages for coded video bitstreams" and be referred to as "versatile supplemental enhancement information" or VSEI. The VSEI standard specifies the syntax and semantics of video usability information (VUI) parameters and supplemental enhancement information (SEI) messages. The VUI parameters and SEI messages defined in the VSEI standard are designed to be conveyed within coded video bitstreams in a manner specified in a video coding specification or to be conveyed by other means determined by the specifications for systems that make use of such coded video bitstreams. The VSEI standard is intended for use with WC coded video bitstreams, although it is drafted in a manner intended to be sufficiently generic that it may also be used with other types of coded video bitstreams. VUI parameters and SEI messages may, for example, assist in processes related to decoding, display or other purposes. In the VSEI standard, the vui_parameters( payloadSize ) syntax structure is specified for the VUI, where payloadSize is an input argument indicating the number of bits in the VUI.

[0087] Some video coding specifications enable metadata OBUs. A metadata OBU comprises a type field, which specifies the type of metadata.

[0088] In some coding formats, a coded video sequence (CVS) may be defined as a sequence of coded pictures in decoding order that is independently decodable and is followed by another coded video sequence or the end of the bitstream.

[0089] In some coding formats, such as AVI, a coded video sequence comprises one or more temporal units. A temporal unit consists of a series of OBUs starting from a temporal delimiter, optional sequence headers, optional metadata OBUs, a sequence of one or more frame headers, each followed by zero or more tile group OBUs as well as optional padding OBUs. A temporal unit may be defined to comprise all the OBUs that are associated with a specific, distinct time instant. A temporal unit may comprise a temporal delimiter OBU, and all the OBUs that follow,up to but not including the next temporal delimiter. A temporal delimiter OBU may be defined as an indication that the following OBUs will have a different presentation / decoding time stamp from the one of the last frame prior to the temporal delimiter.

[0090] A coded layer video sequence (CLVS) may be defined as a sequence of pictures and associated other data within the same scalable layer (e.g., with the same value of nuh layer id in WC) that is decodable independently of other pictures in the same layer.

[0091] Some codecs use a concept of picture order count (POC). A value of POC is derived for each picture and is non-decreasing with increasing picture position in output order. POC therefore indicates the output order of pictures. POC may be used in the decoding process for example for implicit scaling of motion vectors and for reference picture list initialization. Furthermore, POC may be used in the verification of output order conformance. The variable including a POC value of a picture may be referred to as PicOrderCntVal.

[0092] A Decoded Picture Buffer (DPB) may be used in the encoder and / or in the decoder. There may be two reasons to buffer decoded pictures: for references in inter prediction and / or for reordering decoded pictures into output order. Some coding formats, such as HEVC, provide a great deal of flexibility for both reference picture marking and output reordering. Separate buffers for reference picture buffering and output picture buffering may waste memory resources. Hence, the DPB may include a unified decoded picture buffering process for reference pictures and output reordering. A decoded picture may be removed from the DPB when it is no longer used as a reference and is not needed for output.

[0093] Output order may be defined as the order in which the decoded pictures are output from the decoded picture buffer (for the decoded pictures that are to be output from the decoded picture buffer).

[0094] Output time may be defined as a time when a decoded picture is to be output from a decoder or from the DPB of a decoder (for the decoded pictures that are to be output from the DPB), for example as specified by a hypothetical reference decoder specification according to the output timing DPB operation.

[0095] Pictures having the same output order may be defined to mean the same as pictures having the same output time.

[0096] Decoding order may be defined as the order in which syntax elements are processed by the decoding process. It may be required that syntax elements are ordered in a bitstream in their decoding order.

[0097] An identifier may be defined as a syntax element that identifies a syntax structure. A value of the identifier may for example differ in different instances of the same syntax structure, such as a parameter set. A particular instance of the syntax structure may be referenced through its identifier value. For example, a parameter set that is referenced by the (de)coding of a coded video slice may be identified by providing the identifier value of the parameter set in a header of the coded video slice.

[0098] An indicator (ide) may be defined as a syntax element whose value indicates a selection among more than two values (for which semantics have been specified). An indicator syntax element may have _idc postfix in its name.

[0099] An alpha mask (a.k.a. an alpha plane) may be used to provide transparency information for an associated image. A first value of an alpha mask may represent a fully opaque pixel, and a second value may represent a fully transparent pixel. Values between the first and second values may represent different levels of transparency between fully opaque and fully transparent.

[0100] Coded pictures may have different picture types and / or may be categorized according to their picture types into different categories.

[0101] In some coding formats, a picture type may be indicated in a picture header or a frame header. In some coding formats, VCL NAL units may be required to have the same NAL unit type and a picture type may be indicated through the NAL unit type.

[0102] A random access point (RAP) picture (a.k.a. a random access picture) may be defined as a coded picture from which the decoding process can be started.

[0103] An intra random access point (IRAP) picture in an independent layer may comprise only intra-coded data. Provided that the necessary parameter sets and / or all the necessary header syntax structures (e.g., a sequence header) are available, it may be required that an IRAP picture at an independent layer and all subsequent pictures at the independent layer in decoding and output order can be correctly decoded without performing the decoding process of any pictures that precede the IRAP picture in decoding order.

[0104] An IRAP picture belonging to a predicted layer (a.k.a. dependent layer) of a scalable multi-layer bitstream does not use inter prediction from other pictures in the same predicted layer may use inter-layer prediction from its direct reference layers. Provided that the direct reference layers are decoded and the necessary parameter sets and / or all the necessary header syntax structures (e.g., a sequence header) are available, it may be required that an IRAP picture at a dependent layer and all subsequent pictures at the dependent layer in decoding and output order can be correctly decoded without performing the decoding process of any pictures that precede the IRAP picture in decoding order at the dependent layer.

[0105] In some coding formats, an IRAP picture may be referred to as a key frame.

[0106] Gradual Decoding Refresh (GDR) often refers to the ability to start decoding at a non- IRAP picture and to recover decoded pictures that are correct in content after decoding a certain number of pictures. Said otherwise, GDR can be used to achieve random access from non-intra pictures. GDR, which is also known as gradual random access (GRA) or Progressive Intra Refresh (PIR), alleviates the delay issue with intra coded pictures. Instead of coding an intra picture at a random access point, GDR progressively refreshes pictures by spreading intra coded regions (groups of intra coded blocks) over several pictures.

[0107] A GDR picture may be defined as a RAP picture that, when used to start the decoding process, enables recovery of exactly or approximately correct decoded pictures starting from a specific picture, known as the recovery point picture. It is possible to start decoding from a GDR picture.

[0108] In some video coding formats, such as WC, all Video Coding Layer (VCL) Network Abstraction Layer (NAL) units of a GDR picture may have a particular NAL unit type value that indicates a GDR NAL unit.

[0109] In some video coding formats, an SEI message, a metadata OBU or alike with a particular type, such as a recovery point SEI message of HEVC, may be used to indicate a GDR picture and / or a recovery point picture.

[0110] For a single- lay er bitstream, a random access segment may be defined as the sequence of pictures of the bitstream from a RAP picture (inclusive) to the next RAP picture (exclusive) in decoding order. For a multi-layer bitstream, a random access segment may be defined as the sequence of pictures of a CLVS from a RAP picture (inclusive) to the next RAP picture (exclusive) in decoding order.

[0111] A picture unit may be defined as a set of NAL units that are associated with each other according to a specified classification rule, are consecutive in decoding order, and include exactly one coded picture. Some non-VCL NAL units, such as prefix SEI NAL units, are allowed to precede all VCL NAL units of the same picture unit in decoding order and may therefore start a picture unit but are not allowed to follow all VCL NAL units of the same picture unit in decoding order. Some non-VCL NAL units, such as suffix SEI NAL units, are allowed to follow all VCL NAL units of the same picture unit in decoding order but are not allowed to precede all VCL NAL units of the same picture unit in decoding order.

[0112] Having thus introduced one suitable but non-limiting technical context for the practice of the example embodiments of the present disclosure, example embodiments will now be described with greater specificity.

[0113] In this disclosure, the terms “picture”, "image", and "frame" may be used interchangeably. Also, the terms “NN”, "filter", "NN filter", “postfilter”, “post-filter”, “NN postfilter”, “NN post-processing filter”, “NNPF” may be used interchangeably. Moreover, the terms “postfilter”, “post-filter”, "post-processing filter" and “postprocessing filter" may be used interchangeably.

[0114] Features as described herein may generally relate to neural networks (NN). A neural network (NN) may be regarded as a computation graph consisting of two or more layers of computation. Each layer may consist of one or more units, where each unit may perform an elementary computation. The elementary computation may comprise a linear operation or a nonlinear operation. A unit may be connected to one or more other units, and the connection may have a weight associated with it. The weight may be used for scaling the signal passing through the associated connection. Weights may be learnable parameters, i.e., values which can be learned from training data. There may be other learnable parameters, such as those of batch-normalization layers.

[0115] Neural networks may be utilized in an ever increasing number of applications for many different types of device, such as mobile phones, as described above. Examples of applications may include image and video analysis and processing, social media data analysis, device usage data analysis, etc.

[0116] Neural networks, and other machine learning tools, may be able to learn properties of input data or learn to perform tasks given input data, based on a learning algorithm or training algorithm. The learning algorithm may comprise (but may not be limited to) one or more of a supervised learning algorithm, a self-supervised learning algorithm, an unsupervised learning algorithm, etc. Examples of properties of input data that may be learned comprise categories, correlations, causation relations, etc. Examples of tasks comprise classification or categorization, regression, prediction, etc.

[0117] A training algorithm may consist of changing some properties of the neural network so that the output of the neural network is as close as possible to a desired output. Training may comprise changing properties of the neural network so as to minimize or decrease the output's error, also referred to as the loss. Examples of losses include mean squared error (MSE), crossentropy, etc. The training algorithm may comprise an iterative process, where, at each iteration, the algorithm may modify the weights of the neural network to make a gradual improvement of the network's output, i.e., to gradually decrease the loss.

[0118] Training a neural network may be regarded as an optimization process, where the goal is to make the model learn the properties of the data distribution from a limited training dataset. In other words, the training process is additionally used to ensure that the neural network learns to use a limited training dataset in order to learn to generalize to previously unseen data, i.e., data which was not used for training the model. This additional goal is usually referred to as generalization. In practice, data may be split into at least two sets, the training set and the validation set. The training set may be used for training the network, i.e., for modification of its learnable parameters in order to minimize the loss. The validation set may be used for checking the performance of the neural network with data which was not used to minimize the loss (i.e. which was not part of the training set), where the performance of the neural network with the validation set may be an indication of the final performance of the model. The errors on the training set and on the validation set may be monitored during the training process to understand if the neural network is learning at all and if the neural network is learning to generalize. In the case that the network is learning at all, the training set error should decrease. If the network is not learning, the model may be in the regime of underfitting. In the case that the network is learning to generalize, validation set error should decrease and not be much higher than the training set error. If the training set error is low, but the validation set error is much higher than the training set error, or the validation set error does not decrease, or it even increases, the model may be in the regime of overfitting. Overfitting may mean that the model has memorized the training set's properties and performs well only on that set, but performs poorly on a set not used for tuning its parameters. In other words, the model has not learned to generalize.

[0119] Features as described herein may generally relate to NN post-filtering (NNPF). For example, a picture that is decoded by an image or video decoder may be further processed by a post-processing operation. The post-processing operation may comprise one or more NNs, and / or one or more non-NNs operations. The post-processing operation may take as input one or more input pictures, and may output one or more output pictures. Post-processing may be used for several reasons or use cases, including (but not limited to) the following:

[0120] Enhancing the picture with respect to one or more objective metrics, such as peak signal to noise ratio (PSNR), where the objective metrics may be derived based on pixel -wisedistortion metrics (as in the case of PSNR), or may be derived based on perceptual metrics (as in the case of video multi-method assessment fusion (VMAF) or feature-based metrics). Enhancing may comprise removing coding artifacts, denoising, etc.

[0121] - Enhancing the picture with respect to one or more subjective scores or metrics.Enhancing may comprise removing coding artifacts, denoising, etc.

[0122] - Upsampling and super-resolution.

[0123] - Temporal interpolation of one or more output pictures given two or more input pictures.

[0124] - Temporal extrapolation of one or more output pictures given one or more input pictures.

[0125] - Colorization.

[0126] An encoder, such as a video encoder, or a transmitter may signal information to a decoder, such as a video decoder, or a receiver, where the information is indicative of (but not limited to) the following information:

[0127] - Whether to apply a post-processing operation.

[0128] - The characteristics of the post-processing operation, for example in terms of the format of the input and / or output of the post-processing operation.

[0129] - How to apply the post-processing operation.

[0130] Features as described herein may generally relate to the neural-network post-filter characteristics (NNPFC) supplemental enhancement information (SEI) message and / or the neural- network post-filter activation (NNPFA) SEI message. The NNPFC SEI message and the NNPFA SEI message have been described in version 3 of the versatile supplemental enhancement information (VSEI) standard. The syntax structure specifying the NNPFC SEI message may becalled nn_post_filter_characteristics. The syntax structure specifying the NNPFA SEI message may be called nn_post_filter_activation.

[0131] The NNPFC SEI message comprises the nnpfc id syntax element, which contains an identifying number that may be used to identify a post-processing filter. A base post-processing filter is the filter that is contained in or identified by the first NNPFC SEI message, in decoding order, that has a particular nnpfc id value within a coded layer video sequence (CLVS). If there is a second NNPFC SEI message that has the same nnpfc id value that defines the base postprocessing filter, an update relative to the base post-processing filter is applied to obtain a postprocessing filter associated with the nnpfc id value. The update may be obtained by decoding the coded neural network bitstream in the second NNPFC SEI message. Otherwise, the postprocessing filter associated with the nnpfc id value is assigned to be the same as the base postprocessing filter.

[0132] The NNPFC SEI message comprises the nnpfc mode idc syntax element, the semantics of which may be defined as follows:

[0133] - nnpfc mode idc equal to 1 specifies that the base post-processing filter or the update relative to the base post-processing filter associated with the nnpfc id value is a neural network identified by the uniform resource identifier (URI) nnpfc uri with the format identified by the tag URI nnpfc tag uri.

[0134] - nnpfc mode idc equal to 0 indicates that this SEI message contains an ISO / IEC15938-17 bitstream that specifies the base post-processing filter or updates relative to the base post-processing filter with the same nnpfc id value.

[0135] The NNPFC SEI message may also comprise:

[0136] - Purpose of the post-processing filter, which may comprise, but may not be limited to, one or more of the following: visual quality improvement; chroma upsampling from the 4:2:0 chroma format to the 4:2:2 or 4:4:4 chroma format, or from the 4:2:2 chroma format to the 4:4:4chroma format; increasing the width or height of the input picture; frame rate upsampling; bit depth upsampling; and / or colorization.

[0137] - Formatting of the input tensors that are given as input to the neural network inference.

[0138] - Formatting of the output tensors that are resulting from the neural network inference.

[0139] - Characterization of the complexity of the neural network.

[0140] The NNPFC SEI message syntax includes the nnpfc_num_input_pics_minusl syntax element. nnpfc_num_input_pics_minusl plus 1 specifies the number of pictures used as input for the NNPF. The variable numlnputPics may be set equal to nnpfc_num_input_pics_minusl + 1.

[0141] A frame rate upsampling filter may interchangeably be called a picture rate upsampling filter. Such a filter generates or interpolates one or more pictures between a pair of pictures given as input to the filter. It is also possible to have a frame rate upsampling filter where the number of input pictures may be greater than 2. Such a frame rate upsampling filter may generate pictures between more than one pair of input pictures. A frame rate upsampling filter may comprise a neural network, in which case the generation of the interpolated pictures between a pair of input pictures is performed by the inference of the neural network. It is possible to have a frame rate upsampling filter that extrapolates a picture before input picture(s) or after input picture(s), instead of or in addition to between input pictures.

[0142] When the filtering purpose comprises frame rate upsampling, the NNPFC SEI message includes nnpfc_interpolated_pics[ i ] syntax elements for the values of i in the range of 0, inclusive, to nnpfc_num_input_pics_minusl, exclusive. nnpfc_interpolated_pics[ i ] specifies the number of interpolated pictures generated by the NNPF between the i-th and the ( i + 1 )-th picture used as input for the NNPF.

[0143] The NNPFC SEI message syntax may comprise an indication, which may be called nnpfc_absent_input_pic_zero_flag, that indicates how pictures that would not originate from thecurrent bitstream are expected to be replaced in the input tensor. nnpfc_absent_input_pic_zero flag equal to 1 indicates that the NNPF expects an input picture that is not present in the current bitstream to be represented by sample arrays with sample values equal to 0. nnpfc_absent_input_pic_flag equal to 0 indicates that the NNPF expects an input picture that is not present in the current bitstream to be represented by the closest input picture in output order within the current bitstream.

[0144] The NNPFC SEI message syntax may comprise an indication, which may be called nnpfc auxiliary inp idc, that indicates if auxiliary input data in addition to sample array(s) of input picture(s) is present in the input tensor of the NNPF. nnpfc auxiliary inp idc greater than 0 indicates that auxiliary input data is present in the input tensor(s) of the NNPF. Specific semantics may be specified for specific non-zero values of nnpfc auxiliary inp idc. nnpfc auxiliary inp idc equal to 0 indicates that auxiliary input data is not present in the input tensor(s).

[0145] The NNPFA SEI message specifies the neural-network post-processing filter (NNPF) that may be used for post-processing filtering for the current picture, or for post-processing filtering for the current picture and one or more other pictures. The NNPFA SEI message comprises the nnpfa target id syntax element, which indicates that the neural-network postprocessing filter with nnpfc id equal to nnpfa target id may be used for post-processing filtering for the indicated persistence. The indicated persistence may be the current picture only (indicated by nnpfa_persistence_flag equal to 0). Alternatively, the NNPF activation may be indicated to be persistent by nnpfa_persistence_flag equal to 1, in which case the persistence of the NNPF activation may last until the end of the current CLVS or the next picture, in output order, in the current layer associated with a NNPFA SEI message with the same nnpfa target id as the current SEI message.

[0146] The NNPFA SEI message syntax may comprise a syntax element indicative of whether the base post-processing filter or the latest post-processing filter is activated, where the latest post-processing filter is defined by the base post-processing filter relative to which the latest filter update, if any, has been applied. The syntax element may be called nnpfa target base flag.nnpfa_target_base_flag equal to 1 specifies that the target NNPF is the base NNPF with nnpfc_id equal to nnpfa target id. nnpfa target base flag equal to 0 specifies that the target NNPF is the NNPF specified by the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id that precedes the first VCL NAL unit of the current picture in decoding order and is not a repetition of the NNPFC SEI message that contains the base NNPF.

[0147] The NNPFA SEI message syntax may comprise indications which ones of the filtered pictures corresponding to the input pictures are output by the NNPF process. For the i-th input picture that is filtered by the NNPF, the NNPFA SEI message syntax may comprise nnpfa_output_flag[ i ] syntax element, which when equal to 0, specifies that the filtered picture is not output by the NNPF process, and when equal to 1 , specifies that the filtered picture is output by the NNPF process.

[0148] In relation to an NNPFA SEI message, two sets of pictures may be defined, namely nnpfcTargetPictures and nnpfaTargetPictures. nnpfcTargetPictures may be defined to be the set of pictures to which the last NNPFC SEI message with nnpfc id equal to nnpfa target id that precedes the current NNPFA SEI message in decoding order pertains. nnpfaTargetPictures may be defined to be the set of pictures for which the target NNPF is activated by the current NNPFA SEI message. It may be required for a conforming bitstream that any picture included in nnpfaTargetPictures shall also be included in nnpfcTargetPictures.

[0149] An NNPF process comprises performing the NNPF inference for given input pictures. The NNPF inference may be performed in a patch-wise manner so that the entire picture area gets filtered. The NNPF inference may be followed by outputting NNPF -generated pictures in their increasing index order, where all NNPF-generated pictures that were interpolated by the NNPF are output and those NNPF-generated pictures that correspond to any input pictures to the NNPF are output as specified in the semantics of the NNPFA SEI message.

[0150] A general post-processing filtering process using NNPFs may be described as follows. Input to this process is a bitstream BitstreamToFilter. Output of this process is a list of NNPF output pictures ListNnpfOutputPics. First, BitstreamToFilter is decoded, and the listCroppedDecodedPictures is set to be the list of the cropped decoded pictures in output order resulted from decoding BitstreamToFilter. Second, the filtering process for one picture, as described below, is repeatedly invoked, in output order, for each cropped decoded picture that is in CroppedDecodedPictures and for which one or more NNPFs are activated. The order of the pictures in ListNnpfOutputPics is in output order. It may be required that within ListNnpfOutputPics there shall be no more than one picture pertaining to any particular output time instance. When for any particular picture in CroppedDecodedPictures there are multiple NNPFs activated and only one the NNPFs is allowed to be chosen to be applied although any of the NNPFs may be chosen, the above constraint shall apply regardless of which NNPF is chosen to be applied to the particular picture.

[0151] A filtering process for one picture using an NNPF may be described as follows. The filtering process for one picture using an NNPF may be applied to each cropped decoded picture, referred to as the current picture, that is in CroppedDecodedPictures and for which one or more NNPFs are activated. When applying an NNPF to the current picture, the filtered and / or interpolated pictures are generated by the NNPF by applying the NNPF process to the current picture. When applying an NNPF to the current picture, the order of the pictures generated by the NNPF by applying the NNPF process being stored into the output tensor of the NNPF is in output order. When the applied NNPF is the last NNPF that is applied to the current picture, the pictures generated by the NNPF and output by the NNPF process are included into ListNnpfOutputPics, in the same order as when the pictures are stored into the output tensor of the NNPF.

[0152] The use of NNPFC and NNPFA SEI messages for versatile video coding (WC) has been described in version 3 of the WC standard. It is to be understood that NNPFC and NNPFA SEI message may be similarly used for any other video coding specification.

[0153] When NNPFC and NNPFA SEI messages are used for WC, a decoder selects input pictures for the NNPF. The input pictures may be selected in reverse output order starting from a picture for which the NNPF is activated through an NNPFA SEI message. The input pictures may be indexed, starting from index 0 that is assigned for the picture for which the NNPF is activated through an NNPFA SEI message. In an example, the decoder selects the input picture with indexi, where i is greater than 0, to be the latest cropped decoded output picture, in output order, that precedes the input picture with index i-1 in output order. If there is no cropped decoded output picture, in output order, that precedes the input picture with index i-1 in output order as a result of decoding the bitstream, it may be considered that the input picture with index i is not present in the current bitstream (i.e. is missing) and the subsequent input pictures, if any, with index i+1 to numlnputPics-l, inclusive, are likewise missing. A missing input picture may be treated like described above in relation to nnpfc_absent_input_pic_zero_flag syntax element.

[0154] When NNPFC and NNPFA SEI messages are used for WC and a picture rate upsampling NNPF that interpolates pictures between a single pair of input pictures is activated persistently until the end of the bitstream, the NNPF is applied repeatedly at the end of the bitstream for different sets of input pictures up to but excluding a set of input pictures that would cause creation of any interpolated picture after the last picture of the bitstream in output order. In these sets of input pictures, some of the pictures may be missing and may be, for example, replaced by the last picture within the bitstream in output order.

[0155] The neural-network post-filter input (NNPFI) SEI message carries auxiliary input data to be provided to the inference of an associated NNPF.

[0156] An encoder may use the neural-network post-filter input (NNPFI) SEI message to provide auxiliary input data to an inference of a neural-network post-processing filter (NNPF) defined by NNPFC SEI message(s) and activated by NNPFA SEI message(s).

[0157] In an example, when there are multiple NNPFI SEI messages for the same target NNPF with different SEI payload content present in a picture unit, each occurrence of these NNPFI SEI message causes an inference of the target NNPF.

[0158] In an example, the NNPFI SEI message comprises an identifier or a counter value, wherein the identifier or the counter value identifies the NNPFI SEI message for the same target NNPF in the same picture unit. When the identifier or the counter value is the same as for a previous NNPFI SEI message for the same target NNPF in the same picture unit, the currentNNPFI SEI message may be concluded to be a copy of a previous NNPFI SEI message in the same picture unit.

[0159] In an example, an inference of the target NNPF is performed for each unique identifier or a counter value in the NNPFI SEI messages for the same target NNPF in the same picture unit.

[0160] The NNPFI SEI message may, for example, include one or more of the following syntax elements:

[0161] - nnpfi target id that associates the respective NNPF (with the same nnpfc id value) with this SEI message;

[0162] - nnpfi cnt that specifies an NNPFI SEI message instance count value for this nnpfi target id value within a picture unit;

[0163] - nnpfi auxiliary inp idc, which may be required to have the same values as nnpfc_auxiliary_inp_idc, and specifies which auxiliary inputs are provided to the NNPF inference.

[0164] In an example, the NNPF inference for a particular nnpfi target id value is performed in increasing nnpfi cnt order for each unique value of nnpfi cnt within a picture unit.

[0165] Features as described herein may generally relate to visual temporal extrapolation. In at least some embodiments of the present disclosure, the terms visual temporal extrapolation, temporal extrapolation, generative face video, and video prediction may be used interchangeably. Furthermore, the terms temporally extrapolating, generating, predicting may be used interchangeably in at least some embodiments. Visual temporal extrapolation may be defined as a method, algorithm, or process that generates one or more pictures in the future given one or more past pictures as input and one or more other data items. Visual temporal extrapolation may be realized by, but is not necessarily based on or limited to, neural network inference. Use cases for visual temporal extrapolation include (but are not limited to):

[0166] - Very low delay computer vision for domains like robotics and autonomous driving, where extrapolated future pictures facilitate anticipatory decision making.

[0167] - Increase of the rendered picture rate in very low-latency applications, such as cloud gaming, relative to the decoded picture rate.

[0168] - Reduction of the end-to-end delay in low-latency applications through extrapolating and displaying future pictures before they are received.

[0169] - Generative face video for very low bitrate video coding.

[0170] Features as described herein may generally relate to generative artificial intelligence (Al). The term generative Al, or generative modeling, or generative machine learning (and other similar terms), are commonly used to indicate a class of models trained with training data that are capable of generating new data. The new data may be characterized by a probability distribution that is same or different with respect to the probability distribution of the training data. State-of- the-art generative models are based on neural networks.

[0171] One typical example of neural network architecture that allows for generating text data is a Transformer-based “decoder”, where “decoder” may not refer to a decoder that is part of a codec performing compression of input data into a small bitstream. Instead, the decoder of a neural network architecture that processes text may be a neural network that gets / receives / obtains a set of input words, or parts of words, or tokens extracted from input words, and outputs a set of output words, or parts of words, or tokens. At inference time, such a NN is run in auto-regressive mode, where the generated word(s) or token(s) are provided as part of the input word(s) or token(s). In order for such a NN to generate data, it is trained to predict the next word(s) (or an estimate of a probability distribution over the next words) given a set of input words. The NN may be based on the Transformer architecture, which comprises the use of the self-attention mechanism, where an attention score is assigned to each input token or word based on all other input tokens or words, including the previously generated words or tokens. During training of a decoder-style Transformer architecture, the future data items (words or tokens) are masked so not to leak information from the future. In some cases, decoder-style Transformer architectures are referred to as “uni-directional” (because they use or process information from left-to-right), as opposed tosome encoder-style Transformer architectures that are referred to as “bi-directional” (because they use or process information from left-to-right and from right-to-left).

[0172] Another example of generative modeling is visual temporal extrapolation, where a picture is generated by a NN (or a system that comprises a NN) based on one or more previously decoded or generated pictures and on one or more other data items. The one or more previously decoded or generated pictures may be pictures decoded by a process that does not involve generative modeling, such as a traditional codec, e.g., a WC-compliant codec. In one example, the one or more data items may include parameters or features that describe differences between the one or more previously decoded pictures and the current picture to be temporally extrapolated, or between one or more previously generated pictures and the current picture to be temporally extrapolated. In another example, the one or more data items may include parameters or features that describe the position and / or orientation and / or other characteristics of one or more objects that are present in the one or more previously decoded pictures and / or in the current picture to be temporally extrapolated and / or one or more previously generated pictures. Examples of such parameters are facial parameters (such as facial landmarks and their positions or differential positions with respect to the facial landmarks of a previous picture), or parameters of other objects. The one or more data items or data derived from the one or more data items may be signaled from encoder to decoder, for example as part of a versatile supplemental enhancement information (VSEI) message.

[0173] FIG. 6 illustrates a generic block diagram of a generative face video codec. The input base picture 602 is encoded with any video encoder 604, such as a WC encoder, into a coded base picture of a bitstream 606. The input subsequent pictures, e.g., subsequent pictures 608-1, 608-2 are fed into the analysis model 610 for extracting feature parameters. The feature parameters are encoded 612 to generate a feature bitstream 614. Feature parameters may be encoded independently of feature parameters of any other picture, e.g., for the first subsequent picture following the base picture in decoding order. Alternatively, feature parameters may be encoded 612 in a predictive manner with reference to earlier encoded feature parameters in decoding order, wherein feature residual may be encoded relative to predicted feature parameters. Feature residual may undergo quantization and entropy coding as part of feature encoding. Encoded featureparameters are included in or along the feature bitstream 614, e.g., in a generative face video SEI message. For a subsequent picture, the video encoder may encode a dummy picture or a drive picture, as described in the subsequent section. Encoded feature parameters of a subsequent picture may be included in or along the respective coded dummy picture or drive picture, e.g., in the same picture unit.

[0174] The decoder 616 decodes a coded base picture from the bitstream 606. For a subsequent coded dummy or drive picture, the feature decoder 618 decodes encoded feature parameters. The decoded feature parameters and the decoded base picture and optionally a decoded drive picture are input to a generative neural network model inference 620, which generates an output picture sequence 622.

[0175] Features as described herein may generally relate to generative face video (GFV) SEI messages. In JVET, a GFV SEI message has been proposed for a new version of the VSEI standard, in order to support visual temporal extrapolation of faces in videos. One of the functions of the GFV SEI message is to signal facial parameters, which represent an auxiliary input to the NN or system that performs the temporal extrapolation. The latest document describing the GFV SEI message is JVET-AF0234. An approximate summary is provided as follows.

[0176] The GFV SEI message defines an interface between a video decoder and a generator NN that performs visual temporal extrapolation of faces in videos. The following are the types of pictures considered in the GFV SEI message:

[0177] - Base picture: a decoded output picture that may be used by the generative network to generate a novel face picture. It is coded by a traditional codec, such as WC-compliant codecs.

[0178] - Dummy picture: picture unit of the minimum allowed resolution that contains onlySEI messages.

[0179] - (optional) Primary picture (also called drive picture) that can be optionally input to a generative NN to improve background texture and / or facial details. This drive picture may also be coded by a traditional codec such as a WC-compliant codec.

[0180] A second neural network, referred to as Translator NN, is used to convert the facial parameters (signaled within a GFV SEI message) into the following converted parameters: 15 3D- keypoints; one 3x3 matrix; and one 1x3 matrix.

[0181] The input to the generator NN may be one or more of the following: the (previously decoded) base picture; the converted parameters (output of Translator NN); and / or (optionally) the primary / drive picture.

[0182] A technical effect of example embodiments of the present disclosure may be to enable the use of generative Al on a receiver or decoder for the purpose of generating a video based on information that is signaled by a transmitter or encoder.

[0183] A technical effect of example embodiments of the present disclosure may be to support the performance of visual temporal extrapolation or simply temporal extrapolation.

[0184] In an example embodiment, additional information, for example signaling information, may be included in a GFV SEI message. Additionally or alternatively, signaling information may be included in an NNPFC SEI message and / or an NNPFA SEI message. It may be noted that a NNPFC SEI message carries information about the characteristics of a NNPF. It may be noted that a NNPFA SEI message activates a NNPF for one or more pictures. It may be noted that a GFV SEI message activates a NN that performs temporal extrapolation of video frames (e.g. that contain faces).

[0185] In an example embodiment, an encoder may encode signaling information in or along a bitstream, and the decoder may decode signaling information from or along a bitstream, that enables or supports performing visual temporal extrapolation at the decoder side, where the performing of visual temporal extrapolation may comprise using one or more neural networks. In an example embodiment, the bitstream may include an encoded video, such as encoded motion vectors, encoded residual, and / or (optionally encoded) high-level syntax. When an SEI message (or any metadata) is encoded in the bitstream, it may then be part of the bitstream (i.e. in-band signaling). When an SEI message (or any metadata) is encoded along the bitstream, it may be a separate bitstream with respect to the bitstream of the encoded video (i.e. out-of-band signaling).

[0186] For simplicity, we assume that the visual temporal extrapolation may be performed by one neural network that is referred to as Visual Temporal extrapolation Neural Network (VTNN), or as generative NN, or as generator NN. However, it may be understood that example embodiments of the present disclosure may be similarly realized by a system that comprises one or more neural networks and one or more other operations, where the one or more neural networks may comprise an optical flow derivation neural network, followed by a synthesis network that takes an optical flow and one or more pictures as input, and where the one or more other operations may comprise computing a difference of parameters or features (e.g., a difference of facial parameters).

[0187] It may be understood that example embodiments of the present disclosure are not limited to any particular type of neural network. Example embodiments may, for example, be realized with a model that takes one or more seed pictures and one or more driving features as input. In one example, the model may be a generative model, such as a model based on Normalizing Flows. In another example, such a model may be based on Diffusion models (also known as diffusion probabilistic models, or score-based generative models). In yet another example, such a model may be based on a convolutional neural network.

[0188] In at least some example embodiments, the signaling information may be associated to one or more neural networks that perform at least part of the visual temporal extrapolation. In one example, a generative face video (GFV) SEI message may comprise one or more indications of the respective one or more neural networks to be used for temporal extrapolation. In another example, a GFV SEI message may comprise one or more indications of or one or more references to the respective one or more neural network post-filter characteristics (NNPFC) SEI messages that specify the one or more neural networks to be used for temporal extrapolation. In another example, a generative face video characteristics (GF VC) SEI message may comprise one or more indications of the respective one or more neural networks to be used for temporal extrapolation, where the GF VC SEI message may be an SEI message that specifies one or more characteristics or aspects of performing a temporal extrapolation.

[0189] In at least some example embodiments, the signaling information may be associated to one or more pictures of a video. This may be achieved by means of a frame-wise signaling mechanism. Such signaling is referred to as frame-wise (FW) signaling in at least some embodiments of the present disclosure. Examples of FW signaling include a neural network postfilter activation (NNPFA) SEI message, a neural network post-filter input (NNPFI) SEI message, a generative face video (GFV) SEI message, etc.

[0190] In an example embodiment, the signaling information may comprise information indicative of the types of inputs to be used for temporal extrapolation of one or more pictures. Examples of such types include, but are not limited to, the following:

[0191] - Base picture, e.g., a picture that was not generated by means of temporal extrapolation, such as the latest picture that was decoded by a WC-compliant decoder.

[0192] - A previous temporally extrapolated picture.

[0193] - A drive picture.

[0194] - The output of another neural network (such as the Translator NN in the case of GFVSEI messages), where the another neural network may be run at the decoder side and may convert input facial parameters (that were signaled from encoder to decoder) to output converted facial parameters.

[0195] - Facial parameters that may be signaled from encoder to decoder.

[0196] - Audio or speech data that may be signaled from encoder to decoder.

[0197] - Audio or speech data that is extracted at decoder side from an audio track associated to the visual track of a video.

[0198] - Text data that may be signaled from the encoder to the decoder.

[0199] - A tensor, where the tensor may be the output of a neural network that is run at the encoder side, and may be signaled from the encoder to the decoder. In one example, the tensormay represent facial parameters. In another example, the tensor may represent encoded facial parameters. In yet another example, the tensor may represent encoded difference of facial parameters. In yet another example, the tensor may represent parameters that are used to temporally extrapolate a picture or part of a picture. In yet another example, the tensor may represent motion of one or more portions of the picture with respect to another picture.

[0200] For simplicity, in some of the embodiments, the input types that are used to represent information that is specific to the picture to be extrapolated or to a portion of the picture to be extrapolated, such as facial parameters, motion information, audio data, tensor representing facial parameters, tensor representing other information of the picture to be extrapolated (such as motion, or parameters of other objects than faces), may be collectively referred to as “driving information” (which should not be confused with drive picture).

[0201] In one example, an SEI message, such as GFV SEI message, may be signaled from the encoder to the decoder, and may be associated to a certain picture to be temporally extrapolated. The SEI message may comprise a bit-field indicator indicating that the picture is to be temporally extrapolated by using, as inputs, the latest decoded base picture and a tensor representing facial parameters.

[0202] In another example, an SEI message, such as a GFV SEI message, may be signaled from the encoder to the decoder, and may be associated to a certain picture to be temporally extrapolated. The SEI message may comprise a bit-field indicator indicating that the picture is to be temporally extrapolated by using, as inputs, the latest decoded base picture and audio data that is extracted from the audio track and is temporally overlapping with the picture to be temporally extrapolated.

[0203] In yet another example, an SEI message, such as GFV SEI message, may be signaled from the encoder to the decoder, and may be associated to a certain picture to be temporally extrapolated. The SEI message may comprise a bit-field indicator indicating that the picture is to be temporally extrapolated by using, as inputs, the latest decoded base picture, audio data and a tensor representing facial parameters.

[0204] In yet another example, an SEI message, such as GFV SEI message, may be signaled from the encoder to the decoder, and may be associated to a certain picture to be temporally extrapolated. The SEI message may comprise a bit-field indicator indicating that the picture is to be temporally extrapolated by using, as inputs, the latest decoded base picture, a drive picture, and converted facial parameters that are output by a Translator NN.

[0205] In an example embodiment, alternative input types may be signaled. In an example embodiment, the signaling information may comprise information indicative of whether, when one or more information types are missing (i.e. unavailable, not available, etc.) or are not reliable (e.g., based on an indication of reliability of the information type(s)) at decoder side for one or more pictures to be temporally extrapolated, such as facial parameters or tensor representing facial parameters, the one or more pictures are to be temporally extrapolated based on one or more other information types. One or more information types may be missing for one or more pictures for various reasons. In one example, some information type may be missing because it is not signaled from encoder to decoder within a GFV SEI message (e.g. due to failure, corruption, decision to not signal, etc.).

[0206] In one example, the signaling information may be signaled in a GFVC SEI message or in a NNPFC SEI message, and it may comprise information indicating that, when facial parameters or tensors representing facial parameters or tensors representing motion for one or more pictures to be temporally extrapolated are missing (e.g., not signaled from encoder to decoder, or not received, etc.), the temporal extrapolation of the one or more pictures may be performed based on the latest decoded base frame and on audio data extracted from the audio track that temporally overlaps with the one or more pictures.

[0207] In an example embodiment, the signaling information, or a standard specification, may comprise information of which input types are considered as related or alternatives (e.g., comprising approximately similar information), and / or which input types are considered as unrelated or complementary.

[0208] In an example embodiment, when a first input type of two related or alternative input types is missing (i.e. unavailable, not available, etc.) or is not reliable (e.g., based on an indication of reliability), a second input type of the two related or alternative input types may be used for temporal extrapolation if the second input type is not missing (i.e. is available) and / or is reliable, (e.g. as input to a generator NN).

[0209] In an example, the input types facial parameters and tensor representing facial parameters may be considered as related or alternatives. When, for example, facial parameters are missing, a tensor derived from facial parameters may be used as input to a generator NN.

[0210] In another example, the input types facial parameters and audio data may be considered as related or alternatives. When, for example, facial parameters are missing, audio data may be used as input to a generator NN.

[0211] In yet another example, the input types facial parameters and base picture may be considered as unrelated or complementary. Therefore, a base picture may not be used as a replacement for missing facial parameters.

[0212] In an example embodiment, information about multiple generators may be signaled for a same picture. In an example embodiment, a picture may be temporally extrapolated by using two or more neural networks for two or more regions of the picture, where a NN in the two or more NNs may perform temporal extrapolation of one or more of the two or more regions.

[0213] In one example, a picture may be temporally extrapolated by using one neural network for each region of the picture, where region(s) may indicate a face, or at least part of a body of a person. Thus, different neural networks may be used for temporally extrapolating different persons in the picture.

[0214] In an example embodiment, the signaling information may comprise information indicating that, for one or more pictures, two or more indicated regions are to be extrapolated by two or more indicated neural networks.

[0215] In one example, a GFV SEI message may comprise indications of N regions (e.g., in terms of spatial and / or temporal coordinates) and indications for respective N neural networks to be used for temporally extrapolating the N regions. For example, the indications for respective N neural networks may comprise N references to respective N NNPFC SEI messages that specify or define the respective N neural networks.

[0216] In another example, a GFV SEI message may comprise indications for N neural networks and indications for M regions, where M is equal to or less than N, and where the association between each of the N neural networks and the associated region(s) may also be signaled within the GFV SEI message.

[0217] In an example embodiment, two or more SEI messages or other suitable signaling data structures may be signaled from the encoder to the decoder for one or more pictures to be temporally extrapolated, where the two or more SEI messages or other suitable signaling data structures may comprise signaling information associated with two or more regions in the one or more pictures, such as information that identifies a neural network or that identifies another SEI message (such as an NNPFC SEI message) which identifies a neural network.

[0218] In one example, a certain picture to be temporally extrapolated may comprise two faces that are to be temporally extrapolated by two different neural networks. Two GFV SEI messages may be associated with that picture, where a first GFV SEI message may comprise a reference to a first NNPFC SEI message that specifies a first neural network to be used to temporally extrapolate a first face in the picture, and where a second GFV SEI message may comprise a reference to a second NNPFC SEI message that may specify a second neural network to be used to temporally extrapolate a second face in the picture.

[0219] In an example embodiment, two or more regions of a picture to be temporally extrapolated may be temporally extrapolated by two or more neural networks that were updated or finetuned based on one or more initial or base neural networks.

[0220] In one example, two regions of a picture to be temporally extrapolated may be temporally extrapolated by two neural networks that are two finetuned versions of the same base neural network.

[0221] In an example embodiment, two or more regions of a picture to be temporally extrapolated may be temporally extrapolated by using a single neural network.

[0222] In another example embodiment, information about the two or more regions, such as driving information, may be input to respective two or more initial branches of the single neural network, and the single neural network may output respective two or more temporally extrapolated regions.

[0223] In another example embodiment, information about the two or more regions, such as driving information, may be stacked or concatenated across one or more axes of a tensor representing the input to the single neural network, and the single neural network may output respective two or more temporally extrapolated regions. In one example, facial information for two regions may be concatenated across the channel dimension of a tensor.

[0224] In another example embodiment, two or more regions may be temporally extrapolated by a neural network by means of performing multiple inferences, where the multiple inferences may be performed sequentially or in parallel. In one example embodiment, where the multiple inferences may be performed in parallel, information about the two or more regions, such as driving information, may be concatenated across the batch dimension of a tensor, and the single neural network may output respective two or more temporally extrapolated regions. In another example embodiment, where the multiple inferences may be performed sequentially (e.g., one inference after another inference), at each inference only information about one region may be input to the single neural network, and the single neural network may output a respective temporally extrapolated region.

[0225] In an example embodiment, two or more temporally extrapolated regions may be combined or blended by a combination or blending operation.

[0226] In another example embodiment, the combination or blending operation may also use a blending mask or alpha mask or alpha channel.

[0227] In another example embodiment, the blending mask or alpha mask or alpha channel may be generated at the decoder or the receiver side. In one example, the blending mask may be determined and output by the neural network that outputs temporally extrapolated region(s), for example as an auxiliary output with respect to the temporally extrapolated region(s).

[0228] In an example embodiment, uncertainty of facial parameters may be signaled. In an example embodiment, the signaling information may comprise information indicative of whether the driving information may comprise uncertain information. In an example embodiment, the signaling information may comprise an indication of certainty, reliability, and / or unreliability of the driving information. In an example embodiment, the signaling information may comprise information that may be used to determine how reliable or unreliable the driving information is. Driving information may be considered unreliable if the reliability is determined to be below a threshold value. Driving information may be considered to be reliable if the reliability is determined to be at or above a threshold value. The decoder may take further action based on the determined reliability / unreliability of the driving information. For example, if an input type is determined to be unreliable, an alternative input type (that may be determined to be reliable) may be used instead in the temporal extrapolation procedure.

[0229] In an example embodiment, a decoder may use the information about uncertainty of facial parameters to determine whether and / or how to use the facial parameters. In one example, a decoder may use the information about uncertainty of facial parameters to determine whether an input to the generator NN should be the facial parameters (or data derived from the facial parameters). In another example, a decoder may use the information about uncertainty of facial parameters to determine whether another input data type should be used as input to the generator NN, in addition to or as a replacement to the facial parameters.

[0230] In an example embodiment, the signaling information may comprise information indicative of whether the driving information was lossy compressed.

[0231] In another example embodiment, when the driving information was lossy compressed, the signaling information may comprise information indicative of a compression rate of the compressed driving information. In one example, the information indicative of a compression rate comprises information about a quantization step size used to quantize the driving information.

[0232] In an example embodiment, the signaling information may comprise information indicative of whether the driving information was lossless compressed.

[0233] In an example embodiment, the signaling information may comprise information indicative of the compression and / or decompression algorithm for compressing / decompressing the driving information.

[0234] In an example embodiment, the signaling information may comprise information indicative of one or more parameters of a compression and / or decompression algorithm for compressing / decompressing the driving information.

[0235] In an example embodiment, the signaling information may comprise information indicative of whether the driving information was determined based on information that does not comprise the frame to be temporally extrapolated. In one example, an encoder may determine facial parameters by extracting facial parameters from another picture than the picture to be extrapolated and then extrapolating those parameters to the picture to be extrapolated. In another example, an encoder may determine facial parameters by extracting facial parameters from a temporally extrapolated picture.

[0236] In an example, the uncertainty information comprises information that is determined based on an output of a neural network that determines the facial parameters. For example, the neural network that determines the facial parameters may output both the facial parameters and an estimate of the uncertainty of the facial parameters.

[0237] In an example embodiment, the signaling information may comprise information indicative of the type of content in a base or drive picture.

[0238] In an example embodiment, the signaling information may comprise information indicative of whether a base or drive picture comprises only background content.

[0239] In an example embodiment, the signaling information may comprise information indicative of whether a base or drive picture comprises only foreground content.

[0240] In an example embodiment, the signaling information may comprise information indicative of whether a base or a drive picture comprises both background and foreground content.

[0241] In one example, the foreground content may comprise faces, and the background content may comprise other regions or objects that are not faces (e.g. non-face objects).

[0242] In one example, the signaling information may comprise a bit-field indicator that indicates whether a base picture comprises only background content, only foreground content, or both background and foreground content.

[0243] In an example embodiment, an encoder may split or decompose a base picture or a drive picture into a first picture that comprises only background content, and a second picture that comprises only foreground content, and may encode both the first picture and the second picture, or may encode either the first picture or the second picture.

[0244] In an example embodiment, the quality of a base or drive picture may be signaled. In an example embodiment, a base picture or a drive picture may comprise content of different quality in terms of one or more quality metrics.

[0245] In an example embodiment, an encoder may determine a background content and a foreground content of a base picture or a drive picture, and may encode the background content and the foreground content with different qualities in terms of one or more quality metric or with different bitrates. In one example, the foreground content of a drive picture is encoded with higher PSNR than the background content of that drive picture. In another example, the foreground content of a drive picture is encoded with a lower rate-distortion loss than the background content of that drive picture.

[0246] In the present disclosure, the term Translator NN may be used to refer to a neural network that converts signaled facial parameters to converted facial parameters that comprise a format suitable to be input to a neural network (e.g., a generator NN) performing temporal extrapolation of picture(s).

[0247] In an example embodiment, information about the translator NN may be signaled. In an example embodiment, the signaling information may comprise information specifying a Translator NN only once every random access segment of a video sequence. In one example, a GFV SEI message may comprise an MPEG Neural Network compression and Representation (NNR) bitstream that represents an encoded Translator NN only when the GFV SEI message is associated to the first picture of a random access segment, and there may be two or more base pictures in one random access segment.

[0248] In an example embodiment, the signaling information may comprise information indicative of a reference to another SEI message that may identify or specify a Translator NN. In one example, a GFV SEI message may comprise a reference to an NNPFC SEI message that may specify a Translator NN, for example in terms of an NNR bitstream.

[0249] In an example embodiment, the signaling information may comprise information specifying a Translator NN only once for a Coded Layer Video Sequence (CLVS).

[0250] In an example embodiment, the definition of TranslatorNN is gated with a flag. The definition may be required to be present in one or more of the following cases:- when the GFV SEI message is the first GFV SEI message in a CLVS in decoding order;- when the GFV SEI message is the first GFV SEI message in a CLVS in output order;- when the GFV SEI message is included in a random access point (RAP) picture;- when the GFV SEI message is the first GFV SEI message that is present with, or follows (in decoder order or in output order), an RAP picture.

[0251] In an example embodiment, an encoder includes the definition of TranslatorNN in the first GFV SEI message in output order that is present in, or subsequent to, an RAP picture.

[0252] In one example, the GFV SEI message may comprise the following syntax elements:Where gfv_translator_nn_flag is the flag that gates the definition of the Translator NN. gfv translator nn flag equal to 1 specifies that the syntax elements defining the TranslatorNN are present in this SEI message. gfv_translator_nn_flag equal to 0 specifies that the syntax elements defining the TranslatorNN are not present in this SEI message. When gfv translator nn flag is not present, it is inferred to be equal to 0.When a GFV SEI message is the first GFV SEI message with a particular gfv id value within a CLVS, gfv translator nn flag shall be present and equal to 1. When a GFV SEI message is present in an IRAP picture unit, gfv translator nn flag shall be present and equal to 1.When gfv_translator_nn_flag is equal to 0 and TranslatorNN is referenced in the semantics of this SEI message, the applicable TranslatorNN is defined by the previous GFV SEI message with the same gfv id value and gfv translator nn flag equal to 1 in decoding order.

[0253] In an example embodiment, information about the Translator NN is signaled independently of whether the current picture is a base picture.

[0254] In one example, the GFV SEI message comprises the following syntax elements:where gfv_translator_nn_flag is the flag that gates the definition of the Translator NN.

[0255] In an example embodiment, the signaling information may comprise information indicative of whether the current picture is a drive picture.

[0256] In one example, the information indicative of whether the current picture is a drive picture comprises a flag, gfv_drive_pic_flag. The semantics of gfv_drive_pic_fusion_flagcontinue to indicate if the drive picture is used as input to GenerativeNN. The GFV SEI message comprises the following syntax elementsgfv_drive_pic flag equal to 1 specifies that the current decoded picture is a driving picture. gfv_drive_pic_flag equal to 0 specifies that the current decoded picture is not a driving picture. gfv_drive_pic_fusion_flag, when present, equal to 1 indicates the current decoded picture, which corresponds to a driving picture that may be used for fusion, the latest driving picture in output order for this gfv_id value may be input to GenerativeNN( ). gfv_drive_pic_fusion_flag equal to 0 indicates the current decoded picture the driving picture should not be input to GenerativeNN( ). When gfv_drive_pic_flag is equal to 1, gfv_drive_pic_fusion_flag is inferred to be equal to 1. When gfv_drive_pic_fusion_flagis present and equal to 1, it is a requirement of bitstream conformance that there shall be a GFV SEI message that has the same gfv id value and gfv_drive_pic_flag equal to 1 and precedes this SEI message in output order within the same CLVS.

[0257] In an example embodiment, the generator NN may use, as an input, a base picture which is not the latest decoded base picture (i.e. other than the latest decoded base picture).

[0258] In an example embodiment, the signaling information may comprise information indicative of a base picture to be used as an input to the generator NN. In one example, the information indicative of the base picture comprises a POC number that identifies a base picture. In another example, a first GFV SEI message associated to a base picture comprises a first identifier of that base picture; a second GFV SEI message associated to a picture to be extrapolated comprises a second identifier that refers to the first identifier. In yet another example, the information indicative of the base picture refers to an identifier that identifies a GFV SEI message that is associated to a base picture.

[0259] In an example embodiment, the generator NN may use, as an input, a drive picture which is not the latest decoded drive picture (i.e. other than the latest decoded drive picture).

[0260] In an example embodiment, the signaling information may comprise information indicative of a drive picture to be used as an input to the generator NN. In one example, the information indicative of the drive picture comprises a POC number that identifies a drive picture. In another example, a first GFV SEI message associated to a drive picture comprises a first identifier of that drive picture; a second GFV SEI message associated to a picture to be extrapolated comprises a second identifier that refers to the first identifier. In yet another example, the information indicative of the drive picture refers to an identifier that identifies a GFV SEI message that is associated to a drive picture.

[0261] In an example embodiment, the signaling information may comprise an update to one or more neural networks that are part of the generator NN or of the system that performs temporal extrapolation. In one example, the update comprises a parameter update, e.g., an update to one ormore parameters (also referred to as weights) of the one or more neural networks that are part of the generator NN.

[0262] In an example embodiment, the signaling information may comprise one or more adaptation signals or adaptation tensors to be used for adapting respective one or more signals or tensors that are input to the generator NN or that are internally produced by the generator NN or that are output by the generator NN. In one example, one or more adaptation tensors multiply the outputs of respective one or more layers of the generator NN.

[0263] FIG. 4 illustrates the potential steps of an example method 400. The example method 400 may include: generating signaling information for performing a post-processing operation, wherein the post-processing operation comprises use of at least one neural network, 410; and providing the signaling information for performing the post-processing operation, 420. The example method 400 may be performed, for example, with an encoder, a codec, a device including at least one NN, one or more network elements, etc. It may be noted that, instead of generating the signaling information, as at 410, the signaling information may optionally be obtained from another entity, device, module, neural network, task, etc.

[0264] FIG. 5 illustrates the potential steps of an example method 500. The example method 500 may include: receiving signaling information for performing a post-processing operation, wherein the post-processing operation comprises use of at least one neural network, 510; receiving at least one input, 520; and providing the at least one input to the at least one neural network based, at least partially, on the signaling information, 530. The example method 500 may be performed, for example, with a decoder, a codec, a device including at least one NN, one or more network elements, etc.

[0265] Some embodiments have been described in relation to specific SEI messages, such as the GFV SEI message, NNPFC SEI message, and / or NNPFA SEI message. It is to be understood that embodiments are not limited to these specific SEI messages but can be realized with any similar SEI messages.

[0266] Some embodiments have been described in relation to SEI messages. It is to be understood that embodiments are not limited to SEI messages but can be realized with any similar syntax structures, such as metadata OBUs.

[0267] Some embodiments have been described in relation to certain terms, such as picture unit, applicable to some video coding formats, such as WC. It is to be understood that embodiments are not limited to video coding formats where the terms are applicable but can be realized with any video coding formats with their respective terms.

[0268] In the above, some example embodiments have been described with the help of syntax of the bitstream. It needs to be understood, however, that the corresponding structure and / or computer program may reside at the encoder for generating the bitstream and / or at the decoder for decoding the bitstream.

[0269] In the above, where example embodiments have been described with reference to an encoder, it needs to be understood that the resulting bitstream and the decoder have corresponding elements in them. Likewise, where example embodiments have been described with reference to a decoder, it needs to be understood that the encoder has structure and / or computer program for generating the bitstream to be decoded by the decoder.

[0270] In accordance with one example embodiment, an apparatus may comprise: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: generate signaling information for performing a postprocessing operation, wherein the post-processing operation may comprise use of at least one neural network; and provide the signaling information for performing the post-processing operation.

[0271] The post-processing operation may comprise at least one of: temporal extrapolation, or visual temporal extrapolation.

[0272] The signaling information may be associated with one or more pictures of a video.

[0273] The signaling information may be provided with at least one of: at least one neural network post-filter characteristic supplemental enhancement information message, at least one neural network post-filter activation supplemental enhancement information message, at least one generative face video supplemental enhancement information message, at least one generative face video characteristics supplemental enhancement information message, or at least one neural network post-filter input supplemental enhancement information message.

[0274] The signaling information may comprise an indication of at least one of: at least one type of input to the post-processing operation, one or more types of input that are related to the at least one type of input, one or more types of input that are unrelated to the at least one type of input, at least one alternative type of input to the post-processing operation, at least one update for the at least one neural network, at least one adaptation signal or tensor for adapting at least one input to the at least one neural network, a plurality of neural networks the post-processing operation uses, a plurality of neural network post-filter characteristics supplemental enhancement information messages respectively indicating the plurality of neural networks, a plurality of regions of an input for which a same neural network performs the post-processing operation, a neural network post-filter characteristics supplemental enhancement information message configured to indicate to same neural network, an uncertainty associated with the at least one type of input, at least one type of content in at least one base picture, at least one type of content in at least one drive picture, a quality of the at least one type of content in the at least one base picture, a quality of the at least one type of content in the at least one drive picture, whether a current picture is a drive picture, a translator neural network, a reference to a supplemental enhancement information message that comprises the indication of the translator neural network, a neural network post-filter characteristics supplemental enhancement information message configured to indicate the at least one neural network, or at least one picture to be temporally extrapolated using the post-processing operation.

[0275] The at least one alternative type of input to the post-processing operation may comprise one or more types of input for use in response to a determination that the at least one type of input to the post-processing operation is not available or is not reliable.

[0276] The at least one type of input to the post-processing operation may comprise at least one of: a facial parameter, a tensor representing at least one facial parameter, a tensor representing at least one encoded facial parameter, a tensor representing motion of a portion of a first picture relative to a second picture, a tensor representing a difference between facial parameters, a tensor representing a parameter of a non-face object, a tensor representing motion of the non-face object, motion information, audio data, speech data, text data, a base picture, a drive picture, a previous temporally extrapolated picture, or an output of a further neural network.

[0277] The indication of uncertainty associated with the at least one type of input may comprise at least one of: an indication that lossy compression was used for the at least one type of input, an indication that lossless compression was used for the at least one type of input, an indication of a compression algorithm used for compressing the at least one type of input, an indication of a compression rate used for the at least one type of input, an indication of a quantization step size used for the at least one type of input, an indication of a decompression algorithm used for decompressing the at least one type of input, an indication of at least one parameter for compressing the at least one type of input, an indication of at least one parameter for decompressing the at least one type of input, or an indication of a source of the at least one type of input.

[0278] The at least one type of content in the at least one base picture, or the at least one type of content in the at least one drive picture, may comprise at least one of: background content, foreground content, or the background content and the foreground content.

[0279] The background content may comprise one or more non-face regions and / or one or more non-face objects.

[0280] The foreground content may comprise one or more face objects.

[0281] The plurality of neural networks may comprise a neural network associated with a respective region, of two or more regions, of an input picture.

[0282] A number of the plurality of neural networks may be equal to a number of the two or more regions.

[0283] A number of the plurality of neural networks may be greater than a number of the two or more regions.

[0284] The signaling information may comprise, at least, information, associated with respective ones of the two or more regions, identifying respective ones of the plurality of neural networks.

[0285] The indication of the translator neural network may be provided for a random access segment of a video sequence.

[0286] The indication of the translator neural network may be provided for a coded layer video sequence.

[0287] The indication of the translator neural network may be gated with a flag.

[0288] The signaling information may comprise the indication of the translator neural network, wherein the indication of the translator neural network may be provided in: a generative face video supplemental enhancement information message which is the first, in output order, among one or more generative face video supplemental enhancement information messages at or subsequent to a random access point picture.

[0289] The signaling information may comprise the indication of the translator neural network independently of a determination of whether the current picture is a base picture.

[0290] In accordance with one aspect, an example method may be provided comprising: generating, with an encoder, signaling information for performing a post-processing operation, wherein the post-processing operation may comprise use of at least one neural network; and providing the signaling information for performing the post-processing operation.

[0291] The post-processing operation may comprise at least one of: temporal extrapolation, or visual temporal extrapolation.

[0292] The signaling information may be associated with one or more pictures of a video.

[0293] The signaling information may be provided with at least one of: at least one neural network post-filter characteristic supplemental enhancement information message, at least one neural network post-filter activation supplemental enhancement information message, at least one generative face video supplemental enhancement information message, at least one generative face video characteristics supplemental enhancement information message, or at least one neural network post-filter input supplemental enhancement information message.

[0294] The signaling information may comprise an indication of at least one of: at least one type of input to the post-processing operation, one or more types of input that are related to the at least one type of input, one or more types of input that are unrelated to the at least one type of input, at least one alternative type of input to the post-processing operation, at least one update for the at least one neural network, at least one adaptation signal or tensor for adapting at least one input to the at least one neural network, a plurality of neural networks the post-processing operation uses, a plurality of neural network post-filter characteristics supplemental enhancement information messages respectively indicating the plurality of neural networks, a plurality of regions of an input for which a same neural network performs the post-processing operation, a neural network post-filter characteristics supplemental enhancement information message configured to indicate to same neural network, an uncertainty associated with the at least one type of input, at least one type of content in at least one base picture, at least one type of content in at least one drive picture, a quality of the at least one type of content in the at least one base picture, a quality of the at least one type of content in the at least one drive picture, whether a current picture is a drive picture, a translator neural network, a reference to a supplemental enhancement information message that comprises the indication of the translator neural network, a neural network post-filter characteristics supplemental enhancement information message configured to indicate the at least one neural network, or at least one picture to be temporally extrapolated using the post-processing operation.

[0295] The at least one alternative type of input to the post-processing operation may comprise one or more types of input for use in response to a determination that the at least one type of input to the post-processing operation is not available or is not reliable.

[0296] The at least one type of input to the post-processing operation may comprise at least one of: a facial parameter, a tensor representing at least one facial parameter, a tensor representing at least one encoded facial parameter, a tensor representing motion of a portion of a first picture relative to a second picture, a tensor representing a difference between facial parameters, a tensor representing a parameter of a non-face object, a tensor representing motion of the non-face object, motion information, audio data, speech data, text data, a base picture, a drive picture, a previous temporally extrapolated picture, or an output of a further neural network.

[0297] The indication of uncertainty associated with the at least one type of input may comprise at least one of: an indication that lossy compression was used for the at least one type of input, an indication that lossless compression was used for the at least one type of input, an indication of a compression algorithm used for compressing the at least one type of input, an indication of a compression rate used for the at least one type of input, an indication of a quantization step size used for the at least one type of input, an indication of a decompression algorithm used for decompressing the at least one type of input, an indication of at least one parameter for compressing the at least one type of input, an indication of at least one parameter for decompressing the at least one type of input, or an indication of a source of the at least one type of input.

[0298] The at least one type of content in the at least one base picture, or the at least one type of content in the at least one drive picture, may comprise at least one of: background content, foreground content, or the background content and the foreground content.

[0299] The background content may comprise one or more non-face regions and / or one or more non-face objects.

[0300] The foreground content may comprise one or more face objects.

[0301] The plurality of neural networks may comprise a neural network associated with a respective region, of two or more regions, of an input picture.

[0302] A number of the plurality of neural networks may be equal to a number of the two or more regions.

[0303] A number of the plurality of neural networks may be greater than a number of the two or more regions.

[0304] The signaling information may comprise, at least, information, associated with respective ones of the two or more regions, identifying respective ones of the plurality of neural networks.

[0305] The indication of the translator neural network may be provided for a random access segment of a video sequence.

[0306] The indication of the translator neural network may be provided for a coded layer video sequence.

[0307] The indication of the translator neural network may be gated with a flag.

[0308] The signaling information may comprise the indication of the translator neural network, wherein the indication of the translator neural network may be provided in: a generative face video supplemental enhancement information message which is the first, in output order, among one or more generative face video supplemental enhancement information messages at or subsequent to a random access point picture.

[0309] The signaling information may comprise the indication of the translator neural network independently of a determination of whether the current picture is a base picture.

[0310] In accordance with one example embodiment, an apparatus may comprise: circuitry configured to perform: generating, with an encoder, signaling information for performing a postprocessing operation, wherein the post-processing operation may comprise use of at least oneneural network; and circuitry configured to perform: providing the signaling information for performing the post-processing operation.

[0311] In accordance with one example embodiment, an apparatus may comprise: processing circuitry; memory circuitry including computer program code, the memory circuitry and the computer program code configured to, with the processing circuitry, enable the apparatus to: generate signaling information for performing a post-processing operation, wherein the postprocessing operation may comprise use of at least one neural network; and provide the signaling information for performing the post-processing operation.

[0312] As used in this application, the term “circuitry” or “means” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry) and (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.” This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.

[0313] In accordance with one example embodiment, an apparatus may comprise means for: generating signaling information for performing a post-processing operation, wherein the postprocessing operation may comprise use of at least one neural network; and providing the signaling information for performing the post-processing operation.

[0314] The post-processing operation may comprise at least one of: temporal extrapolation, or visual temporal extrapolation.

[0315] The signaling information may be associated with one or more pictures of a video.

[0316] The signaling information may be provided with at least one of: at least one neural network post-filter characteristic supplemental enhancement information message, at least one neural network post-filter activation supplemental enhancement information message, at least one generative face video supplemental enhancement information message, at least one generative face video characteristics supplemental enhancement information message, or at least one neural network post-filter input supplemental enhancement information message.

[0317] The signaling information may comprise an indication of at least one of: at least one type of input to the post-processing operation, one or more types of input that are related to the at least one type of input, one or more types of input that are unrelated to the at least one type of input, at least one alternative type of input to the post-processing operation, at least one update for the at least one neural network, at least one adaptation signal or tensor for adapting at least one input to the at least one neural network, a plurality of neural networks the post-processing operation uses, a plurality of neural network post-filter characteristics supplemental enhancement information messages respectively indicating the plurality of neural networks, a plurality of regions of an input for which a same neural network performs the post-processing operation, a neural network post-filter characteristics supplemental enhancement information message configured to indicate to same neural network, an uncertainty associated with the at least one type of input, at least one type of content in at least one base picture, at least one type of content in at least one drive picture, a quality of the at least one type of content in the at least one base picture, a quality of the at least one type of content in the at least one drive picture, whether a current picture is a drive picture, a translator neural network, a reference to a supplemental enhancement information message that comprises the indication of the translator neural network, a neural network post-filter characteristics supplemental enhancement information message configured to indicate the at least one neural network, or at least one picture to be temporally extrapolated using the post-processing operation.

[0318] The at least one alternative type of input to the post-processing operation may comprise one or more types of input for use in response to a determination that the at least one type of input to the post-processing operation is not available or is not reliable.

[0319] The at least one type of input to the post-processing operation may comprise at least one of: a facial parameter, a tensor representing at least one facial parameter, a tensor representing at least one encoded facial parameter, a tensor representing motion of a portion of a first picture relative to a second picture, a tensor representing a difference between facial parameters, a tensor representing a parameter of a non-face object, a tensor representing motion of the non-face object, motion information, audio data, speech data, text data, a base picture, a drive picture, a previous temporally extrapolated picture, or an output of a further neural network.

[0320] The indication of uncertainty associated with the at least one type of input may comprise at least one of: an indication that lossy compression was used for the at least one type of input, an indication that lossless compression was used for the at least one type of input, an indication of a compression algorithm used for compressing the at least one type of input, an indication of a compression rate used for the at least one type of input, an indication of a quantization step size used for the at least one type of input, an indication of a decompression algorithm used for decompressing the at least one type of input, an indication of at least one parameter for compressing the at least one type of input, an indication of at least one parameter for decompressing the at least one type of input, or an indication of a source of the at least one type of input.

[0321] The at least one type of content in the at least one base picture, or the at least one type of content in the at least one drive picture, may comprise at least one of: background content, foreground content, or the background content and the foreground content.

[0322] The background content may comprise one or more non-face regions and / or one or more non-face objects.

[0323] The foreground content may comprise one or more face objects.

[0324] The plurality of neural networks may comprise a neural network associated with a respective region, of two or more regions, of an input picture.

[0325] A number of the plurality of neural networks may be equal to a number of the two or more regions.

[0326] A number of the plurality of neural networks may be greater than a number of the two or more regions.

[0327] The signaling information may comprise, at least, information, associated with respective ones of the two or more regions, identifying respective ones of the plurality of neural networks.

[0328] The indication of the translator neural network may be provided for a random access segment of a video sequence.

[0329] The indication of the translator neural network may be provided for a coded layer video sequence.

[0330] The indication of the translator neural network may be gated with a flag.

[0331] The signaling information may comprise the indication of the translator neural network, wherein the indication of the translator neural network may be provided in: a generative face video supplemental enhancement information message which is the first, in output order, among one or more generative face video supplemental enhancement information messages at or subsequent to a random access point picture.

[0332] The signaling information may comprise the indication of the translator neural network independently of a determination of whether the current picture is a base picture.

[0333] A processor, memory, and / or example algorithms (which may be encoded as instructions, program, or code) may be provided as example means for providing or causing performance of operation.

[0334] In accordance with one example embodiment, a non-transitory computer-readable medium comprising instructions stored thereon which, when executed with at least one processor, cause the at least one processor to: generate signaling information for performing a post-processing operation, wherein the post-processing operation may comprise use of at least one neural network; and cause providing of the signaling information for performing the post-processing operation.

[0335] In accordance with one example embodiment, a non-transitory computer-readable medium comprising program instructions stored thereon for performing at least the following: generating signaling information for performing a post-processing operation, wherein the postprocessing operation may comprise use of at least one neural network; and causing providing of the signaling information for performing the post-processing operation.

[0336] In accordance with another example embodiment, a non-transitory program storage device readable by a machine may be provided, tangibly embodying instructions executable by the machine for performing operations, the operations comprising: generating signaling information for performing a post-processing operation, wherein the post-processing operation may comprise use of at least one neural network; and causing providing of the signaling information for performing the post-processing operation.

[0337] In accordance with another example embodiment, a non-transitory computer-readable medium comprising instructions that, when executed by an apparatus, cause the apparatus to perform at least the following: generating signaling information for performing a post-processing operation, wherein the post-processing operation may comprise use of at least one neural network; and causing providing of the signaling information for performing the post-processing operation.

[0338] A computer implemented system comprising: at least one processor and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the system at least to perform: generating signaling information for performing a post-processing operation, wherein the post-processing operation may comprise use of at least one neural network; and causing providing of the signaling information for performing the post-processing operation.

[0339] A computer implemented system comprising: means for generating signaling information for performing a post-processing operation, wherein the post-processing operation may comprise use of at least one neural network; and means for causing providing of the signaling information for performing the post-processing operation.

[0340] In accordance with one example embodiment, an apparatus may comprise: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: receive signaling information for performing a postprocessing operation, wherein the post-processing operation may comprise use of at least one neural network; receive at least one input; and provide the at least one input to the at least one neural network based, at least partially, on the signaling information.

[0341] The post-processing operation may comprise at least one of: temporal extrapolation, or visual temporal extrapolation.

[0342] The at least one input may comprise, at least, one or more pictures of a video.

[0343] The signaling information may be received via at least one of: at least one neural network post-filter characteristic supplemental enhancement information message, at least one neural network post-filter activation supplemental enhancement information message, at least one generative face video supplemental enhancement information message, at least one generative face video characteristics supplemental enhancement information message, or at least one neural network post-filter input supplemental enhancement information message.

[0344] The signaling information may comprise an indication of at least one of: at least one type of input to the post-processing operation, one or more types of input that are related to the at least one type of input, one or more types of input that are unrelated to the at least one type of input, at least one alternative type of input to the post-processing operation, at least one update for the at least one neural network, at least one adaptation signal or tensor for adapting at least one input to the at least one neural network, a plurality of neural networks the post-processing operation uses, a plurality of neural network post-filter characteristics supplemental enhancement information messages respectively indicating the plurality of neural networks, a plurality ofregions of an input for which a same neural network performs the post-processing operation, a neural network post-filter characteristics supplemental enhancement information message configured to indicate to same neural network, an uncertainty associated with the at least one type of input, at least one type of content in at least one base picture, at least one type of content in at least one drive picture, a quality of the at least one type of content in the at least one base picture, a quality of the at least one type of content in the at least one drive picture, whether a current picture is a drive picture, a translator neural network, a reference to a supplemental enhancement information message that comprises the indication of the translator neural network, a neural network post-filter characteristics supplemental enhancement information message configured to indicate the at least one neural network, or at least one picture to be temporally extrapolated using the post-processing operation.

[0345] The at least one alternative type of input to the post-processing operation may comprise one or more types of input for use in response to a determination that the at least one type of input to the post-processing operation is not available or is not reliable.

[0346] The at least one type of input to the post-processing operation may comprise at least one of: a facial parameter, a tensor representing at least one facial parameter, a tensor representing at least one encoded facial parameter, a tensor representing motion of a portion of a first picture relative to a second picture, a tensor representing a difference between facial parameters, a tensor representing a parameter of a non-face object, a tensor representing motion of the non-face object, motion information, audio data, speech data, text data, a base picture, a drive picture, a previous temporally extrapolated picture, or an output of a further neural network.

[0347] The example apparatus may be further configured to: provide one or more input facial parameters to the further neural network; and obtain the output of the further neural network, wherein the output of the further neural network may comprise one or more converted facial parameters.

[0348] The further neural network may comprise the translator neural network.

[0349] The example apparatus may be further configured to: extract the audio data or the speech data from an audio track of the at least one input.

[0350] The indication of uncertainty associated with the at least one type of input may comprise at least one of: an indication that lossy compression was used for the at least one type of input, an indication that lossless compression was used for the at least one type of input, an indication of a compression algorithm used for compressing the at least one type of input, an indication of a compression rate used for the at least one type of input, an indication of a quantization step size used for the at least one type of input, an indication of a decompression algorithm used for decompressing the at least one type of input, an indication of at least one parameter for compressing the at least one type of input, an indication of at least one parameter for decompressing the at least one type of input, or an indication of a source of the at least one type of input.

[0351] The at least one type of content in the at least one base picture, or the at least one type of content in the at least one drive picture, may comprise at least one of: background content, foreground content, or the background content and the foreground content.

[0352] The background content may comprise one or more non-face regions and / or one or more non-face objects.

[0353] The foreground content may comprise one or more face objects.

[0354] The plurality of neural networks may comprise a neural network associated with a respective region, of two or more regions, of an input picture.

[0355] A number of the plurality of neural networks may be equal to a number of the two or more regions.

[0356] A number of the plurality of neural networks may be greater than a number of the two or more regions.

[0357] The signaling information may comprise, at least, information, associated with respective ones of the two or more regions, identifying respective ones of the plurality of neural networks.

[0358] The indication of the translator neural network may be received for a random access segment of a video sequence.

[0359] The indication of the translator neural network may be received for a coded layer video sequence.

[0360] The signaling information may be received in or a long a bitstream.

[0361] The example apparatus may be further configured to: determine that one or more of the at least one type of input is unavailable; determine one or more of the at least one alternative type of input that can be used instead of the one or more types of input that are unavailable; and provide, at least, the one or more alternative types of input to the at least one neural network.

[0362] The indication of the translator neural network may be gated with a flag.

[0363] The signaling information may comprise the indication of the translator neural network, wherein the indication of the translator neural network may be received in: a generative face video supplemental enhancement information message which is the first, in output order, among one or more generative face video supplemental enhancement information messages at or subsequent to a random access point picture.

[0364] The signaling information may comprise the indication of the translator neural network independently of a determination of whether the current picture is a base picture.

[0365] The at least one neural network may comprise, at least, a generator neural network, wherein a base picture, other than a latest decoded base picture, may be provided to the generator neural network.

[0366] The signaling information may comprise, at least, an indication of the base picture.

[0367] The at least one neural network may comprise, at least, a generator neural network, wherein a drive picture, other than a latest decoded drive picture, may be provided to the generator neural network.

[0368] The signaling information may comprise, at least, an indication of the drive picture.

[0369] In accordance with one aspect, an example method may be provided comprising: receiving, with a decoder, signaling information for performing a post-processing operation, wherein the post-processing operation may comprise use of at least one neural network; receiving at least one input; and providing the at least one input to the at least one neural network based, at least partially, on the signaling information.

[0370] The post-processing operation may comprise at least one of: temporal extrapolation, or visual temporal extrapolation.

[0371] The at least one input may comprise, at least, one or more pictures of a video.

[0372] The signaling information may be received via at least one of: at least one neural network post-filter characteristic supplemental enhancement information message, at least one neural network post-filter activation supplemental enhancement information message, at least one generative face video supplemental enhancement information message, at least one generative face video characteristics supplemental enhancement information message, or at least one neural network post-filter input supplemental enhancement information message.

[0373] The signaling information may comprise an indication of at least one of: at least one type of input to the post-processing operation, one or more types of input that are related to the at least one type of input, one or more types of input that are unrelated to the at least one type of input, at least one alternative type of input to the post-processing operation, at least one update for the at least one neural network, at least one adaptation signal or tensor for adapting at least one input to the at least one neural network, a plurality of neural networks the post-processing operation uses, a plurality of neural network post-filter characteristics supplemental enhancement information messages respectively indicating the plurality of neural networks, a plurality ofregions of an input for which a same neural network performs the post-processing operation, a neural network post-filter characteristics supplemental enhancement information message configured to indicate to same neural network, an uncertainty associated with the at least one type of input, at least one type of content in at least one base picture, at least one type of content in at least one drive picture, a quality of the at least one type of content in the at least one base picture, a quality of the at least one type of content in the at least one drive picture, whether a current picture is a drive picture, a translator neural network, a reference to a supplemental enhancement information message that comprises the indication of the translator neural network, a neural network post-filter characteristics supplemental enhancement information message configured to indicate the at least one neural network, or at least one picture to be temporally extrapolated using the post-processing operation.

[0374] The at least one alternative type of input to the post-processing operation comprises one or more types of input for use in response to a determination that the at least one type of input to the post-processing operation is not available or is not reliable.

[0375] The at least one type of input to the post-processing operation may comprise at least one of: a facial parameter, a tensor representing at least one facial parameter, a tensor representing at least one encoded facial parameter, a tensor representing motion of a portion of a first picture relative to a second picture, a tensor representing a difference between facial parameters, a tensor representing a parameter of a non-face object, a tensor representing motion of the non-face object, motion information, audio data, speech data, text data, a base picture, a drive picture, a previous temporally extrapolated picture, or an output of a further neural network.

[0376] The example method may further comprise: providing one or more input facial parameters to the further neural network; and obtaining the output of the further neural network, wherein the output of the further neural network may comprise one or more converted facial parameters.

[0377] The further neural network may comprise the translator neural network.

[0378] The example method may further comprise: extracting the audio data or the speech data from an audio track of the at least one input.

[0379] The indication of uncertainty associated with the at least one type of input may comprise at least one of: an indication that lossy compression was used for the at least one type of input, an indication that lossless compression was used for the at least one type of input, an indication of a compression algorithm used for compressing the at least one type of input, an indication of a compression rate used for the at least one type of input, an indication of a quantization step size used for the at least one type of input, an indication of a decompression algorithm used for decompressing the at least one type of input, an indication of at least one parameter for compressing the at least one type of input, an indication of at least one parameter for decompressing the at least one type of input, or an indication of a source of the at least one type of input.

[0380] The at least one type of content in the at least one base picture, or the at least one type of content in the at least one drive picture, may comprise at least one of: background content, foreground content, or the background content and the foreground content.

[0381] The background content may comprise one or more non-face regions and / or one or more non-face objects.

[0382] The foreground content may comprise one or more face objects.

[0383] The plurality of neural networks may comprise a neural network associated with a respective region, of two or more regions, of an input picture.

[0384] A number of the plurality of neural networks may be equal to a number of the two or more regions.

[0385] A number of the plurality of neural networks may be greater than a number of the two or more regions.

[0386] The signaling information may comprise, at least, information, associated with respective ones of the two or more regions, identifying respective ones of the plurality of neural networks.

[0387] The indication of the translator neural network may be received for a random access segment of a video sequence.

[0388] The indication of the translator neural network may be received for a coded layer video sequence.

[0389] The signaling information may be received in or along a bitstream.

[0390] The example method may further comprise: determining that one or more of the at least one type of input is unavailable; determining one or more of the at least one alternative type of input that can be used instead of the one or more types of input that are unavailable; and providing, at least, the one or more alternative types of input to the at least one neural network.

[0391] The indication of the translator neural network may be gated with a flag.

[0392] The signaling information may comprise the indication of the translator neural network, wherein the indication of the translator neural network may be received in: a generative face video supplemental enhancement information message which is the first, in output order, among one or more generative face video supplemental enhancement information messages at or subsequent to a random access point picture.

[0393] The signaling information may comprise the indication of the translator neural network independently of a determination of whether the current picture is a base picture.

[0394] The at least one neural network may comprise, at least, a generator neural network, wherein a base picture, other than a latest decoded base picture, may be provided to the generator neural network.

[0395] The signaling information may comprise, at least, an indication of the base picture.

[0396] The at least one neural network may comprise, at least, a generator neural network, wherein a drive picture, other than a latest decoded drive picture, may be provided to the generator neural network.

[0397] The signaling information may comprise, at least, an indication of the drive picture.

[0398] In accordance with one example embodiment, an apparatus may comprise: circuitry configured to perform: receiving, with a decoder, signaling information for performing a postprocessing operation, wherein the post-processing operation may comprise use of at least one neural network; circuitry configured to perform: receiving at least one input; and circuitry configured to perform: providing the at least one input to the at least one neural network based, at least partially, on the signaling information.

[0399] In accordance with one example embodiment, an apparatus may comprise: processing circuitry; memory circuitry including computer program code, the memory circuitry and the computer program code configured to, with the processing circuitry, enable the apparatus to: receive signaling information for performing a post-processing operation, wherein the postprocessing operation may comprise use of at least one neural network; receive at least one input; and provide the at least one input to the at least one neural network based, at least partially, on the signaling information.

[0400] In accordance with one example embodiment, an apparatus may comprise means for: receiving signaling information for performing a post-processing operation, wherein the postprocessing operation may comprise use of at least one neural network; receiving at least one input; and providing the at least one input to the at least one neural network based, at least partially, on the signaling information.

[0401] The post-processing operation may comprise at least one of: temporal extrapolation, or visual temporal extrapolation.

[0402] The at least one input may comprise, at least, one or more pictures of a video.

[0403] The signaling information may be received via at least one of: at least one neural network post-filter characteristic supplemental enhancement information message, at least one neural network post-filter activation supplemental enhancement information message, at least one generative face video supplemental enhancement information message, at least one generative face video characteristics supplemental enhancement information message, or at least one neural network post-filter input supplemental enhancement information message.

[0404] The signaling information may comprise an indication of at least one of: at least one type of input to the post-processing operation, one or more types of input that are related to the at least one type of input, one or more types of input that are unrelated to the at least one type of input, at least one alternative type of input to the post-processing operation, at least one update for the at least one neural network, at least one adaptation signal or tensor for adapting at least one input to the at least one neural network, a plurality of neural networks the post-processing operation uses, a plurality of neural network post-filter characteristics supplemental enhancement information messages respectively indicating the plurality of neural networks, a plurality of regions of an input for which a same neural network performs the post-processing operation, a neural network post-filter characteristics supplemental enhancement information message configured to indicate to same neural network, an uncertainty associated with the at least one type of input, at least one type of content in at least one base picture, at least one type of content in at least one drive picture, a quality of the at least one type of content in the at least one base picture, a quality of the at least one type of content in the at least one drive picture, whether a current picture is a drive picture, a translator neural network, a reference to a supplemental enhancement information message that comprises the indication of the translator neural network, a neural network post-filter characteristics supplemental enhancement information message configured to indicate the at least one neural network, or at least one picture to be temporally extrapolated using the post-processing operation.

[0405] The at least one alternative type of input to the post-processing operation may comprise one or more types of input for use in response to a determination that the at least one type of input to the post-processing operation is not available or is not reliable.

[0406] The at least one type of input to the post-processing operation may comprise at least one of: a facial parameter, a tensor representing at least one facial parameter, a tensor representing at least one encoded facial parameter, a tensor representing motion of a portion of a first picture relative to a second picture, a tensor representing a difference between facial parameters, a tensor representing a parameter of a non-face object, a tensor representing motion of the non-face object, motion information, audio data, speech data, text data, a base picture, a drive picture, a previous temporally extrapolated picture, or an output of a further neural network.

[0407] The means may be further configured for: providing one or more input facial parameters to the further neural network; and obtaining the output of the further neural network, wherein the output of the further neural network may comprise one or more converted facial parameters.

[0408] The further neural network may comprise the translator neural network.

[0409] The means may be further configured for: extracting the audio data or the speech data from an audio track of the at least one input.

[0410] The indication of uncertainty associated with the at least one type of input may comprise at least one of: an indication that lossy compression was used for the at least one type of input, an indication that lossless compression was used for the at least one type of input, an indication of a compression algorithm used for compressing the at least one type of input, an indication of a compression rate used for the at least one type of input, an indication of a quantization step size used for the at least one type of input, an indication of a decompression algorithm used for decompressing the at least one type of input, an indication of at least one parameter for compressing the at least one type of input, an indication of at least one parameter for decompressing the at least one type of input, or an indication of a source of the at least one type of input.

[0411] The at least one type of content in the at least one base picture, or the at least one type of content in the at least one drive picture, may comprise at least one of: background content, foreground content, or the background content and the foreground content.

[0412] The background content may comprise one or more non-face regions and / or one or more non-face objects.

[0413] The foreground content may comprise one or more face objects.

[0414] The plurality of neural networks may comprise a neural network associated with a respective region, of two or more regions, of an input picture.

[0415] A number of the plurality of neural networks may be equal to a number of the two or more regions.

[0416] A number of the plurality of neural networks may be greater than a number of the two or more regions.

[0417] The signaling information may comprise, at least, information, associated with respective ones of the two or more regions, identifying respective ones of the plurality of neural networks.

[0418] The indication of the translator neural network may be received for a random access segment of a video sequence.

[0419] The indication of the translator neural network may be received for a coded layer video sequence.

[0420] The signaling information may be received in or along a bitstream.

[0421] The means may be further configured for: determining that one or more of the at least one type of input is unavailable; determining one or more of the at least one alternative type of input that can be used instead of the one or more types of input that are unavailable; and providing, at least, the one or more alternative types of input to the at least one neural network.

[0422] The indication of the translator neural network may be gated with a flag.

[0423] The signaling information may comprise the indication of the translator neural network, wherein the indication of the translator neural network may be received in: a generative face video supplemental enhancement information message which is the first, in output order, among one or more generative face video supplemental enhancement information messages at or subsequent to a random access point picture.

[0424] The signaling information may comprise the indication of the translator neural network independently of a determination of whether the current picture is a base picture.

[0425] The at least one neural network may comprise, at least, a generator neural network, wherein a base picture, other than a latest decoded base picture, may be provided to the generator neural network.

[0426] The signaling information may comprise, at least, an indication of the base picture.

[0427] The at least one neural network may comprise, at least, a generator neural network, wherein a drive picture, other than a latest decoded drive picture, may be provided to the generator neural network.

[0428] The signaling information may comprise, at least, an indication of the drive picture.

[0429] In accordance with one example embodiment, a non-transitory computer-readable medium comprising instructions stored thereon which, when executed with at least one processor, cause the at least one processor to: cause receiving of signaling information for performing a postprocessing operation, wherein the post-processing operation may comprise use of at least one neural network; cause receiving of at least one input; and cause providing of the at least one input to the at least one neural network based, at least partially, on the signaling information.

[0430] In accordance with one example embodiment, a non-transitory computer-readable medium comprising program instructions stored thereon for performing at least the following: causing receiving of signaling information for performing a post-processing operation, wherein the post-processing operation may comprise use of at least one neural network; causing receivingof at least one input; and causing providing of the at least one input to the at least one neural network based, at least partially, on the signaling information.

[0431] In accordance with another example embodiment, a non-transitory program storage device readable by a machine may be provided, tangibly embodying instructions executable by the machine for performing operations, the operations comprising: causing receiving of signaling information for performing a post-processing operation, wherein the post-processing operation may comprise use of at least one neural network; causing receiving of at least one input; and causing providing of the at least one input to the at least one neural network based, at least partially, on the signaling information.

[0432] In accordance with another example embodiment, a non-transitory computer-readable medium comprising instructions that, when executed by an apparatus, cause the apparatus to perform at least the following: causing receiving of signaling information for performing a postprocessing operation, wherein the post-processing operation may comprise use of at least one neural network; causing receiving of at least one input; and causing providing of the at least one input to the at least one neural network based, at least partially, on the signaling information.

[0433] A computer implemented system comprising: at least one processor and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the system at least to perform: causing receiving of signaling information for performing a postprocessing operation, wherein the post-processing operation may comprise use of at least one neural network; causing receiving of at least one input; and causing providing of the at least one input to the at least one neural network based, at least partially, on the signaling information.

[0434] A computer implemented system comprising: means for causing receiving of signaling information for performing a post-processing operation, wherein the post-processing operation may comprise use of at least one neural network; means for causing receiving of at least one input; and means for causing providing of the at least one input to the at least one neural network based, at least partially, on the signaling information.

[0435] The term “non-transitory,” as used herein, is a limitation of the medium itself (i.e. tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM vs. ROM).

[0436] It should be understood that the foregoing description is only illustrative. Various alternatives and modifications can be devised by those skilled in the art. For example, features recited in the various dependent claims could be combined with each other in any suitable combination(s). In addition, features from different embodiments described above could be selectively combined into a new embodiment. Accordingly, the description is intended to embrace all such alternatives, modification and variances which fall within the scope of the appended claims.

Claims

CLAIMSWhat is claimed is:

1. An apparatus comprising means for: generating signaling information for performing a post-processing operation, wherein the post-processing operation comprises use of at least one neural network; and providing the signaling information for performing the post-processing operation.

2. The apparatus of claim 1, wherein the post-processing operation comprises at least one of: temporal extrapolation, or visual temporal extrapolation.

3. The apparatus of claim 1 or 2, wherein the signaling information is associated with one or more pictures of a video.

4. The apparatus of any of claims 1 through 3, wherein the signaling information is provided with at least one of: at least one neural network post-filter characteristic supplemental enhancement information message, at least one neural network post-filter activation supplemental enhancement information message, at least one generative face video supplemental enhancement information message, at least one generative face video characteristics supplemental enhancement information message, orat least one neural network post-filter input supplemental enhancement information message.

5. The apparatus of any of claims 1 through 4, wherein the signaling information comprises an indication of at least one of: at least one type of input to the post-processing operation, one or more types of input that are related to the at least one type of input, one or more types of input that are unrelated to the at least one type of input, at least one alternative type of input to the post-processing operation, at least one update for the at least one neural network, at least one adaptation signal or tensor for adapting at least one input to the at least one neural network, a plurality of neural networks the post-processing operation uses, a plurality of neural network post-filter characteristics supplemental enhancement information messages respectively indicating the plurality of neural networks, a plurality of regions of an input for which a same neural network performs the post-processing operation, a neural network post-filter characteristics supplemental enhancement information message configured to indicate to same neural network, an uncertainty associated with the at least one type of input, at least one type of content in at least one base picture, at least one type of content in at least one drive picture,a quality of the at least one type of content in the at least one base picture, a quality of the at least one type of content in the at least one drive picture, whether a current picture is a drive picture, a translator neural network, a reference to a supplemental enhancement information message that comprises the indication of the translator neural network, a neural network post-filter characteristics supplemental enhancement information message configured to indicate the at least one neural network, or at least one picture to be temporally extrapolated using the post-processing operation.

6. The apparatus of claim 5, wherein the at least one alternative type of input to the postprocessing operation comprises one or more types of input for use in response to a determination that the at least one type of input to the post-processing operation is not available or is not reliable.

7. The apparatus of claim 5 or 6, wherein the at least one type of input to the post-processing operation comprises at least one of: a facial parameter, a tensor representing at least one facial parameter, a tensor representing at least one encoded facial parameter, a tensor representing motion of a portion of a first picture relative to a second picture, a tensor representing a difference between facial parameters,a tensor representing a parameter of a non-face object, a tensor representing motion of the non-face object, motion information, audio data, speech data, text data, a base picture, a drive picture, a previous temporally extrapolated picture, or an output of a further neural network.

8. The apparatus of any of claims 5 through 7, wherein the indication of uncertainty associated with the at least one type of input comprises at least one of: an indication that lossy compression was used for the at least one type of input, an indication that lossless compression was used for the at least one type of input, an indication of a compression algorithm used for compressing the at least one type of input, an indication of a compression rate used for the at least one type of input, an indication of a quantization step size used for the at least one type of input, an indication of a decompression algorithm used for decompressing the at least one type of input,an indication of at least one parameter for compressing the at least one type of input, an indication of at least one parameter for decompressing the at least one type of input, or an indication of a source of the at least one type of input.

9. The apparatus of any of claims 5 through 8, wherein the at least one type of content in the at least one base picture, or the at least one type of content in the at least one drive picture, comprises at least one of: background content, foreground content, or the background content and the foreground content.

10. The apparatus of claim 9, wherein the background content comprises one or more nonface regions and / or one or more non-face objects.

11. The apparatus of claim 9 or 10, wherein the foreground content comprises one or more face objects.

12. The apparatus of any of claims 5 through 11, wherein the signaling information comprises the indication of the translator neural network, wherein the indication of the translator neural network is provided in: a generative face video supplemental enhancement information message which is the first, in output order, among one or more generative face video supplemental enhancement information messages at or subsequent to a random access point picture.

13. The apparatus of any of claims 5 through 12, wherein the signaling information comprises the indication of the translator neural network independently of a determination of whether the current picture is a base picture.

14. A method comprising: generating, with an encoder, signaling information for performing a postprocessing operation, wherein the post-processing operation comprises use of at least one neural network; and providing the signaling information for performing the post-processing operation.

15. The method of claim 14, wherein the post-processing operation comprises at least one of: temporal extrapolation, or visual temporal extrapolation.

16. The method of claim 14 or 15, wherein the signaling information is associated with one or more pictures of a video.

17. The method of any of claims 14 through 16, wherein the signaling information is provided with at least one of: at least one neural network post-filter characteristic supplemental enhancement information message, at least one neural network post-filter activation supplemental enhancement information message, at least one generative face video supplemental enhancement information message, at least one generative face video characteristics supplemental enhancement information message, orat least one neural network post-filter input supplemental enhancement information message.

18. The method of any of claims 14 through 17, wherein the signaling information comprises an indication of at least one of: at least one type of input to the post-processing operation, one or more types of input that are related to the at least one type of input, one or more types of input that are unrelated to the at least one type of input, at least one alternative type of input to the post-processing operation, at least one update for the at least one neural network, at least one adaptation signal or tensor for adapting at least one input to the at least one neural network, a plurality of neural networks the post-processing operation uses, a plurality of neural network post-filter characteristics supplemental enhancement information messages respectively indicating the plurality of neural networks, a plurality of regions of an input for which a same neural network performs the post-processing operation, a neural network post-filter characteristics supplemental enhancement information message configured to indicate to same neural network, an uncertainty associated with the at least one type of input, at least one type of content in at least one base picture, at least one type of content in at least one drive picture,a quality of the at least one type of content in the at least one base picture, a quality of the at least one type of content in the at least one drive picture, whether a current picture is a drive picture, a translator neural network, a reference to a supplemental enhancement information message that comprises the indication of the translator neural network, a neural network post-filter characteristics supplemental enhancement information message configured to indicate the at least one neural network, or at least one picture to be temporally extrapolated using the post-processing operation.

19. The method of claim 18, wherein the at least one alternative type of input to the postprocessing operation comprises one or more types of input for use in response to a determination that the at least one type of input to the post-processing operation is not available or is not reliable.

20. The method of claim 18 or 19, wherein the at least one type of input to the postprocessing operation comprises at least one of: a facial parameter, a tensor representing at least one facial parameter, a tensor representing at least one encoded facial parameter, a tensor representing motion of a portion of a first picture relative to a second picture, a tensor representing a difference between facial parameters,a tensor representing a parameter of a non-face object, a tensor representing motion of the non-face object, motion information, audio data, speech data, text data, a base picture, a drive picture, a previous temporally extrapolated picture, or an output of a further neural network.

21. The method of any of claims 18 through 20, wherein the indication of uncertainty associated with the at least one type of input comprises at least one of: an indication that lossy compression was used for the at least one type of input, an indication that lossless compression was used for the at least one type of input, an indication of a compression algorithm used for compressing the at least one type of input, an indication of a compression rate used for the at least one type of input, an indication of a quantization step size used for the at least one type of input, an indication of a decompression algorithm used for decompressing the at least one type of input,an indication of at least one parameter for compressing the at least one type of input, an indication of at least one parameter for decompressing the at least one type of input, or an indication of a source of the at least one type of input.

22. The method of any of claims 18 through 21, wherein the at least one type of content in the at least one base picture, or the at least one type of content in the at least one drive picture, comprises at least one of: background content, foreground content, or the background content and the foreground content.

23. The method of claim 22, wherein the background content comprises one or more nonface regions and / or one or more non-face objects.

24. The method of claim 22 or 23, wherein the foreground content comprises one or more face objects.

25. The method of any of claims 18 through 24, wherein the signaling information comprises the indication of the translator neural network, wherein the indication of the translator neural network is provided in: a generative face video supplemental enhancement information message which is the first, in output order, among one or more generative face video supplemental enhancement information messages at or subsequent to a random access point picture.

26. The method of any of claims 18 through 25, wherein the signaling information comprises the indication of the translator neural network independently of a determination of whether the current picture is a base picture.

27. An apparatus comprising means for: receiving signaling information for performing a post-processing operation, wherein the post-processing operation comprises use of at least one neural network; receiving at least one input; and providing the at least one input to the at least one neural network based, at least partially, on the signaling information.

28. The apparatus of claim 27, wherein the post-processing operation comprises at least one of: temporal extrapolation, or visual temporal extrapolation.

29. The apparatus of claim 27 or 28, wherein the at least one input comprises, at least, one or more pictures of a video.

30. The apparatus of any of claims 27 through 29, wherein the signaling information is received via at least one of: at least one neural network post-filter characteristic supplemental enhancement information message, at least one neural network post-filter activation supplemental enhancement information message, at least one generative face video supplemental enhancement information message,at least one generative face video characteristics supplemental enhancement information message, or at least one neural network post-filter input supplemental enhancement information message.

31. The apparatus of any of claims 27 through 30, wherein the signaling information comprises an indication of at least one of: at least one type of input to the post-processing operation, one or more types of input that are related to the at least one type of input, one or more types of input that are unrelated to the at least one type of input, at least one alternative type of input to the post-processing operation, at least one update for the at least one neural network, at least one adaptation signal or tensor for adapting at least one input to the at least one neural network, a plurality of neural networks the post-processing operation uses, a plurality of neural network post-filter characteristics supplemental enhancement information messages respectively indicating the plurality of neural networks, a plurality of regions of an input for which a same neural network performs the post-processing operation, a neural network post-filter characteristics supplemental enhancement information message configured to indicate to same neural network, an uncertainty associated with the at least one type of input, at least one type of content in at least one base picture,at least one type of content in at least one drive picture, a quality of the at least one type of content in the at least one base picture, a quality of the at least one type of content in the at least one drive picture, whether a current picture is a drive picture, a translator neural network, a reference to a supplemental enhancement information message that comprises the indication of the translator neural network, a neural network post-filter characteristics supplemental enhancement information message configured to indicate the at least one neural network, or at least one picture to be temporally extrapolated using the post-processing operation.

32. The apparatus of claim 31, wherein the at least one alternative type of input to the postprocessing operation comprises one or more types of input for use in response to a determination that the at least one type of input to the post-processing operation is not available or is not reliable.

33. The apparatus of claim 31 or 32, wherein the at least one type of input to the postprocessing operation comprises at least one of: a facial parameter, a tensor representing at least one facial parameter, a tensor representing at least one encoded facial parameter, a tensor representing motion of a portion of a first picture relative to a second picture,a tensor representing a difference between facial parameters, a tensor representing a parameter of a non-face object, a tensor representing motion of the non-face object, motion information, audio data, speech data, text data, a base picture, a drive picture, a previous temporally extrapolated picture, or an output of a further neural network.

34. The apparatus of claim 33, wherein the means are further configured for: providing one or more input facial parameters to the further neural network; and obtaining the output of the further neural network, wherein the output of the further neural network comprises one or more converted facial parameters.

35. The apparatus of claim 33 or 34, wherein the further neural network comprises the translator neural network.

36. The apparatus of any of claims 33 through 35, wherein the means are further configured for:extracting the audio data or the speech data from an audio track of the at least one input.

37. The apparatus of any of claims 31 through 36, wherein the indication of uncertainty associated with the at least one type of input comprises at least one of: an indication that lossy compression was used for the at least one type of input, an indication that lossless compression was used for the at least one type of input, an indication of a compression algorithm used for compressing the at least one type of input, an indication of a compression rate used for the at least one type of input, an indication of a quantization step size used for the at least one type of input, an indication of a decompression algorithm used for decompressing the at least one type of input, an indication of at least one parameter for compressing the at least one type of input, an indication of at least one parameter for decompressing the at least one type of input, or an indication of a source of the at least one type of input.

38. The apparatus of any of claims 31 through 37, wherein the at least one type of content in the at least one base picture, or the at least one type of content in the at least one drive picture, comprises at least one of: background content, foreground content, or the background content and the foreground content.

39. The apparatus of claim 38, wherein the background content comprises one or more non-face regions and / or one or more non-face objects.

40. The apparatus of claim 38 or 39, wherein the foreground content comprises one or more face objects.

41. A method comprising: receiving, with a decoder, signaling information for performing a post-processing operation, wherein the post-processing operation comprises use of at least one neural network; receiving at least one input; and providing the at least one input to the at least one neural network based, at least partially, on the signaling information.

42. The method of claim 41, wherein the post-processing operation comprises at least one of: temporal extrapolation, or visual temporal extrapolation.

43. The method of claim 41 or 42, wherein the at least one input comprises, at least, one or more pictures of a video.

44. The apparatus of any of claims 41 through 43, wherein the signaling information is received via at least one of: at least one neural network post-filter characteristic supplemental enhancement information message, at least one neural network post-filter activation supplemental enhancement information message,at least one generative face video supplemental enhancement information message, at least one generative face video characteristics supplemental enhancement information message, or at least one neural network post-filter input supplemental enhancement information message.

45. The method of any of claims 41 through 44, wherein the signaling information comprises an indication of at least one of: at least one type of input to the post-processing operation, one or more types of input that are related to the at least one type of input, one or more types of input that are unrelated to the at least one type of input, at least one alternative type of input to the post-processing operation, at least one update for the at least one neural network, at least one adaptation signal or tensor for adapting at least one input to the at least one neural network, a plurality of neural networks the post-processing operation uses, a plurality of neural network post-filter characteristics supplemental enhancement information messages respectively indicating the plurality of neural networks, a plurality of regions of an input for which a same neural network performs the post-processing operation, a neural network post-filter characteristics supplemental enhancement information message configured to indicate to same neural network, an uncertainty associated with the at least one type of input,at least one type of content in at least one base picture, at least one type of content in at least one drive picture, a quality of the at least one type of content in the at least one base picture, a quality of the at least one type of content in the at least one drive picture, whether a current picture is a drive picture, a translator neural network, a reference to a supplemental enhancement information message that comprises the indication of the translator neural network, a neural network post-filter characteristics supplemental enhancement information message configured to indicate the at least one neural network, or at least one picture to be temporally extrapolated using the post-processing operation.

46. The method of claim 45, wherein the at least one alternative type of input to the postprocessing operation comprises one or more types of input for use in response to a determination that the at least one type of input to the post-processing operation is not available or is not reliable.

47. The method of claim 45 or 46, wherein the at least one type of input to the postprocessing operation comprises at least one of: a facial parameter, a tensor representing at least one facial parameter, a tensor representing at least one encoded facial parameter,a tensor representing motion of a portion of a first picture relative to a second picture, a tensor representing a difference between facial parameters, a tensor representing a parameter of a non-face object, a tensor representing motion of the non-face object, motion information, audio data, speech data, text data, a base picture, a drive picture, a previous temporally extrapolated picture, or an output of a further neural network.

48. The method of claim 47, further comprising: providing one or more input facial parameters to the further neural network; and obtaining the output of the further neural network, wherein the output of the further neural network comprises one or more converted facial parameters.

49. The apparatus of claim 47 or 48, wherein the further neural network comprises the translator neural network.

50. The apparatus of any of claims 47 through 49, further comprising:extracting the audio data or the speech data from an audio track of the at least one input.

51. The method of any of claims 45 through 50, wherein the indication of uncertainty associated with the at least one type of input comprises at least one of: an indication that lossy compression was used for the at least one type of input, an indication that lossless compression was used for the at least one type of input, an indication of a compression algorithm used for compressing the at least one type of input, an indication of a compression rate used for the at least one type of input, an indication of a quantization step size used for the at least one type of input, an indication of a decompression algorithm used for decompressing the at least one type of input, an indication of at least one parameter for compressing the at least one type of input, an indication of at least one parameter for decompressing the at least one type of input, or an indication of a source of the at least one type of input.

52. The method of any of claims 45 through 51, wherein the at least one type of content in the at least one base picture, or the at least one type of content in the at least one drive picture, comprises at least one of: background content, foreground content, or the background content and the foreground content.

53. The method of claim 52, wherein the background content comprises one or more nonface regions and / or one or more non-face objects.

54. The method of claim 52 or 53, wherein the foreground content comprises one or more face objects.

Citation Information

Patent Citations

  • Coding techniques and metadata for video communications using generative face video

    WO2024263644A2

Cited By

  • Devices and methods for signaling multiple extrapolations in video coding via a neural-network post-filter characteristics supplemental enhancement information message

    US12666066B2

  • Systems and methods for signaling multiple spatial extrapolations in video coding

    US20260012624A1