Using generative artifical intelligence for decoding media data
By integrating neural networks for visual temporal extrapolation within multimedia encoding and decoding, generative artificial intelligence is utilized to enhance media data processing, addressing limitations in existing technologies and enabling advanced visual enhancements.
Patent Information
- Application Number
- PCT/IB2025/050144
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-08
- Filing Date
- 2025-01-07
- Publication Date
- 2025-07-17
AI Technical Summary
Existing multimedia encoding and decoding technologies lack the ability to effectively perform visual temporal extrapolation using generative artificial intelligence, limiting the enhancement and manipulation of media data.
Incorporating neural networks for visual temporal extrapolation, utilizing signaling information in bitstreams to enable generative artificial intelligence for decoding media data, including neural network post-filter characteristics and generative face video supplemental enhancement information messages.
Enables advanced visual enhancements such as picture improvement, temporal interpolation, and extrapolation, enhancing the quality and flexibility of media decoding processes.
Smart Images

Figure IB2025050144_17072025_PF_FP_ABST
Abstract
Description
USING GENERATIVE ARTIFICAL INTELLIGENCE FOR DECODING MEDIA DATATECHNICAL FIELD
[0001] The examples and non-limiting embodiments relate generally to multimedia transport and, more particularly, enabling the use of generative artificial intelligence for encoding, decoding or signaling media data.BACKGROUND
[0002] It is known to provide standardized formats for encoding, signaling, or decoding of media data.SUMMARY
[0003] Example 1 : An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: encoding signaling information in or along a bitstream, wherein the signaling information enables or supports performing visual temporal extrapolation; and signaling the signaling information.
[0004] Example 2: An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving signaling information, that enables or supports performing visual temporal extrapolation, in or along a bitstream; decoding the signaling information to obtain a decoded signaling information (simply also referred to as signaling information); and performing the visual temporal extrapolation, based at least on the decoded signaling information.
[0005] Example 3: The apparatus of any of the examples 1 or 2, wherein one or more neural networks (VTNN) are used to perform the visual temporal extrapolation .
[0006] Example 4: he apparatus of any of the examples 1 to 2, wherein the signaling information is comprised in a neural network post-filter characteristics (NNPFC) supplemental enhancement information (SEI) message; a generative face video (GFV) SEI message; or in a neural-network postfilter characteristics (NNPFC) SEI message and a GFV SEI message.
[0007] Example 5 : The apparatus of any of the examples 1 to 4, wherein the signaling information comprises one or more of the following: a purpose of performing the visual temporal extrapolation at a decoder side; or a temporal extrapolation characteristics (TEC) information that characterizes a neural network performing temporal extrapolation; or a purpose of generative modelling or generative artificial intelligence (Al) at decoder side.
[0008] Example 6: The apparatus of any of the examples 1 to 5, wherein the TEC information is signaled based on one or more of the following: when the signaling information indicates that the purpose of the VENN associated with the signaling information comprises visual temporal extrapolation, generative modeling, or generative artificial intelligence (Al), wherein the generative modeling or the generative Al refers to using the VTNN for generating data conditional to one or more inputs; or when the purpose comprises generative face video or generative video.
[0009] Example 7 : The apparatus of example 6, wherein the purpose of the generative face video comprises generating or improving a picture or at least a part of a video comprising at least one or more faces, given one or more of the following: coordinates of the at least one or more faces; coordinates of landmarks or key-points of the at least one or more faces; transformation matrices that describe a motion or temporal change of the one or more key-points of the at least one or more faces; features of facial expressions; indication of blinking; texts associated to the video; audio associated to the video; or one or more other pictures.
[0010] Example 8: The apparatus of any of the examples 5 to 7, wherein the TEC information comprises one or more of the following: information indicative of number of pictures that are extrapolated in one activation or inference of the VTNN; information indicative of time interval of pictures being extrapolated relative to a latest picture used as input for the inference; information indicative of whether the VTNN takes an auxiliary input that indicates output instances of one or more pictures being temporally extrapolated; information indicative of whether the VTNN takes an auxiliary input that indicates or is derived from facial parameters, wherein the auxiliary input is used by the VTNN in order to perform the temporal extrapolation of one or more faces in a picture or a video; information indicative of whether the VTNN takes an auxiliary input that is derived from the output of another VTNN, wherein the another VTNN is run or executed at the decoder side; information indicative of whether the VTNN takes an auxiliary input that is derived from the output of the another VTNN that converts or translates facial parameters to a format that is accepted or required by the VTNN; information indicative of whether the VTNN takes an auxiliary input that represents or is derived from a tensor that is generated at an encoder side; information indicative of whether the VTNNtakes an auxiliary input that represents or is derived from an audio or a speech; information about characteristics of the audio or the speech, when the auxiliary input represents the audio or the speech; information indicative of whether the VTNN takes an auxiliary input that represents or is derived from text; information indicative of whether the VTNN produces different outputs for different inferences or executions of the VTNN with same input; information indicative of type or cause of output variability of the VTNN; information indicative of whether the VTNN is instructed or operated in a way that the VTNN does not yield output variability; information indicative of whether the VTNN is allowed to have output variability; information that makes the network not to have output variability; information indicative of a type of extrapolated outputs; information indicative of whether the VTNN takes an auxiliary input that indicates picture order count (POC) differences, timestamp differences, relative output position, or relative output differences of output pictures are to be generated or extrapolated; information indicative of format for the facial parameters which is accepted or required by the VTNN; information about a format of the tensor; information indicative of whether the VTNN takes an auxiliary input that represents text associated to the speech of a person for whom the face is to be temporally extrapolated; information indicative of whether the VTNN takes an auxiliary input that represents a text; information indicative of whether the VTNN produces different outputs for different inferences or executions of the VTNN with one or more same inputs; information indicative of whether an input to the VTNN consists of background information; information indicative of whether an input to the VTNN consists of foreground information; information indicative of whether an input to the VTNN comprises a drive picture; information indicative of whether an auxiliary input to the VTNN is derived from a lossy coding process; information indicative of whether the VTNN accepts an input that controls one or more aspects of how a final output is determined based on an output of the VTNN; one or more indications for indicating which facial parameters are present and / or used by the VTNN; indicating that the number of extrapolated pictures of a neural network post-filter (NNPF) vary for each inference and the number of extrapolated pictures for each inference is indicated by using another signaling mechanism; an indication indicative of whether a decoder side is to generate a random seed value to the VTNN or derive a pseudo-random seed value based on other information that is signaled by an encoder; an indication indicative of whether a pseudo-random seed value is derived based on a value of a particular syntax element that is comprised in the TEC information or that is signaled to the decoder by other means; an indication indicative of whether the VTNN specified by the TEC information takes a seed value as an input or auxiliary input; an indication indicative of whether the VTNN specified by the TEC information takes a pseudo-random seed as the input or the auxiliary input; or a characteristics of randomness for the VTNN.
[0011] Example 9: The apparatus of example 8, wherein the other value comprises one of following: a sequence-wise quantization parameter (QP) value, and wherein the pseudo-random seedvalue is set equal to that sequence-wise QP value; a slice-wise QP value, and wherein the pseudorandom seed value is set equal to that slice-wise QP value; a value of a syntax element in the TEC information; a value of a syntax element in a neural -network post-fdter characteristics (NNPFC) supplemental enhancement information (SEI) message that specifies the VENN; or a value of a syntax element in a generative face video (GFV) SEI message that specifies the VENN.
[0012] Example 10: The apparatus of any of the examples 8, wherein the facial parameters that are used as auxiliary input or that are used to derive the auxiliary input are carried within one of the following: a newly defined SEI message; a generative face video (GFV) SEI message; a neural network post-filter activation (NNPFA) SEI message; or a neural-network post-filter input (NNPFI) SEI message.
[0013] Example 11: The apparatus of any of the examples 8, wherein the tensor is an output of or is derived from another neural network that is run or executed at encoder side.
[0014] Example 12: The apparatus of example 11, wherein the tensor is signaled within the NNPFA SEI message, the GFV SEI message, or the NNPFI SEI message.
[0015] Example 13: The apparatus of example 8, wherein information about characteristics of the audio or speech comprises one or more of the following: sampling rate of the audio; bit depth and / or dynamic range of the audio; number of audio channels and / or channel configuration of the audio; duration of the audio and / or number of audio samples to be included in an input tensor; information on whether the audio is processed with source separation algorithm; information on whether the audio is spatial audio; or timing of the audio provided as input in relation to one or more pictures to be extrapolated.
[0016] Example 14: The apparatus of example 8, wherein the type or cause of output variability comprises one of the following: a pseudo-randomness; a deliberate variability; or floating-point numerical format that is used to represent value of one or more parameters of the VENN and / the value of one or more signals flowing through the VENN.
[0017] Example 15: The apparatus of example 8, wherein the information indicative of whether the VENN takes an auxiliary input that indicates output instances of one or more pictures being temporally extrapolated is carried within the NNPFC SEI message or a neural network post-filter activation (NNPFA) SEI message.
[0018] Example 16: The apparatus of example 8, wherein the information indicative of whether the VTNN takes the auxiliary input that indicates the POC differences, timestamp differences, relative output position, or relative output differences is signaled as part of a neural-network post-fdter characteristics (NNPFC) supplemental enhancement information (SEI) message.
[0019] Example 17: The apparatus of any of the examples 8 or 16, wherein the auxiliary inputthat indicates the POC differences, timestamp differences, relative output position, or relative output differences is carried within a neural-network post-filter activation(NNPFA) supplemental enhancement information (SEI) message or a neural-network post-filter input (NPFI) SEI message.
[0020] Example 18: The apparatus of any of the examples 8 or 10, wherein the facial parameters comprise one or more of the following: 2D key-points; 3D key-points; or one or more matrices.
[0021] Example 19: The apparatus of example 10, wherein the facial parameters obtained from the NNPFA SEI message, the GFV SEI message, or the NNPFI SEI message are converted to a required format that is indicated in the TEC information.
[0022] Example 20: The apparatus any of the examples 8 or 13, wherein information indicative of whether the VTNN takes the auxiliary input that represents or is derived from the audio or the speech is used by the VTNN to generate or synthesize one or more features of a face.
[0023] Example 21: The apparatus of example 8, wherein when the VTNN takes the auxiliary input that represents the text, the TEC information further comprise information indicative of characteristics of the text comprising one or more of the following: whether the VTNN accepts or requires raw text; whether the VTNN accepts or requires tokens; whether the VTNN accepts or requires text embeddings; which tokenizer to use; which text embedder to use; or timing of the text provided as input in relation to the one or more pictures to be extrapolated.
[0024] Example 22: The apparatus of example 8, wherein the information indicative of the type of extrapolated outputs comprises indicating one or more categories of visual content that the VTNN is able to extrapolate.
[0025] Example 23: The apparatus of example 22, wherein the one or more categories comprise a face, hands, or a body.
[0026] Example 24: The apparatus of any of the examples 3 to 23, wherein the signaling information further comprises indicating whether the VTNN is adapted for all the pictures to which the VTNN is applied, activated, or the VTNN is sequence-level adapted.
[0027] Example 25: The apparatus is example 24, wherein the signaling information further comprises indicating a purpose of the sequence-level adaptation.
[0028] Example 26: The apparatus of example 24 or 25, wherein a persistent auxiliary input (PAI) signal comprised in the signaling information is used for indicating whether the VTNN is adapted for all pictures to which the VTNN is applied.
[0029] Example 27: The apparatus of example 26, wherein the signaling information further comprises indicating a type of the PAI signal.
[0030] Example 28: The apparatus of example 27, wherein the type of the PAI comprises, a tensor, a text, or an image.
[0031] Example 29: The apparatus of any of the examples 3 to 23, wherein the signaling information further comprises: information indicative of a purpose for a neural network that converts input facial parameters to an output that represents converted facial parameters; information indicative of a purpose for a neural network that converts input parameters to an output that represents converted parameters; or information indicative of a purpose for a neural network that converts a format of input parameters to another format of output parameters.
[0032] Example 30: The apparatus of example 8, wherein a neural-network post-fdter input (NNPFI) SEI message is used to provide the auxiliary input to a neural-network post-processing filter defined by a (NNPFC) SEI message and activated by a neural network post-filter activation (NNPFA) SEI message.
[0033] Example 31: An apparatus comprising at least one processor; and at least one non- transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: concluding that a temporal interleaving frame packing arrangement is in use in a bitstream; concluding a constituent frame parity that a current frame where a temporal extrapolation post-filter is activated comprises a temporal interleaving frame packing arrangement; selecting one or more input frames, for the temporal extrapolation post-filter, that comprise theconcluded constituent frame parity; and applying the temporal extrapolation post-filter with the one or more input frames as input.
[0034] Example 32: A method comprising: encoding signaling information in or along a bitstream, wherein the signaling information enables or supports performing visual temporal extrapolation; and signaling the signaling information.
[0035] Example 33: A method comprising: receiving signaling information, that enables or supports performing visual temporal extrapolation, in or along a bitstream; decoding the signaling information to obtain a decoded signaling information (signaling information) to obtain a decoded signaling information (signaling information); and performing the visual temporal extrapolation, based at least on the decoded signaling information.
[0036] Example 34: The method of any of the examples 32 or 33, wherein one or more neural networks (VENN) are used to perform the visual temporal extrapolation.
[0037] Example 35: The method of any of the examples 32 to 33, wherein the signaling information is comprised in a neural network post-fdter characteristics (NNPFC) supplemental enhancement information (SEI) message, a generative face video (GFV) SEI message; or in a neural- network post-fdter characteristics (NNPFC) SEI message and a GFV SEI message.
[0038] Example 36: The method of any of the examples 32 to 35, wherein the signaling information comprises one or more of the following: a purpose of performing the visual temporal extrapolation at a decoder side; or a temporal extrapolation characteristics (TEC) information that characterizes a neural network performing temporal extrapolation; or a purpose of generative modelling or generative artificial intelligence (Al) at decoder side.
[0039] Example 37: The method of any of the examples 32 to 36, wherein the TEC information is signaled based on one or more of the following: when the signaling information indicates that the purpose of the VENN associated with the signaling information comprises visual temporal extrapolation, generative modeling, or generative artificial intelligence (Al), wherein the generative modeling or the generative Al refers to using the VENN for generating data conditional to one or more inputs; or when the purpose comprises generative face video or generative video.
[0040] Example 38: The method of example 37, wherein the purpose of the generative face video comprises generating or improving a picture or at least a part of a video comprising at least one or morefaces, given one or more of the following: coordinates of the at least one or more faces; coordinates of landmarks or key-points of the at least one or more faces; transformation matrices that describe a motion or temporal change of the one or more key-points of the at least one or more faces; features of facial expressions; indication of blinking; texts associated to the video; audio associated to the video; or one or more other pictures.
[0041] Example 39: The method of any of the examples 36 to 38, wherein the TEC information comprises one or more of the following: information indicative of number of pictures that are extrapolated in one activation or inference of the VTNN; information indicative of time interval of pictures being extrapolated relative to a latest picture used as input for the inference; information indicative of whether the VTNN takes an auxiliary input that indicates output instances of one or more pictures being temporally extrapolated; information indicative of whether the VTNN takes an auxiliary input that indicates or is derived from facial parameters, wherein the auxiliary input is used by the VTNN in order to perform the temporal extrapolation of one or more faces in a picture or a video; information indicative of whether the VTNN takes an auxiliary input that is derived from the output of another VTNN, wherein the another VTNN is run or executed at the decoder side; information indicative of whether the VTNN takes an auxiliary input that is derived from the output of the another VTNN that converts or translates facial parameters to a format that is accepted or required by the VTNN; information indicative of whether the VTNN takes an auxiliary input that represents or is derived from a tensor that is generated at an encoder side; information indicative of whether the VTNN takes an auxiliary input that represents or is derived from an audio or a speech; information about characteristics of the audio or the speech, when the auxiliary input represents the audio or the speech; information indicative of whether the VTNN takes an auxiliary input that represents or is derived from text; information indicative of whether the VTNN produces different outputs for different inferences or executions of the VTNN with same input; information indicative of type or cause of output variability of the VTNN; information indicative of whether the VTNN is instructed or operated in a way that the VTNN does not yield output variability; information indicative of whether the VTNN is allowed to have output variability; information that makes the network not to have output variability; information indicative of a type of extrapolated outputs; information indicative of whether the VTNN takes an auxiliary input that indicates picture order count (POC) differences, timestamp differences, relative output position, or relative output differences of output pictures are to be generated or extrapolated; information indicative of format for the facial parameters which is accepted or required by the VTNN; information about a format of the tensor; information indicative of whether the VTNN takes an auxiliary input that represents text associated to the speech of a person for whom the face is to be temporally extrapolated; information indicative of whether the VTNN takes an auxiliary input that represents a text; information indicative of whether the VTNN produces different outputs for different inferences orexecutions of the VTNN with one or more same inputs; information indicative of whether an input to the VTNN consists of background information; information indicative of whether an input to the VTNN consists of foreground information; information indicative of whether an input to the VTNN comprises a drive picture; information indicative of whether an auxiliary input to the VTNN is derived from a lossy coding process; information indicative of whether the VTNN accepts an input that controls one or more aspects of how a final output is determined based on an output of the VTNN one or more indications for indicating which facial parameters are present and / or used by the VTNN; indicating that the number of extrapolated pictures of a neural network post-filter (NNPF) vary for each inference and the number of extrapolated pictures for each inference is indicated by using another signaling mechanism; an indication indicative of whether a decoder side is to generate a random seed value to the VTNN or derive a pseudo-random seed value based on other information that is signaled by an encoder; an indication indicative of whether a pseudo-random seed value is derived based on a value of a particular syntax element that is comprised in the TEC information or that is signaled to the decoder by other means; an indication indicative of whether the VTNN specified by the TEC information takes a seed value as an input or auxiliary input; an indication indicative of whether the VTNN specified by the TEC information takes a pseudo-random seed as the input or the auxiliary input; or a characteristics of randomness for the VTNN.
[0042] Example 40: The method of example 39, wherein the other value comprises one of following: a sequence-wise quantization parameter (QP) value, and wherein the pseudo-random seed value is set equal to that sequence-wise QP value; a slice-wise QP value, and wherein the pseudorandom seed value is set equal to that slice-wise QP value; a value of a syntax element in the TEC information; a value of a syntax element in a neural -network post-filter characteristics (NNPFC) supplemental enhancement information (SEI) message that specifies the VTNN; or a value of a syntax element in a generative face video (GFV) SEI message that specifies the VTNN.
[0043] Example 41: The method of example 39, wherein the facial parameters that are used as auxiliary input or that are used to derive the auxiliary input are carried within one of the following: a newly defined SEI message; a generative face video (GFV) SEI message; a neural network post-filter activation (NNPF A) SEI message; or a neural -network post-filter input (NNPFI) SEI message.
[0044] Example 42: The method of example 39, wherein the tensor is an output of or is derived from another neural network that is run or executed at encoder side.
[0045] Example 41 : The method of example 42, wherein the tensor is signaled within the NNPFA SEI message, the GFV SEI message, or the NNPFI SEI message.
[0046] Example 44: The method of example 39, wherein information about characteristics of the audio or speech comprises one or more of the following: sampling rate of the audio; bit depth and / or dynamic range of the audio; number of audio channels and / or channel configuration of the audio; duration of the audio and / or number of audio samples to be included in an input tensor; information on whether the audio is processed with source separation algorithm; information on whether the audio is spatial audio; or timing of the audio provided as input in relation to one or more pictures to be extrapolated.
[0047] Example 45: The method of example 39, wherein the type or cause of output variability comprises one of the following: a pseudo-randomness; a deliberate variability; or floating-point numerical format that is used to represent value of one or more parameters of the VTNN and / the value of one or more signals flowing through the VTNN.
[0048] Example 46: The method of example 39, wherein the information indicative of whether the VTNN takes an auxiliary input that indicates output instances of one or more pictures being temporally extrapolated is carried within the NNPFC SEI message or a neural network post-fdter activation (NNPFA) SEI message.
[0049] Example 47: The method of example 39, wherein the information indicative of whether the VTNN takes the auxiliary input that indicates the POC differences, timestamp differences, relative output position, or relative output differences is signaled as part of a neural-network post-fdter characteristics (NNPFC) supplemental enhancement information (SEI).
[0050] Example 48: The method of any of the examples 39 or 47, wherein the auxiliary input that indicates the POC differences, timestamp differences, relative output position, or relative output differences is carried within a neural-network post-fdter activation (NNPFA) supplemental enhancement information (SEI) message or a neural-network post-fdter input (NPFI) SEI message.
[0051] Example 49: The method of any of the examples 39 or 41, wherein the facial parameters comprise one or more of the following: 2D key-points; 3D key-points; or one or more matrices.
[0052] Example 50: The method of example 41, wherein the facial parameters obtained from the NNPFA SEI message, the GFV SEI message, or the NNPFI SEI message are converted to a required format that is indicated in the TEC information.
[0053] Example 51: The method any of the examples 39 or 44, wherein information indicative of whether the VTNN takes the auxiliary input that represents or is derived from the audio or the speech is used by the VTNN to generate or synthesize one or more features of a face.
[0054] Example 52: The method of example 39, wherein when the VTNN takes the auxiliary input that represents the text, the TEC information further comprise information indicative of characteristics of the text comprising one or more of the following: whether the VTNN accepts or requires raw text; whether the VTNN accepts or requires tokens; whether the VTNN accepts or requires text embeddings; which tokenizer to use; which text embedder to use; or timing of the text provided as input in relation to the one or more pictures to be extrapolated.
[0055] Example 53: The method of example 39, wherein the information indicative of the type of extrapolated outputs comprises indicating one or more categories of visual content that the VTNN is able to extrapolate.
[0056] Example 54: The method of example 53, wherein the one or more categories comprise a face, hands, or a body.
[0057] Example 55: The method of any of the examples 34 to 54, wherein the signaling information further comprises indicating whether the VTNN is adapted for all the pictures to which the VTNN is applied, activated, or the VTNN is sequence-level adapted.
[0058] Example 56: The method is example 55, wherein the signaling information further comprises indicating a purpose of the sequence -level adaptation.
[0059] Example 57: The method of example 55 or 56, wherein a persistent auxiliary input (PAI) signal comprised in the signaling information is used for indicating whether the VTNN is adapted for all pictures to which the VTNN is applied.
[0060] Example 58: The method of example 57, wherein the signaling information further comprises indicating a type of the PAI signal.
[0061] Example 59: The method of example 58, wherein the type of the PAI comprises, a tensor, a text, or an image.
[0062] Example 60: The method of any of the examples 34 to 54, wherein the signaling information further comprises: information indicative of a purpose for a neural network that converts input facial parameters to an output that represents converted facial parameters; information indicative of a purpose for a neural network that converts input parameters to an output that represents converted parameters; or information indicative of a purpose for a neural network that converts a format of input parameters to another format of output parameters.
[0063] Example 61: The method of example 39, wherein a neural-network post-filter input (NNPFI) SEI message is used to provide the auxiliary input to a neural-network post-processing filter defined by a (NNPFC) SEI message and activated by a neural network post-filter activation (NNPFA) SEI message.
[0064] Example 62: A method comprising: concluding that a temporal interleaving frame packing arrangement is in use in a bitstream; concluding a constituent frame parity that a current frame where a temporal extrapolation post-filter is activated comprises a temporal interleaving frame packing arrangement; selecting one or more input frames, for the temporal extrapolation post-filter, that comprise the concluded constituent frame parity; and applying the temporal extrapolation post-filter with the one or more input frames as input.
[0065] Example 63 : An apparatus comprising means for performing the methods as described in any of the examples 32 to 62.
[0066] Example 64: A computer readable medium comprising program instructions which, when executed by an apparatus, cause the apparatus to perform the methods as described in any of the examples 32 to 62.
[0067] Example 65: The computer readable medium of example 64, wherein the computer readable medium comprises a non-transitory computer readable medium.BRIEF DESCRIPTION OF THE DRAWINGS
[0068] The foregoing embodiments and other features are explained in the following description, taken in connection with the accompanying drawings, wherein:
[0069] FIG. 1 shows schematically an electronic device employing embodiments of the examples described herein.
[0070] FIG. 2 shows schematically a user equipment suitable for employing embodiments of the examples described herein.
[0071] FIG. 3 further shows schematically electronic devices employing embodiments of the examples described herein connected using wireless and wired network connections.
[0072] FIG. 4 shows a block diagram of a general structure of a video encoder.
[0073] FIG. 5 is an example apparatus, which may be implemented in hardware, and is caused to, implement examples described herein.
[0074] FIG. 6 shows a representation of an example of non-volatile memory media used to store instructions that implement the examples described herein.
[0075] FIG. 7 is an example method to implement the embodiments described herein, in accordance with another embodiment.
[0076] FIG. 8 is another example method to implement the embodiments described herein, in accordance with another embodiment.
[0077] FIG. 9 is still another example method to implement the embodiments described herein, in accordance with another embodiment.
[0078] FIG. 10 is a block diagram of one possible and non-limiting system in which the example embodiments may be practiced.DETAILED DESCRIPTION OF EXAMPLE EMBODIMENTS
[0079] The following acronyms and abbreviations that may be found in the specification and / or the drawing figures are defined as follows:4CC four character code5G fifth generation cellular network technology5GC 5G core network a.k.a. also known asAVC advanced video codingCU central unitDSP digital signal processorDU distributed unit eNB (or eNodeB) evolved Node B (for example, an LTE base station)EN-DC E-UTRA-NR dual connectivity en-gNB or En-gNB node providing NR user plane and control plane protocol terminations towards the UE, and acting as secondary node in EN-DCE-UTRA evolved universal terrestrial radio access, for example, the LTE radio access technologyFl or Fl-C interface between CU and DU control interface gNB (or gNodeB) base station for 5G / NR, for example, a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GCIEC International Electrotechnical Commission loT internet of thingsISO International Organization for StandardizationISOBMFF ISO base media fde formatJPEG joint photographic experts groupLTE long-term evolution mdat MediaDataBoxMIME Multipurpose Internet Mail ExtensionMME mobility management entity moov MovieBoxMP4 file format for MPEG-4 Part 14 filesMPEG moving picture experts groupMPEG-2 H.222 / H.262 as defined by the ITUMPEG-4 audio and video coding standard for ISO / IEC 14496 ng or NG new generation ng-eNB or NG-eNB new generation eNBNR new radio (5G radio)N / W or NW networkPDCP packet data convergence protocolPHY physical layerPNG portable network graphicsRAN radio access networkRFC request for commentsRLC radio link controlRRC radio resource controlRRH remote radio headRU radio unitRx receiverSDAP service data adaptation protocolSGW serving gatewaySMF session management functionSPS sequence parameter setSVC scalable video codingS 1 interface between eNodeBs and the EPC trak TrackBoxTx transmitterUE user equipmentUICC Universal Integrated Circuit CardUPF user plane functionURL uniform resource locatorX2 interconnecting interface between two eNodeBs in LTE networkXn interface between two NG-RAN nodes
[0080] Some embodiments will now be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all, embodiments are shown. Indeed, various embodiments of the invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Like reference numerals refer to like elements throughout. As used herein, the terms ‘data,’ ‘content,’ ‘information,’ and similar terms may be used interchangeably to refer to data capable of being transmitted, received and / or stored in accordance with embodiments of the present invention. Thus, use of any such terms should not be taken to limit the spirit and scope of embodiments.
[0081] Additionally, as used herein, the term ‘circuitry’ refers to (a) hardware-only circuit implementations (e.g., implementations in analog circuitry and / or digital circuitry); (b) combinations of circuits and computer program product(s) comprising software and / or firmware instructions stored on one or more computer readable memories that work together to cause an apparatus to perform oneor more functions described herein; and (c) circuits, such as, for example, a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation even when the software or firmware is not physically present. This definition of ‘circuitry’ applies to all uses of this term herein, including in any claims. As a further example, as used herein, the term ‘circuitry’ also includes an implementation comprising one or more processors and / or portion(s) thereof and accompanying software and / or firmware. As another example, the term ‘circuitry’ as used herein also includes, for example, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, other network device, and / or other computing device.
[0082] As defined herein, a ‘computer-readable storage medium,’ which refers to a non- transitory physical storage medium (e.g., volatile or non-volatile memory device), can be differentiated from a ‘computer-readable transmission medium,’ which refers to an electromagnetic signal.
[0083] A method, apparatus and computer program product are provided in accordance with example embodiments for enabling the use of generative artificial intelligence for encoding, decoding or signaling media data.
[0084] In an example, the following describes in detail suitable apparatus and possible mechanisms for enabling the use of generative artificial intelligence for encoding, decoding or signaling media data. In this regard reference is first made to FIG. 1 and FIG. 2, where FIG. 1 shows an example block diagram of an apparatus 50. The apparatus may be an internet of things (loT) apparatus configured to perform various functions, for example, gathering information by one or more sensors, receiving or transmitting information, analyzing information gathered or received by the apparatus, or the like. The apparatus may comprise a video coding system, which may incorporate a codec. FIG. 2 shows a layout of an apparatus according to an example embodiment. The elements of FIG. 1 and FIG. 2 will be explained next.
[0085] The apparatus 50, may for example be, a mobile terminal or user equipment of a wireless communication system, a sensor device, a tag, or a lower power device. However, it would be appreciated that embodiments of the examples described herein may be implemented within any electronic device or apparatus which may process data by neural networks.
[0086] The apparatus 50 may comprise a housing 30 for incorporating and protecting the device. The apparatus 50 may further comprise a display 32, for example, in the form of a liquid crystal display, light emitting diode display, organic light emitting diode display, and the like. In other embodiments ofthe examples described herein the display may be any suitable display technology suitable to display media or multimedia content, for example, an image or a video. The apparatus 50 may further comprise a keypad 34. In other embodiments of the examples described herein any suitable data or user interface mechanism may be employed. For example, the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.
[0087] The apparatus may comprise a microphone 36 or any suitable audio input which may be a digital or analogue signal input. The apparatus 50 may further comprise an audio output device which in embodiments of the examples described herein may be any one of: an earpiece 38, speaker, or an analogue audio or digital audio output connection. The apparatus 50 may also comprise a battery (or in other embodiments of the examples described herein the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator). The apparatus may further comprise a camera 42 capable of recording or capturing images and / or video. The apparatus 50 may further comprise an infrared port for short range line of sight communication to other devices. In other embodiments the apparatus 50 may further comprise any suitable short range communication solution such as for example a Bluetooth wireless connection or a USB / firewire wired connection.
[0088] The apparatus 50 may comprise a controller 56, a processor or a processor circuitry for controlling the apparatus 50. The controller 56 may be connected to a memory 58 which in embodiments of the examples described herein may store both data in the form of an image, audio data and video data, and / or may also store instructions for implementation on the controller 56. The controller 56 may further be connected to codec circuitry 54 suitable for carrying out coding and / or decoding of audio, image and / or video data or assisting in coding and / or decoding carried out by the controller.
[0089] The apparatus 50 may further comprise a card reader 48 and a smart card 46, for example, a universal integrated circuit card (UICC) and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network.
[0090] The apparatus 50 may comprise radio interface circuitry 52 connected to the controller and suitable for generating wireless communication signals, for example, for communication with a cellular communications network, a wireless communications system or a wireless local area network. The apparatus 50 may further comprise an antenna 44 connected to the radio interface circuitry 52 for transmitting radio frequency signals generated at the radio interface circuitry 52 to other apparatus(es) and / or for receiving radio frequency signals from other apparatus(es).
[0091] The apparatus 50 may comprise a camera 42 capable of recording or detecting individual frames which are then passed to the codec circuitry 54 or the controller for processing. The apparatus may receive the video image data for processing from another device prior to transmission and / or storage. The apparatus 50 may also receive either wirelessly or by a wired connection the image for coding / decoding. The structural elements of apparatus 50 described above represent examples of means for performing a corresponding function.
[0092] With respect to FIG. 3, an example of a system within which embodiments of the examples described herein can be utilized is shown. The system 10 comprises multiple communication devices which can communicate through one or more networks. The system 10 may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network (such as a GSM, UMTS, CDMA, LTE, 4G, 5G network etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a Bluetooth personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and the Internet.
[0093] The system 10 may include both wired and wireless communication devices and / or apparatus 50 suitable for implementing embodiments of the examples described herein.
[0094] For example, the system shown in FIG. 3 shows a mobile telephone network 11 and a representation of the internet 28. Connectivity to the internet 28 may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.
[0095] The example communication devices shown in the system 10 may include, but are not limited to, an electronic device or apparatus 50, a combination of a personal digital assistant (PDA) and a mobile telephone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, a notebook computer 22, or a head-mounted apparatus 21, which head-mounted apparatus 21 may be a head-mounted display (HMD), or glasses having a camera or other device used for processing images and / or video. The apparatus 50 may be stationary or mobile when carried by an individual who is moving. The apparatus 50 may also be located in a mode of transport including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle or any similar suitable mode of transport.
[0096] The embodiments may also be implemented in a set-top box; e.g. a digital TV receiver, which may / may not have a display or wireless capabilities, in tablets or (laptop) personal computers (PC), which have hardware and / or software to process neural network data, in various operatingsystems, and in chipsets, processors, DSPs and / or embedded systems offering hardware / software based coding.
[0097] Some or further apparatus may send and receive calls and messages and communicate with service providers through a wireless connection 25 to a base station 24. The base station 24 may be connected to a network server 26 that allows communication between the mobile telephone network 11 and the internet 28. The system may include additional communication devices and communication devices of various types.
[0098] The communication devices may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocolinternet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11, 3GPP Narrowband loT and any similar wireless communication technology. A communications device involved in implementing various embodiments of the examples described herein may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.
[0099] In telecommunications and data networks, a channel may refer either to a physical channel or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire, whereas a logical channel may refer to a logical connection over a multiplexed medium, capable of conveying several logical channels. A channel may be used for conveying an information signal, for example a bitstream, from one or several senders (or transmitters) to one or several receivers.
[0100] The embodiments may also be implemented in so-called loT devices. The Internet of Things (loT) may be defined, for example, as an interconnection of uniquely identifiable embedded computing devices within the existing Internet infrastructure. The convergence of various technologies has and may enable many fields of embedded systems, such as wireless sensor networks, control systems, home / building automation, etc. to be included in the Internet of Things (loT). In order to utilize the Internet loT devices are provided with an IP address as a unique identifier. loT devices may be provided with a radio transmitter, such as a WLAN or Bluetooth transmitter or a RFID tag. Alternatively, loT devices may have access to an IP -based network via a wired network, such as an Ethernet-based network or a power-line connection (PLC).
[0101] FIG. 4 shows a block diagram of a general structure of a video encoder. FIG. 4 presents an encoder for two layers, but it would be appreciated that presented encoder could be similarly extended to encode more than two layers. FIG. 4 illustrates a video encoder comprising a first encoder section 401 for a base layer and a second encoder section 451 for an enhancement layer. Each of the first encoder section 401 and the second encoder section 451 may comprise similar elements for encoding incoming pictures. The encoder sections 401, 451 may comprise a pixel predictor 402, 452, prediction error encoder 403, 453 and prediction error decoder 404, 454. FIG. 4 also shows an embodiment of the pixel predictor 402, 452 as comprising an inter-predictor 406, 456, an intra-predictor 408, 458, a mode selector 410, 460, a filter 416, 466, and a reference frame memory 418, 468. The pixel predictor 402 of the first encoder section 401 receives base layer picture(s) / image(s) 400 of a video stream to be encoded at both the inter-predictor 406 (which determines the difference between the image and a motion compensated reference frame) and the intra-predictor 408 (which determines a prediction for an image block based only on the already processed parts of current frame or picture). The output of both the inter-predictor and the intra-predictor are passed to the mode selector 410. The intra-predictor 408 may have more than one intra-prediction modes. Hence, each mode may perform the intra-prediction and provide the predicted signal to the mode selector 410. The mode selector 410 also receives a copy of the base layer image(s) 400. Correspondingly, the pixel predictor 452 of the second encoder section 451 receives enhancement layer picture(s) / images(s) 450 of a video stream to be encoded at both the interpredictor 456 (which determines the difference between the image and a motion compensated reference frame) and the intra-predictor 458 (which determines a prediction for an image block based only on the already processed parts of current frame or picture). The output of both the inter-predictor and the intra- predictor are passed to the mode selector 460. The intra-predictor 458 may have more than one intraprediction modes. Hence, each mode may perform the intra-prediction and provide the predicted signal to the mode selector 460. The mode selector 460 also receives a copy of the enhancement layer pictures 450.
[0102] Depending on which encoding mode is selected to encode the current block, the output of the inter-predictor 406, 456 or the output of one of the optional intra-predictor modes or the output of a surface encoder within the mode selector is passed to the output of the mode selector 410, 460. The output of the mode selector 410, 460 is passed to a first summing device 421, 471. The first summing device may subtract the output of the pixel predictor 402, 452 from the base layer image(s) 400 / enhancement layer image(s) 450 to produce a first prediction error signal 420, 470 which is input to the prediction error encoder 403, 453.
[0103] The pixel predictor 402, 452 further receives from a preliminary reconstructor 439, 489 the combination of the prediction representation of the image block 412, 462 and the output 438, 488of the prediction error decoder 404, 454. The preliminary reconstructed image 414, 464 may be passed to the intra-predictor 408, 458 and to the fdter 416, 466. The filter 416, 466 receiving the preliminary representation may filter the preliminary representation and output a final reconstructed image 440, 490 which may be saved in the reference frame memory 418, 468. The reference frame memory 418 may be connected to the inter-predictor 406 to be used as the reference image against which a future base layer image 400 is compared in inter-prediction operations. Subject to the base layer being selected and indicated to be source for inter-layer sample prediction and / or inter-layer motion information prediction of the enhancement layer according to some embodiments, the reference frame memory 418 may also be connected to the inter-predictor 456 to be used as the reference image against which a future enhancement layer image(s) 450 is compared in inter-prediction operations. Moreover, the reference frame memory 468 may be connected to the inter-predictor 456 to be used as the reference image against which the future enhancement layer image(s) 450 is compared in inter-prediction operations.
[0104] Filtering parameters from the filter 416 of the first encoder section 401 may be provided to the second encoder section 451 subject to the base layer being selected and indicated to be source for predicting the filtering parameters of the enhancement layer according to some embodiments.
[0105] The prediction error encoder 403, 453 comprises a transform unit 442, 492 and a quantizer 444, 494. The transform unit 442, 492 transforms the first prediction error signal 420, 470 to a transform domain. The transform is, for example, the DCT transform. The quantizer 444, 494 quantizes the transform domain signal, for example, the DCT coefficients, to form quantized coefficients.
[0106] The prediction error decoder 404, 454 receives the output from the prediction error encoder 403, 453 and performs the opposite processes of the prediction error encoder 403, 453 to produce a decoded prediction error signal 438, 488 which, when combined with the prediction representation of the image block 412, 462 at the second summing device 439, 489, produces the preliminary reconstructed image 414, 464. The prediction error decoder may be considered to comprise a dequantizer 446, 496, which dequantizes the quantized coefficient values, for example, DCT coefficients, to reconstruct the transform signal and an inverse transformation unit 448, 498, which performs the inverse transformation to the reconstructed transform signal wherein the output of the inverse transformation unit 448, 498 includes reconstructed block(s). The prediction error decoder may also comprise a block filter which may filter the reconstructed block(s) according to further decoded information and filter parameters.
[0107] The entropy encoder 430, 480 receives the output of the prediction error encoder 403, 453 and may perform a suitable entropy encoding / variable length encoding on the signal to provide acompressed signal. The outputs of the entropy encoders 430, 480 may be inserted into a bitstream, for example, by a multiplexer 465.
[0108] The Advanced Video Coding standard (which may be abbreviated H.264, AVC or H.264 / AVC) was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunications Standardization Sector of International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). The H.264 / AVC standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, also known as MPEG-4 Part 10 Advanced Video Coding (AVC). There have been multiple versions of the H.264 / AVC standard, each integrating new extensions or features to the specification. These extensions include Scalable Video Coding (SVC) and Multiview Video Coding (MVC).
[0109] The High Efficiency Video Coding standard (which may be abbreviated H.265, HEVC or H.265 / HEVC) was developed by the Joint Collaborative Team - Video Coding (JCT-VC) of VCEG and MPEG. The standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2, also known as MPEG- H Part 2 High Efficiency Video Coding (HEVC). Extensions to H.265 / HEVC include scalable, multiview, three-dimensional, and fidelity range extensions, which may be referred to as SHVC, MV- HEVC, 3D-HEVC, and REXT, respectively. The references in this description to H.265 / HEVC, SHVC, MV-HEVC, 3D-HEVC and REXT that have been made for the purpose of understanding definitions, structures or concepts of these standard specifications are to be understood to be references to the latest versions of these standards that were available before the date of this application, unless otherwise indicated.
[0110] Versatile Video Coding (which may be abbreviated VVC, H.266, or H.266 / VVC) is a video compression standard developed as the successor to HEVC. VVC is specified in ITU-T Recommendation H.266 and equivalently in ISO / IEC 23090-3, which is also referred to as MPEG-I Part 3.
[0111] A specification of the AV 1 bitstream format and decoding process were developed by the Alliance of Open Media (AOM). The AVI specification was published in 2018. AOM is reportedly working on the AV2 specification.
[0112] An elementary unit for the input to an encoder and the output of a decoder, respectively, in most cases is a picture. A picture given as an input to an encoder may also be referred to as a sourcepicture, and a picture decoded by a decoded may be referred to as a decoded picture or a reconstructed picture.
[0113] The source and decoded pictures are each comprised of one or more sample arrays, such as one of the following sets of sample arrays:- Luma (Y) only (monochrome).- Luma and two chroma (Y CbCr or Y CgCo).- Green, Blue and Red (GBR, also known as RGB).- Arrays representing other unspecified monochrome or tri-stimulus color samplings (for example, YZX, also known as XYZ).
[0114] In the following, these arrays may be referred to as luma (or L or Y) and chroma, where the two chroma arrays may be referred to as Cb and Cr or Cg and Co; regardless of the actual color representation method in use. The actual color representation method in use may be indicated, e.g., in a coded bitstream e.g., using the Video Usability Information (VUI) syntax. A component may be defined as an array or single sample from one of the three sample arrays (luma and two chroma) or the array or a single sample of the array that compose a picture in monochrome format.
[0115] A picture may be defined to be either a frame or a field. A frame comprises a matrix of luma samples and possibly the corresponding chroma samples. A field is a set of alternate sample rows of a frame and may be used as encoder input, when the source signal is interlaced. Chroma sample arrays may be absent (and hence monochrome sampling may be in use) or chroma sample arrays may be subsampled when compared to luma sample arrays.
[0116] Some chroma formats may be summarized as follows:- In monochrome sampling there is only one sample array, which may be nominally considered the luma array.- In 4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array.- In 4:2:2 sampling, each of the two chroma arrays has the same height and half the width of the luma array.- In 4:4:4 sampling when no separate color planes are in use, each of the two chroma arrays has the same height and width as the luma array.
[0117] Coding formats or standards may allow to code sample arrays as separate color planes into the bitstream and respectively decode separately coded color planes from the bitstream. Whenseparate color planes are in use, each one of them is separately processed (by the encoder and / or the decoder) as a picture with monochrome sampling.
[0118] Video codec consists of an encoder that transforms the input video into a compressed representation suited for storage / transmission and a decoder that can decompress the compressed video representation back into a viewable form. Typically encoder discards some information in the original video sequence in order to represent the video in a more compact form (that is, at lower bitrate).
[0119] Typical hybrid video codecs, for example ITU-T H.263 and H.264, encode the video information in two phases. Firstly pixel values in a certain picture area (or “block”) are predicted for example by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Secondly the prediction error, i.e. the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the difference in pixel values using a specified transform (e.g. Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate).
[0120] Inter prediction, which may also be referred to as temporal prediction, motion compensation, or motion-compensated prediction, exploits temporal redundancy. In inter prediction the sources of prediction are previously decoded pictures (a.k.a. reference pictures).
[0121] In temporal inter prediction, the sources of prediction are previously decoded pictures in the same scalable layer. In intra block copy (IBC; a.k.a. intra-block-copy prediction), prediction may be applied similarly to temporal inter prediction but the reference picture is the current picture and only previously decoded samples can be referred in the prediction process. Inter-layer or inter-view prediction may be applied similarly to temporal inter prediction, but the reference picture is a decoded picture from another scalable layer or from another view, respectively. In some cases, inter prediction may refer to temporal inter prediction only, while in other cases inter prediction may refer collectively to temporal inter prediction and any of intra block copy, inter-layer prediction, and inter-view prediction provided that they are performed with the same or similar process than temporal prediction. Inter prediction, temporal inter prediction, or temporal prediction may sometimes be referred to as motion compensation or motion-compensated prediction.
[0122] Intra prediction utilizes the fact that adjacent pixels within the same picture are likely tobe correlated. Intra prediction can be performed in spatial or transform domain, i.e., either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.
[0123] One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy-coded more efficiently if they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors and only the difference relative to the motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.
[0124] The decoder reconstructs the output video by applying prediction means similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in spatial pixel domain). After applying prediction and prediction error decoding means the decoder sums up the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) can also apply additional filtering means to improve the quality of the output video before passing it for display and / or storing it as prediction reference for the forthcoming frames in the video sequence.
[0125] Image and video codecs may use a set of filters, which may enhance the visual quality of the predicted visual content. Filters may be applied either in-loop or out-of-loop, or both. In-loop filters (which may be also called loop filters) are used in reconstructing prediction reference that may be used for predicting forthcoming video signal. In other words, in the case of in-loop filters, the filter applied on one block in the currently encoded frame may affect the encoding of another block in the same frame and / or in another frame which is predicted from the current frame. An in-loop filter may affect the bitrate and / or the visual quality. In fact, an enhanced block may cause a smaller residual (difference between original block and predicted-and-filtered block), thus requiring less bits to be encoded. An out- of-the loop filter (which may also be called a post-processing filter or a post-filter) may be applied on a frame or part of a frame after it has been reconstructed, the filtered visual content may not be used as a source for prediction, and thus it may only impact the visual quality of the frames that are output by the decoder.
[0126] In typical video codecs the motion information is indicated with motion vectors associated with each motion compensated image block. Each of these motion vectors represents the displacementof the image block in the picture to be coded (in the encoder side) or decoded (in the decoder side) and the prediction source block in one of the previously coded or decoded pictures.
[0127] In order to represent motion vectors efficiently those are typically coded differentially with respect to block specific predicted motion vectors. In typical video codecs the predicted motion vectors are created in a predefined way, for example calculating the median of the encoded or decoded motion vectors of the adjacent blocks.
[0128] Another way to create motion vector predictions is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures and signalling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index is typically predicted from adjacent blocks and / or or co-located blocks in temporal reference picture.
[0129] Moreover, typical high efficiency video codecs employ an additional motion information coding / decoding mechanism, often called merging / merge mode, where all the motion field information, which includes motion vector and corresponding reference picture index for each available reference picture list, is predicted and used without any modification / correction. Similarly, predicting the motion field information is carried out using the motion field information of adjacent blocks and / or co-located blocks in temporal reference pictures and the used motion field information is signalled among a list of motion field candidate list filled with motion field information of available adjacent / co-located blocks.
[0130] In typical video codecs the prediction residual after motion compensation is first transformed with a transform kernel (like DCT) and then coded. The reason for this is that often there still exists some correlation among the residual and transform can in many cases help reduce this correlation and provide more efficient coding.
[0131] Typical video encoders utilize Lagrangian cost functions to find optimal coding modes, e.g. the desired Macroblock mode and associated motion vectors. This kind of cost function uses a weighting factor X to tie together the (exact or estimated) image distortion due to lossy coding methods and the (exact or estimated) amount of information that is required to represent the pixel values in an image area:C = D + XR where C is the Lagrangian cost to be minimized, D is the image distortion (e.g. Mean Squared Error) with the mode and motion vectors considered, and R the number of bits needed to represent the requireddata to reconstruct the image block in the decoder (including the amount of data to represent the candidate motion vectors).
[0132] An out-of-band transmission, signaling, or storage may refer to the capability of transmitting, signaling, or storing information in a manner that associates the information with a video bitstream. The out-of-band transmission may use a more reliable transmission mechanism compared to the protocols used for carrying coded video data, such as slices. The out-of-band transmission, signaling or storage may additionally or alternatively be used, e.g., for ease of access or session negotiation. For example, a sample entry of a track in a fde conforming to the ISO Base Media File Format may comprise parameter sets, while the coded data in the bitstream is stored elsewhere in the fde or in another fde. Another example of out-of-band transmission, signaling, or storage comprises including information, such as NN and / or NN updates in a fde format track that is separate from track(s) including coded video data.
[0133] The phrase along the bitstream (e.g., indicating along the bitstream) or along a coded unit of a bitstream (e.g., indicating along a coded tile) may be used in claims and described embodiments to refer to transmission, signaling, or storage in a manner that the ‘out-of-band’ data is associated with, but not included within, the bitstream or the coded unit, respectively. The phrase decoding along the bitstream or along a coded unit of a bitstream or alike may refer to decoding the referred out-of-band data (which may be obtained from out-of-band transmission, signaling, or storage) that is associated with the bitstream or the coded unit, respectively. For example, the phrase along the bitstream may be used when the bitstream is included in a container fde, such as a fde conforming to the ISO Base Media File Format, and certain fde metadata is stored in the fde in a manner that associates the metadata to the bitstream, such as boxes in the sample entry for a track including the bitstream, a sample group for the track including the bitstream, or a timed metadata track associated with the track including the bitstream. In another example, the phrase along the bitstream may be used when the bitstream is made available as a stream over a communication protocol and a media description, such as a streaming manifest, is provided to describe the stream.
[0134] A bitstream may be defined as a sequence of bits or a sequence of syntax structures. A bitstream format may constrain the order of syntax structures in the bitstream.
[0135] A syntax element may be defined as an element of data represented in a bitstream. A syntax structure may be defined as zero or more syntax elements present together in a bitstream in a specified order.
[0136] Syntax structures may be specified, for example, using arithmetic, logical, relational, bitwise, and assignment operators similar to those available in many programming languages. For example, & may indicate a bit-wise ‘AND’ operation. Furthermore, syntax structures may be specified with reference to mathematical functions.
[0137] Syntax structures and semantics may use the values of variables derived from the values of syntax elements. Naming conventions may be defined for variables. For example, variables may be named by a mixture of lower case and upper case letter and without any underscore characters. Variables starting with an upper case letter may be derived for the decoding of the current syntax structure and all depending syntax structures. Variables starting with an upper case letter may, in some cases, be used in the decoding process for later syntax structures without mentioning the originating syntax structure of the variable. Variables starting with a lower case letter may only be used in relation to the syntax structure or function they have been defined for.
[0138] An elementary unit for the output of a video encoder and the input of a video decoder, respectively, may be a network abstraction layer (NAL) unit. For transport over packet-oriented networks or storage into structured files, NAL units may be encapsulated into packets or similar structures. A bytestream format encapsulating NAL units may be used for transmission or storage environments that do not provide framing structures. The bytestream format may separate NAL units from each other by attaching a start code in front of each NAL unit. To avoid false detection of NAL unit boundaries, encoders may run a byte-oriented start code emulation prevention algorithm, which may add an emulation prevention byte to the NAL unit payload, when a start code would have occurred otherwise. In order to enable straightforward gateway operation between packet and stream-oriented systems, start code emulation prevention may be performed regardless of whether the bytestream format is in use or not. A NAL unit may be defined as a syntax structure including an indication of the type of data to follow and bytes including that data in the form of a raw byte sequence payload interspersed as necessary with emulation prevention bytes. A raw byte sequence payload (RBSP) may be defined as a syntax structure including an integer number of bytes that is encapsulated in a NAL unit. An RBSP is either empty or has the form of a string of data bits including syntax elements followed by an RBSP stop bit and followed by zero or more subsequent bits equal to 0.
[0139] A bitstream may be defined to logically include a syntax structure, such as a NAL unit, when the syntax structure is transmitted along the bitstream but may be included in the bitstream according to the bitstream format. A bitstream may be defined to natively comprise a syntax structure, when the bitstream includes the syntax structure.
[0140] In some coding formats or standards, a bitstream may be in the form of a network abstraction layer (NAL) unit stream or a byte stream, that forms the representation of coded pictures and associated data forming one or more coded video sequences.
[0141] In some formats or standards, a first bitstream may be followed by a second bitstream in the same logical channel, such as in the same file or in the same connection of a communication protocol. An elementary stream (in the context of video coding) may be defined as a sequence of one or more bitstreams.
[0142] In some coding formats, such as AVI, a bitstream may comprise a sequence of open bitstream units (OBUs). An OBU comprises a header and a payload, wherein the header identifies a type of the OBU. Furthermore, the header may comprise a size of the payload in bytes.
[0143] In some coding standards, NAL units include a header and payload. The NAL unit header indicates the type of the NAL unit. In some coding standards, the NAL unit header indicates a scalability layer identifier (e.g., called nuh_layer_id in H.265 / HEVC and H.266 / VVC), which may be used, e.g., for indicating spatial or quality layers, views of a multiview video, or auxiliary layers (such as depth maps or alpha planes). In some coding standards, the NAL unit header includes a temporal sublayer identifier, which may be used for indicating temporal subsets of the bitstream, such as a 30-frames-per- second subset of a 60-frames-per-second bitstream.
[0144] Bitstreams or coded video sequences may be encoded to be temporally scalable as follows. Each picture may be assigned to a particular temporal sub-layer. A temporal sub-layer may be equivalently called a sub-layer, temporal sublayer, sublayer, or temporal level. Temporal sub-layers may be enumerated, e.g., from 0 upwards. The lowest temporal sub-layer, sub-layer 0, may be decoded independently. Pictures at temporal sub-layer 1 may be predicted from reconstructed pictures at temporal sub-layers 0 and 1. Pictures at temporal sub-layer 2 may be predicted from reconstructed pictures at temporal sub-layers 0, 1, and 2, and so on. In other words, a picture at temporal sub-layer N does not use any picture at temporal sub-layer greater than N as a reference for inter prediction. The bitstream created by excluding all pictures greater than or equal to a selected sub-layer value and including pictures remains conforming.
[0145] Each picture of a temporally scalable bitstream may be assigned with a temporal identifier (also known as temporal layer identifier, temporal sublayer identifier, or temporal layer ID), which may be, for example, assigned to a variable Temporalld. The temporal identifier may, for example, be indicated in a NAL unit header or in an OBU extension header. Temporalld equal to 0 corresponds tothe lowest temporal level. The bitstream created by excluding all coded pictures having a Temporalld greater than or equal to a selected value and including all other coded pictures remains conforming. Consequently, a picture having Temporalld equal to tid value does not use any picture having a Temporalld greater than tid value as a prediction reference.
[0146] NAL units may be categorized into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units are typically coded slice NAL units.
[0147] A non-VCL NAL unit may be, for example, one of the following types: a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a supplemental enhancement information (SEI) NAL unit, an access unit delimiter, an end of sequence NAL unit, an end of bitstream NAL unit, or a filler data NAL unit. Parameter sets may be needed for the reconstruction of decoded pictures, whereas many of the other non-VCL NAL units are not necessary for the reconstruction of decoded sample values.
[0148] Video coding specifications may enable the use of supplemental enhancement information (SEI) messages or alike. Some video coding specifications include SEI NAL units, and some video coding specifications include both prefix SEI NAL units and suffix SEI NAL units, where the former type can start a picture unit (PU) or alike and the latter type can end a picture unit or alike. An SEI NAL unit includes one or more SEI messages, which are not required for the decoding of output pictures but may assist in related processes, such as picture output timing, post-processing of decoded pictures, rendering, error detection, error concealment, and resource reservation. Several SEI messages are specified in H.264 / AVC, H.265 / HEVC, H.266 / VVC, and H.274 / VSEI standards, and the user data SEI messages enable organizations and companies to specify SEI messages for their own use. The standards may include the syntax and semantics for the specified SEI messages but a process for handling the messages in the recipient might not be defined. Consequently, encoders may be required to follow the standard specifying a SEI message when they create SEI message(s), and decoders might not be required to process SEI messages for output order conformance. One of the reasons to include the syntax and semantics of SEI messages in standards is to allow different system specifications to interpret the supplemental information identically and hence interoperate. It is intended that system specifications can require the use of particular SEI messages both in the encoding end and in the decoding end, and additionally the process for handling particular SEI messages in the recipient can be specified.
[0149] Some video coding specifications enable metadata OBUs. A metadata OBU comprises a type field, which specifies the type of metadata.
[0150] In some coding formats, a coded video sequence (CVS) may be defined as a sequence of coded pictures in decoding order that is independently decodable and is followed by another coded video sequence or the end of the bitstream.
[0151] In some coding formats, such as AVI, a coded video sequence comprises one or more temporal units. A temporal unit consists of a series of OBUs starting from a temporal delimiter, optional sequence headers, optional metadata OBUs, a sequence of one or more frame headers, each followed by zero or more tile group OBUs as well as optional padding OBUs. A temporal unit may be defined to comprise all the OBUs that are associated with a specific, distinct time instant. A temporal unit may comprise a temporal delimiter OBU, and all the OBUs that follow, up to but not including the next temporal delimiter. A temporal delimiter OBU may be defined as an indication that the following OBUs will have a different presentation / decoding time stamp from the one of the last frame prior to the temporal delimiter.
[0152] A coded layer video sequence (CUVS) may be defined as a sequence of pictures and associated other data within the same scalable layer (e.g., with the same value of nuh_layer_id in VVC) that is decodable independently of other pictures in the same layer.
[0153] Some codecs use a concept of picture order count (POC). A value of POC is derived for each picture and is non-decreasing with increasing picture position in output order. POC therefore indicates the output order of pictures. POC may be used in the decoding process for example for implicit scaling of motion vectors and for reference picture list initialization. Furthermore, POC may be used in the verification of output order conformance. The variable including a POC value of a picture may be referred to as PicOrderCntVal.
[0154] A Decoded Picture Buffer (DPB) may be used in the encoder and / or in the decoder. There may be two reasons to buffer decoded pictures, for references in inter prediction and for reordering decoded pictures into output order. Some coding formats, such as HEVC, provide a great deal of flexibility for both reference picture marking and output reordering, separate buffers for reference picture buffering and output picture buffering may waste memory resources. Hence, the DPB may include a unified decoded picture buffering process for reference pictures and output reordering. A decoded picture may be removed from the DPB when it is no longer used as a reference and is not needed for output.
[0155] Output order may be defined as the order in which the decoded pictures are output fromthe decoded picture buffer (for the decoded pictures that are to be output from the decoded picture buffer).
[0156] Output time may be defined as a time when a decoded picture is to be output from a decoder or from the DPB of a decoder (for the decoded pictures that are to be output from the DPB), for example as specified by a hypothetical reference decoder specification according to the output timing DPB operation.
[0157] Pictures having the same output order may be defined to mean the same as pictures having the same output time.
[0158] Decoding order may be defined as the order in which syntax elements are processed by the decoding process. It may be required that syntax elements are ordered in a bitstream in their decoding order.
[0159] An identifier may be defined as a syntax element that identifies a syntax structure . A value of the identifier may for example differ in different instances of the same syntax structure, such as a parameter set. A particular instance of the syntax structure may be referenced through its identifier value. For example, a parameter set that is referenced by the (de)coding of a coded video slice may be identified by providing the identifier value of the parameter set in a header of the coded video slice.
[0160] An indicator (ide) may be defined as a syntax element whose value indicates a selection among more than two values (for which semantics have been specified). An indicator syntax element may have _idc postfix in its name.
[0161] Video coding specifications may enable the use of supplemental enhancement information (SEI) messages or alike. Several SEI messages are specified in H.264 / AVC, H.265 / HEVC, H.266 / VVC, and H.274 / VSEI standards, and the user data SEI messages enable organizations and companies to specify SEI messages for their own use. The standards may contain the syntax and semantics for the specified SEI messages, which the encoders are required to follow, but the decoders might not be required to process SEI messages for output order conformance.
[0162] Video usability information (VUI) may be defined as a syntax structure that identifies properties of interpretation of decoded pictures for display purposes, particularly including color representation information. VUI parameters may include, but may not be limited to, one or more of the following: progressive source indication flag, interlaced source indication flag, frame packing indication flag, projected content indication flag, sample aspect ratio information, overscan information,color primaries, transfer characteristics, matrix coefficients for color conversion, sample value range indication (e.g., indicative of full range or a studio range), chroma sample location information. VUI parameters may be carried as part of a parameter set, such as a sequence parameter set or a video parameter set.
[0163] ITU-T Recommendation H.274, which is equivalent to ISO / IEC 23002-7, may be called "versatile supplemental enhancement information messages for coded video bitstreams" and be referred to as "versatile supplemental enhancement information" or VSEI. The VSEI standard specifies the syntax and semantics of video usability information (VUI) parameters and supplemental enhancement information (SEI) messages. The VUI parameters and SEI messages defined in the VSEI standard are designed to be conveyed within coded video bitstreams in a manner specified in a video coding specification or to be conveyed by other means determined by the specifications for systems that make use of such coded video bitstreams. The VSEI standard is intended for use with VVC coded video bitstreams, although it is drafted in a manner intended to be sufficiently generic that it may also be used with other types of coded video bitstreams. VUI parameters and SEI messages may, for example, assist in processes related to decoding, display or other purposes. In the VSEI standard, the vui_parameters( payloadSize ) syntax structure is specified for the VUI, where payloadSize is an input argument indicating the number of bits in the VUI.
[0164] Frame packing may be defined to comprise arranging more than one input picture, which may be referred to as (input) constituent frames, into an output picture, or arranging the input pictures as a temporal interleaving of alternating first and second constituent frames.
[0165] A constituent frame parity may be defined as a first constituent frame or a second constituent frame, or equivalent as constituent frame 0 or constituent frame 1.
[0166] In general, frame packing is not limited to any particular type of constituent frames or the constituent frames need not have a particular relation with each other. In many cases, frame packing is used for arranging constituent frames of a stereoscopic video clip into a single picture sequence, as explained in more details in the next paragraph. The arranging may include placing the input pictures in spatially non-overlapping areas within the output picture . For example, in a side-by-side arrangement, two input pictures are placed within an output picture horizontally adjacently to each other. The arranging may also include partitioning of one or more input pictures into two or more constituent frame partitions and placing the constituent frame partitions in spatially non-overlapping areas within the output picture. The output picture or a sequence of frame-packed output pictures may be encoded into a bitstream e.g., by a video encoder. The bitstream may be decoded e.g., by a video decoder. Thedecoder or a post-processing operation after decoding may extract the decoded constituent frames from the decoded picture(s) e.g., for displaying.
[0167] In frame-compatible stereoscopic video (a.k.a. frame packing of stereoscopic video), a spatial packing of a stereo pair into a single frame is performed at the encoder side as a pre-processing step for encoding and then the frame-packed frames are encoded with a conventional 2D video coding scheme. The output frames produced by the decoder includes constituent frames of a stereo pair.
[0168] In a typical operation mode, the spatial resolution of the original frames of each view and the packaged single frame have the same resolution. In this case the encoder downsamples the two views of the stereoscopic video before the packing operation. The spatial packing may use for example a side-by-side or top-bottom format, and the downsampling need to be performed accordingly.
[0169] An encoder may indicate the use of frame packing by including one or more frame packing arrangement SEI messages, e.g., as defined in VSEI, in the bitstream. Likewise, a decoder may conclude the use of frame packing by decoding one or more frame packing arrangement SEI messages from the bitstream. When a frame packing arrangement SEI message applies to the CLVS, a cropped decoded picture includes samples of multiple distinct spatially packed constituent frames that are packed into one frame, or that the output cropped decoded pictures in output order form a temporal interleaving of alternating first and second constituent frames, using an indicated frame packing arrangement scheme. This information may be used by the decoder to appropriately rearrange the samples and process the samples of the constituent frames appropriately for display or other purposes.
[0170] In some video codecs, video usability information (VUI) may be included in a sequence parameter set (SPS). VUI specified in VSEI comprises the following: vui_non_packed_constraint_flag equal to 1 specifies that there may not be any frame packing arrangement SEI messages present in the bitstream that apply to the CLVS. vui_non_packed_constraint_flag equal to 0 does not impose such a constraint.
[0171] Information on neural network (NN) post-filtering
[0172] A picture that is decoded by an image or video decoder may be further processed by a postprocessing operation. The post-processing operation may comprise one or more neural networks (NNs), and / or one or more non-NNs operations. The post-processing operation may take as input one or more input pictures, and may output one or more output pictures. Post-processing may be used for several reasons or use cases, including, but not limited to, the following:- Enhancing the picture with respect to one or more objective metrics, such as the peak signal-to-noise ratio (PSNR), where the objective metrics may be derived based onpixel-wise distortion metrics (as in the case of PSNR), or may be derived based on perceptual metrics (as in the case of the video multi -method assessment fusion, VMAF, or feature-based metrics). Enhancing may comprise removing coding artifacts, denoising, and the like;- Enhancing the picture with respect to one or more subjective scores or metrics. Enhancing may comprise removing coding artifacts, denoising, and the like;- Upsampling and super-resolution;- Temporal interpolation of one or more output pictures given two or more input pictures;- Temporal extrapolation of one or more output pictures given one or more input pictures; and- Colorization.
[0173] An encoder, such as a video encoder, or a transmitter may signal information to a decoder, such as a video decoder, or a receiver, where the information is indicative of, but not limited to, the following information:- Whether to apply a post-processing operation;- The characteristics of the post-processing operation, for example in terms of the format of the input and output of the post-processing operation; and- How to apply the post-processing operation.
[0174] Information on neural-network post-fdter characteristics (NNPFC) and neural-network post-filter activation (NNPFA) supplemental enhancement information (SEI) messages
[0175] The NNPFC SEI message and the NNPFA SEI message have been described in version 3 of the versatile supplemental enhancement information (VSEI) standard.
[0176] The syntax structure specifying the NNPFC SEI message may be called nn_post_fdter_characteristics. The syntax structure specifying the NNPFA SEI message may be called nn_post_filter_activation.
[0177] The NNPFC SEI message comprises the nnpfc_id syntax element, which includes an identifying number that may be used to identify a post-processing filter. A base post-processing filter is the filter that is included in or identified by the first NNPFC SEI message, in decoding order, that has a particular nnpfc_id value within a coded layer video sequence (CLVS). When there is a second NNPFC SEI message that has the same nnpfc_id value that defines the base post-processing filter, an update relative to the base post-processing filter is applied to obtain a post-processing filter associatedwith the nnpfc id value. The update may be obtained by decoding the coded neural network bitstream in the second NNPFC SEI message. Otherwise, the post-processing filter associated with the nnpfc_id value is assigned to be the same as the base post-processing filter.
[0178] The NNPFC SEI message comprises the nnpfc_mode_idc syntax element, the semantics of which may be defined as follows:
[0179] nnpfc_mode_idc equal to 1 specifies that the base post-processing filter or the update relative to the base post-processing filter associated with the nnpfc_id value is a neural network identified by the Uniform Resource Identifier (URI) nnpfc uri with the format identified by the tag URI nnpfc tag uri.
[0180] nnpfc_mode_idc equal to 0 indicates that this SEI message contains an ISO / IEC 15938-17 bitstream that specifies the base post-processing filter or updates relative to the base post-processing filter with the same nnpfc_id value.
[0181] The NNPFC SEI message may also comprise:- Purpose of the post-processing filter, which may comprise, but may not be limited to, one or more of the following:■ Visual quality improvement;■ Chroma upsampling from the 4:2:0 chroma format to the 4:2:2 or 4:4:4 chroma format, or from the 4:2:2 chroma format to the 4:4:4 chroma format;■ Increasing the width or height of the input picture;■ Frame rate upsampling;■ Bit depth upsampling; or■ Colorization.- Formatting of the input tensors that are given as input to the neural network inference- Formatting of the output tensors that are resulting from the neural network inference; and- Characterization of the complexity of the neural network.
[0182] The NNPFC SEI message syntax comprises nnpfc_base_flag. nnpfc_base_flag equal to 1 specifies that the SEI message specifies the base NNPF. nnpfc_base_flag equal to 0 specifies that the SEI message specifies an update relative to the base NNPF.
[0183] When nnpfc_base_flag is equal to 0, the following applies:- This SEI message defines an update relative to the preceding base NNPF in decoding order with the same nnpfc id value. Updates are not cumulative but rather each update is applied on the base NNPF, which is the NNPF specified by the first NNPFC SEI message, in decoding order, that has a particular nnpfc id value within the current CLVS. The NNPF defined by this SEI message is obtained by applying the update defined by this SEI message relative to the base NNPF with the same nnpfc_id value.- This SEI message pertains to the current decoded picture and all subsequent decoded pictures of the current layer, in output order, until the end of the current CLVS or up to but excluding the decoded picture that follows the current decoded picture in output order within the current CLVS and is associated with a subsequent NNPFC SEI message, in decoding order, having nnpfc_base_flag equal to 0 and that particular nnpfc_id value within the current CLVS, whichever is earlier.
[0184] The NNPFC SEI message syntax includes the nnpfc num input _pics_minus 1 syntax element. nnpfc_num_input_pics_minus 1 plus 1 specifies the number of pictures used as input for the NNPF. The variable numlnputPics may be set equal to nnpfc_num_input_pics_minus 1 + 1.
[0185] A frame rate upsampling filter may interchangeably be called a picture rate upsampling filter. Such a filter generates or interpolates one or more pictures between a pair of pictures given as input to the filter. It is also possible to have a frame rate upsampling filter where the number of input pictures may be greater than 2. Such a frame rate upsampling filter may generate pictures between more than one pair of input pictures. A frame rate upsampling filter may comprise a neural network, in which case the generation of the interpolated pictures between a pair of input pictures is performed by the inference of the neural network. It is possible to have a frame rate upsampling filter that extrapolates a picture before input picture(s) or after input picture(s), instead of or in addition to between input pictures.
[0186] When the filtering purpose comprises frame rate upsampling, the NNPFC SEI message includes nnpfc_interpolated_pics[ i ] syntax elements for the values of i in the range of 0, inclusive, to nnpfc_num_input_pics_minusl, exclusive. nnpfc_interpolated_pics[ i ] specifies the number of interpolated pictures generated by the NNPF between the i-th and the ( i + 1 )-th picture used as input for the NNPF.
[0187] The NNPFC SEI message syntax may comprise an indication, which may be called nnpfc_absent_input_pic_zero_flag, that indicates how pictures that would not originate from the current bitstream are expected to be replaced in the input tensor. nnpfc_absent_input_pic_zero_flag equal to 1indicates that the NNPF expects an input picture that is not present in the current bitstream to be represented sample arrays with sample values equal to 0. nnpfc_absent_input_pic_flag equal to 0 indicates that the NNPF expects an input picture that is not present in the current bitstream to be represented by the closest input picture in output order within the current bitstream.
[0188] The NNPFC SEI message syntax may comprise an indication, which may be called nnpfc auxiliary inp idc, that indicates if auxiliary input data in addition to sample array(s) of input picture(s) is present in the input tensor of the NNPF. nnpfc_auxiliary_inp_idc greater than 0 indicates that auxiliary input data is present in the input tensor of the NNPF. Specific semantics may be specified for specific non-zero values of nnpfc auxiliary inp idc. nnpfc auxiliary inp idc equal to 0 indicates that auxiliary input data is not present in the input tensor.
[0189] The NNPFA SEI message specifies the neural -network post-processing filter (NNPF) that may be used for post-processing filtering for the current picture, or for post-processing filtering for the current picture and one or more other pictures. The NNPFA SEI message comprises the nnpfa_target_id syntax element, which indicates that the neural -network post-processing filter with nnpfc id equal to nnpfa_target_id may be used for post-processing filtering for the indicated persistence. The indicated persistence may be the current picture only (indicated by nnpfa_persistence_flag equal to 0). Alternatively, the NNPF activation may be indicated to be persistent by nnpfa_persistence_flag equal to 1, in which case the persistence of the NNPF activation may last until the end of the current CLVS or the next picture, in output order, in the current layer associated with a NNPFA SEI message with the same nnpfa_target_id as the current SEI message.
[0190] The NNPFA SEI message syntax may comprise a syntax element indicative if the base post-processing filter or the latest post-processing filter is activated, where the latest post-processing filter is defined by the base post-processing filter relative to which the latest filter update, if any, has been applied. The syntax element may be called nnpfa_target_base_flag. nnpfa_target_base_flag equal to 1 specifies that the target NNPF is the base NNPF with nnpfc_id equal to nnpfa_target_id. nnpfa_target_base_flag equal to 0 specifies that the target NNPF is the NNPF specified by the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id that precedes the first VCL NAL unit of the current picture in decoding order and is not a repetition of the NNPFC SEI message that contains the base NNPF.
[0191] The NNPFA SEI message syntax may comprise indications which ones of the filtered pictures corresponding to the input pictures are output by the NNPF process. For the i-th input picture that is filtered by the NNPF, the NNPFA SEI message syntax may comprise nnpfa_output_flag[ i ]syntax element, which when equal to 0, specifies that the filtered picture is not output by the NNPF process, and when equal to 1, specifies that the filtered picture is output by the NNPF process.
[0192] In relation to an NNPFA SEI message, two sets of pictures may be defined, namely nnpfcTargetPictures and nnpfaTargetPictures. nnpfcTargetPictures may be defined to be the set of pictures to which the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id that precedes the current NNPFA SEI message in decoding order pertains. nnpfaTargetPictures may be defined to be the set of pictures for which the target NNPF is activated by the current NNPFA SEI message. It may be required for a conforming bitstream that any picture included in nnpfaTargetPictures shall also be included in nnpfcTargetPictures.
[0193] An NNPF process comprises performing the NNPF inference forgiven input pictures. The NNPF inference may be performed in a patch-wise manner so that the entire picture area gets filtered. The NNPF inference may be followed by outputting NNPF -generated pictures in their increasing index order, where all NNPF -generated pictures that were interpolated by the NNPF are output and those NNPF -generated pictures that correspond to any input pictures to the NNPF are output as specified in the semantics of the NNPFA SEI message.
[0194] A general post-processing filtering process using NNPFs may be described as follows. Input to this process is a bitstream BitstreamToFilter. Output of this process is a list of NNPF output pictures ListNnpfOutputPics. First, BitstreamToFilter is decoded, and the list CroppedDecodedPictures is set to be the list of the cropped decoded pictures in output order resulted from decoding BitstreamToFilter. Second, the filtering process for one picture, as described below, is repeatedly invoked, in output order, for each cropped decoded picture that is in CroppedDecodedPictures and for which one or more NNPFs are activated. The order of the pictures in ListNnpfOutputPics is in output order. It may be required that within ListNnpfOutputPics there shall be no more than one picture pertaining to any particular output time instance. When for any particular picture in CroppedDecodedPictures there are multiple NNPFs activated and only one the NNPFs is allowed to be chosen to be applied although any of the NNPFs may be chosen, the above constraint shall apply regardless of which NNPF is chosen to be applied to the particular picture.
[0195] A filtering process for one picture using an NNPF may be described as follows. The filtering process for one picture using an NNPF may be applied to each cropped decoded picture, referred to as the current picture, that is in CroppedDecodedPictures and for which one or more NNPFs are activated. When applying an NNPF to the current picture, the filtered and / or interpolated pictures are generated by the NNPF by applying the NNPF process to the current picture. When applying anNNPF to the current picture, the order of the pictures generated by the NNPF by applying the NNPF process being stored into the output tensor of the NNPF is in output order. When the applied NNPF is the last NNPF that is applied to the current picture, the pictures generated by the NNPF and output by the NNPF process are included into ListNnpfOutputPics, in the same order as when the pictures are stored into the output tensor of the NNPF.
[0196] The use of NNPFC and NNPFA SEI messages for VVC has been described in version 3 of the versatile video coding (VVC) standard. It is to be understood that NNPFC and NNPFA SEI message may be similarly used for any other video coding specification.
[0197] When NNPFC and NNPFA SEI messages are used for VVC, a decoder selects input pictures for the NNPF. The input pictures may be selected in reverse output order starting from a picture for which the NNPF is activated through an NNPFA SEI message. The input pictures may be indexed, starting from index 0 that is assigned for the picture for which the NNPF is activated through an NNPFA SEI message. In an example, the decoder selects the input picture with index i, where i is greater than 0, to be the latest cropped decoded output picture, in output order, that precedes the input picture with index i-1 in output order. When there is no cropped decoded output picture, in output order, that precedes the input picture with index i-1 in output order as a result of decoding the bitstream, it may be considered that the input picture with index i is not present in the current bitstream (e.g., missing) and the subsequent input pictures, when any, with index i+1 to numlnputPics-l, inclusive, are likewise missing. A missing input picture may be treated like described above in relation to nnpfc_absent_input_pic_zero_flag syntax element.
[0198] When NNPFC and NNPFA SEI messages are used for VVC and a picture rate upsampling NNPF that interpolates pictures between a single pair of input pictures is activated persistently until the end of the bitstream, the NNPF is applied repeatedly at the end of the bitstream for different sets of input pictures up to but excluding a set of input pictures that would cause creation of any interpolated picture after the last picture of the bitstream in output order. In these sets of input pictures, some of the pictures may be missing and may be, for example, replaced by the last picture within the bitstream in output order.
[0199] Visual temporal extrapolation
[0200] It is to be understood that, in various embodiments, the terms visual temporal extrapolation, temporal extrapolation, and video prediction may be used interchangeably. Visual temporal extrapolation may be defined as a method, algorithm, or process that generates one or morepictures in the future given one or more past pictures as input. Visual temporal extrapolation may be realized by, but is not necessarily based on or limited to, neural network inference.
[0201] Use cases for visual temporal extrapolation
[0202] Use cases for visual temporal extrapolation include, but are not limited to:- Very low delay computer vision for domains like robotics and autonomous driving, where extrapolated future pictures facilitate anticipatory decision making;- Increase of the rendered picture rate in very low-latency applications, such as cloud gaming, relative to the decoded picture rate;- Reduction of the end-to-end delay in low-latency applications through extrapolating and displaying future pictures before they are received; and- Generative face video for very low bitrate video coding.
[0203] Information on generative artificial intelligence (Al), including temporal extrapolation and generative face video
[0204] The term generative artificial intelligence (Al), or generative modeling, or generative machine learning (and other similar terms), are commonly used to indicate a class of models learned from data that are capable of generating new data. State-of-the-art generative models are based on neural networks. The basic components or layers of a generative NN are usually not different from the components or layers of a non-generative NN. Example of such components are convolution layers, non-linear layers, fully-connected layers, normalization layers, attention layers, etc.
[0205] One typical example of neural network architecture that allows for generating text data is a Trasformer-based “decoder”, where “decoder” may not refer to a decoder that is part of a codec performing compression of input data into a small bitstream. Instead, the decoder is a neural network that gets a set of input words or parts of words or tokens extracted from input words, and outputs a set of output words or parts of words or tokens. At inference time, such a NN is run in auto-regressive mode, where the generated word(s) or token(s) is provided as part of the input word(s) or token(s). In order for such a NN to generate data, it is trained to predict the next word(s) (or an estimate of a probability distribution over the next words) given a set of input words. The NN may be based on the Transformer architecture, which comprises the use of the self-attention mechanism, where an attention score is assigned to each input token or word based on all other input tokens or words, including the previously generated words or tokens. During training of a decoder-style Transformer architecture, the future data items (words or tokens) are masked so not to leak information from the future. In somecases, decoder-style Transformer architectures are referred to as “uni-directional”, because they use or process information from left-to-right, as opposed to some encoder-style Transformer architectures that are referred to as “bi-directional” (because they use or process information from left-to-right and from right-to-left).
[0206] Another example of generative modeling is visual temporal extrapolation, where a picture is generated by a NN based on one or more previously decoded or generated pictures and on one or more other data items. The one or more previously decoded or generated pictures may be pictures decoded by a process that does not involve generative modeling, such as a traditional codec, e.g., a VVC-compliant codec. The one or more data items may include parameters or features that describe the differences between the one or more previously decoded pictures and the current picture to be temporally extrapolated. Examples of such parameters are facial parameters (such as facial keypoints or facial landmarks and their positions or differential positions with respect to the facial landmarks of a previous picture), or parameters of other objects. The one or more data items may be signaled from encoder to decoder.
[0207] Another example of parameters or features that describe the differences between the one or more previously decoded pictures and the current picture to be temporally extrapolated is a facial matrix that may represent the transformation between the one or more previously decoded pictures and the current picture. The matrix may, for example, be one of the following (where * is a multiplication sign):- Affine translation matrix with the size of 2*2 or 3*3;- Covariance matrix with size of 2*2 or 3*3;- Mouth matrix representing mouth motion;- Eye matrix representing the open-close status and level of eyes;- Head rotation paramters with the size of 2*2 or 3*3 representing the head rotation in 2D space or 3D space;- Head translation matrix with the size of 1*2 or 1*3 representing head translationin 2D space or 3D space; or- Head location matrix with size of 1 * 2 or 1 * 3 representing the head location in 2D space or 3D space.
[0208] Facial parameters may include, but need not be limited to, either or both of: facial keypoints and / or one or more facial matrix(es).
[0209] An example of an encoder utilizing generative Al may be described as follows: The input base picture is encoded with any video encoding method, such as VVC, into a coded base picture of abitstream. The input subsequent pictures are fed into the analysis model for extracting feature parameters. The feature parameters are encoded. Feature parameters may be encoded independently of feature parameters of any other picture, e.g., for the first subsequent picture following the base picture in decoding order. Alternatively, feature parameters may be encoded in a predictive manner with reference to earlier encoded feature parameters in decoding order, wherein feature residual may be encoded relative to predicted feature parameters. Feature residual may undergo quantization and entropy coding as part of feature encoding. Encoded feature parameters are included in or along the bitstream, e.g., in a generative face video SEI message. For a subsequent picture, the video encoder may encode a dummy picture or a drive picture, as described subsequently. Encoded feature parameters of a subsequent picture may be included in or along the respective coded dummy picture or drive picture, e.g., in the same picture unit.
[0210] An example of a decoder utilizing generative Al may be described as follows: The decoder decodes a coded base picture from a bitstream. For a subsequent coded dummy or drive picture, the decoder decodes encoded feature parameters. The decoded feature parameters and the decoded base picture; and optionally a decoded drive picture are input to a generative neural network model inference, which generates an output picture.
[0211] Information on generative face video (GFV) SEI message
[0212] In JVET, a generative face video (GFV) SEI message has been proposed for a new version of the Versatile Supplemental Enhancement Information (VSEI) standard, in order to support visual temporal extrapolation of faces in videos. One of the functions of the GFV SEI message is to signal facial parameters, which represent an auxiliary input to the NN that performs the temporal extrapolation. The latest document describing the GFV SEI message is JVET-AF0234 available from [https: / / jvet-(last accessed on December 21, 2023)]. An example summary is provided below.
[0213] The GFV SEI message defines an interface between a video decoder and a generator NN that performs visual temporal extrapolation of faces in videos.
[0214] The following are the types of pictures considered in the GFV SEI message:- Base picture: a decoded output picture that may be used by the generative network to generate a novel face picture. It is coded by a traditional codec such as VVC-compliant codecs;- Dummy picture: picture unit of the minimum allowed resolution that contains only SEI messages; and(optional) Primary picture (also called drive picture) that can be optionally input to a generative NN to improve background texture and / or facial details. This drive picture may also be coded by a traditional codec such as a VVC-compliant codec.
[0215] A second neural network, referred to as Translator NN, is used to convert the facial parameters (signaled within a GFV SEI message) into the following converted parameters:15 3D-keypoints;- One 3x3 matrix; and- One 1x3 matrix.
[0216] The input to the generator NN (which may also be referred to as GenerativeNN) may be one or more of the following:- The (previously decoded) base picture.- The converted parameters (output of Translator NN)- (optionally) The primary / drive picture.
[0217] Various embodiments aims at solving, for example, the problem of enabling the use of generative Al on a receiver or decoder for the purpose of generating a video based on information that is signaled by a transmitter or encoder.
[0218] At least some of the embodiments describe methods and apparatuses for supporting the performing of visual temporal extrapolation.
[0219] Furthermore, at least some of the embodiments describe methods and apparatuses for enabling a Neural Network Post-Filter (NNPFC) SEI message to signal information that is needed at decoder or receiver side for performing visual temporal extrapolation by using one or more postprocessing neural networks.
[0220] In an embodiment, an encoder encodes signaling information in or along a bitstream, and the decoder decodes signaling information from or along a bitstream, that enables or supports performing visual temporal extrapolation at decoder side, where the performing of visual temporal extrapolation may comprise using one or more neural networks (NNs), where the one or more NNs are referred to as VENN for simplicity. In an example, the signaling information may be comprised in a Neural Network Post-Filter Characteristics (NNPFC) Supplemental Enhancement Information (SEI) message.
[0221] In an embodiment, the signaling information may comprise a purpose of performing visual temporal extrapolation at decoder side.
[0222] In an embodiment, the signaling information may comprise a purpose of generative modelling or generative artificial intelligence (Al) at decoder side.
[0223] In an embodiment, the signaling information may comprise information that characterizes a neural network performing temporal extrapolation, which may include, but may not be limited to, information characterizing inputs to and / or outputs from the neural network. This information may be referred to as temporal extrapolation characteristics (TEC) information.
[0224] In an embodiment, the TEC information may be signaled when the purpose comprises visual temporal extrapolation.
[0225] In another embodiment, at least some of the TEC information is signaled when the signaling information indicates that the purpose of one or more neural networks associated with the signaling information comprises one or more of temporal extrapolation, generative modeling, or generative artificial intelligence (Al), where generative modeling or generative Al may refer to using the one or more neural networks for generating data conditional to one or more inputs.
[0226] In an embodiment, the TEC information may be signaled when the purpose comprises generative face video or generative video.
[0227] In an embodiment, the TEC information may comprise information indicative of whether the VENN is capable of extrapolating a variable number of pictures or a fixed number of pictures.
[0228] In an embodiment, the TEC information may comprise information indicative of the maximum number of pictures that the VENN can extrapolate.
[0229] In an embodiment, when the TEC information indicates that the VENN is capable of extrapolating a variable number of pictures, for each inference of the VENN the encoder may indicate to the decoder the number of pictures to be extrapolated.
[0230] In an embodiment, the TEC information may comprise information indicative of the number of pictures that are extrapolated in one activation or inference of the VENN.
[0231] In an embodiment, the TEC information may comprise information indicative of the time interval of pictures being extrapolated relative to the (latest) picture used as input for the inference.
[0232] In an embodiment, the TEC information may comprise information indicative of whether the VENN takes an auxiliary input that indicates output instances of the picture(s) being temporally extrapolated.
[0233] In an embodiment, the TEC information may comprise information indicative of whether the VENN takes an auxiliary input that indicates or is derived from facial parameters, where this auxiliary input may be used by the VENN in order to perform temporal extrapolation of one or more faces in a picture.
[0234] In an embodiment, the TEC information may comprise indication(s) for indicating which facial parameters are present and / or used by the VENN. For example, the TEC information may comprise a first flag for indicating when facial keypoints are in use and a second flag for indicating when a facial matrix is in use.
[0235] In an embodiment, the facial parameters that are used as auxiliary input or that are used to derive the auxiliary input may be carried within a newly defined SEI message, which may, for example, be referred to as the neural -network post-filter input (NNPFI) SEI message.
[0236] In another embodiment, the facial parameters that are used as auxiliary input or that are used to derive the auxiliary input may be carried within a generative face video (GFV) SEI message.
[0237] In an embodiment, the TEC information may comprise information indicative of whether an auxiliary input to the VENN, such as an auxiliary input that indicates or is derived from facial parameters, can be derived from a lossy coding process. In one example, an encoder performs lossy (and, optionally, lossless) compression of facial parameters; a decoder performs lossless and / or lossy decompression of encoded facial parameters; the decompressed facial parameters are input to a VENN.
[0238] In an embodiment, the TEC information may comprise information indicative of whether the VENN takes an auxiliary input that is derived from the output of another neural network, where the another neural network is run or executed at decoder side.
[0239] In an embodiment, the TEC information may comprise information indicative of whether the VENN takes an auxiliary input that is derived from the output of another neural network that converts or translates facial parameters to a format that is accepted or required by the VENN. The facialparameters may be carried within an NNPFA SEI message or a GFV SEI message or a NNPFI SEI message (or, more generally, a frame-wise (FW) signaling) or in any other suitable way.
[0240] In an embodiment, the TEC information may comprise information indicative of whether the VENN takes an auxiliary input that may represent or is derived from a tensor that is generated at encoder side.
[0241] In an embodiment, the tensor is (or is derived from) an output of another neural network that is run or executed at encoder side, and the tensor may be signaled within an NNPFA SEI message or a GFV SEI message or NNPFI SEI message (or, more generally, a FW signaling).
[0242] In an embodiment, the TEC information may comprise information indicative of whether the VENN takes an auxiliary input that represents or is derived from audio or speech.
[0243] This information may be used by the VENN to generate or synthesize several aspects of a face, such as mouth, lips, face expression, blinking, and the like.
[0244] In an embodiment, when the VTNN takes an auxiliary input that represents audio or speech, the TEC information may comprise information about characteristics of the audio or speech.
[0245] In an embodiment, the TEC information may comprise information indicative of whether the VTNN takes an auxiliary input that represents or is derived from text.
[0246] This information may be used by the VTNN to generate or synthesize several features of a face, such as mouth, lips, face expression, blinking, etc.
[0247] In an embodiment, the TEC information may comprise information indicative of whether an input to the VTNN can be an input that comprises only background information.
[0248] In an embodiment, the TEC information may comprise information indicative of whether an input to the VTNN can be an input that comprises only foreground information, such as faces.
[0249] In an embodiment, the TEC information may comprise information indicative of whether an input to the VTNN can be a drive picture. In one example, an encoder indicates to a decoder that the VTNN supports two types of input pictures, where a first input picture is a base picture and a second picture is a drive picture.
[0250] In an embodiment, the TEC information may comprise information indicative of whether the VTNN may produce different outputs for different inferences or executions of the VTNN with the same input(s). Such a behavior or characteristic of the VTNN is referred to as output variability of the VTNN.
[0251] In an additional embodiment, the TEC information may comprise information indicative of the type or cause of output variability of the VTNN.
[0252] In an embodiment, the TEC information may comprise information indicative of whether the VTNN can be instructed or operated in a way that it does not yield output variability.
[0253] In an embodiment, the TEC information may comprise information indicative of whether the VTNN is allowed to have output variability.
[0254] In an embodiment, the TEC information may comprise information that makes the VTNN not to have output variability.
[0255] In an embodiment, the TEC information may comprise information indicative of the type of extrapolated output.
[0256] In an embodiment, the signaling information may comprise indicating whether the VTNN is adapted for all the pictures to which it is applied or activated.
[0257] In an embodiment, a VTNN may be adapted in order to specialize it with respect to some aspects of the generated output, such as specializing to a specific face.
[0258] In an embodiment, a VTNN may be adapted in order to specialize it to a specific task or purpose.
[0259] In an embodiment, the signaling information may comprise indicating that the NN is sequence-level adapted.
[0260] In an embodiment, the signaling information may comprise indicating the purpose of the sequence-level adaptation.
[0261] In an embodiment, the signaling information may comprise indicating whether the VTNN is adapted for all pictures to which it is applied by means of inputting a persistent auxiliary input (PAI) signal.
[0262] In an embodiment, an encoder encodes signaling information in or along a bitstream, and the decoder decodes signaling information from or along a bitstream, that enables or supports performing visual temporal extrapolation at decoder side, where the performing of visual temporal extrapolation may comprise using one or more neural networks. For simplicity, it is assumed that the visual temporal extrapolation is performed by one neural network that is referred to as VTNN. It is to be understood that embodiments may be similarly realized by a combination of neural networks, such as an optical flow derivation neural network, followed by a synthesis network that takes an optical flow and one or more pictures as input.
[0263] It is to be understood that embodiments are not limited to any particular type of a neural network. Embodiments may, for example, be realized with a model that takes one or more seed pictures and one or more driving features as input. In an example, the model is a generative model, such as a model based on Normalizing Flows. In another example, such a model is based on Diffusion models (also known as diffusion probabilistic models, or score-based generative models). In yet another example, such a model is based on a convolutional neural network.
[0264] In at least some embodiments, the signaling information is associated to one or more neural networks that perform at least part of the visual temporal extrapolation task.
[0265] In an example, the signaling information may be comprised in a Neural Network PostFilter Characteristics (NNPFC) Supplemental Enhancement Information (SEI) message. In another example, the signaling information may be comprised in a Generative Face Video (GFV) SEI message. In yet another example, the signaling information may be comprised in an NNPFC SEI message and a GFV SEI message.
[0266] In some embodiments, some features may refer to signaling information associated to one or more pictures of a video. This may be achieved by means of a frame-wise signaling mechanism. Such signaling is referred to as frame-wise (FW) signaling. Examples ofFW signaling include a Neural Network Post-Filter Activation (NNPFA) SEI message and a generative face video (GFV) SEI message.
[0267] Purpose
[0268] In an embodiment, the signaling information may comprise a purpose of performing visual temporal extrapolation at decoder side.
[0269] In an additional embodiment, the purpose may be associated with a defined neural network.
[0270] In an additional embodiment, the purpose may be indicated as a bit field wherein a particular bit position is assigned to visual temporal extrapolation and a value of the bit in the particular bit position indicates whether visual temporal extrapolation is among the purposes of the defined neural network. Other bit positions may be assigned to other purposes and may be indicated with visual temporal extrapolation.
[0271] An example embodiment of the signaling information comprising a purpose of performing visual temporal extrapolation.
[0272] Addition of temporal extrapolation as one of the purposes of nnpfc jiurposc. This addition includes the following features:- A new bit in nnpfc_purpose is dedicated for temporal extrapolation.- When nnpfC purposc indicates temporal extrapolation, nnpfc_extrapolated_pics_minus 1 plus 1 indicates the number of extrapolated pictures generated by the NNPF.
[0273] Adding temporal extrapolation purpose
[0274] NNPFC SEI message syntax
[0275] NNPFC SEI message semantics
[0276] The following table provides an example of a definition of nnpfc_purpose. It is to be understood that the presented bitmask assignment is an example and embodiments may be similarly realized with any other bitmask assignment.
[0277] The variables ChromaUpsamplingFlag, ResolutionResamplingFlag, PictureRateUpsamplingFlag, BitDepthUpsamplingFlag, ColourizationFlag, and TemporalExtrapolationFlag, specifying whether nnpfc_purpose indicates the purpose of the NNPF to include chroma upsampling, resolution resampling, picture rate upsampling, bit depth upsampling, colourization, and temporal extrapolation, respectively, are derived as follows:ChromaUpsamplingFlag = ( ( nnpfc_purpose & 0x02 ) > 0 ) ? 1 : 0 ResolutionResamplingFlag = ( ( nnpfc_purpose & 0x04 ) > 0 ) ? 1 : 0 PictureRateUpsamplingFlag = ( ( nnpfc_purpose & 0x08 ) > 0 ) ? 1 : 0 BitDepthUpsamplingFlag = ( ( nnpfc_purpose & 0xl0 ) > 0 ) ? l : 0 ColourizationFlag = ( ( nnpfc_purpose & 0x20 ) > 0 ) ? 1 : 0TemporalExtrapolationFlag = ( ( nnpfc_purpose & 0x40 ) > 0 ) ? 1 : 0
[0278] nnpfc_interpolated_pics[ i ] specifies the number of interpolated pictures generated by the NNPF between the i-th and the ( i + 1 )-th input picture for the NNPF.
[0279] When PictureRateUpsamplingFlag is equal to 1 for an NNPF and the NNPFA SEI message that activated this NNPF has nnpfa_persistence_flag equal to 1, only for a single value of i in the range of 0 to numlnputPics - 1, inclusive, the value of nnpfc_interpolated_pics[ i ] is greater than 0.
[0280] nnpfc_extrapolated_pics_minus 1 plus 1 specifies the number of extrapolated pictures generated by the NNPF subsequent to all input pictures for the NNPF in output order.
[0281] The variables NumlnpPicsInOutputTensor, specifying the number of pictures that have a corresponding input picture and are present in the output tensor of the NNPF, Inpldx[ idx ], specifying the input picture index, to the list of input pictures in reverse output order, of the idx-th picture that is present in the output tensor of the NNPF and has a corresponding input picture, and numPicsInOutputTensor, specifying the total number of pictures present in the output tensor of the NNPF, are derived as follows: for( i = 0, numPicsInOutputTensor = 0; i < numlnputPics; i++ ) if( nnpfc_input_pic_filtering_flag[ i ] ) { Inpldx[ numPicsInOutputTensor ] = i numPicsInOutputTensor++}NumlnpPicsInOutputTensor = numPicsInOutputTensor if( PictureRateUpsamplingFlag ) for( i = 0; i <= numlnputPics - 2; i++ ) numPicsInOutputTensor += nnpfc_interpolated_pics[ i ] if( TemporalExtrapolationFlag ) numPicsInOutputTensor += nnpfc_extrapolated_pics + 1
[0282] When TemporalExtrapolationFlag is equal to 1, the extrapolated pictures generated by the NNPF follow all input pictures of the NNPF in output order.
[0283] In an embodiment, when TemporalExtrapolationFlag is equal to 1 and there is a decoded output picture that follows, in output order, the current picture for which the NNPF is activated, the extrapolated pictures generated by the NNPF precede that decoded output picture in output order.
[0284] In an embodiment, when temporal frame packing arrangement is in use, a decoder may select decoded pictures of the same constituent frame (e.g., only left-view pictures or only right-view pictures of stereoscopic video that is temporally frame-packed) as input pictures to an NNPF performing temporal extrapolation. For example, when PictureRateUpsamplingFlag or TemporalExtrapolationFlag is equal to 1 and the input picture with index 0 is associated with a frame packing arrangement SEI message with fp_arrangement_type equal to 5, all input pictures are associated with a frame packing arrangement SEI message with fp_arrangement_type equal to 5 and the same value of fp_current_frame_is_frameO_flag .
[0285] In an embodiment, a decoder: concludes that temporal interleaving frame packing arrangement is in use in a bitstream, e.g., from a frame packing arrangement SEI message or alike;concludes a constituent frame parity that a current frame where a temporal extrapolation postfilter is activated has for the temporal interleaving frame packing arrangement; selects input frames for the temporal extrapolation post-filter that have the concluded constituent frame parity; and applies the temporal extrapolation post-filter with the input frames as input.
[0286] Constraint to avoid output order conflicts.
[0287] In an embodiment, the constraints to avoid conflicting or ambiguous output order in relation to the general post-processing filtering process are updated to concern extrapolated pictures in addition to interpolated pictures. Within the process, for any particular pair of pictures inputPicA and inputPicB consecutive in output order in CroppedDecodedPictures, when there are one or more pictures intermediatePicSetA in ListNnpfOutputPics between inputPicA and inputPicB in output order, it may be required that one and only one of the following applies:- The pictures in intermediatePicSetA shall be among the pictures that were output by applying a particular NNPF nnpfA with PictureRateUpsamplingFlag equal to 1 when a particular picture currPicA in CroppedDecodedPictures was the current picture.- The pictures in intermediatePicSetA shall be among the pictures that were output by applying a particular NNPF nnpfA with TemporalExtrapolationFlag equal to 1 when a particular picture currPicA in CroppedDecodedPictures was the current picture.
[0288] It may be required that the application of any other NNPF that was used in the filtering process for one picture when currPicA was the current picture or the application of any NNPF (including nnpfA) that was used in the filtering process for one picture when any other picture currPicB in CroppedDecodedPictures was the current picture does not output any picture between the inputPicA and inputPicB in output order.
[0289] Adding generative modelling purpose
[0290] The following table provides an example of a definition of nnpfc_purpose. It is to be understood that the presented bitmask assignment is an example and embodiments may be similarly realized with any other bitmask assignment.
[0291] The variables GeneralVisualQualitylmprovementFlag and GenerativeModellingFlag, specifying whether nnpfc_purpose indicates the purpose of the NNPF to include general visual quality improvement and generative modelling, respectively, are derived as follows:GeneralVisualQualitylmprovementFlag = ( nnpfc_purpose & 0x01 ) GenerativeModellingFlag = ( ( nnpfc_purpose & 0x40 ) > 0 ) ? 1 : 0
[0292] GenerativeModellingFlag equal to 1 may indicate that the NNPF output picture(s) do not represent the same visual content as the input picture(s) but are generated by the NNPF based on the input pictures and auxiliary inputs, if any.
[0293] It may be required that when GeneralVisualQualitylmprovementFlag is equal to 1, GenerativeModellingFlag shall be equal to 0.
[0294] Example characteristics
[0295] In one embodiment, the signaling information may comprise information that characterizes a neural network performing temporal extrapolation, which may include, but may not be limited to, information characterizing inputs to and / or outputs from the neural network. This information may be referred to as temporal extrapolation characteristics (TEC) information.
[0296] In an embodiment, the TEC information may be signaled when the purpose comprises visual temporal extrapolation. In an example, the TEC information is signaled when the signaling information indicates that the purpose of one or more neural networks associated with the signaling information comprises visual temporal extrapolation.
[0297] In another embodiment, at least some of the TEC information may be signaled when the purpose comprises other purposes than visual temporal extrapolation. In an example, some of the TEC information is signaled when the signaling information indicates that the purpose of one or more neuralnetworks associated with the signaling information comprises generative modeling, or generative Artificial Intelligence (Al), where generative modeling or generative Al refers to using the one or more neural networks for generating data conditional to one or more inputs. An example of generative modeling or generative Al is using a neural network for generating an image conditional to an input description in natural language. Another example of generative modeling or generative Al is using a neural network for generating additional data for a given image, such as pixels in areas not captured by the imaging process that produced the given image (e.g., a 360 degrees view).
[0298] In an embodiment, the TEC information may be signaled when the purpose comprises generative face video or generative video. In an example, the purpose of generative face video comprises generating or improving, at decoder or receiver side, a picture (part of a video) comprising at least one or more faces, given one or more of the following:- Coordinates of the at least one or more faces;- Coordinates of landmarks or keypoints of the at least one or more faces;- Transformation matrices that describe the motion or temporal change of one or more keypoints of the at least one or more faces;- Features of facial expressions;- Indication of blinking;- Texts associated to the video;- Audio associated to the video; or- One or more other pictures, such as a key picture or a drive picture, where a key picture may be a picture coded based on other techniques than generative models, and a drive picture may be a picture used to refresh or enhance some parts of the picture.
[0299] It is to be noted that at least some of the characteristics in the TEC information may be relevant or valid also for neural networks that perform other tasks or purposes than visual temporal extrapolation, such as tasks or purposes performed by other generative NNs or generative Al models, such as spatial extrapolation generative NN.
[0300] Characteristic: number of extrapolated pictures
[0301] In an embodiment, the TEC information may comprise information indicative of the number of pictures that are extrapolated in one activation or inference of the VENN.
[0302] An example embodiment of the information indicative of the number of pictures that are extrapolated in one activation or inference of the VTNN described above in nnpfc_extrapolated_pics_minus 1 syntax element.
[0303] In an embodiment, the TEC information may comprise information indicating that the number of extrapolated pictures of the NNPF may vary for each inference and the number of extrapolated pictures for each inference is indicated by means of another signaling mechanism, such as another SEI message (e.g., a GFV SEI message or a NNPFA SEI message). For example, the TEC information may comprise a flag indicating when the TEC information comprises the number of pictures that are extrapolated in one activation or inference of the VTNN. When the flag indicates that the TEC information does not comprise the number of pictures that are extrapolated in one activation or inference of the VTNN, it indicates that the number of extrapolated pictures of the NNPF may vary for each inference and the number of extrapolated pictures for each inference is indicated by means of another signaling mechanism, such as another SEI message (e.g., a GFV SEI message or a NNPFA SEI message). In another example, the TEC information may comprise a syntax element indicating the number of pictures that are extrapolated in one activation or inference of the VTNN, and when the syntax element is equal to 0, it indicates that the number of extrapolated pictures of the NNPF can vary for each inference and the number of extrapolated pictures for each inference is indicated by means of another signaling mechanism, such as another SEI message (e.g., a GFV SEI message or a NNPFA SEI message).
[0304] Characteristic: time interval of extrapolated pictures
[0305] In an embodiment, the TEC information may comprise information indicative of the time interval of pictures being extrapolated relative to the (latest) picture used as input for the inference.
[0306] Characteristic: auxiliary input indicating output position / POC / timestamp differences
[0307] In an embodiment, the TEC information may comprise information indicative of whether the VTNN takes an auxiliary input that indicates output instances of the picture(s) being temporally extrapolated. The output instances may be defined in, but may not be limited to, one or more of the following ways: relative output position provided in units of a given or pre-defined fraction of a nominal picture interval, POC differences or timestamp differences between the picture(s) to be temporally extrapolated and input pictures (e.g., relative to the latest input picture, i.e., input picture with index 0). The considered timestamps may, for example, be output timestamps or display timestamps.
[0308] In an embodiment, the information indicative of whether the VTNN takes an auxiliary input that indicates output instances of the picture(s) being temporally extrapolated may be carried within an NNPFC SEI message, for example in a syntax element indicating the type of auxiliary input(s), which may be referred to as nnpfc auxiliary inp idc.
[0309] In another embodiment, the information indicative of whether the VTNN takes an auxiliary input that indicates output instances of the picture(s) being temporally extrapolated may be carried within a Neural Network Post-Filter Activation (NNPFA) SEI message.
[0310] In an embodiment, when such auxiliary input is not present, or the VTNN is not designed to be input such an auxiliary input, a fixed picture rate may be assumed by the VTNN.
[0311] In an embodiment, the TEC information may comprise information indicative of whether the VTNN takes an auxiliary input that indicates the picture order count (POC) differences, timestamp differences, relative output position, or relative output differences (relative to interval of input pictures) of output pictures are to be generated or extrapolated.
[0312] In an example, TEC information may comprise information indicative of whether the VTNN takes an auxiliary input that indicates the POC differences, timestamp differences, relative output position, or relative output differences is signaled as part of a NNPFC SEI message.
[0313] In an example embodiment, such an auxiliary input may be carried within a Neural Network Post-Filter Activation (NNPFA) SEI message.
[0314] In another example embodiment, such an auxiliary input may be carried within a newly defined SEI message, which may, for example, be referred to as the neural-network post-filter input (NNPFI) SEI message.
[0315] In an example, an NNPF may take two pictures as input and extrapolate one picture as output. In a NNPFC SEI message, the syntax element nnpfc auxiliary input idc indicates that an auxiliary input to the NN associated to this NNPFC SEI message indicates a relative output difference (relative to the interval of input pictures). In an NNPFA SEI message, an auxiliary input parameter can indicate 0.5 in order to instruct the VTNN to generate an extrapolated picture at half of the interval between the input pictures. E.g., when input pictures have timestamps 0 and 2, indicating 0.5 would cause the VTNN to generate an extrapolated picture at timestamp 3. In another example, when inputpictures have timestamps 0 and 1, indicating 0.5 would cause the VENN to generate an extrapolated picture at time stamp 1.5, e.g., a picture in the middle of time stamp 1 and 2.
[0316] An example embodiment of the auxiliary input characteristic is described as following:
[0317] Enabling multiple neural network inferences for the same set of input pictures, which includes either or both of the following features:- An nnpfc auxiliary inp idc value is specified to indicate that the relative output positions of the pictures to be extrapolated is provided as input to the temporal extrapolation NN.- The neural-network post-filter input (NNPFI) SEI message is proposed. The NNPFI SEI message carries auxiliary input data to be provided to the inference of an associated NNPF. The proposed NNPFI SEI message includes the following ingredients: o nnpfi_target_id that associates the respective NNPF (with the same nnpfc id value) with this SEI message; o nnpfi_cancel_flag and nnpfi_persistence_flag specified similarly to many other SEI messages; o nnpfi_cnt that specifies an NNPFI SEI message instance count value for this nnpfi target id value within a picture unit; o nnpfi auxiliary inp idc and nnpfi_extrapolated_pics_minus 1 that shall have the same values as nnpfc auxiliary inp idc and nnpfc_extrapolated_pics_minus 1 of the associated NNPF, respectively, and are repeated in the NNPFI SEI message to avoid parsing dependency; and o When nnpfi_auxiliary_inp_idc indicates the relative output position of the picture to be extrapolated, the relative output position is provided in units of a given fraction of a nominal picture interval. The output position is required to increase in ascending order of nnpfi_cnt in the NNPFI SEI messages of the same picture unit.- The NNPF process is performed in increasing nnpfi_cnt order for each unique value of nnpfi cnt within a picture unit.
[0318] Neural-network post-filter input SEI message
[0319] In an embodiment, the neural -network post-filter input (NNPFI) SEI message is used to provide auxiliary input data to a neural-network post-processing filter defined by NNPFC SEI message(s) and activated by NNPFA SEI message(s).
[0320] In an embodiment, when there are multiple NNPFI SEI messages for the same target NNPF with different SEI payload content present in a picture unit, each occurrence of these NNPFI SEI message causes an inference of the target NNPF.
[0321] In an embodiment, the NNPFI SEI message comprises an identifier or a counter value, wherein the identifier or the counter value identifies the NNPFI SEI message for the same target NNPF in the same picture unit. When the identifier or the counter value is the same as for a previous NNPFI SEI message for the same target NNPF in the same picture unit, the current NNPFI SEI message may be concluded to be a copy of a previous NNPFI SEI message in the same picture unit.
[0322] In an embodiment, an inference of the target NNPF is performed for each unique identifier or a counter value in the NNPFI SEI messages for the same target NNPF in the same picture unit.
[0323] In an example embodiment, the syntax of the NNPFI SEI message payload may be specified as follows:
[0324] Semantics of the syntax elements of the NNPFI SEI message may be specified as follows.
[0325] nnpfi_target_id indicates the nnpfc_id of the target NNPF, which is specified by one or more NNPFC SEI messages that pertain to the current picture and have nnpfc_id equal to nnpfi target id. It may be required that an NNPFI SEI message with a particular value of nnpfi target id shall not be present in a current PU unless one or both of the following conditions are true: 1) Within the current CLVS there is an NNPFC SEI message with nnpfc_id equal to the particular value of nnpfa target id present in a PU preceding the current PU in decoding order. 2) There is an NNPFC SEI message with nnpfc_id equal to the particular value of nnpfi_target_id in the current PU.
[0326] When a PU contains both an NNPFA SEI message with a particular value of nnpfa target id and an NNPFI SEI message with nnpfi target id equal to the particular value of nnpfa_target_id, it may be required that the NNPFA SEI message precedes the NNPFI SEI message in decoding order.
[0327] nnpfi_cancel_flag equal to 1 indicates that the persistence of the NNPF auxiliary input data established by any previous NNPFI SEI message with the same nnpfi_target_id as the current SEI message is cancelled, i.e., the NNPF auxiliary input data is no longer used unless it is activated by another NNPFI SEI message with the same nnpfi_target_id as the current SEI message and nnpfi cancel flag equal to 0. nnpfi_cancel_flag equal to 0 indicates that syntax elements specifying the NNPF auxiliary input data follow.
[0328] nnpfi _persistence_flag specifies the persistence of the NNPF auxiliary input data for the current layer. nnpfi_persistence_flag equal to 0 specifies that the NNPF auxiliary input data may be used for post-processing filtering for the current picture only. nnpfi_persistence_flag equal to 1 specifies that the NNPF auxiliary input data may be used for post-processing filtering for the current picture and all subsequent pictures of the current layer in output order until one or more of the following conditions are true:- A new CLVS of the current layer begins.- The bitstream ends.- A picture in the current layer associated with an NNPFI SEI message with the same nnpfi target id as the current SEI message is output that follows the current picture in output order.- The NNPF auxiliary input data is not applied for this subsequent picture in the current layer associated with an NNPFI SEI message with the same nnpfi_target_id as the current SEI message.
[0329] Let nnpfiTargetPictures be the set of pictures for which the NNPF auxiliary input data persist according to the current NNPFI SEI message. Let nnpfaTargetPictures be the set of pictures for which an NNPF with nnpfc_id equal to nnpfi_target_id is activated by any NNPFA SEI message withnnpfa target id equal to nnpfi target id. It may be required that any picture included in nnpfiTargetPictures shall also be included in nnpfaTargetPictures.
[0330] nnpfi_cnt specifies an NNPFI SEI message instance count value for this nnpfi_target_id value within a picture unit. It may be required that the nnpfi cnt of the first NNPFI SEI message, in decoding order, with a particular value of nnpfi target id within picture unit is equal to 0. It may be required that when nnpfi_cnt assigned to currNnpfiCnt is greater than 0, an NNPFI SEI message with the same nnpfi_target_id value and nnpfi_cnt equal to currNnpfiCnt - 1 shall precede the current NNPFI SEI message in decoding order in the same picture unit.
[0331] nnpfi auxiliary inp idc has the same semantics of nnpfc auxiliary inp idc and shall be set equal to equal to the value of nnpfc_auxiliary_inp_idc in the NNPFC SEI message with nnpfc_id equal to nnpfi target id applying to the current picture.
[0332] nnpfi_extrapolated_pics_minus 1 has the same semantics of nnpfc_extrapolated_pics_minus 1 and shall be set equal to equal to the value of nnpfc_extrapolated_pics_minus 1 in the NNPFC SEI message with nnpfc_id equal to nnpfi_target_id applying to the current picture.
[0333] nnpfi_output_interval_div_minus2 is used to derive the units for specifying the relative output position.
[0334] nnpfi output interval mul minus 1 [ i ] + 1 specifies the output position of the picture in units of 1 ( nnpfi_output_interval_div_minus2 + 2 ) nominal picture intervals. If there are at least two input pictures to the target NNPF, the nominal picture interval is the interval between each consecutive pair of input pictures in output order. Otherwise, the nominal picture interval is an arbitrary interval that shall be consistently used for this value of nnpfi target id in the same CLVS. The output position Output Pos | nnpfi cnt ][ i ] is set equal to ( nnpfi output interval mul minus 1 [ i ] + 1 ) ( nnpfi_output_interval_div_minus2 + 2 ).
[0335] When nnpfi_cnt is greater than 0 and i is equal to 0, it may be required that Output Pos | nnpfi cnt ]
[0000] is greater thanOutputPosf nnpfi cnt - 1 ][ nnpfi_extrapolated_pics_minusl ] derived from NNPFI SEI messages having the same nnpfi_target_id in the same picture unit.
[0336] It may be required that when i is greater than 0, OutputPosf nnpfi cnt ] [ i ] shall be greater than Output Pos | nnpfi cnt ] [ i - 1 ] .
[0337] When a picture is present in both nnpfaTargetPictures and nnpfiTargetPictures, the auxiliary input tensor AuxInputTensoro is a vector of real values equal to Output Pos | nnpfi cnt |. Otherwise, the auxiliary input tensor AuxInputTensoro is a vector of real values equal to OutputPos[ nnpfi cnt ] derived by setting nnpfi cnt equal to 0, nnpfi_extrapolated_pics_minus 1 equal to nnpfc_extrapolated_pics_minusl, nnpfi_output_interval_div_minus2 equal to nnpfi_extrapolated_pics_minusl, and nnpfi_output_interval_mul_minusl[ i ] equal to i for each value of i.
[0338] Characteristic: auxiliary input indicating facial parameters
[0339] In an embodiment, the TEC information may comprise information indicative of whether the VTNN takes an auxiliary input that indicates or is derived from facial parameters, where this auxiliary input may be used by the VTNN in order to perform temporal extrapolation of one or more faces in a picture.
[0340] In an embodiment, the facial parameters that are used as auxiliary input, or that are used to derive the auxiliary input, may be carried from encoder to decoder within an SEI message or other suitable signaling mechanism.
[0341] In an embodiment, the information indicative of whether the VTNN takes an auxiliary input that indicates or is derived from facial parameters may comprise indicating that the auxiliary input is based on an SEI message of a given payload type value and further indicating the SEI message payload type value of the Generative Face Video (GFV) SEI message.
[0342] In an embodiment, the information indicative of whether the VTNN takes an auxiliary input that indicates or is derived from facial parameters may additionally comprise an identifier value that is included in or derived from the SEI message(s) of a given SEI payload type value, where an indication of the given SEI payload type value may also be comprised in the TEC information. The auxiliary input is to be derived based on those SEI messages that are associated with the identifier value.
[0343] In an example, facial parameters comprise one or more of the following:- 2D key-points.- 3D key-points.- One or more matrices.
[0344] In an embodiment, the facial parameters that are used as auxiliary input or that are used to derive the auxiliary input may be carried within an NNPFA SEI message.
[0345] In another embodiment, the facial parameters that are used as auxiliary input or that are used to derive the auxiliary input may be carried within a newly defined SEI message, which may, for example, be referred to as the neural -network post-filter input (NNPFI) SEI message.
[0346] In another embodiment, the facial parameters that are used as auxiliary input or that are used to derive the auxiliary input may be carried within a Generative Face Video (GFV) SEI message.
[0347] In an embodiment, the TEC information may comprise information indicative of a format for the facial parameters which is accepted or required by the VENN. In an example, there could be two or more possible formats defined in a standard, each associated with a format identifier, and the TEC information would comprise a value for the format identifier.
[0348] Facial parameters obtained from a NNPFA SEI message or from a GFV SEI message or from a NNPFI SEI message (or, more generally, from a FW signaling) may be converted to the required format that is indicated in the TEC information.
[0349] Characteristic: auxiliary input comprises an output of another network run at decoder side
[0350] In an embodiment, the TEC information may comprise information indicative of whether the VENN takes an auxiliary input that is derived from the output of another neural network, where the another neural network is run or executed at decoder side.
[0351] In an embodiment, the TEC information may comprise information indicative of whether the VENN takes an auxiliary input that is derived from the output of another neural network that converts or translates facial parameters to a format that is accepted or required by the VENN. The facial parameters may be carried within an NNPFA SEI message or a GFV SEI message or a NNPFI SEI message (or, more generally, a FW signaling) or in any other suitable way.
[0352] It is to be understood that many of the embodiments for the characteristic "auxiliary input indicating facial parameters" are likewise applicable for the present characteristic.
[0353] In an embodiment, the auxiliary input type (nnpfc auxiliary inp idc) indicated in the NNPFC SEI message identifies the specific auxiliary input or the specific SEI message to derive the auxiliary input. For example, an auxiliary input type may be defined for the GFV SEI message for the facial parameters and facial matrix derived based on the GFV SEI message(s).
[0354] In an embodiment, the auxiliary input type (nnpfc auxiliary inp idc) indicated in the NNPFC SEI message specifies that the auxiliary input is derived from SEI message(s) of an SEI payload type value, which is provided in a syntax element (e.g., called nnpfc_sei_aux_payload_type) within the NNPFC SEI message.
[0355] In an embodiment, the syntax element for defining auxiliary input types (nnpfc_auxiliary_inp_idc) indicated in the NNPFC SEI message is regarded as a bit field, where each bit position defines the use of a specific auxiliary input type. In an example, ( nnpfc auxiliary inp idc & 2 ) greater than 1 specifies that the auxiliary input is derived from SEI message(s) of a given payload type value (nnpfc_sei_aux_payload_type).
[0356] In an embodiment, when nnpfc_sei_aux_payload_type is present in the NNPFC SEI message, a syntax element nnpfc_sei_aux_id is also present. nnpfc_sei_aux_id specifies an identifying number that may be used to identify which one(s) of the SEI messages of payload type equal to nnpfc_sei_aux_payload_type are used to derive auxiliary input data or auxiliary input tensor(s) to the NNPF.
[0357] In an embodiment, when the NNPFC SEI message indicates that auxiliary input is derived from the GFV SEI message, nnpfc sei aux id is present and indicates the gfv id value of the GFV SEI message for which the NNPF defined in the NNPFC SEI message serves as the GenerativeNN.
[0358] In an embodiment, when the NNPFC SEI message indicates that auxiliary input is derived from the GFV SEI message, the NNPFC SEI message includes a flag, where one value indicates that the nnpfc_id value serves as the gfv_id value for which the NNPF defined in the NNPFC SEI message serves as the GenerativeNN, and another value indicates that nnpfc_sei_aux_id is present and indicates the gfv_id value of the GFV SEI message for which the NNPF defined in the NNPFC SEI message serves as the GenerativeNN.
[0359] An example embodiment that combines features of this characteristic and the previous characteristic ("auxiliary input indicating facial parameters") are described in following paragraphs.
[0360] One possible syntax included in the NNPFC SEI message may comprise the following:
[0361] Another possible syntax included in the NNPFC SEI message may comprise the following:
[0362] When ( nnpfc_auxiliary_inp_idc & 2 ) is equal to 2 and nnpfc_sei_aux_payload_type is equal to SEI GFV PAYLOAD TYPE, nnpfc_num_input_pics_minus 1 shall be equal to 1.
[0363] nnpfc_auxiliary_inp_idc greater than 0 indicates that auxiliary input data is present in the input tensor of the NNPF or auxiliary input tensor(s) are provided to the NNPF. nnpfc auxiliary inp idc equal to 0 indicates that auxiliary input data is not present in the input tensor and no auxiliary input tensors are provided to the NNPF.
[0364] ( nnpfc auxiliary inp idx & 2 ) equal to 2 specifies that auxiliary input data or auxiliary input tensor(s) are derived based on a given payload type value of an SEI message and a given identifier value.
[0365] nnpfc_sei_aux_payload_type specifies the SEI message payload type value that is used to derive auxiliary input data or auxiliary input tensor(s) to the NNPF.
[0366] nnpfc_sei_aux_id specifies an identifying number that may be used to identify which one(s) of the SEI messages of payload type equal to nnpfc sci auxjiayload typc are used to derive auxiliary input data or auxiliary input tensor(s) to the NNPF.
[0367] In an example, when nnpfc_sei_aux_payload_type is equal to the payload type value of the GFV SEI message (hereafter, SEI GFV PAYLOAD TYPE), it may be required that when a generative face video SEI message with gfv_id equal to nnpfc_sei_aux_id is present in a picture unit, the picture unit includes an NNPFA SEI message with nnpfa_target_id equal to nnpfc_id.
[0368] When ( nnpfc_auxiliary_inp_idc & 2 ) is equal to 2 and nnpfc_sei_aux_payload_type is equal to SEI GFV PAYLOAD TYPE, the following applies:- inputBaseY, inputBaseCb (when applicable), inputBaseCr (when applicable), inputBaseKeyPoint, and inputBase Matrix are derived as specified for the semantics of the generative face video SEI message that has gfv_id equal to nnpfc_sei_aux_id and gfv_base_pic_flag equal to 1. If there are multiple such generative face video SEI messages present, the derivation is based on the last one preceding, in decoding order, the picture unit containing an NNPFA SEI message with nnpfa_target_id equal to nnpfc_id used for activating the NNPF.- The associated generative face video SEI message is the generative face video SEI message that has gfv_id equal to nnpfc_sei_aux_id and is present in the picture unit that contains an NNPFA SEI message with nnpfa_target_id equal to nnpfc_id. Subsequently, the syntax elements of the generative face video SEI message refer to the associated generative face video SEI message.- inputDriveKeyPoint and inputDriveMatrix are derived as specified for the semantics of the associated generative face video SEI message.- When gfv_drive_pic_fusion_flag is equal to 0, samples values of CroppedYPic
[0000] , CroppedCbPic
[0000] (when applicable), and CroppedCrPic
[0000] (when applicable) are all set equal to 0.- An input picture to the NNPF that has sample arrays with sample value 0 is a dummy picture as defined in the generative face video SEI message.- The input picture with index 1 is overwritten with the base picture as defined in the generative face video SEI message as follows: CroppedYPic
[0001] is set equal to inputBaseY, CroppedCbPic
[0001] is set equal to inputBaseCb (when applicable), and CroppedCrPic
[0001] is set equal to inputBaseCr.
[0369] The process Derive InputTensors( ), for deriving the input tensor inputTensor for a given vertical sample coordinate cTop and a horizontal sample coordinate cLeft specifying the top-left sample location for the patch of samples included in the input tensor, the number of auxiliary input tensors numAuxInputTensors, and the auxiliary input tensors aux InputTensorf j ] , if any, for each value of j in the range of 0 to numAuxInputTensors - 1, inclusive, is specified as follows: if( ( nnpfc auxiliary inp idx & 2 ) &&( nnpfc_sei_aux_payload_type = = SEI GFV PAYLOAD TYPE ) ) numAuxInputTensors = 2 elsenumAuxInputTensors = 0 for( i = 0; i < numlnputPics; i++ ) { if( nnpfc inp order idc = = 0 ) for( yP = -nnpfc overlap; yP < inpPatchHeight + nnpfc overlap; yP++) for( xP = -nnpfc overlap; xP < inpPatchWidth + nnpfc o verlap; xP++ ) { inpVal = InpY( InpSampleVal( cTop + yP, cLeft + xP, CroppedHeight, CroppedWidth, CroppedYPic[ i ], 0 ) ) yPovlp = yP + nnpfc overlap xPovlp = xP + nnpfc overlap if( !nnpfc_component_last_flag ) inputTensor
[0000] [ i ]
[0000] [ yPovlp ] [ xPovlp ] = inpVal else inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0000] = inpVal if( nnpfc auxiliary inp idc = = 1 ) if( !nnpfc_component_last_flag ) inputTensor
[0000] [ i ]
[0001] [ yPovlp ][ xPovlp ] = strengthControlScaledVal[ i ] else inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0001] = strengthControlScaledVal[ i ] } else if( nnpfc inp order idc = = 1 ) for( yP = -nnpfc overlap; yP < inpPatchHeight + nnpfc overlap; yP++) for( xP = -nnpfc overlap; xP < inpPatchWidth + nnpfc o verlap; xP++ ) { inpCbVal = InpC( InpSampleVal( cTop + yP, cLeft + xP, CroppedHeight / SubHeightC,CroppedWidth / SubWidthC, CroppedCbPic[ i ], 1 ) ) inpCrVal = InpC( InpSampleVal( cTop + yP, cLeft + xP, CroppedHeight / SubHeightC,CroppedWidth / SubWidthC, CroppedCrPic[ i ], 2 ) ) yPovlp = yP + nnpfc overlap xPovlp = xP + nnpfc overlap if( !nnpfc_component_last_flag ) { inputTensor
[0000] [ i ]
[0000] [ yPovlp ] [ xPovlp ] = inpCbVal inputTensor
[0000] [ i ]
[0001] [ yPovlp ][ xPovlp ] = inpCrVal} else { inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0000] = inpCbVal inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0001] = inpCrVal} if( nnpfc auxiliary inp idc = = 1 ) if( !nnpfc_component_last flag ) inputTensor
[0000] [ i ]
[0002] [ yPovlp ][ xPovlp ] = strengthControlScaledVal[ i ] else inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0002] = strengthControlScaledVal[ i ] } else if( nnpfc_inp_order_idc = = 2 ) for( yP = -nnpfc overlap; yP < inpPatchHeight + nnpfc overlap; yP++) for( xP = -nnpfc overlap; xP < inpPatchWidth + nnpfc o verlap; xP++ ) { yY = cTop + yP xY = cLeft + xP yC = yY / SubHeightC xC = xY / SubWidthC inpYVal = InpY( InpSampleVal( yY, xY, CroppedHeight,CroppedWidth, CroppedYPic[ i ], 0 ) ) inpCbVal = InpC( InpSampleVal( yC, xC, CroppedHeight / SubHeightC,CroppedWidth / SubWidthC, CroppedCbPic[ i ], 1 ) ) inpCrVal = InpC( InpSampleVal( yC, xC, CroppedHeight / SubHeightC,CroppedWidth / SubWidthC, CroppedCrPic[ i ], 2 ) ) yPovlp = yP + nnpfc overlap xPovlp = xP + nnpfc overlap if( !nnpfc_component_last_flag ) { inputTensor
[0000] [ i ]
[0000] [ yPovlp ][ xPovlp ] = inpYVal inputTensor
[0000] [ i ]
[0001] [ yPovlp ][ xPovlp ] = inpCbVal inputTensor
[0000] [ i ]
[0002] [ yPovlp ][ xPovlp ] = inpCrVal} else { inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0000] = inpYVal inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0001] = inpCbVal inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0002] = inpCrVal} if( nnpfc auxiliary inp idc = = 1 ) if( !nnpfc_component_last_flag ) inputTensor
[0000] [ i ]
[0003] [ yPovlp ][ xPovlp ] = strengthControlScaledVal[ i ] else inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0003] = strengthControlScaledVal[ i ]} else if( nnpfc_inp_order_idc = = 3 ) for( yP = -nnpfc overlap; yP < inpPatchHeight + nnpfc overlap; yP++) for( xP = -nnpfc overlap; xP < inpPatchWidth + nnpfc o verlap; xP++ ) { yTL = cTop + yP * 2 xTL = cLeft + xP * 2 yBR = yTL + 1 xBR = xTL + 1 yC = cTop / 2 + yP xC = cLeft / 2 + xP inpTLVal = InpY( InpSampleVal( yTL, xTL, CroppedHeight, CroppedWidth, CroppedYPic[ i ], 0 ) ) inpTRVal = InpY( InpSampleVal( yTL, xBR, CroppedHeight, CroppedWidth, CroppedYPic[ i ], 0 ) ) inpBLVal = InpY( InpSampleVal( yBR, xTL, CroppedHeight, CroppedWidth, CroppedYPic[ i ], 0 ) ) inpBRVal = InpY( InpSampleVal( yBR, xBR, CroppedHeight, CroppedWidth, CroppedYPic[ i ], 0 ) ) inpCbVal = InpC( InpSampleVal( yC, xC, CroppedHeight / 2, CroppedWidth / 2, CroppedCbPic[ i ], 1 ) ) inpCrVal = InpC( InpSampleVal( yC, xC, CroppedHeight / 2, CroppedWidth / 2, CroppedCrPic[ i ], 2 ) ) yPovlp = yP + nnpfc overlap xPovlp = xP + nnpfc overlap if( !nnpfc_component_last_flag ) { inputTensor
[0000] [ i ]
[0000] [ yPovlp ][ xPovlp ] = inpTLVal inputTensor
[0000] [ i ]
[0001] [ yPovlp ] [ xPovlp ] = inpTRVal inputTensor
[0000] [ i ]
[0002] [ yPovlp ][ xPovlp ] = inpBLVal inputTensor
[0000] [ i ]
[0003] [ yPovlp ][ xPovlp ] = inpBRVal inputTensor
[0000] [ i ]
[0004] [ yPovlp ] [ xPovlp ] = inpCbVal inputTensor
[0000] [ i ]
[0005] [ yPovlp ][ xPovlp ] = inpCrVal} else { inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0000] = inpTLVal inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0001] = inpTRVal inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0002] = inpBLVal inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0003] = inpBRValinputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0004] = inpCbVal inputTensor
[0000] [ i ] [ yPovlp ] [ xPovlp ]
[0005] = inpCrVal} if( nnpfc auxiliary inp idc = = 1 ) if( !nnpfc_component_last flag ) inputTensor
[0000] [ i ]
[0006] [ yPovlp ][ xPovlp ] = strengthControlScaledVal[ i ] else inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0006] = strengthControlScaledVal[ i ]} if( ( nnpfc auxiliary inp idx & 2 ) &&( nnpfc_sei_aux_payload_type = = SEI GFV PAYLOAD TYPE ) ) { if( i = = 0 ) { / / drive picture aux!nputTensoro
[0000] [ i ] = inputDriveKeyPoint aux!nputTensori
[0000] [ i ] = inputDriveMatrix} else { / / i = = 1, base picture aux!nputTensoro
[0000] [ i ] = inputBaseKeyPoint aux!nputTensori
[0000] [ i ] = inputBaseMatrix}}}
[0370] An NNPF PostProcessingFilter( ) is the target NNPF as derived in the semantics of the NNPFA SEI message. The NNPF PostProcessingFilter( ) inputs inputTensor, followed by numAuxInputTensors tensors auxInputTensorj where j is in the range of 0 to numAuxInputTensors - 1, inclusive. The following example process may be used, with the NNPF PostProcessingFilter( ), to generate, in a patch-wise manner, the fdtered and / or interpolated picture(s), which contain Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated by nnpfc out order idc : if( nnpfc_inp_order_idc = = 0 | | nnpfc_inp_order_idc = = 2 ) for( cTop = 0; cTop < CroppedHeight; cTop += inpPatchHeight ) for( cLeft = 0; cLeft < CroppedWidth; cLeft += inpPatchWidth ) { DeriveInputTensors( ) outputTensor = PostProcessingFilter( ) StoreOutputTensors( outputTensor ) } else if( nnpfc inp order idc = = 1 ) for( cTop = 0; cTop < CroppedHeight / SubHeightC; cTop += inpPatchHeight )for( cLeft = 0; cLeft < CroppedWidth / SubWidthC; cLeft += inpPatchWidth ) { DeriveInputTensors( ) outputTensor = PostProcessingFilter( )StoreOutputTensors( outputTensor )} else if( nnpfc_inp_order_idc = = 3 ) for( cTop = 0; cTop < CroppedHeight; cTop += inpPatchHeight * 2 ) for( cLeft = 0; cLeft < CroppedWidth; cLeft += inpPatchWidth * 2 ) {DeriveInputTensors( ) outputTensor = PostProcessingFilter( )StoreOutputTensors( outputTensor )}
[0371] In an embodiment, the signaling information may comprise information indicative of a purpose for a neural network that converts input facial parameters to an output that represents converted facial parameters, where the converted facial parameters or data derived from converted facial parameters are comprised in an input to a VTNN.
[0372] In an embodiment, the signaling information may comprise information indicative of a purpose for a neural network that converts input parameters to an output that represents converted parameters.
[0373] In an embodiment, the signaling information may comprise information indicative of a purpose for a neural network that converts a format of input parameters to another format of output parameters.
[0374] Characteristic: identifying an NNPF as generator NN and activation of the generator NN
[0375] In an embodiment, an encoder indicates that an NNPF is a generator NN and / or a decoder concludes that an NNPF is a generator NN when an NNPFC SEI message indicates that the NNPF uses an auxiliary input based on an SEI message payload type value that indicates generative Al, such as the GFV SEI message.
[0376] In an embodiment, an encoder indicates that an NNPF is a generator NN and / or a decoder concludes that an NNPF is a generator NN when a generative Al SEI message, such as the GFV SEI message, includes an identifier of the NNPF. In an example embodiment, an encoder indicates that an NNPF is a generator NN and / or a decoder concludes that an NNPF is a generator NN, when the gfv_id of the GFV SEI message is equal to nnpfc_id of an NNPFC SEI message defining the NNPF. In anotherexample embodiment, an encoder indicates that an NNPF is a generator NN and / or a decoder concludes that an NNPF is a generator NN through an identifier syntax element in the GFV SEI message other than gfv_id, where the identifier syntax element is equal to nnpfc_id of an NNPFC SEI message defining the NNPF. In an embodiment, the GFV SEI message may include two identifiers, a first identifier identifying a face and a second identifier identifying a generator NN. In yet another example embodiment, an encoder indicates that an NNPF is a generator NN and / or a decoder concludes that an NNPF is a generator NN when the second identifier is equal to nnpfc_id of an NNPFC SEI message defining the NNPF.
[0377] In an embodiment, the generator NN is activated in the decoder side as signaled with the NNPFA SEI message for an NNPF for which an auxiliary input derived from a generative Al SEI message, such as the GFV SEI message, is indicated. The NNPF with nnpfc_id equal to nnpfa_target_id is activated, and auxiliary input is derived from a generative Al SEI message, such as the GFV SEI message, as discussed in other embodiments.
[0378] In an embodiment, the generator NN is activated in the decoder side as signaled with a generative Al SEI message, such as a GFV SEI message. The NNPF with nnpfc_id indicated in or for a generative Al SEI message, such as the GFV SEI message, is activated and auxiliary input is derived from the generative Al SEI message, such as the GFV SEI message, as discussed in other embodiments.
[0379] Characteristic: auxiliary input comprises a tensor generated at encoder side
[0380] In one embodiment, the TEC information may comprise information indicative of whether the VENN takes an auxiliary input that may represent or is derived from a tensor that is generated at encoder side.
[0381] A tensor refers to an N-dimensional array. Typically, the tensor would be a 3D or 4D tensor, where two dimensions represent the spatial dimensions (height and width), one dimension represents the channels, and one dimension (in the case of 4D tensor) represents the batch size in the case of batch processing. However, even ID tensor (e.g., an array of M elements) is possible, when the batch size is assumed to be equal to 1, or a 2D tensor (e.g., a matrix of Ml by M2 elements) is possible, where one axis or dimension represents the batch size.
[0382] In an example, the tensor is (or is derived from) an output of another neural network that is run or executed at encoder side, and the tensor may be signaled within an NNPFA SEI message or a GFV SEI message or NNPFI SEI message (or, more generally, a FW signaling).
[0383] In an example, the tensor represents or is derived from facial parameters. At encoder side, the facial parameters may be input to a neural network (different from the VTNN performing temporal extrapolation), the NN is run or executed, and the output of such a NN would be the tensor. The tensor may be lossy or lossless coded at encoder side and lossy or lossless decoded at decoder side. At decoder side, the lossy or lossless decoded tensor, or a signal derived from the lossy or lossless decoded tensor, is input to a VTNN.
[0384] In an embodiment, the TEC information may comprise information about the format of the tensor, such as number of dimensions, size of each dimension, numerical format of the values in the tensor (e.g., floating-point 32 bits, or integer 16 bits, and the like.), serialization order or format of the tensor. In an example, such information may be signaled by using a standard for signaling tensors, such as a standard for a tensor serialization format.
[0385] In an embodiment, the VTNN may take more than one tensor.
[0386] In an embodiment, the number of auxiliary tensors to be input to the VTNN may be a characteristic of a VTNN, and information about such number may be signaled in an NNPFC SEI message. In another embodiment, the number of auxiliary tensors to be input to the VTNN may vary and information about such a number may be signaled in an NNPFA SEI message or a GFV SEI message or a newly defined SEI message, which may be referred to as neural-network post-filter input SEI message (or, more generally, a FW signaling).
[0387] In an example, a VTNN accepts more than one auxiliary tensors, where different auxiliary tensors represent different features of a face to be temporally extrapolated. The auxiliary tensors may comprise, but need not be limited to, one or more of the following: a first auxiliary tensor representing face expression, a second auxiliary tensor representing head pose, a third auxiliary tensor representing finer details of the face.
[0388] Characteristic: auxiliary input comprises audio or speech
[0389] In one embodiment, the TEC information may comprise information indicative of whether the VTNN takes an auxiliary input that represents or is derived from audio or speech.
[0390] This information may be used by the VTNN to generate or synthesize several feature of a face, such as mouth, lips, face expression, blinking, and the like.
[0391] In an embodiment, when the VTNN takes an auxiliary input that represents audio or speech, the TEC information may comprise information about characteristics of the audio or speech, such as one or more of the following:Sampling rate of the expected audio signal;- Bit depth and / or dynamic range;- Number of audio channels and / or channel configuration;- Duration of the audio signal and / or number of audio samples to be included in the input tensor;- Information on whether the audio is processed with source separation algorithm. E.g., Is the voice isolated? And (in case of multiple faces in the frame) are different voices separated?- Information on whether it is spatial audio; or- Timing of the audio signal provided as input in relation to picture(s) to be extrapolated.
[0392] In an embodiment, this auxiliary input may be derived from the audio track associated with the video.
[0393] In an additional embodiment, this auxiliary input may be derived by processing the audio track with a neural network or other processing algorithm, where the neural network or other processing algorithm are defined in a standard or are associated to an identifier that can be indicted by an encoder. In an example, an encoder may signal an identifier of a NN to be used for processing the audio track and deriving the auxiliary input to the VTNN.
[0394] In an embodiment, the auxiliary input may be (or may be derived from) a signal that is carried within another signaling mechanism, such as an NNPFA SEI or a GFV SEI message or a newly defined SEI message, which may be referred to as neural -network post-filter input SEI message (or, more generally, a FW signaling).
[0395] Characteristic: auxiliary input comprises text
[0396] In an embodiment, the TEC information may comprise information indicative of whether the VTNN takes an auxiliary input that represents or is derived from text.
[0397] This information may be used by the VTNN to generate or synthesize several features of a face, such as mouth, lips, face expression, blinking, and the like.
[0398] In an embodiment, the text may comprise text that is derived from the audio track associated with the pictures of the video, for example, by speech-to-text techniques.
[0399] In an embodiment, the text may comprise instructions for the VTNN, related for example to face expression, or mood, emotions, or whether the eyes are closed, and the like. The instructions may or may not be in the form of natural language.
[0400] In an embodiment, the TEC information may comprise information indicative of whether the VTNN takes an auxiliary input that represents text associated to the speech of a person for whom the face is to be temporally extrapolated.
[0401] In an embodiment, the TEC information may comprise information indicative of whether the VTNN takes an auxiliary input that represents text representing instructions for high-level or abstract concepts related to the subject of temporal extrapolation, such as emotions and mood and face expressions in the case where the subject of temporal extrapolation is faces.
[0402] In an embodiment, when the VTNN takes an auxiliary input that represents text, the TEC information may comprise information indicative of characteristics of the text, such as one or more of the following:- Whether the VTNN accepts or requires raw text;- Whether the VTNN accepts or requires tokens (e.g., words, or parts of words, or indexes of a dictionary where an index is associated to a respective word or part of word);- Whether the VTNN accepts or requires text embeddings, e.g., features extracted from the raw text or from tokens;- Which tokenizer to use;- Which text embedder to use; or- Timing of the text provided as input in relation to picture(s) to be extrapolated.
[0403] In an embodiment, this auxiliary input may be derived from the timed text track associated with the video.
[0404] In an embodiment, this auxiliary input may be derived from the subtitles or captions included in or along the video bitstream. For example, the subtitles or captions may be included in one or more SEI messages.
[0405] In an embodiment, the text may be obtained at decoder side. For example, it may be derived based on the audio track.
[0406] In an alternative embodiment, the text or data derived from the text may be signaled from encoder to decoder via a FW signaling mechanism (e.g., NNPFA SEI message or GFV SEI message of NNPFI SEI message).
[0407] Characteristic: temperature
[0408] In an embodiment, the TEC information may comprise information indicative of whether the VENN accepts an input that controls one or more aspects of how a final output is determined based on the output of the VENN. In one example, a VENN may accept an input parameter referred to as “temperature”, where the temperature parameter is used to control the probabilities that are output by the VENN, where a higher value of the temperature parameter allows for more randomness in the final output and a lower value allows for a more predictable final output.
[0409] Characteristic: output variability
[0410] In an embodiment, the TEC information may comprise information indicative of whether the VENN may produce different outputs for different inferences or executions of the VENN with the same input(s). Such a behavior or characteristic of the VENN is referred to as output variability of the VENN.
[0411] In an additional embodiment, the TEC information may comprise information indicative of the type or cause of output variability of the VENN. For example, the type or cause of output variability may be one of the following:- Pseudo-randomness;- Deliberate variability, e.g., the VENN is characterized by output variability by design; or- Floating-point numerical format that is used to represent the value of one or more parameters of the VENN and / or the value of one or more signals flowing through the VENN.
[0412] In an embodiment, the TEC information may comprise information indicative of whether the VENN can be instructed or operated in a way that it does not yield output variability. In an additional embodiment, this information is comprised in the TEC information when the TEC information indicates that the VENN is characterized by output variability.
[0413] In an embodiment, the TEC information may comprise information indicative of whether the VTNN is allowed to have output variability. In other words, the VTNN associated to the TEC information may be characterized by or support output variability, but an encoder may signal to a decoder whether the VTNN is to be run or executed in a way that it does not yield output variability. In an additional embodiment, this information is comprised in the TEC information when the TEC information indicates that the VTNN is characterized by output variability. In one example, the TEC information comprises a flag that indicates whether the VTNN is allowed to have output variability.
[0414] In an embodiment, the TEC information may comprise information that makes the VTNN not to have output variability. This information is referred to as Deterministic Inference (DI) information. In one example, the DI information comprises a pseudo-random seed value.
[0415] In an embodiment, the TEC information comprises an indication, such as a flag, indicative whether a decoder side is to generate a random seed to the VTNN by itself or derive a pseudo-random seed based on other information that is signaled by an encoder.
[0416] A hash value may be defined to be a result of a hash function. A hash value may be alternatively or additionally referred to as a checksum or a hash sum. A hash function may be defined as any function that can be used to map digital data of arbitrary size to digital data of fixed size, with slight differences in input data possibly producing big differences in output data. A cryptographic hash function may be defined as a hash function that is intended to be practically impossible to invert, i.e., to create the input data based on the hash value alone. Cryptographic hash function may comprise, e.g., the MD5 function. An MD5 value may be a null -terminated string of UTF-8 characters containing a base64 encoded MD5 digest of the input data. One method of calculating the string is specified in IETF RFC 1864. It should be understood that instead of or in addition to MD5, other types of integrity check schemes could be used in various embodiments, such as different forms of the cyclic redundancy check (CRC), such as the CRC scheme used in ITU-T Recommendation H.271.
[0417] In an embodiment, a pseudo-random seed value may be derived at a decoder side or receiver side based on other information that is signaled by an encoder or that is present or determined within the decoder, and where the other information may be used also for other purposes than to derive the pseudo-random seed value. In one example, the other information is a sequence-wise Quantization Parameter (QP) value, and the pseudo-random seed is set equal to that sequence-wise QP value. In another example, the other information is a slice-wise QP value, and the pseudo-random seed is set equal to that slice-wise QP value. In yet another example, the other information is a value of a syntax element in the TEC information. In yet another example, the other information is a value of a syntax element in an NNPFC SEI message that specifies the VTNN, such as nnpfc_id, and the pseudo-randomseed is set equal to the value of nnpfc id. In yet another example, the other information is a value of a syntax element in a Generative Face Video (GFV) SEI message that specifies the VTNN, such as gfv id, and the pseudo-random seed is set equal to the value of gfv id. In a further example, a hash value is derived from coded data, such as the NNPFC SEI message, the GFV SEI message, the SEI NAL unit comprising the NNPFC SEI message or the GFV SEI message, or a coded picture associated with the VTNN, or a picture unit associated with the VTNN. The hash value is used as the pseudorandom seed. In another further example, a hash value is derived from decoded data, such as one or more decoded pictures or one or more input pictures to the VTNN, and the hash value is used as the pseudo-random seed.
[0418] In an embodiment, a pseudo-random seed for the VTNN is derived at a decoder side or receiver side in two phases: In a first phase, a first pseudo-random seed is derived as described in any embodiment of the previous paragraph. In a second phase, the first pseudo-random seed is provided as initialization to a pseudo-random number generator that returns a second pseudo-random seed. The second pseudo-random seed is provided to the VTNN. The pseudo-random number generator may be pre-defined, for example in a standard, or it may be freely chosen by a decoder side or receiver side.
[0419] In an embodiment, the TEC information may comprise an indication, such as a flag, of whether a pseudo-random seed value may be derived based on a value of a particular syntax element that is comprised in the TEC information or that is signaled to a decoder by other means. In one example, an NNPFC SEI message comprises a flag nnpfc_seed_from_nnpfc_id that, when it is equal to 1, it indicates that a pseudo-random seed value is equal to the value of the syntax element nnpfc id. In another example, an NNPFC SEI message comprises a flag nnpfc_seed_from_nnpfc_id that, when it is equal to 1, it indicates that a pseudo-random seed value should be input to the NNPF that is specified by the NNPFC SEI message and it is equal to the value of the syntax element nnpfc_id, and when it is equal to 0 it indicates that no pseudo-random seed value should be input to the NNPF that is specified by the NNPFC SEI message.
[0420] In an embodiment, the TEC information may comprise an indication that the VTNN specified by the TEC information takes as an input or auxiliary input a seed value. The seed value may be a random seed generated in the decoder side or a pseudo-random seed based on other information that is signaled by an encoder.
[0421] In an embodiment, the TEC information may comprise an indication that the VTNN specified by the TEC information takes as an input or auxiliary input a pseudo-random seed.
[0422] In an embodiment, the TEC information may comprise an indication that the VTNN specified by the TEC information takes as an input or auxiliary input a pseudo-random seed, where thepseudo-random seed may be determined or obtained in one of the following ways: the pseudo-random seed is specified in a standard specification; the pseudo-random seed is signaled as part of the TEC information; the pseudo-random seed is set equal to or is determined based on a value of another syntax element that is comprised in the TEC information or that is signaled to a decoder by other means; the pseudo-random seed is provided to the decoder or the receiver by external means, such as by an application.
[0423] In an embodiment, a particular bit of nnpfc auxiliary inp idc being equal to 1 may be used as the indication that the VENN specified by the TEC information takes as an input or auxiliary input a seed value. For example, when nnpfc auxiliary inp idc & 4 is greater than 0, the auxiliary input data includes the seed value for the NNPF.
[0424] In one example, when a particular NNPFC SEI message specifying an NNPF that performs spatial or temporal extrapolation (or other generative task) comprises a syntax element ( nnpfc auxiliary inp idc & 4 ) that is equal to 4, the NNPF takes as an input or as an auxiliary input a pseudo-random seed that is equal to a value of the syntax element nnpfc_id of the particular NNPFC SEI message.
[0425] In another example, when a particular NNPFC SEI message specifying an NNPF that performs spatial or temporal extrapolation (or other generative task) comprises a syntax element (nnpfc_auxiliary_inp_idc & 4) that is equal to 4, the NNPFC SEI message also comprises a syntax element nnpfc_seed that specifies a value of a pseudo-random seed to be provided to the NNPF as an input or as an auxiliary input. nnpfc_seed, when present, specifies a seed value that is used as auxiliary input for the NNPF. The following syntax included in an NNPFC SEI message may be used in this example:
[0426] In an embodiment, an auxiliary input, such as a seed value, may be provided to the NNPF in one of the following ways:- A channel or matrix of the input tensor comprises one or more occurrences of the seed value.- A scalar seed value is provided separately from the input tensor to the NNPF process.
[0427] In an embodiment, an NNPF process or alike for the VTNN comprises a stage of adding pseudo-random noise to the input picture(s) provided to the NNPF process. The stage of adding pseudorandom noise may be performed prior to the neural network inference of the VTNN. The stage may comprise initializing a pseudo-random number generator with the seed value, which has been determined as described in other embodiments. Furthermore, the stage may comprise one or more of:- deriving a pseudo-random number for a sample of an input picture,- scaling the pseudo-random number to a value range that may depend on the sample value range (which may, for example, be an integer value range of a given bit depth or a floating point value from 0 to 1),- adding the scaled pseudo-random number to the sample value, and- clipping the resulting sample value to the sample value range.
[0428] In an embodiment, the pseudo-random noise addition and the neural network inference of the VTNN are performed iteratively so that the NNPF-generated pictures of the previous iteration are used as the input to the next iteration.
[0429] In an embodiment, the pseudo-random noise is first added to the input picture(s) for the VTNN. Subsequently, the neural network inference of the VTNN is performed iteratively so that the NNPF-generated pictures of the previous iteration are used as the input to the next iteration.
[0430] In an embodiment, the DI information or the TEC information comprises characteristics of randomness for the VTNN, which may comprise, but may not be limited to, one or more of the following:- Value range(s) of the pseudo-random noise added to samples of the input picture(s). A value range may be provided jointly for all color components, or separate value ranges may be provided for luma and chroma, or a separate value range may be provided for each color component, such as Y, Cb, and Cr.- A distribution function of the pseudo-random noise.- A number of iterations of the neural network inference.
[0431] In an embodiment, at least part of the signaling information, or at least part of the TEC information, or at least part of the DI information may be comprised in a Neural Network Post-Filter Activation (NNPFA) SEI message.
[0432] In an example, information indicative of a seed value may be comprised in an NNPFA SEI message, where the seed value or data derived from the seed value is used as an input or auxiliary input to the NNPF. In another example, information indicative of whether an NNPF takes as an input or auxiliary input a seed value may be comprised in an NNPFA SEI message.
[0433] In another example, the DI information comprises a flag indicating the use of a quantized version of the VTNN.
[0434] In another example, the DI information comprises one or more quantization parameters to be used for quantizing the VTNN.
[0435] In an additional or alternative embodiment, the DI information may be carried in a NNPFA SEI message (or, more generally, a FW signaling).
[0436] In an example, a VTNN may be tasked to generate two or more parts or aspects of a picture, such as two faces. An NNPFA SEI message or a GFV SEI message may carry respective two or more sets of DI information for generating the two or more parts or aspects of a picture, such as two or more pseudo-random seeds.
[0437] Characteristic: type of extrapolated output
[0438] In an embodiment, the TEC information may comprise information indicative of the type of extrapolated output.
[0439] In an additional embodiment, the information indicative of the type of extrapolated output may comprise indicating one or more categories of visual content that the NN is able to extrapolate.
[0440] In an example, the one or more categories comprise the following categories:Face;Hand(s); and / or Body.
[0441] In an example, a VTNN performing visual temporal extrapolation supports extrapolating only faces. In another example, a VTNN performing visual temporal extrapolation supports extrapolating faces and hands.
[0442] Sequence-Level Adaptation
[0443] In an embodiment, the signaling information may comprise indicating whether the VTNN is adapted for all the pictures to which it is applied or activated.
[0444] Example use case 1 : specialization with respect to generated content
[0445] In an embodiment, a VTNN may be adapted in order to specialize it with respect to some aspects of the generated output, such as specializing to a specific face. In an example, a VTNN may be trained on a big dataset and may not perform well for all possible faces; the VTNN could be adapted by providing as auxiliary input some extra information, such as peculiar face features of a person whose face is to be temporally extrapolated. In another example, a VTNN may be adapted to generate faces under a certain illumination condition by providing as auxiliary input extra information about characteristic of illumination and color tone of the environment.
[0446] Example use case 2: task specialization
[0447] In an embodiment, a VTNN may be adapted in order to specialize it to a specific task or purpose. In one example, when the purpose of the VTNN is visual temporal extrapolation, the VTNN may be specialized to extrapolate with a certain visual or artistic style. In another example, when the purpose of a NN is visual generative Al (e.g., a NN that is capable of generating visual content), the NN may be specialized to add eyeglasses to faces in the pictures. In another example, when the purpose of a NN is visual generative Al (e.g., a NN that is capable of generating visual content), the NN may be specialized to perform visual temporal extrapolation.
[0448] Example use case 3: specialization with respect to what to generate and what not to generate
[0449] In an embodiment, the signaling information may comprise indicating that the VTNN is adapted for all pictures to which it is applied by inputting a signal that instructs the VTNN not to modify or generate one or more aspects. In an example, the signaling instructs the VTNN not to modify the background information. In another example, the signaling instructs the VTNN not to modify faces.
[0450] In an embodiment, the signaling information may comprise indicating that the VTNN is adapted for all pictures to which it is applied by inputting a signal that instructs the VTNN to modify or generate only one or more aspects. In an example, the signaling instructs the VTNN to modify only the background information. In another example, the signaling instructs the VTNN to modify only faces. In another example, the signaling instructs the VTNN to only extend a frame horizontally (e.g., from a normal picture to a 360 degrees view).
[0451] Signaling that sequence-level adaptation is performed
[0452] In an embodiment, the signaling information may comprise indicating that the NN is sequence-level adapted.
[0453] In an example embodiment, a new nnpfc_mode_idc value (e.g., 2) may be used in the NNPFC SEI message. When nnpfc_mode_idc is equal to 2, the NN is sequence-level adapted.
[0454] Signaling the purpose of adaptation
[0455] In an embodiment, the signaling information may comprise indicating the purpose of the sequence-level adaptation. A standard specification may specify several purposes of sequence-level adaptations, that may be referred to as adaptation purposes, where each adaptation purpose may be associated to an identifier value; an encoder may signal an identifier value for identifying an adaptation purpose. In an example, an adaptation purpose may comprise specializing the NN to one or more specific faces that appear in the video. In another example, an adaptation purpose may comprise specializing the NN to a certain more specific task than what indicated by the general purpose of the NN (e.g., by nnpfc_purpose in the NNPFC SEI message).
[0456] For example, the syntax element name may be nnpfc_adaptation_purpose.
[0457] In an example embodiment, the adaptation purpose may be indicated as a bit field wherein particular bit positions are assigned to particular adaptation purposes and a value of the bit in the particular bit position indicates whether the respective particular adaptation purpose is among the adaptation purposes of the defined neural network.
[0458] Signaling how the adaptation is performed
[0459] In an embodiment, the signaling information may comprise indicating whether the VTNN is adapted for all pictures to which it is applied by means of inputting a persistent auxiliary input (PAI) signal.
[0460] In an embodiment, the signaling information may comprise indicating the type of PAI signal, such as:- A tensor;- Text (e .g . , a text prompt; or- Image (e.g., an image prompt).
[0461] For example, the syntax element name may be nnpfc_persistent_aux_input_type.
[0462] Example on the number of required NNPFC SEI messages for performing adaptation
[0463] In an embodiment, sequence -level adaptation of a NN may be indicated by using a single NNPFC SEI message associated to that NN. The syntax element nnpfc_base_flag is set to 1.
[0464] In an alternative embodiment, sequence -level adaptation of a NN may be indicated in a second NNPFC SEI message associated to that NN, where the second NNPFC SEI message has same nnpfc_id as a first NNPFC SEI message present in the same CLVS. The first NNPFC SEI message has nnpfc_base_flag equal to 1, and the second NNPFC SEI message has nnpfc_base_flag equal to 0. There may be other NNPFC SEI messages which indicate sequence-level adaptation of the same NN, in addition to the second NNPFC SEI message. The adapted NN is applied to all pictures that activate the NN with this nnpfc_id value and nnpfa_target_base_flag equal to 0 and where the particular NNPFC SEI message with nnpfc_base_flag equal to 0 persists. This way, it may be possible to indicate more than one sequence-level adaptation or specialization of a base NN.
[0465] In an embodiment, the signaling information may comprise a purpose of foundational model. In an example, a foundational model is a machine -learned model, such as a neural network, that has been trained for a generic task such as next-word prediction, or next-pixel prediction, or next-frame prediction, inpainting, interpolation, and the like, on a large set of data. A foundational model may be specialized to perform a more specific task, for example by means of one or more of the following: further training, or finetuning, or adaptation based on an input instruction or prompt.
[0466] In an embodiment, when the purpose signaled by a first SEI message, such as a first NNPFC SEI message, comprises one of foundational model, or visual temporal extrapolation, or generative Al, it may be required that the encoder signals a second SEI message, such as a second NNPFC SEI message, that indicates a sequence-level adaptation. In one example, the NN whose characteristics (and eventually also the weights of the NN) are signaled in a first NNPFC SEI message cannot be used in an inference until a second NNPFC SEI message that indicates a sequence-level adaptation of that NN is signaled.
[0467] FIG. 5 is an example apparatus, which may be implemented in hardware, caused to enable the use of generative artificial intelligence for encoding, decoding or signaling media data. The apparatus 500 comprises at least one processor 502 (e.g., an FPGA and / or CPU), one or more memories504 including computer program code 505, the computer program code 505 having instructions to carry out the methods described herein, wherein the at least one memory 504 and the computer program code505 are configured to, with the at least one processor 502, cause the apparatus 500 to implement circuitry, a process, component, module, or function (implemented with control module 506) to implement the examples described herein, including enabling the use of generative artificial intelligence for encoding, decoding or signaling media data. Optionally included encoder 580 of the control module506 performs encoding, and optionally included decoder 590 implements decoding. The memory 504 may be a non-transitory memory, atransitory memory, a volatile memory (e.g., RAM), or a non-volatile memory (e.g., ROM).
[0468] The apparatus 500 includes a display and / or I / O interface 508, which includes user interface (UI) circuitry and elements, that may be used to display features or a status of the methods described herein (e.g., as one of the methods is being performed or at a subsequent time), or to receive input from a user such as with using a keypad, camera, touchscreen, touch area, microphone, biometric recognition, one or more sensors, and the like. The apparatus 500 includes one or more communication, e.g., network (N / W) interfaces (I / F(s)) 510. The communication I / F(s) 510 may be wired and / or wireless and communicate over the Intemet / other network(s) via any communication technique including via one or more links 524. The communication I / F(s) 510 may comprise one or more transmitters or one or more receivers.
[0469] The transceiver 516 comprises one or more transmitters 518 and one or more receivers 520. The transceiver 516 and / or communication I / F(s) 510 may comprise standard well-known components such as an amplifier, filter, frequency-converter, (de)modulator, and encoder / decoder circuitries and one or more antennas, such as antennas 514 used for communication over wireless link 522.
[0470] The control module 506 of the apparatus 500 comprises one of or both parts 506-1 and / or 506-2, which may be implemented in a number of ways. The control module 506 may be implemented in hardware as control module 506-1, such as being implemented as part of the one or more processors 502. The control module 506-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the control module 506 may be implemented as control module 506-2, which is implemented as computer program code (having corresponding instructions) 505 and is executed by the one or more processors 502. For instance, theone or more memories 504 store instructions that, when executed by the one or more processors 502, cause the apparatus 500 to perform one or more of the operations as described herein. Furthermore, the one or more processors 502, one or more memories 504, and example algorithms (e.g., as flowcharts and / or signaling diagrams), encoded as instructions, programs, or code, are means for causing performance of the operations described herein.
[0471] The apparatus 500 to implement the functionality of control 506 may correspond to any of the apparatuses depicted herein. Alternatively, apparatus 500 and its elements may not correspond to any of the other apparatuses depicted herein, as apparatus 500 may be part of a self- organizing / optimizing network (SON) node or other node, such as a node in a cloud.
[0472] The apparatus 500 may also be distributed throughout the network (e.g., internet 28) including within and between apparatus 500 and any network element (such as a base station 24 and / or apparatus 90).
[0473] Interface 512 enables data communication and signaling between the various items of apparatus 500, as shown in FIG. 5. For example, the interface 512 may be one or more buses such as address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. Computer program code (e.g., instructions) 505, including control 506 may comprise object- oriented software configured to pass data or messages between objects within computer program code 505. The apparatus 500 need not comprise each of the features mentioned, or may comprise other features as well. The various components of apparatus 500 may at least partially reside in a common housing 528, or a subset of the various components of apparatus 500 may at least partially be located in different housings, which different housings may include housing 528.
[0474] FIG. 6 shows a schematic representation of non-volatile memory media 600a (e.g. computer / compact disc (CD) or digital versatile disc (DVD)) and 600b (e.g. universal serial bus (USB) memory stick) and 600c (e.g. cloud storage for downloading instructions and / or parameters 602 or receiving emailed instructions and / or parameters 602) storing instructions and / or parameters 602 which when executed by a processor allows the processor to perform one or more of the operations of the methods described herein.
[0475] FIG. 7 is an example method 700 to implement the embodiments described herein, in accordance with an embodiment. At 702 the method 700 includes encoding signaling information in oralong a bitstream. The signaling information enables or supports performing visual temporal extrapolation. At 704 the method 700 includes signaling the signaling information.
[0476] The method 700 may be performed with an apparatus described herein, for example, the any apparatus of FIG. 1 to FIG. 5, any apparatus of FIG. 10, or any other apparatus described herein.
[0477] FIG. 8 is another example method 800 to implement the embodiments described herein, in accordance with an embodiment. At 802 the method 800 includes receiving signaling information, that enables or supports performing visual temporal extrapolation, in or along a bitstream. At 804 the method 800 includes decoding the signaling information to obtain a decoded signaling information. The decoded signaling information may simply be referred to as signaling information. At 806 the method 800 includes performing the visual temporal extrapolation, based at least on the decoded signaling information.
[0478] The method 800 may be performed with an apparatus described herein, for example, the any apparatus of FIG. 1 to FIG. 5, any apparatus of FIG. 10, or any other apparatus described herein.
[0479] FIG. 9 is still another example method 900 to implement the embodiments described herein, in accordance with an embodiment. At 902 the method 900 includes concluding that a temporal interleaving frame packing arrangement is in use in a bitstream. At 904 the method 900 includes concluding a constituent frame parity that a current frame where a temporal extrapolation post-filter is activated comprises a temporal interleaving frame packing arrangement. At 906 the method 900 includes selecting one or more input frames, for the temporal extrapolation post-filter, that comprise the concluded constituent frame parity. At 908 the method 900 includes applying the temporal extrapolation post-filter with the one or more input frames as input.
[0480] The method 900 may be performed with an apparatus described herein, for example, the any apparatus of FIG. 1 to FIG. 5, any apparatus of FIG. 10, or any other apparatus described herein.
[0481] Referring to FIG. 10, this figure shows a block diagram of one possible and non-limiting example in which the examples may be practiced. A user equipment (UE) 110, radio access network (RAN) node 170, and network element(s) 190 are illustrated. In the example of FIG. 1, the user equipment (UE) 110 is in wireless communication with a wireless network 100. A UE is a wireless device that can access the wireless network 100. The UE 110 includes one or more processors 120, one or more memories 125, and one or more transceivers 130 interconnected through one or more buses 127. Each of the one or more transceivers 130 includes a receiver, Rx, 132 and a transmitter, Tx, 133.The one or more buses 127 may be address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like . The one or more transceivers 130 are connected to one or more antennas 128. The one or more memories 125 include computer program code 123. The UE 110 includes a module 140, comprising one of or both parts 140-1 and / or 140-2, which may be implemented in a number of ways. The module 140 may be implemented in hardware as module 140-1, such as being implemented as part of the one or more processors 120. The module 140-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the module 140 may be implemented as module 140-2, which is implemented as computer program code 123 and is executed by the one or more processors 120. For instance, the one or more memories 125 and the computer program code 123 may be configured to, with the one or more processors 120, cause the user equipment 110 to perform one or more of the operations as described herein. The UE 110 communicates with a radio access network (RAN) node 170 via a wireless link 111.
[0482] The RAN node 170 in this example is a base station that provides access by wireless devices such as the UE 110 to the wireless network 100. The RAN node 170 may be, for example, a base station for fifth generation cellular network technology (5G), also called New Radio (NR). In 5G, the RAN node 170 may be a NG-RAN node, which is defined as either a gNB (e.g., base station for 5G / NR, for example, a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GC) or an ng (new generation)-eNB. A gNB is a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to a 5G core network (5GC) (such as, for example, the network element(s) 190). The ng-eNB, is a node providing evolved universal terrestrial radio access (E-UTRA), for example, the LTE radio access technology, user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GC. The NG-RAN node may include multiple gNBs, which may also include a central unit (CU) (gNB-CU) 196 and distributed unit(s) (DUs) (gNB-DUs), of which DU 195 is shown. Note that the DU may include or be coupled to and control a radio unit (RU). The gNB- CU is a logical node hosting radio resource control (RRC), service data adaptation protocol (SDAP) and PDCP protocols of the gNB or RRC and packet data convergence protocol (PDCP) protocols of the en-gNB (e.g., node providing NR user plane and control plane protocol terminations towards the UE, and acting as secondary node in E-UTRA-NR dual connectivity (EN-DC)) that controls the operation of one or more gNB-DUs. The gNB-CU terminates the interface between CU and DU control interface (Fl or Fl-C) interface connected with the gNB-DU. The Fl interface is illustrated as reference 198, although reference 198 also illustrates a link between remote elements of the RAN node 170 and centralized elements of the RAN node 170, such as between the gNB-CU 196 and the gNB-DU 195.The gNB-DU is a logical node hosting radio link control (RLC), MAC and physical layer (PHY) layers of the gNB or en-gNB, and its operation is partly controlled by gNB-CU. One gNB-CU supports one or multiple cells. One cell is supported by only one gNB-DU. The gNB-DU terminates the Fl interface 198 connected with the gNB-CU. Note that the DU 195 is considered to include the transceiver 160, for example, as part of a RU, but some examples of this may have the transceiver 160 as part of a separate RU, for example, under control of and connected to the DU 195. The RAN node 170 may also be an eNB (evolved NodeB) base station, for example, long term evolution (UTE), or any other suitable base station or node.
[0483] The RAN node 170 includes one or more processors 152, one or more memories 155, one or more network interfaces (N / W I / F(s)) 161, and one or more transceivers 160 interconnected through one or more buses 157. Each of the one or more transceivers 160 includes a receiver, Rx, 162 and a transmitter, Tx, 163. The one or more transceivers 160 are connected to one or more antennas 158. The one or more memories 155 include computer program code 153. The CU 196 may include the processor(s) 152, memories 155, and network interfaces 161. Note that the DU 195 may also include its own memory / memories and processor(s), and / or other hardware, but these are not shown.
[0484] The RAN node 170 includes a module 150, comprising one of or both parts 150-1 and / or 150-2, which may be implemented in a number of ways. The module 150 may be implemented in hardware as module 150-1, such as being implemented as part of the one or more processors 152. The module 150-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the module 150 may be implemented as module 150-2, which is implemented as computer program code 153 and is executed by the one or more processors 152. For instance, the one or more memories 155 and the computer program code 153 are configured to, with the one or more processors 152, cause the RAN node 170 to perform one or more of the operations as described herein. Note that the functionality of the module 150 may be distributed, such as being distributed between the DU 195 and the CU 196, or be implemented solely in the DU 195.
[0485] The one or more network interfaces 161 communicate over a network such as via the links 176 and 131. Two or more gNBs 170 may communicate using, for example, link 176. The link 176 may be wired or wireless or both and may implement, for example, an Xn interface for 5G, an X2 interface for LTE, or other suitable interface for other standards.
[0486] The one or more buses 157 may be address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, wireless channels, and the like. For example, the one ormore transceivers 160 may be implemented as a remote radio head (RRH) 195 for LTE or a distributed unit (DU) 195 for gNB implementation for 5G, with the other elements of the RAN node 170 possibly being physically in a different location from the RRH / DU, and the one or more buses 157 could be implemented in part as, for example, fiber optic cable or other suitable network connection to connect the other elements (for example, a central unit (CU), gNB-CU) of the RAN node 170 to the RRH / DU 195. Reference 198 also indicates those suitable network link(s).
[0487] It is noted that description herein indicates that ‘cells’ perform functions, but it should be clear that equipment which forms the cell may perform the functions. The cell makes up part of a base station. That is, there can be multiple cells per base station. For example, there could be three cells for a single carrier frequency and associated bandwidth, each cell covering one-third of a 360 degree area so that the single base station’s coverage area covers an approximate oval or circle. Furthermore, each cell can correspond to a single carrier and a base station may use multiple carriers. So when there are three 120 degree cells per carrier and two carriers, then the base station has a total of 6 cells.
[0488] The wireless network 100 may include a network element or elements 190 that may include core network functionality, and which provides connectivity via a link or links 181 with a further network, such as a telephone network and / or a data communications network (for example, the Internet). Such core network functionality for 5G may include access and mobility management fimction(s) (AMF(S)) and / or user plane functions (UPF(s)) and / or session management function(s) (SMF(s)). Such core network functionality for UTE may include MME (Mobility Management Entity) / SGW (Serving Gateway) functionality. These are merely example functions that may be supported by the network element(s) 190, and note that both 5G and LTE functions might be supported. The RAN node 170 is coupled via a link 131 to the network element 190. The link 131 may be implemented as, for example, an NG interface for 5G, or an SI interface for LTE, or other suitable interface for other standards. The network element 190 includes one or more processors 175, one or more memories 171, and one or more network interfaces (N / W I / F(s)) 180, interconnected through one or more buses 185. The one or more memories 171 include computer program code 173. The one or more memories 171 and the computer program code 173 are configured to, with the one or more processors 175, cause the network element 190 to perform one or more operations.
[0489] The wireless network 100 may implement network virtualization, which is the process of combining hardware and software network resources and network functionality into a single, softwarebased administrative entity, a virtual network. Network virtualization involves platform virtualization, often combined with resource virtualization. Network virtualization is categorized as either external, combining many networks, or parts of networks, into a virtual unit, or internal, providing network-likefunctionality to software containers on a single system. Note that the virtualized entities that result from the network virtualization are still implemented, at some level, using hardware such as processors 152 or 175 and memories 155 and 171, and also such virtualized entities create technical effects.
[0490] The computer readable memories 125, 155, and 171 may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The computer readable memories 125, 155, and 171 may be means for performing storage functions. The processors 120, 152, and 175 may be of any type suitable to the local technical environment, and may include one or more of general purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on a multi-core processor architecture, as non-limiting examples. The processors 120, 152, and 175 may be means for performing functions, such as controlling the UE 110, RAN node 170, network element(s) 190, and other functions as described herein.
[0491] In general, the various embodiments of the user equipment 110 can include, but are not limited to, cellular telephones such as smart phones, tablets, personal digital assistants (PDAs) having wireless communication capabilities, portable computers having wireless communication capabilities, image capture devices such as digital cameras having wireless communication capabilities, gaming devices having wireless communication capabilities, music storage and playback appliances having wireless communication capabilities, Internet appliances permitting wireless Internet access and browsing, tablets with wireless communication capabilities, as well as portable units or terminals that incorporate combinations of such functions.
[0492] One or more of modules 140-1, 140-2, 150-1, and 150-2 may be caused to enable the use of generative artificial intelligence for encoding, decoding or signaling media data. Computer program code 173 may also be caused to enable the use of generative artificial intelligence for encoding, decoding or signaling media data.
[0493] As described above, FIGs. 7 to 9 include flowcharts of an apparatus (e.g. 50, 500, or any other apparatuses described herein), method, and computer program product according to certain example embodiments. It will be understood that each block of the flowcharts, and combinations of blocks in the flowcharts, may be implemented by various means, such as hardware, firmware, processor, circuitry, and / or other devices associated with execution of software including one or more computer program instructions. For example, one or more of the procedures described above may be embodied by computer program instructions. In this regard, the computer program instructions which embody theprocedures described above may be stored by a memory (e.g. 58, 125, or 504) of an apparatus employing an embodiment of the present invention and executed by processing circuitry (e.g. 56, 120, or 502) of the apparatus. As will be appreciated, any such computer program instructions may be loaded onto a computer or other programmable apparatus (e.g., hardware) to produce a machine, such that the resulting computer or other programmable apparatus implements the functions specified in the flowchart blocks. These computer program instructions may also be stored in a computer-readable memory that may direct a computer or other programmable apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture, the execution of which implements the function specified in the flowchart blocks. The computer program instructions may also be loaded onto a computer or other programmable apparatus to cause a series of operations to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide operations for implementing the functions specified in the flowchart blocks.
[0494] A computer program product is therefore defined in those instances in which the computer program instructions, such as computer-readable program code portions, are stored by at least one non- transitory computer-readable storage medium with the computer program instructions, such as the computer-readable program code portions, being configured, upon execution, to perform the functions described above, such as in conjunction with the flowchart(s) of FIGs. 7 to 9. In other embodiments, the computer program instructions, such as the computer-readable program code portions, need not be stored or otherwise embodied by a non-transitory computer-readable storage medium, but may, instead, be embodied by a transitory medium with the computer program instructions, such as the computer- readable program code portions, still being configured, upon execution, to perform the functions described above.
[0495] Accordingly, blocks of the flowcharts support combinations of means for performing the specified functions and combinations of operations for performing the specified functions for performing the specified functions. It will also be understood that one or more blocks of the flowcharts, and combinations of blocks in the flowcharts, may be implemented by special purpose hardware -based computer systems which perform the specified functions, or combinations of special purpose hardware and computer instructions.
[0496] In some embodiments, certain ones of the operations above may be modified or further amplified. Furthermore, in some embodiments, additional optional operations may be included.Modifications, additions, or amplifications to the operations above may be performed in any order and in any combination.
[0497] Some embodiments have been described in relation to one or more neural networks performing visual temporal extrapolation. It is to be understood that embodiments can be realized with any generative modelling neural networks.
[0498] Some embodiments have been described in relation to specific SEI messages, such as the GFV SEI message, NNPFC SEI message, and / or NNPFA SEI message. It is to be understood that embodiments are not limited to these specific SEI messages but can be realized with any similar SEI messages.
[0499] Some embodiments have been described in relation to SEI messages. It is to be understood that embodiments are not limited to SEI messages but can be realized with any similar syntax structures, such as metadata OBUs.
[0500] Some embodiments have been described in relation to certain terms, such as picture unit, applicable to some video coding formats, such as VVC. It is to be understood that embodiments are not limited to video coding formats where the terms are applicable but can be realized with any video coding formats with their respective terms.
[0501] In the above, some example embodiments have been described with the help of syntax of the bitstream. It needs to be understood, however, that the corresponding structure and / or computer program may reside at the encoder for generating the bitstream and / or at the decoder for decoding the bitstream.
[0502] In the above, where example embodiments have been described with reference to an encoder, it needs to be understood that the resulting bitstream and the decoder have corresponding elements in them. Likewise, where example embodiments have been described with reference to a decoder, it needs to be understood that the encoder has structure and / or computer program for generating the bitstream to be decoded by the decoder.
[0503] Many modifications and other embodiments of the inventions set forth herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the inventions are not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims. Moreover, although the foregoing descriptions and the associated drawings describe example embodiments in thecontext of certain example combinations of elements and / or functions, it should be appreciated that different combinations of elements and / or functions may be provided by alternative embodiments without departing from the scope of the appended claims. In this regard, for example, different combinations of elements and / or functions than those explicitly described above are also contemplated as may be set forth in some of the appended claims. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
[0504] It should be understood that the foregoing description is only illustrative. Various alternatives and modifications may be devised by those skilled in the art. For example, features recited in the various dependent claims could be combined with each other in any suitable combination(s). In addition, features from different embodiments described above could be selectively combined into a new embodiment. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims.
[0505] References to a ‘computer’, ‘processor’, etc. should be understood to encompass not only computers having different architectures such as single / multi-processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device such as instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device, and the like.
[0506] As used herein, the term ‘circuitry’ may refer to any of the following: (a) hardware circuit implementations, such as implementations in analog and / or digital circuitry, and (b) combinations of circuits and software (and / or firmware), such as (as applicable): (i) a combination of processor(s) or (ii) portions of processor(s) / software including digital signal processor(s), software, and memory(ies) that work together to cause an apparatus to perform various functions, and (c) circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even when the software or firmware is not physically present. This description of ‘circuitry’ applies to uses of this term in this application. As a further example, as used herein, the term ‘circuitry’ would also cover an implementation of merely a processor (or multiple processors) or a portion of a processor and its (or their) accompanying software and / or firmware. The term ‘circuitry’ would also cover, for example and when applicable to the particular element, a baseband integrated circuit or applicationsprocessor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, or another network device.
[0507] Circuitry or Circuit: As used in this application, the term ‘circuitry’ or ‘circuit’ may refer to one or more or all of the following:(a) hardware -only circuit implementations (such as implementations in only analog and / or digital circuitry); and(b) combinations of hardware circuits and software, such as (as applicable):(i) a combination of analog and / or digital hardware circuit(s) with software / firmware; and(ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and(c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.
[0508] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example, and when applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
Claims
CLAIMSWhat is claimed is:
1. An apparatus comprising at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: encoding signaling information in or along a bitstream, wherein the signaling information enables or supports performing visual temporal extrapolation; and signaling the signaling information.
2. An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving signaling information, that enables or supports performing visual temporal extrapolation, in or along a bitstream; decoding the signaling information to obtain a decoded signaling information (signaling information); and performing the visual temporal extrapolation, based at least on the decoded signaling information.
3. The apparatus of any of the claims 1 or 2, wherein one or more neural networks (VTNN) are used to perform the visual temporal extrapolation .
4. The apparatus of any of the claims 1 to 2, wherein the signaling information is comprised in a neural network post-fdter characteristics (NNPFC) supplemental enhancement information (SEI) message; a generative face video (GFV) SEI message; or in a neural-network post-fdter characteristics (NNPFC) SEI message and a GFV SEI message.
5. The apparatus of any of the claims 1 to 4, wherein the signaling information comprises one or more of the following: a purpose of performing the visual temporal extrapolation at a decoder side; a temporal extrapolation characteristics (TEC) information that characterizes a neural network performing temporal extrapolation; or a purpose of generative modelling or generative artificial intelligence (Al) at decoder side.
6. The apparatus of any of the claims 1 to 5, wherein the TEC information is signaled based on one or more of the following: when the signaling information indicates that the purpose of the VTNN associated with the signaling information comprises visual temporal extrapolation, generative modeling, or generative artificial intelligence (Al), wherein the generative modeling or the generative Al refers to using the VTNN for generating data conditional to one or more inputs; or when the purpose comprises generative face video or generative video.
7. The apparatus of claim 6, wherein the purpose of the generative face video comprises generating or improving a picture or at least a part of a video comprising at least one or more faces, given one or more of the following: coordinates of the at least one or more faces; coordinates of landmarks or key-points of the at least one or more faces; transformation matrices that describe a motion or temporal change of the one or more key-points of the at least one or more faces; features of facial expressions; indication of blinking; texts associated to the video; audio associated to the video; or one or more other pictures.
8. The apparatus of any of the claims 5 to 7, wherein the TEC information comprises one or more of the following: information indicative of number of pictures that are extrapolated in one activation or inference of the VTNN; information indicative of time interval of pictures being extrapolated relative to a latest picture used as input for the inference; information indicative of whether the VTNN takes an auxiliary input that indicates output instances of one or more pictures being temporally extrapolated; information indicative of whether the VTNN takes an auxiliary input that indicates or is derived from facial parameters, wherein the auxiliary input is used by the VTNN in order to perform the temporal extrapolation of one or more faces in a picture or a video; information indicative of whether the VTNN takes an auxiliary input that is derived from the output of another VTNN, wherein the another VTNN is run or executed at the decoder side;information indicative of whether the VTNN takes an auxiliary input that is derived from the output of the another VTNN that converts or translates facial parameters to a format that is accepted or required by the VTNN; information indicative of whether the VTNN takes an auxiliary inputthat represents or is derived from a tensor that is generated at an encoder side; information indicative of whether the VTNN takes an auxiliary inputthat represents or is derived from an audio or a speech; information about characteristics of the audio or the speech, when the auxiliary input represents the audio or the speech; information indicative of whether the VTNN takes an auxiliary inputthat represents or is derived from text; information indicative of whether the VTNN produces different outputs for different inferences or executions of the VTNN with same input; information indicative of type or cause of output variability of the VTNN; information indicative of whether the VTNN is instructed or operated in a way that the VTNN does not yield output variability; information indicative of whether the VTNN is allowed to have output variability; information that makes the network not to have output variability; information indicative of a type of extrapolated outputs; information indicative of whether the VTNN takes an auxiliary input that indicates picture order count (POC) differences, timestamp differences, relative output position, or relative output differences of output pictures are to be generated or extrapolated; information indicative of format for the facial parameters which is accepted or required by the VTNN; information about a format of the tensor; information indicative of whether the VTNN takes an auxiliary input that represents text associated to the speech of a person for whom the face is to be temporally extrapolated; information indicative of whether the VTNN takes an auxiliary input that represents a text; information indicative of whether the VTNN produces different outputs for different inferences or executions of the VTNN with one or more same inputs; information indicative of whether an input to the VTNN consists of background information; information indicative of whether an input to the VTNN consists of foreground information; information indicative of whether an input to the VTNN comprises a drive picture;information indicative of whether an auxiliary input to the VTNN is derived from a lossy coding process; information indicative of whether the VTNN accepts an input that controls one or more aspects of how a final output is determined based on an output of the VTNN; one or more indications for indicating which facial parameters are present and / or used by the VTNN; information indicating that the number of extrapolated pictures of a neural network post-filter (NNPF) vary for each inference and the number of extrapolated pictures for each inference is indicated by using another signaling mechanism; an indication indicative of whether a decoder side is to generate a random seed value to the VTNN or derive a pseudo-random seed value based on other information that is signaled by an encoder; an indication indicative of whether a pseudo-random seed value is derived based on a value of a particular syntax element that is comprised in the TEC information or that is signaled to the decoder by other means; an indication indicative of whether the VTNN specified by the TEC information takes a seed value as an input or auxiliary input; an indication indicative of whether the VTNN specified by the TEC information takes a pseudo-random seed as the input or the auxiliary input; or a characteristics of randomness for the VTNN.
9. The apparatus of claim 8, wherein the other value comprises one of following: a sequence-wise quantization parameter (QP) value, and wherein the pseudo-random seed value is set equal to that sequence-wise QP value; a slice-wise QP value, and wherein the pseudo-random seed value is set equal to that slice-wise QP value; a value of a syntax element in the TEC information; a value of a syntax element in a neural-network post-filter characteristics (NNPFC) supplemental enhancement information (SEI) message that specifies the VTNN; or a value of a syntax element in a generative face video (GFV) SEI message that specifies the VTNN.
10. The apparatus of claim 8, wherein the facial parameters that are used as auxiliary input or that are used to derive the auxiliary input are carried within one of the following: a newly defined SEI message; a generative face video (GFV) SEI message; a neural network post-filter activation (NNPFA) SEI message; or a neural -network post-filter input (NNPFI) SEI message.
11. The apparatus of claim 8, wherein the tensor is an output of or is derived from another neural network that is run or executed at encoder side.
12. The apparatus of claim 11, wherein the tensor is signaled within the NNPFA SEI message, the GFV SEI message, or the NNPFI SEI message.
13. The apparatus of claim 8, wherein information about characteristics of the audio or speech comprises one or more of the following: sampling rate of the audio; bit depth and / or dynamic range of the audio; number of audio channels and / or channel configuration of the audio; duration of the audio and / or number of audio samples to be included in an input tensor; information on whether the audio is processed with source separation algorithm; information on whether the audio is spatial audio; or timing of the audio provided as input in relation to one or more pictures to be extrapolated.
14. The apparatus of claim 8, wherein the type or cause of output variability comprises one of the following: a pseudo-randomness; a deliberate variability; or floating-point numerical format that is used to represent value of one or more parameters of the VTNN and / the value of one or more signals flowing through the VTNN.
15. The apparatus of claim 8, wherein the information indicative of whether the VTNN takes an auxiliary input that indicates output instances of one or more pictures being temporally extrapolated is carried within a the NNPFC SEI message or a neural network post-filter activation (NNPFA) SEI message.
16. The apparatus of claim 8, wherein the information indicative of whether the VTNN takes the auxiliary input that indicates the POC differences, timestamp differences, relative output position, or relative output differences is signaled as part of a neural-network post-fdter characteristics (NNPFC) supplemental enhancement information (SEI) message.
17. The apparatus of any of the claims 8 or 16, wherein the auxiliary input that indicates the POC differences, timestamp differences, relative output position, or relative output differences is carried within a neural-network post-fdter activation (NNPFA) supplemental enhancement information (SEI) message or a neural-network post-fdter input (NPFI) SEI message.
18. The apparatus of any of the claims 8 or 10, wherein the facial parameters comprise one or more of the following:2D key-points;3D key-points; or one or more matrices.
19. The apparatus of claim 10, wherein the facial parameters obtained from the NNPFA SEI message, the GFV SEI message, or the NNPFI SEI message are converted to a required format that is indicated in the TEC information.
20. The apparatus any of the claims 8 or 13, wherein information indicative of whether the VTNN takes the auxiliary input that represents or is derived from the audio or the speech is used by the VTNN to generate or synthesize one or more features of a face.
21. The apparatus of claim 8, wherein when the VTNN takes the auxiliary input that represents the text, the TEC information further comprise information indicative of characteristics of the text comprising one or more of the following: whether the VTNN accepts or requires raw text; whether the VTNN accepts or requires tokens; whether the VTNN accepts or requires text embeddings; which tokenizer to use; which text embedder to use; or timing of the text provided as input in relation to the one or more pictures to be extrapolated.
22. The apparatus of claim 8, wherein the information indicative of the type of extrapolated outputs comprises indicating one or more categories of visual content that the VTNN is able to extrapolate.
23. The apparatus of claim 22, wherein the one or more categories comprise a face, hands, or a body.
24. The apparatus of any of the claims 3 to 23, wherein the signaling information further comprises indicating whether the VTNN is adapted for all the pictures to which the VTNN is applied, activated, or the VTNN is sequence -level adapted.
25. The apparatus is claim 24, wherein the signaling information further comprises indicating a purpose of the sequence-level adaptation.
26. The apparatus of claim 24 or 25, wherein a persistent auxiliary input (PAI) signal comprised in the signaling information is used for indicating whether the VTNN is adapted for all pictures to which the VTNN is applied.
27. The apparatus of claim 26, wherein the signaling information further comprises indicating a type of the PAI signal.
28. The apparatus of claim 27, wherein the type of the PAI comprises, a tensor, a text, or an image.
29. The apparatus of any of the claims 3 to 23, wherein the signaling information further comprises: information indicative of a purpose for a neural network that converts input facial parameters to an output that represents converted facial parameters; information indicative of a purpose for a neural network that converts input parameters to an output that represents converted parameters; or information indicative of a purpose for a neural network that converts a format of input parameters to another format of output parameters.
30. The apparatus of claim 8, wherein a neural -network post-fdter input (NNPFI) SEI message is used to provide the auxiliary input to a neural-network post-processing fdter defined by a (NNPFC) SEI message and activated by a neural network post-filter activation (NNPFA) SEI message.
31. An apparatus comprising at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform:concluding that a temporal interleaving frame packing arrangement is in use in a bitstream; concluding a constituent frame parity that a current frame where a temporal extrapolation post-filter is activated comprises a temporal interleaving frame packing arrangement; selecting one or more input frames, for the temporal extrapolation post-filter, that comprise the concluded constituent frame parity; and applying the temporal extrapolation post-filter with the one or more input frames as input.
32. A method comprising: encoding signaling information in or along a bitstream, wherein the signaling information enables or supports performing visual temporal extrapolation; and signaling the signaling information.
33. A method comprising : receiving signaling information, that enables or supports performing visual temporal extrapolation, in or along a bitstream; decoding the signaling information to obtain a decoded signaling information (signaling information); and performing the visual temporal extrapolation, based at least on the decoded signaling information.
34. The method of any of the claims 31 or 32, wherein one or more neural networks (VTNN) are used to perform the visual temporal extrapolation .
35. The method of any of the claims 31 to 32, wherein the signaling information is comprised in a neural network post-filter characteristics (NNPFC) supplemental enhancement information (SEI) message, a generative face video (GFV) SEI message; or in a neural-network post-filter characteristics (NNPFC) SEI message and a GFV SEI message.
36. The method of any of the claims 31 to 34, wherein the signaling information comprises one or more of the following: a purpose of performing the visual temporal extrapolation at a decoder side; or a temporal extrapolation characteristics (TEC) information that characterizes a neural network performing temporal extrapolation; ora purpose of generative modelling or generative artificial intelligence (Al) at decoder side.
37. The method of any of the claims 31 to 35, wherein the TEC information is signaled based on one or more of the following: when the signaling information indicates that the purpose of the VTNN associated with the signaling information comprises visual temporal extrapolation, generative modeling, or generative artificial intelligence (Al), wherein the generative modeling or the generative Al refers to using the VTNN for generating data conditional to one or more inputs; or when the purpose comprises generative face video or generative video.
38. The method of claim 36, wherein the purpose of the generative face video comprises generating or improving a picture or at least a part of a video comprising at least one or more faces, given one or more of the following: coordinates of the at least one or more faces; coordinates of landmarks or key-points of the at least one or more faces; transformation matrices that describe a motion or temporal change of the one or more key-points of the at least one or more faces; features of facial expressions; indication of blinking; texts associated to the video; audio associated to the video; or one or more other pictures.
39. The method of any of the claims 35 to 37, wherein the TEC information comprises one or more of the following: information indicative of number of pictures that are extrapolated in one activation or inference of the VTNN; information indicative of time interval of pictures being extrapolated relative to a latest picture used as input for the inference; information indicative of whether the VTNN takes an auxiliary input that indicates output instances of one or more pictures being temporally extrapolated; information indicative of whether the VTNN takes an auxiliary input that indicates or is derived from facial parameters, wherein the auxiliary input is used by the VTNN in order to perform the temporal extrapolation of one or more faces in a picture or a video;information indicative of whether the VTNN takes an auxiliary input that is derived from the output of another VTNN, wherein the another VTNN is run or executed at the decoder side; information indicative of whether the VTNN takes an auxiliary input that is derived from the output of the another VTNN that converts or translates facial parameters to a format that is accepted or required by the VTNN; information indicative of whether the VTNN takes an auxiliary input that represents or is derived from a tensor that is generated at an encoder side; information indicative of whether the VTNN takes an auxiliary input that represents or is derived from an audio or a speech; information about characteristics of the audio or the speech, when the auxiliary input represents the audio or the speech; information indicative of whether the VTNN takes an auxiliary input that represents or is derived from text; information indicative of whether the VTNN produces different outputs for different inferences or executions of the VTNN with same input; information indicative of type or cause of output variability of the VTNN; information indicative of whether the VTNN is instructed or operated in a way that the VTNN does not yield output variability; information indicative of whether the VTNN is allowed to have output variability; information that makes the network not to have output variability; information indicative of a type of extrapolated outputs; information indicative of whether the VTNN takes an auxiliary input that indicates picture order count (POC) differences, timestamp differences, relative output position, or relative output differences of output pictures are to be generated or extrapolated; information indicative of format for the facial parameters which is accepted or required by the VTNN; information about a format of the tensor; information indicative of whether the VTNN takes an auxiliary input that represents text associated to the speech of a person for whom the face is to be temporally extrapolated; information indicative of whether the VTNN takes an auxiliary input that represents a text; information indicative of whether the VTNN produces different outputs for different inferences or executions of the VTNN with one or more same inputs; information indicative of whether an input to the VTNN consists of background information;information indicative of whether an input to the VTNN consists of foreground information; information indicative of whether an input to the VTNN comprises a drive picture; information indicative of whether an auxiliary input to the VTNN is derived from a lossy coding process; information indicative of whether the VTNN accepts an input that controls one or more aspects of how a final output is determined based on an output of the VTNN; one or more indications for indicating which facial parameters are present and / or used by the VTNN; information indicating that the number of extrapolated pictures of a neural network post-filter (NNPF) vary for each inference and the number of extrapolated pictures for each inference is indicated by using another signaling mechanism; an indication indicative of whether a decoder side is to generate a random seed value to the VTNN or derive a pseudo-random seed value based on other information that is signaled by an encoder; an indication indicative of whether a pseudo-random seed value is derived based on a value of a particular syntax element that is comprised in the TEC information or that is signaled to the decoder by other means; an indication indicative of whether the VTNN specified by the TEC information takes a seed value as an input or auxiliary input; an indication indicative of whether the VTNN specified by the TEC information takes a pseudo-random seed as the input or the auxiliary input; or a characteristics of randomness for the VTNN.
40. The method of claim 39, wherein the other value comprises one of following: a sequence-wise quantization parameter (QP) value, and wherein the pseudo-random seed value is set equal to that sequence-wise QP value; a slice-wise QP value, and wherein the pseudo-random seed value is set equal to that slice-wise QP value; a value of a syntax element in the TEC information; a value of a syntax element in a neural-network post-filter characteristics (NNPFC) supplemental enhancement information (SEI) message that specifies the VTNN; or a value of a syntax element in a generative face video (GFV) SEI message that specifies the VTNN.
41. The method of claim 39, wherein the facial parameters that are used as auxiliary input or that are used to derive the auxiliary input are carried within one of the following: a newly defined SEI message; a generative face video (GFV) SEI message; a neural network post-filter activation (NNPFA) SEI message; or a neural -network post-filter input (NNPFI) SEI message.
42. The method of claim 39, wherein the tensor is an output of or is derived from another neural network that is run or executed at encoder side.
43. The method of claim 42, wherein the tensor is signaled within the NNPFA SEI message, the GFV SEI message, or the NNPFI SEI message.
44. The method of claim 39, wherein information about characteristics of the audio or speech comprises one or more of the following: sampling rate of the audio; bit depth and / or dynamic range of the audio; number of audio channels and / or channel configuration of the audio; duration of the audio and / or number of audio samples to be included in an input tensor; information on whether the audio is processed with source separation algorithm; information on whether the audio is spatial audio; or timing of the audio provided as input in relation to one or more pictures to be extrapolated.
45. The method of claim 39, wherein the type or cause of output variability comprises one of the following: a pseudo-randomness; a deliberate variability; or floating-point numerical format that is used to represent value of one or more parameters of the VTNN and / the value of one or more signals flowing through the VTNN.
46. The method of claim 39, wherein the information indicative of whether the VTNN takes an auxiliary input that indicates output instances of one or more pictures being temporally extrapolated is carried within the NNPFC SEI message or a neural network post-filter activation (NNPFA) SEI message.
47. The method of claim 39, wherein the information indicative of whether the VTNN takes the auxiliary input that indicates the POC differences, timestamp differences, relative output position, or relative output differences is signaled as part of a neural-network post-fdter characteristics (NNPFC) supplemental enhancement information (SEI).
48. The method of any of the claims 39 or 47, wherein the auxiliary input that indicates the POC differences, timestamp differences, relative output position, or relative output differences is carried within a neural-network post-fdter activation (NNPFA) supplemental enhancement information (SEI) message or a neural -network post-fdter input (NPFI) SEI message.
49. The method of any of the claims 39 or 41, wherein the facial parameters comprise one or more of the following:2D key-points;3D key-points; or one or more matrices.
50. The method of claim 41, wherein the facial parameters obtained from the NNPFA SEI message, the GFV SEI message, or the NNPFI SEI message are converted to a required format that is indicated in the TEC information.
51. The method any of the claims 39 or 44, wherein information indicative of whether the VTNN takes the auxiliary input that represents or is derived from the audio or the speech is used by the VTNN to generate or synthesize one or more features of a face.
52. The method of claim 39, wherein when the VTNN takes the auxiliary input that represents the text, the TEC information further comprise information indicative of characteristics of the text comprising one or more of the following: whether the VTNN accepts or requires raw text; whether the VTNN accepts or requires tokens; whether the VTNN accepts or requires text embeddings; which tokenizer to use; which text embedder to use; or timing of the text provided as input in relation to the one or more pictures to be extrapolated.
53. The method of claim 39, wherein the information indicative of the type of extrapolated outputs comprises indicating one or more categories of visual content that the VTNN is able to extrapolate.
54. The method of claim 53, wherein the one or more categories comprise a face, hands, or a body.
55. The method of any of the claims 34 to 54, wherein the signaling information further comprises indicating whether the VTNN is adapted for all the pictures to which the VTNN is applied, activated, or the VTNN is sequence -level adapted.
56. The method is claim 55, wherein the signaling information further comprises indicating a purpose of the sequence-level adaptation.
57. The method of claim 55 or 56, wherein a persistent auxiliary input (PAI) signal comprised in the signaling information is used for indicating whether the VTNN is adapted for all pictures to which the VTNN is applied.
58. The method of claim 57, wherein the signaling information further comprises indicating a type of the PAI signal.
59. The method of claim 58, wherein the type of the PAI comprises, a tensor, a text, or an image.
60. The method of any of the claims 34 to 54, wherein the signaling information further comprises: information indicative of a purpose for a neural network that converts input facial parameters to an output that represents converted facial parameters; information indicative of a purpose for a neural network that converts input parameters to an output that represents converted parameters; or information indicative of a purpose for a neural network that converts a format of input parameters to another format of output parameters.
61. The method of claim 39, wherein a neural -network post-fdter input (NNPFI) SEI message is used to provide the auxiliary input to a neural-network post-processing fdter defined by a (NNPFC) SEI message and activated by a neural network post-filter activation (NNPFA) SEI message.
62. A method comprising:concluding that a temporal interleaving frame packing arrangement is in use in a bitstream; concluding a constituent frame parity that a current frame where a temporal extrapolation post-filter is activated comprises a temporal interleaving frame packing arrangement; selecting one or more input frames, for the temporal extrapolation post-filter, that comprise the concluded constituent frame parity; and applying the temporal extrapolation post-filter with the one or more input frames as input.
63. An apparatus comprising means for performing the methods as claimed in any of the claims 32 to 62.
64. A computer readable medium comprising program instructions which, when executed by an apparatus, cause the apparatus to perform the methods as claimed in any of the claims 32 to 62.
65. The computer readable medium of claim 64, wherein the computer readable medium comprises a non-transitory computer readable medium.
Citation Information
Cited By
Devices and methods for signaling multiple extrapolations in video coding via a neural-network post-filter characteristics supplemental enhancement information message
US12666066B2
Systems and methods for signaling multiple spatial extrapolations in video coding
US20260012624A1