Integrating generative neural network in video coding
Integrating generative neural networks into multimedia coding through face video enhancement information and encoded feature parameters improves video encoding and decoding efficiency and quality.
Patent Information
- Application Number
- PCT/IB2025/050263
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-10
- Filing Date
- 2025-01-09
- Publication Date
- 2025-07-17
AI Technical Summary
Existing multimedia coding technologies lack integration of generative neural networks, limiting the enhancement and efficiency of video encoding and decoding processes.
Integrate generative neural networks into multimedia coding by encoding and signaling face video enhancement information using generative video supplemental enhancement information messages within picture units, and incorporating encoded feature parameters to invoke neural network inference for improved decoding.
Enhances video encoding and decoding processes by leveraging generative neural networks for improved quality and efficiency, enabling advanced video reconstruction and prediction techniques.
Smart Images

Figure IB2025050263_17072025_PF_FP_ABST
Abstract
Description
INTEGRATING GENERATIVE NEURAL NETWORK IN VIDEO CODINGTECHNICAL FIELD
[0001] The examples and non-limiting embodiments relate generally to multimedia coding and, more particularly to integrating generative neural network in multimedia coding.BACKGROUND
[0002] It is known to provide standardized formats for encoding, signaling, or decoding of media data.SUMMARY
[0003] Example 1 : An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: encoding, a suffix enhancement information unit comprising a face video enhancement information message in the picture unit that comprises a base or drive picture referenced by the face video enhancement information, to generate an encoded picture; and signaling the encoded picture.
[0004] Example 2: The apparatus of example 1, wherein the face video enhancement information message comprises a generative video supplemental enhancement information message.
[0005] Example 3: The apparatus of any of the examples 1 or 2, wherein the suffix enhancement information unit suffix supplemental enhancement information network abstraction layer unit.
[0006] Example 4: An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving a first input picture; encoding the first input picture into a first coded picture, wherein the first coded picture comprises a coded base picture or the coded drive picture, and wherein a first picture unit comprises the first coded picture; receiving a second input picture; extracting feature parameters from the second input picture; encoding the feature parameters into encoded feature parameters; and including the encoded feature parameters in the first picture unit or in a subsequent picture unit, wherein the subsequent picture unit comprises a subsequent coded picture, wherein the subsequent coded picture comprises a second coded drive picture or a second coded base picture.
[0007] Example 5: The apparatus of example 4, wherein the apparatus is further caused to perform: including the encoded feature parameters in a generative video (GV) supplemental enhancement information (SEI) message that follows the first coded picture in decoding order.
[0008] Example 6: The apparatus of example 5, wherein the apparatus is further caused to perform: including includes the generative video SEI message in a suffix SEI network abstraction layer (NAL) unit within the first picture unit.
[0009] Example 7 : The apparatus of any of the examples 4 to 6, wherein when the encoded feature parameters are decoded, the encoded feature parameters invoke a generative neural network inference.
[0010] Example 8: The apparatus of any of the examples 4 to 7, wherein the apparatus is further caused to perform: including an activation signal to invoke the generative neural network inference along the encoded feature parameters.
[0011] Example 9: The apparatus of any of the examples 4 to 8, wherein apparatus is further caused to perform: including an identifier or a counter value in the generative video SEI message, wherein the identifier or the counter value identifies the generative video SEI message from other generative video SEI messages in the first picture unit or the subsequent picture unit.
[0012] Example 10: The apparatus of 4, wherein the apparatus is further caused to perform: generating multiple generative face video (GFV) supplemental enhancement information (SEI) messages with different payload in the first picture unit or the subsequent picture unit, wherein each instance of a GFV SEI message invokes a generative neural network (NN) inference.
[0013] Example 11: The apparatus of example 10, wherein the apparatus is caused to perform: using a suffix SEI NAL unit comprising a GFV SEI message in the first picture unit that comprises a base picture or a drive picture referenced by the GFV SEI message.
[0014] Example 12: The apparatus of example 11, wherein the apparatus is caused to perform: adding a counter syntax element in the GFV SEI message.
[0015] Example 13: An apparatus comprising at least one processor; and at least one non- transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving a first coded picture from a first picture unit, wherein the first coded picture comprises a coded base picture or a coded drive picture; decoding the first coded picture to generate a first decoded picture; receiving encoded feature parameters from the first picture unit or from a subsequent picture unit; creating a feature input tensor from the encoded feature parameters; andinvoking a generative neural network inference with the first decoded picture and the feature input tensor as inputs to generate a picture.
[0016] Example: 14 The apparatus of example 13, wherein the apparatus is caused to perform accepting, by an interface, less than a picture unit to be passed for decoding, and wherein decoding of the first coded picture is invoked prior to receiving or creating the feature input tensor.
[0017] Example: 15 The apparatus of example 13, wherein the apparatus is caused to perform: accepting, by an interface , less than a picture unit to be passed for decoding, and wherein creating the feature input tensor is performed prior to receiving or decoding a subsequent coded picture from the subsequent picture unit.
[0018] Example 16: The apparatus of any of the examples 13 to 15, wherein the apparatus is further caused to perform: decoding the encoded feature parameters; and invoking the generative neural network inference based on decoding the encoded feature parameters.
[0019] Example 17: The apparatus of any of the examples 13 to 16, wherein apparatus is further caused to perform: decoding an identifier or a counter value from the generative video SEI message, wherein the identifier or the counter value identifies the generative video SEI message from other generative video SEI messages in the first picture unit or the subsequent picture unit.
[0020] Example 18: The apparatus of example 17, wherein when the identifier or the counter value is the same as for a previous generative video SEI message in the first picture unit or the subsequent picture unit, the current generative video SEI message is concluded to be a copy of the previous generative video SEI message.
[0021] Example 19: The apparatus of 13, wherein the apparatus is further caused to perform: decoding multiple generative face video (GFV) supplemental enhancement information (SEI) messages with different payload from the first picture unit or the subsequent picture unit, wherein each instance of a GFV SEI message invokes a generative neural network (NN) inference.
[0022] Example 20: The apparatus of example 19, wherein the apparatus is caused to perform: decoding a suffix SEI NAL unit comprising a GFV SEI message from the first picture unit that comprises a base picture or a drive picture referenced by the GFV SEI message.
[0023] Example 21: The apparatus of example 20, wherein the apparatus is caused to perform: decoding a counter syntax element from the GFV SEI message.
[0024] Example 22: An apparatus comprising at least one processor; and at least one non- transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving a first input picture; encoding the first input picture into a first coded picture, wherein the first coded picture is a coded base picture, and wherein the first coded picture resides in an independent layer, and wherein a first access unit comprises the first coded picture; receiving a second input picture; encoding a second coded picture, wherein the second coded picture comprises a coded drive picture based on the second input picture or a coded dummy picture, and wherein the second coded picture resides in a dependent layer, and wherein a second picture unit comprises the second coded picture, and wherein a second access unit comprises the second coded picture; extracting feature parameters from the second input picture; encoding the feature parameters into encoded feature parameters; and including the encoded feature parameters in the second picture unit.
[0025] Example 23: The apparatus of example 22, wherein the apparatus is further caused to perform: encoding a coded replica picture in the first access unit, wherein the coded replica picture resides in the dependent layer and is predicted from the independent layer, and wherein the coded replica picture comprises no prediction residual.
[0026] Example 24: The apparatus of any of the examples 22 or 23, wherein the apparatus is further caused to perform: including the encoded feature parameters in a generative video supplemental enhancement information (SEI) message in the second picture unit.
[0027] Example 25: The apparatus of any of the examples 22 to 24, wherein the apparatus is further caused to perform: predicting a coded dummy picture from a previous picture, in decoding order, within the dependent layer.
[0028] Example 26: An apparatus comprising at least one processor; and at least one non- transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving a first input picture; encoding the first input picture into a first coded base-layer picture, wherein the first coded base-layer picture is a coded base picture, and wherein the first coded base-layer picture resides in an independent layer, and wherein a first access unit comprises the first coded base-layer picture; encoding a first coded dependent-layer picture, wherein the first coded dependent-layer picture resides in a dependent layer, and wherein the first access unit comprises the first coded dependent-layer picture; receiving a second input picture; encoding a second coded base-layer picture, wherein the second coded base-layer picture comprises a coded drive picturebased on the second input picture or a coded dummy picture, and wherein the second coded base-layer picture resides in the independent layer, and wherein a second access unit comprises the second coded base-layer picture; encoding a second coded dependent-layer picture, wherein the second coded dependent-layer picture resides in a dependent layer, and the second access unit comprises the second coded dependent-layer picture; indicating in or along the second coded dependent-layer picture that the second coded dependent-layer picture uses the first coded dependent-layer picture and the second coded base-layer picture as reference for prediction; extracting feature parameters from the second input picture; encoding the feature parameters into encoded feature parameters; and indicating that the encoded feature parameters are to be used for decoding of the second coded dependent-layer picture.
[0029] Example 27: The apparatus of example 26, wherein the apparatus is further caused to perform encoding the first dependent-layer picture as a coded replica picture, wherein the coded replica picture is predicted from the independent layer, and wherein the coded replica picture no prediction residual.
[0030] Example 28: The apparatus of any of the examples 26 or 27, wherein the apparatus is further caused to perform: encoding an indication in or along the second coded base-layer picture to omit output of a reconstructed picture decoded from the second coded base-layer picture.
[0031] Example 29: The apparatus of any of the examples 26 or 27, wherein the apparatus is further caused to perform: indicating that a first reconstructed base-layer picture decoded from the first coded base-layer picture is output; indicating that a first reconstructed dependent-layer picture decoded from the first coded dependent-layer picture is output; indicating that a second reconstructed dependentlayer picture decoded from the second coded dependent-layer picture is output; indicating that an independent layer and a dependent layer form an output layer set and the dependent layer is an output layer of the output layer set; or indicating a profile to the output layer set, wherein the profile includes the generative neural network capability.
[0032] Example 30: The apparatus of any of the examples 26 to 28, wherein the apparatus is further caused to perform: creating a feature input tensor from the encoded feature parameters.
[0033] Example 31: The apparatus of any of the examples 26 to 30, wherein the apparatus is further caused to perform: including the encoded feature parameters or the feature input tensor in an adaptation parameter set with a type indicating generative video; or including the encoded feature parameters or the feature input tensor in a picture header of the second coded dependent-layer picture.
[0034] Example 32: An apparatus comprising at least one processor; and at least one non- transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: reconstructing a first decoded dependent-layer picture from a first coded dependent-layer picture; and reconstructing a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependent-layer picture and a feature input tensor as inputs to generate a picture.
[0035] Example 33: The apparatus of example 32, wherein a syntax of the second coded dependent-layer picture excludes syntax elements that are not relevant to reconstructing the second decoded dependent-layer picture with the generative neural network.
[0036] Example 34: The apparatus of example 32, a syntax of one or more syntax structures of the second dependent-layer excludes syntax elements that are not relevant for reconstructing the second decoded dependent-layer picture with the generative neural network.
[0037] Example 35: An apparatus comprising at least one processor; and at least one non- transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: reconstructing a first decoded dependent-layer picture from a first coded dependent-layer picture; reconstructing one or more prediction units of a second decoded dependentlayer picture by invoking a generative neural network inference with the first decoded dependent-layer picture and a feature input tensor as inputs; and encoding a difference between the one or more prediction units and collocated units of the second input picture as prediction residual of the second coded dependent-layer picture.
[0038] Example 36: An apparatus comprising at least one processor; and at least one non- transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: reconstructing a first decoded dependent-layer picture from a first coded dependent-layer picture; invoking a generative neural network inference with the first decoded dependent-layer picture and a feature input tensor as inputs to generate an inter-layer reference picture; and encoding a second coded dependent-layer picture with reference to the inter-layer reference picture.
[0039] Example 37: The apparatus of example 36, wherein the apparatus is further caused to perform: encoding difference between the second input picture and the inter-layer reference picture as prediction residual of the second coded dependent-layer picture.
[0040] Example 38: An apparatus comprising at least one processor; and at least one non- transitory memory storing instructions that, when executed by the at least one processor, cause theapparatus at least to perform: receiving a first coded base-layer picture, wherein the first coded baselayer picture is a coded base picture, the first coded base-layer picture resides in an independent layer, and a first access unit comprises the first coded base-layer picture; decoding the first coded base-layer picture into a first decoded base-layer picture; receiving a first coded dependent-layer picture, wherein the first coded dependent-layer picture resides in a dependent layer, and the first access unit comprises the first coded dependent-layer picture; decoding the first coded dependent-layer picture into a first decoded dependent-layer picture; receiving a second coded base-layer picture, wherein the second coded base-layer picture may be a coded drive picture based on the second input picture or a coded dummy picture, the second coded base-layer picture resides in the independent layer, and a second access unit comprises the second coded base-layer picture; in response to the second coded base-layer picture being a coded drive picture, decoding the second coded base-layer picture into a second decoded base-layer picture; receiving a second coded dependent-layer picture, wherein the second coded dependent-layer picture resides in a dependent layer, and the second access unit comprises the second coded dependent-layer picture; decoding from or along the second coded dependent-layer picture that the second coded dependent-layer picture may use the first coded dependent-layer picture and the second coded base-layer picture as reference for prediction; obtaining a feature input tensor that is in use for the decoding of the second coded dependent-layer picture; and reconstructing a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependent-layer picture and the feature input tensor as inputs to generate a picture.
[0041] Example 39: An apparatus comprising at least one processor; and at least one non- transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving a first input picture; encoding the first input picture into a first coded picture, wherein the first coded picture is a coded base picture or a coded drive picture, and wherein the first coded picture resides in an independent layer, and wherein a first access unit comprises the first coded picture; receiving a second input picture; extracting feature parameters from the second input picture; encoding the feature parameters into encoded feature parameters; encoding a second coded picture, wherein the second coded picture resides in a dependent layer, and wherein the second access unit comprises the second coded picture; and indicating that the encoded feature parameters are used for decoding the second coded dependent-layer picture.
[0042] Example 40: The apparatus of example 39, wherein the apparatus is further caused to perform: reconstructing a first decoded picture from the first coded picture; creating a feature input tensor from the encoded feature parameters; invoking a generative neural network inference with the first decoded picture and the feature input tensor as inputs to generate a first inter-layer reference picture; and encoding the second coded picture with reference to the first inter-layer reference picture.
[0043] Example 41: The apparatus of any of the examples 39 or 40, wherein the apparatus is further caused to perform: omitting encoding prediction residual for the second coded picture.
[0044] Example 42: The apparatus of any of the examples 39 to 41, wherein the apparatus is further caused to perform: encoding difference between the second input picture and the first inter-layer reference picture as prediction residual of the second coded picture.
[0045] Example 43: The apparatus of any of the examples 39 to 41, wherein the apparatus is further caused to perform: in response to the first coded picture being a coded base picture, indicating that the first coded picture is an output; in response to the first coded picture being a coded drive picture, indicating that the first coded picture is not the output; indicating that the independent layer and the dependent layer form an output layer set and the dependent layer is an output layer of the output layer set; or indicating that the second coded picture is output.
[0046] Example 44: The apparatus of any of the examples 39 to 43, wherein the apparatus is further caused to perform: indicating in or along the first coded picture an indication indicating that the first coded picture is the coded base picture or that the first coded picture is the coded drive picture; or indicating with a picture header extra bit that the first coded picture is the coded base picture or that the first coded picture is the coded drive picture.
[0047] Example 45: The apparatus of any of the examples 39 to 44, wherein the apparatus is further caused to perform: indicating that the second coded picture uses inter-layer prediction from the independent layer.
[0048] Example 46: The apparatus of any of the examples 39 to 44, wherein the apparatus is further caused to perform: inferring that the inter-layer prediction from the independent layer refers to the first inter-layer reference picture.
[0049] Example 47: The apparatus of any of the examples 39 to 44, wherein the apparatus is further caused to perform: receiving a third input picture; and encoding the third input picture into a third coded picture, wherein the third coded picture is the coded drive picture, and wherein the third coded picture resides in the independent layer, and the third coded picture follows the first coded picture in decoding order and precedes the second coded picture in the decoding order.
[0050] Example 48: The apparatus of example 47, wherein the apparatus is further caused to perform: in response to the third coded picture residing in the second access unit, omitting indicating that the third coded picture is the coded drive picture and infers that the third coded picture is the coded drive picture.
[0051] Example 49: The apparatus of example 47, wherein the apparatus is further caused to perform: indicating in or along the third coded picture an indication indicating that the third coded picture is the coded drive picture.
[0052] Example 50: The apparatus of example 47, wherein the apparatus is further caused to perform: indicating using a picture header extra bit that the third coded picture is the coded drive picture.
[0053] Example 51: The apparatus of any of the examples 47 to 50, wherein the apparatus is further caused to perform: reconstructing the third decoded picture from the third coded picture; and using the third decoded picture in the invocation of the generative neural network inference as an additional input.
[0054] Example 52: The apparatus of any of the examples 47 to 51, wherein the apparatus is further caused to perform: inferring when the inter-layer prediction from the independent layer refers to the first inter-layer reference picture or the third decoded picture.
[0055] Example 53: The apparatus of any of the examples 47 to 51, wherein the apparatus is further caused to perform: indicating in a reference picture list syntax structure information indicating that the first coded picture is used as input for generating the inter-layer reference picture.
[0056] Example 54: The apparatus of example 53, wherein the information indicating that the first coded picture is used as input for generating the inter-layer reference picture comprises a picture order count difference between a current picture and the first coded picture.
[0057] Example 55: A method comprising: encoding, a suffix enhancement information unit comprising a face video enhancement information message in the picture unit that comprises a base or drive picture referenced by the face video enhancement information, to generate an encoded picture; and signaling the encoded picture.
[0058] Example 56: The method of example 55, wherein the face video enhancement information message comprises a generative video supplemental enhancement information message.
[0059] Example 57: The method of any of the examples 55 or 56, wherein the suffix enhancement information unit suffix supplemental enhancement information network abstraction layer unit.
[0060] Example 58: A method comprising: receiving a first input picture; encoding the first input picture into a first coded picture, wherein the first coded picture comprises a coded base picture or a coded drive picture, and wherein a first picture unit comprises the first coded picture; receiving a second input picture; extracting feature parameters from the second input picture; encoding the feature parameters into encoded feature parameters; and including the encoded feature parameters in the first picture unit or in a subsequent picture unit, wherein the subsequent picture unit comprises a subsequent coded picture, wherein the subsequent coded picture comprises a second coded drive picture or a second coded base picture.
[0061] Example 59: The method of example 58 further comprising: including the encoded feature parameters in a generative video (GV) supplemental enhancement information (SEI) message that follows the first coded picture in decoding order.
[0062] Example 60: The method of example 59 further comprising: including includes the generative video SEI message in a suffix SEI network abstraction layer (NAL) unit within the first picture unit.
[0063] Example 61: The method of any of the examples 58 to 60, wherein when the encoded feature parameters are decoded, the encoded feature parameters invoke a generative neural network inference.
[0064] Example 62: The method of any of the examples 58 to 61 further comprising: including an activation signal to invoke the generative neural network inference along the encoded feature parameters.
[0065] Example 63: The method of any of the examples 58 to 62 further comprising: including an identifier or a counter value in the generative video SEI message, wherein the identifier or the counter value identifies the generative video SEI message from other generative video SEI messages in the first picture unit or the subsequent picture unit.
[0066] Example 64: The method of 58 further comprising: generating multiple generative face video (GFV) supplemental enhancement information (SEI) messages with different payload in the first picture unit or the subsequent picture unit, wherein each instance of a GFV SEI message invokes a generative neural network (NN) inference.
[0067] Example 65: The method of example 64 further comprising: using a suffix SEI NAL unit comprising a GFV SEI message in the first picture unit that comprises a base picture or a drive picture referenced by the GFV SEI message.
[0068] Example 66: The method of example 65 further comprising: adding a counter syntax element in the GFV SEI message.
[0069] Example 67: The method comprising: receiving a first coded picture from a first picture unit, wherein the first coded picture comprises a coded base picture or a coded drive picture; decoding the first coded picture to generate a first decoded picture; receiving encoded feature parameters from the first picture unit or from a subsequent picture unit; creating a feature input tensor from the encoded feature parameters; and invoking a generative neural network inference with the first decoded picture and the feature input tensor as inputs to generate a picture.
[0070] Example 68: The method of example 67 further comprising: accepting, by an interface, less than a picture unit to be passed for decoding, and wherein decoding of the first coded picture is invoked prior to receiving or creating the feature input tensor.
[0071] Example 69: The method of example 67 accepting, by an interface, less than a picture unit to be passed for decoding, and wherein creating the feature input tensor is performed prior to receiving or decoding a subsequent coded picture from the subsequent picture unit.
[0072] Example 70: The method of any of the examples 67 to 69 further comprising: decoding the encoded feature parameters; and invoking the generative neural network inference based on decoding the encoded feature parameters.
[0073] Example 71: The method of any of the examples 67 to 70 further comprising: decoding an identifier or a counter value from the generative video SEI message, wherein the identifier or the counter value identifies the generative video SEI message from other generative video SEI messages in the first picture unit or the subsequent picture unit.
[0074] Example 72: The method of example 71, wherein when the identifier or the counter value is the same as for a previous generative video SEI message in the first picture unit or the subsequent picture unit, the current generative video SEI message is concluded to be a copy of the previous generative video SEI message.
[0075] Example 73: The method of 67 further comprising: decoding multiple generative face video (GFV) supplemental enhancement information (SEI) messages with different payload from the first picture unit or the subsequent picture unit, wherein each instance of a GFV SEI message invokes a generative neural network (NN) inference.
[0076] Example 74: The method of example 73 further comprising: decoding a suffix SEI NAL unit comprising a GFV SEI message from the first picture unit that comprises a base picture or a drive picture referenced by the GFV SEI message.
[0077] Example 75: The method of example 74 further comprising: decoding a counter syntax element from the GFV SEI message.
[0078] Example 76: A method comprising: receiving a first input picture; encoding the first input picture into a first coded picture, wherein the first coded picture is a coded base picture, and wherein the first coded picture resides in an independent layer, and wherein a first access unit comprises the first coded picture; receiving a second input picture; encoding a second coded picture, wherein the second coded picture comprises a coded drive picture based on the second input picture or a coded dummy picture, and wherein the second coded picture resides in a dependent layer, and wherein a second picture unit comprises the second coded picture, and wherein a second access unit comprises the second coded picture; extracting feature parameters from the second input picture; encoding the feature parameters into encoded feature parameters; and including the encoded feature parameters in the second picture unit.
[0079] Example 77: The method of example 76 further comprising: encoding a coded replica picture in the first access unit, wherein the coded replica picture resides in the dependent layer and is predicted from the independent layer, and wherein the coded replica picture comprises no prediction residual.
[0080] Example 78: The method of any of the examples 76 or 77 further comprising: including the encoded feature parameters in a generative video supplemental enhancement information (SEI) message in the second picture unit.
[0081] Example 79: The method of any of the examples 76 to 78 further comprising: predicting a coded dummy picture from a previous picture, in decoding order, within the dependent layer.
[0082] Example 80: A method comprising: receiving a first input picture; encoding the first input picture into a first coded base-layer picture, wherein the first coded base-layer picture is a coded base picture, and wherein the first coded base-layer picture resides in an independent layer, and wherein a first access unit comprises the first coded base-layer picture; encoding a first coded dependent-layer picture, wherein the first coded dependent-layer picture resides in a dependent layer, and wherein the first access unit comprises the first coded dependent-layer picture; receiving a second input picture;encoding a second coded base-layer picture, wherein the second coded base-layer picture comprises a coded drive picture based on the second input picture or a coded dummy picture, and wherein the second coded base-layer picture resides in the independent layer, and wherein a second access unit comprises the second coded base-layer picture; encoding a second coded dependent-layer picture, wherein the second coded dependent-layer picture resides in a dependent layer, and the second access unit comprises the second coded dependent-layer picture; indicating in or along the second coded dependent-layer picture that the second coded dependent-layer picture uses the first coded dependent-layer picture and the second coded base-layer picture as reference for prediction; extracting feature parameters from the second input picture; encoding the feature parameters into encoded feature parameters; and indicating that the encoded feature parameters are to be used for decoding of the second coded dependent-layer picture.
[0083] Example 81: The method of example 80 further comprising: encoding the first dependentlayer picture as a coded replica picture, wherein the coded replica picture is predicted from the independent layer, and wherein the coded replica picture no prediction residual.
[0084] Example 82: The method of any of the examples 80 or 81 further comprising: encoding an indication in or along the second coded base-layer picture to omit output of a reconstructed picture decoded from the second coded base-layer picture.
[0085] Example 83: The method of any of the examples 80 or 81 further comprising: indicating that a first reconstructed base-layer picture decoded from the first coded base-layer picture is output; indicating that a first reconstructed dependent-layer picture decoded from the first coded dependentlayer picture is output; indicating that a second reconstructed dependent-layer picture decoded from the second coded dependent-layer picture is output; indicating that an independent layer and a dependent layer form an output layer set and the dependent layer is an output layer of the output layer set; or indicating a profile to the output layer set, wherein the profile includes the generative neural network capability.
[0086] Example 84: The method of any of the examples 80 to 82 further comprising: creating a feature input tensor from the encoded feature parameters.
[0087] Example 85: The method of any of the examples 80 to 84 further comprising: including the encoded feature parameters or the feature input tensor in an adaptation parameter set with a type indicating generative video; or including the encoded feature parameters or the feature input tensor in a picture header of the second coded dependent-layer picture.
[0088] Example 86: A method comprising: reconstructing a first decoded dependent-layer picture from a first coded dependent-layer picture; and reconstructing a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependent-layer picture and a feature input tensor as inputs to generate a picture.
[0089] Example 87 : The method of example 86, wherein a syntax of the second coded dependentlayer picture excludes syntax elements that are not relevant to reconstructing the second decoded dependent-layer picture with the generative neural network.
[0090] Example 88: The method of example 86, a syntax of one or more syntax structures of the second dependent-layer excludes syntax elements that are not relevant for reconstructing the second decoded dependent-layer picture with the generative neural network.
[0091] Example 89: A method comprising: reconstructing a first decoded dependent-layer picture from a first coded dependent-layer picture; reconstructing one or more prediction units of a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependent-layer picture and a feature input tensor as inputs; and encoding a difference between the one or more prediction units and collocated units of the second input picture as prediction residual of the second coded dependent-layer picture.
[0092] Example 90: A method comprising: reconstructing a first decoded dependent-layer picture from a first coded dependent-layer picture; invoking a generative neural network inference with the first decoded dependent-layer picture and a feature input tensor as inputs to generate an inter-layer reference picture; and encoding a second coded dependent-layer picture with reference to the inter-layer reference picture.
[0093] Example 91: The method of example 90 further comprising: encoding difference between the second input picture and the inter-layer reference picture as prediction residual of the second coded dependent-layer picture.
[0094] Example 92: A method comprising: receiving a first coded base-layer picture, wherein the first coded base-layer picture is a coded base picture, the first coded base-layer picture resides in an independent layer, and a first access unit comprises the first coded base-layer picture; decoding the first coded base-layer picture into a first decoded base-layer picture; receiving a first coded dependent-layer picture, wherein the first coded dependent-layer picture resides in a dependent layer, and the first access unit comprises the first coded dependent-layer picture; decoding the first coded dependent-layer picture into a first decoded dependent-layer picture; receiving a second coded base-layer picture, wherein the second coded base-layer picture may be a coded drive picture based on the second input picture or acoded dummy picture, the second coded base-layer picture resides in the independent layer, and a second access unit comprises the second coded base- layer picture; in response to the second coded baselayer picture being a coded drive picture, decoding the second coded base-layer picture into a second decoded base-layer picture; receiving a second coded dependent-layer picture, wherein the second coded dependent-layer picture resides in a dependent layer, and the second access unit comprises the second coded dependent-layer picture; decoding from or along the second coded dependent-layer picture that the second coded dependent-layer picture may use the first coded dependent-layer picture and the second coded base-layer picture as reference for prediction; obtaining a feature input tensor that is in use for the decoding of the second coded dependent-layer picture; and reconstructing a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependent-layer picture and the feature input tensor as inputs to generate a picture.
[0095] Example 93: A method comprising: receiving a first input picture; encoding the first input picture into a first coded picture, wherein the first coded picture is a coded base picture or a coded drive picture, and wherein the first coded picture resides in an independent layer, and wherein a first access unit comprises the first coded picture; receiving a second input picture; extracting feature parameters from the second input picture; encoding the feature parameters into encoded feature parameters, encoding a second coded picture, wherein the second coded picture resides in a dependent layer, and wherein the second access unit comprises the second coded picture; and indicating that the encoded feature parameters are used for decoding the second coded dependent-layer picture.
[0096] Example 94: The method of example 93 further comprising: reconstructing a first decoded picture from the first coded picture; creating a feature input tensor from the encoded feature parameters; invoking a generative neural network inference with the first decoded picture and the feature input tensor as inputs to generate a first inter-layer reference picture; and encoding the second coded picture with reference to the first inter-layer reference picture.
[0097] Example 95: The method of any of the examples 93 or 94 further comprising: omitting encoding prediction residual for the second coded picture.
[0098] Example 96: The method of any of the examples 93 to 95 further comprising: encoding difference between the second input picture and the first inter-layer reference picture as prediction residual of the second coded picture.
[0099] Example 97: The method of any of the examples 93 to 95 further comprising: in response to the first coded picture being a coded base picture, indicating that the first coded picture is an output; in response to the first coded picture being a coded drive picture, indicating that the first coded picture is not the output; indicating that the independent layer and the dependent layer form an output layer setand the dependent layer is an output layer of the output layer set; or indicating that the second coded picture is output.
[0100] Example 98: The method of any of the examples 93 to 97 further comprising: indicating in or along the first coded picture an indication indicating that the first coded picture is the coded base picture or that the first coded picture is the coded drive picture; or indicating with a picture header extra bit that the first coded picture is the coded base picture or that the first coded picture is the coded drive picture.
[0101] Example 99: The method of any of the examples 93 to 98 further comprising: indicating that the second coded picture uses inter-layer prediction from the independent layer.
[0102] Example 100: The method of any of the examples 93 to 98 further comprising: inferring that the inter-layer prediction from the independent layer refers to the first inter-layer reference picture.
[0103] Example 101: The method of any of the examples 93 to 98 further comprising: receiving a third input picture; and encoding the third input picture into a third coded picture, wherein the third coded picture is the coded drive picture, and wherein the third coded picture resides in the independent layer, and the third coded picture follows the first coded picture in decoding order and precedes the second coded picture in the decoding order.
[0104] Example 102: The method of example 101 further comprising: in response to the third coded picture residing in the second access unit, omitting indicating that the third coded picture is the coded drive picture and infers that the third coded picture is the coded drive picture.
[0105] Example 103: The method of example 101 further comprising: indicating in or along the third coded picture an indication indicating that the third coded picture is the coded drive picture.
[0106] Example 104: The method of example 101 further comprising: indicating using a picture header extra bit that the third coded picture is the coded drive picture.
[0107] Example 105: The method of any of the examples 101 to 104 further comprising: reconstructing the third decoded picture from the third coded picture; and using the third decoded picture in the invocation of the generative neural network inference as an additional input.
[0108] Example 106: The method of any of the examples 101 to 105 further comprising: inferring when the inter-layer prediction from the independent layer refers to the first inter-layer reference picture or the third decoded picture.
[0109] Example 107: The method of any of the examples 101 to 105 further comprising: indicating in a reference picture list syntax structure information indicating that the first coded picture is used as input for generating the inter-layer reference picture.
[0110] Example 108: The method of example 107, wherein the information indicating that the first coded picture is used as input for generating the inter-layer reference picture comprises a picture order count difference between a current picture and the first coded picture.
[0111] Example 109: An apparatus comprising means for performing the methods as described in any of the examples 55 to 108.
[0112] Example 110: A computer readable medium comprising program instructions which, when executed by an apparatus, cause the apparatus to perform the methods as described in any of the examples 55 to 108.
[0113] Example 111: The computer readable medium of example 110, wherein the computer readable medium comprises a non-transitory computer readable medium.BRIEF DESCRIPTION OF THE DRAWINGS
[0114] The foregoing embodiments and other features are explained in the following description, taken in connection with the accompanying drawings, wherein:
[0115] FIG. 1 shows schematically an electronic device employing embodiments of the examples described herein.
[0116] FIG. 2 shows schematically a user equipment suitable for employing embodiments of the examples described herein.
[0117] FIG. 3 further shows schematically electronic devices employing embodiments of the examples described herein connected using wireless and wired network connections.
[0118] FIG. 4 shows a block diagram of a general structure of a video encoder.
[0119] FIG. 5 illustrates a picture unit.
[0120] FIG. 6 illustrates a generic block diagram of a generative face video codec.
[0121] FIG. 7 illustrates a picture unit generated by an encoder.
[0122] FIG. 8 illustrates of a bitstream generated by an encoder, in accordance with an embodiment.
[0123] FIG. 9 illustrates a bitstream generated by an encoder, in accordance with another embodiment.
[0124] FIG. 10 illustrates of an example bitstream generated by an encoder, in accordance with yet another embodiment.
[0125] FIG. 11 is an example apparatus, which may be implemented in hardware, and is caused to, implement examples described herein.
[0126] FIG. 12 shows a representation of an example of non-volatile memory media used to store instructions that implement the examples described herein.
[0127] FIG. 13 is an example method to implement the embodiments described herein, in accordance with an embodiment.
[0128] FIG. 14 is another example method to implement the embodiments described herein, in accordance with an embodiment.
[0129] FIG. 15 is still another example method to implement the embodiments described herein, in accordance with an embodiment.
[0130] FIG. 16 is still another example method to implement the embodiments described herein, in accordance with an embodiment.
[0131] FIG. 17 is still another example method to implement the embodiments described herein, in accordance with an embodiment.
[0132] FIG. 18 is still another example method to implement the embodiments described herein, in accordance with an embodiment.
[0133] FIG. 19 is still another example method to implement the embodiments described herein, in accordance with an embodiment.
[0134] FIG. 20 is still another example method to implement the embodiments described herein, in accordance with an embodiment.
[0135] FIG. 21 is still another example method to implement the embodiments described herein, in accordance with an embodiment.
[0136] FIG. 22 is still another example method to implement the embodiments described herein, in accordance with an embodiment.DETAILED DESCRIPTION OF EXAMPLE EMBODIMENTS
[0137] The following acronyms and abbreviations that may be found in the specification and / or the drawing figures are defined as follows:4CC four character code5G fifth generation cellular network technology5GC 5G core network a.k.a. also known asAVC advanced video codingCU central unitDSP digital signal processorDU distributed unit eNB (or eNodeB) evolved Node B (for example, an LTE base station)EN-DC E-UTRA-NR dual connectivity en-gNB or En-gNB node providing NR user plane and control plane protocol terminations towards the UE, and acting as secondary node in EN-DCE-UTRA evolved universal terrestrial radio access, for example, the LTE radio access technologyFl or Fl-C interface between CU and DU control interfacegNB (or gNodeB) base station for 5G / NR, for example, a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GCIEC International Electrotechnical Commission loT internet of thingsISO International Organization for StandardizationISOBMFF ISO base media file formatJPEG joint photographic experts groupLTE long-term evolution mdat MediaDataBoxMIME Multipurpose Internet Mail ExtensionMME mobility management entity moov MovieBoxMP4 file format for MPEG-4 Part 14 filesMPEG moving picture experts groupMPEG-2 H.222 / H.262 as defined by the ITUMPEG-4 audio and video coding standard for ISO / IEC 14496 ng or NG new generation ng-eNB or NG-eNB new generation eNBNR new radio (5G radio)N / W or NW networkPDCP packet data convergence protocolPHY physical layerPNG portable network graphicsRAN radio access networkRFC request for commentsRLC radio link controlRRC radio resource controlRRH remote radio headRU radio unitRx receiverSDAP service data adaptation protocolSGW serving gatewaySMF session management functionSPS sequence parameter setSVC scalable video codingSI interface between eNodeBs and the EPC trak TrackBoxTx transmitterUE user equipmentUICC Universal Integrated Circuit CardUPF user plane functionURL uniform resource locatorX2 interconnecting interface between two eNodeBs in LTE networkXn interface between two NG-RAN nodes
[0138] Some embodiments will now be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all, embodiments are shown. Indeed, various embodiments of the invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Like reference numerals refer to like elements throughout. As used herein, the terms ‘data,’ ‘content,’ ‘information,’ and similar terms may be used interchangeably to refer to data capable of being transmitted, received and / or stored in accordance with embodiments of the present invention. Thus, use of any such terms should not be taken to limit the spirit and scope of embodiments.
[0139] Additionally, as used herein, the term ‘circuitry’ refers to (a) hardware-only circuit implementations (e.g., implementations in analog circuitry and / or digital circuitry); (b) combinations of circuits and computer program product(s) comprising software and / or firmware instructions stored on one or more computer readable memories that work together to cause an apparatus to perform one or more functions described herein; and (c) circuits, such as, for example, a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation even when the software or firmware is not physically present. This definition of ‘circuitry’ applies to all uses of this term herein, including in any claims. As a further example, as used herein, the term ‘circuitry’ also includes an implementation comprising one or more processors and / or portion(s) thereof and accompanying software and / or firmware. As another example, the term ‘circuitry’ as used herein also includes, for example, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, other network device, and / or other computing device.
[0140] As defined herein, a ‘computer-readable storage medium,’ which refers to a non-transitory physical storage medium (e.g., volatile or non-volatile memory device), can be differentiated from a ‘computer-readable transmission medium,’ which refers to an electromagnetic signal.
[0141] A method, apparatus and computer program product are provided in accordance with example embodiments for integrating generative neural network in multimedia coding.
[0142] In an example, the following describes in detail suitable apparatus and possible mechanisms for integrating generative neural network in multimedia coding. In this regard reference is first made to FIG. 1 and FIG. 2, where FIG. 1 shows an example block diagram of an apparatus 50. The apparatus may be an internet of things (loT) apparatus configured to perform various functions, for example, gathering information by one or more sensors, receiving or transmitting information, analyzing information gathered or received by the apparatus, or the like. The apparatus may comprise a video coding system, which may incorporate a codec. FIG. 2 shows a layout of an apparatus according to an example embodiment. The elements of FIG. 1 and FIG. 2 will be explained next.
[0143] The apparatus 50, may for example be, a mobile terminal or user equipment of a wireless communication system, a sensor device, a tag, or a lower power device. However, it would be appreciated that embodiments of the examples described herein may be implemented within any electronic device or apparatus which may process data by neural networks.
[0144] The apparatus 50 may comprise a housing 30 for incorporating and protecting the device. The apparatus 50 may further comprise a display 32, for example, in the form of a liquid crystal display, light emitting diode display, organic light emitting diode display, and the like. In other embodiments of the examples described herein the display may be any suitable display technology suitable to display media or multimedia content, for example, an image or a video. The apparatus 50 may further comprise a keypad 34. In other embodiments of the examples described herein any suitable data or user interface mechanism may be employed. For example, the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.
[0145] The apparatus may comprise a microphone 36 or any suitable audio input which may be a digital or analogue signal input. The apparatus 50 may further comprise an audio output device which in embodiments of the examples described herein may be any one of: an earpiece 38, speaker, or an analogue audio or digital audio output connection. The apparatus 50 may also comprise a battery (or in other embodiments of the examples described herein the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator). The apparatus may further comprise a camera 42 capable of recording or capturing images and / or video. The apparatus 50 may further comprise an infrared port for short range line of sight communication to other devices. In other embodiments the apparatus 50 may further comprise any suitable short range communication solution such as for example a Bluetooth wireless connection or a USB / firewire wired connection.
[0146] The apparatus 50 may comprise a controller 56, a processor or a processor circuitry for controlling the apparatus 50. The controller 56 may be connected to a memory 58 which in embodiments of the examples described herein may store both data in the form of an image, audio data and video data, and / or may also store instructions for implementation on the controller 56. The controller 56 may further be connected to codec circuitry 54 suitable for carrying out coding and / or decoding of audio, image and / or video data or assisting in coding and / or decoding carried out by the controller.
[0147] The apparatus 50 may further comprise a card reader 48 and a smart card 46, for example, a universal integrated circuit card (UICC) and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network.
[0148] The apparatus 50 may comprise radio interface circuitry 52 connected to the controller and suitable for generating wireless communication signals, for example, for communication with a cellular communications network, a wireless communications system or a wireless local area network. The apparatus 50 may further comprise an antenna 44 connected to the radio interface circuitry 52 for transmitting radio frequency signals generated at the radio interface circuitry 52 to other apparatus(es) and / or for receiving radio frequency signals from other apparatus(es).
[0149] The apparatus 50 may comprise a camera 42 capable of recording or detecting individual frames which are then passed to the codec circuitry 54 or the controller for processing. The apparatus may receive the video image data for processing from another device prior to transmission and / or storage. The apparatus 50 may also receive either wirelessly or by a wired connection the image for coding / decoding. The structural elements of apparatus 50 described above represent examples of means for performing a corresponding function.
[0150] With respect to FIG. 3, an example of a system within which embodiments of the examples described herein can be utilized is shown. The system 10 comprises multiple communication devices which can communicate through one or more networks. The system 10 may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network (such as a GSM, UMTS, CDMA, LTE, 4G, 5G network etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a Bluetooth personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and the Internet.
[0151] The system 10 may include both wired and wireless communication devices and / or apparatus 50 suitable for implementing embodiments of the examples described herein.
[0152] For example, the system shown in FIG. 3 shows a mobile telephone network 11 and a representation of the internet 28. Connectivity to the internet 28 may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.
[0153] The example communication devices shown in the system 10 may include, but are not limited to, an electronic device or apparatus 50, a combination of a personal digital assistant (PDA) and a mobile telephone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, a notebook computer 22, or a head-mounted apparatus 21, which head-mounted apparatus 21 may be a head-mounted display (HMD), or glasses having a camera or other device used for processing images and / or video. The apparatus 50 may be stationary or mobile when carried by an individual who is moving. The apparatus 50 may also be located in a mode of transport including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle or any similar suitable mode of transport.
[0154] The embodiments may also be implemented in a set-top box; e.g. a digital TV receiver, which may / may not have a display or wireless capabilities, in tablets or (laptop) personal computers (PC), which have hardware and / or software to process neural network data, in various operating systems, and in chipsets, processors, DSPs and / or embedded systems offering hardware / software based coding.
[0155] Some or further apparatus may send and receive calls and messages and communicate with service providers through a wireless connection 25 to a base station 24. The base station 24 may be connected to a network server 26 that allows communication between the mobile telephone network 11 and the internet 28. The system may include additional communication devices and communication devices of various types.
[0156] The communication devices may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocolinternet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11, 3GPP Narrowband loT and any similar wireless communication technology. A communications device involved in implementing various embodiments of the examples described herein may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.
[0157] In telecommunications and data networks, a channel may refer either to a physical channel or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire, whereas a logical channel may refer to a logical connection over a multiplexed medium, capable of conveying several logical channels. A channel may be used for conveying an information signal, for example a bitstream, from one or several senders (or transmitters) to one or several receivers.
[0158] The embodiments may also be implemented in so-called loT devices. The Internet of Things (loT) may be defined, for example, as an interconnection of uniquely identifiable embedded computing devices within the existing Internet infrastructure. The convergence of various technologies has and may enable many fields of embedded systems, such as wireless sensor networks, control systems, home / building automation, etc. to be included in the Internet of Things (loT). In order to utilize the Internet loT devices are provided with an IP address as a unique identifier. loT devices may be provided with a radio transmitter, such as a WLAN or Bluetooth transmitter or a RFID tag. Alternatively, loT devices may have access to an IP-based network via a wired network, such as an Ethernet-based network or a power-line connection (PLC).
[0159] FIG. 4 shows a block diagram of a general structure of a video encoder. FIG. 4 presents an encoder for two layers, but it would be appreciated that presented encoder could be similarly extended to encode more than two layers. FIG. 4 illustrates a video encoder comprising a first encoder section 401 for a base layer and a second encoder section 451 for an enhancement layer. Each of the first encoder section 401 and the second encoder section 451 may comprise similar elements for encoding incoming pictures. The encoder sections 401, 451 may comprise a pixel predictor 402, 452, prediction error encoder 403, 453 and prediction error decoder 404, 454. FIG. 4 also shows an embodiment of the pixel predictor 402, 452 as comprising an inter-predictor 406, 456, an intra-predictor 408, 458, a mode selector 410, 460, a filter 416, 466, and a reference frame memory 418, 468. The pixel predictor 402 of the first encoder section 401 receives base layer picture(s) / image(s) 400 of a video stream to be encoded at both the inter-predictor 406 (which determines the difference between the image and a motion compensated reference frame) and the intra-predictor 408 (which determines a prediction for an image block based only on the already processed parts of current frame or picture). The output of both the inter-predictor and the intra-predictor are passed to the mode selector 410. The intra-predictor 408 may have more than one intra-prediction modes. Hence, each mode may perform the intra-prediction and provide the predicted signal to the mode selector 410. The mode selector 410 also receives a copy of the base layer image(s) 400. Correspondingly, the pixel predictor 452 of the second encoder section 451 receives enhancement layer picture(s) / images(s) 450 of a video stream to be encoded at both the interpredictor 456 (which determines the difference between the image and a motion compensated reference frame) and the intra-predictor 458 (which determines a prediction for an image block based only on thealready processed parts of current frame or picture). The output of both the inter-predictor and the intrapredictor are passed to the mode selector 460. The intra-predictor 458 may have more than one intraprediction modes. Hence, each mode may perform the intra-prediction and provide the predicted signal to the mode selector 460. The mode selector 460 also receives a copy of the enhancement layer pictures 450.
[0160] Depending on which encoding mode is selected to encode the current block, the output of the inter-predictor 406, 456 or the output of one of the optional intra-predictor modes or the output of a surface encoder within the mode selector is passed to the output of the mode selector 410, 460. The output of the mode selector 410, 460 is passed to a first summing device 421, 471. The first summing device may subtract the output of the pixel predictor 402, 452 from the base layer image(s) 400 / enhancement layer image(s) 450 to produce a first prediction error signal 420, 470 which is input to the prediction error encoder 403, 453.
[0161] The pixel predictor 402, 452 further receives from a preliminary reconstructor 439, 489 the combination of the prediction representation of the image block 412, 462 and the output 438, 488 of the prediction error decoder 404, 454. The preliminary reconstructed image 414, 464 may be passed to the intra-predictor 408, 458 and to the filter 416, 466. The filter 416, 466 receiving the preliminary representation may filter the preliminary representation and output a final reconstructed image 440, 490 which may be saved in the reference frame memory 418, 468. The reference frame memory 418 may be connected to the inter-predictor 406 to be used as the reference image against which a future base layer image 400 is compared in inter-prediction operations. Subject to the base layer being selected and indicated to be source for inter-layer sample prediction and / or inter-layer motion information prediction of the enhancement layer according to some embodiments, the reference frame memory 418 may also be connected to the inter-predictor 456 to be used as the reference image against which a future enhancement layer image(s) 450 is compared in inter-prediction operations. Moreover, the reference frame memory 468 may be connected to the inter-predictor 456 to be used as the reference image against which the future enhancement layer image(s) 450 is compared in inter -prediction operations.
[0162] Filtering parameters from the filter 416 of the first encoder section 401 may be provided to the second encoder section 451 subject to the base layer being selected and indicated to be source for predicting the filtering parameters of the enhancement layer according to some embodiments.
[0163] The prediction error encoder 403, 453 comprises a transform unit 442, 492 and a quantizer 444, 494. The transform unit 442, 492 transforms the first prediction error signal 420, 470 to a transform domain. The transform is, for example, the DCT transform. The quantizer 444, 494 quantizes the transform domain signal, for example, the DCT coefficients, to form quantized coefficients.
[0164] The prediction error decoder 404, 454 receives the output from the prediction error encoder 403, 453 and performs the opposite processes of the prediction error encoder 403, 453 to produce a decoded prediction error signal 438, 488 which, when combined with the prediction representation of the image block 412, 462 at the second summing device 439, 489, produces the preliminary reconstructed image 414, 464. The prediction error decoder may be considered to comprise a dequantizer 446, 496, which dequantizes the quantized coefficient values, for example, DCT coefficients, to reconstruct the transform signal and an inverse transformation unit 448, 498, which performs the inverse transformation to the reconstructed transform signal wherein the output of the inverse transformation unit 448, 498 includes reconstructed block(s). The prediction error decoder may also comprise a block filter which may filter the reconstructed block(s) according to further decoded information and filter parameters.
[0165] The entropy encoder 430, 480 receives the output of the prediction error encoder 403, 453 and may perform a suitable entropy encoding / variable length encoding on the signal to provide a compressed signal. The outputs of the entropy encoders 430, 480 may be inserted into a bitstream, for example, by a multiplexer 465.
[0166] Some video coding and video metadata specifications
[0167] The Advanced Video Coding standard (which may be abbreviated H.264, AVC or H.264 / AVC) was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunications Standardization Sector of International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). The H.264 / AVC standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, also known as MPEG-4 Part 10 Advanced Video Coding (AVC). There have been multiple versions of the H.264 / AVC standard, each integrating new extensions or features to the specification. These extensions include Scalable Video Coding (SVC) and Multiview Video Coding (MVC).
[0168] The High Efficiency Video Coding standard (which may be abbreviated H.265, HEVC or H.265 / HEVC) was developed by the Joint Collaborative Team - Video Coding (JCT-VC) of VCEG and MPEG. The standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2, also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC). Extensions to H.265 / HEVC include scalable, multiview, three-dimensional, and fidelity range extensions, which may be referred to as SHVC, MV-HEVC, 3D-HEVC, and REXT, respectively. The references in this description to H.265 / HEVC, SHVC, MV- HEVC, 3D-HEVC and REXT that have been made for the purpose of understanding definitions, structures or concepts of these standard specifications are to be understood to be references to the latest versions of these standards that were available before the date of this application, unless otherwise indicated.
[0169] Versatile Video Coding (which may be abbreviated VVC, H.266, or H.266 / VVC) is a video compression standard developed as the successor to HEVC. VVC is specified in ITU-T Recommendation H.266 and equivalently in ISO / IEC 23090-3, which is also referred to as MPEG-I Part 3.
[0170] A specification of the AVI bitstream format and decoding process were developed by the Alliance for Open Media (AOM). The AVI specification was published in 2018. AOM is reportedly working on the AV2 specification.
[0171] ITU-T Recommendation H.274, which is equivalent to ISO / IEC 23002-7, may be called "versatile supplemental enhancement information messages for coded video bitstreams" and be referred to as "versatile supplemental enhancement information" or VSEI. The VSEI standard specifies the syntax and semantics of video usability information (VUI) parameters and supplemental enhancement information (SEI) messages. The VUI parameters and SEI messages defined in the VSEI standard are designed to be conveyed within coded video bitstreams in a manner specified in a video coding specification or to be conveyed by other means determined by the specifications for systems that make use of such coded video bitstreams. The VSEI standard is intended for use with VVC coded video bitstreams, although it is drafted in a manner intended to be sufficiently generic that it may also be used with other types of coded video bitstreams. VUI parameters and SEI messages may, for example, assist in processes related to decoding, display or other purposes.
[0172] Video coding
[0173] An elementary unit for the input to an encoder and the output of a decoder, respectively, in most cases is a picture. A picture given as an input to an encoder may also be referred to as a source picture, and a picture decoded by a decoded may be referred to as a decoded picture or a reconstructed picture.
[0174] The source and decoded pictures are each comprised of one or more sample arrays, such as one of the following sets of sample arrays:- Luma (Y) only (monochrome).- Luma and two chroma (YCbCr or YCgCo).- Green, Blue and Red (GBR, also known as RGB).- Arrays representing other unspecified monochrome or tri-stimulus color samplings (for example, YZX, also known as XYZ).
[0175] In the following, these arrays may be referred to as luma (or L or Y) and chroma, where the two chroma arrays may be referred to as Cb and Cr or Cg and Co; regardless of the actual color representation method in use. The actual color representation method in use may be indicated, e.g., in a coded bitstream e.g., using the Video Usability Information (VUI) syntax. A component may be defined as an array or single sample from one of the three sample arrays (luma and two chroma) or the array or a single sample of the array that compose a picture in monochrome format.
[0176] A picture may be defined to be either a frame or a field. A frame comprises a matrix of luma samples and possibly the corresponding chroma samples. A field is a set of alternate sample rows of a frame and may be used as encoder input, when the source signal is interlaced. Chroma sample arrays may be absent (and hence monochrome sampling may be in use) or chroma sample arrays may be subsampled when compared to luma sample arrays.
[0177] Some chroma formats may be summarized as follows:- In monochrome sampling there is only one sample array, which may be nominally considered the luma array.- In 4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array.- In 4:2:2 sampling, each of the two chroma arrays has the same height and half the width of the luma array.- In 4:4:4 sampling when no separate color planes are in use, each of the two chroma arrays has the same height and width as the luma array.
[0178] Coding formats or standards may allow to code sample arrays as separate color planes into the bitstream and respectively decode separately coded color planes from the bitstream. When separate color planes are in use, each one of them is separately processed (by the encoder and / or the decoder) as a picture with monochrome sampling.
[0179] A video codec may comprise an encoder that transforms the input video into a compressed representation suited for storage / transmission and a decoder that can uncompress the compressed video representation back into a viewable form. The compressed representation may be referred to as a bitstream or a video bitstream. A video encoder and / or a video decoder may also be separate from eachother, i.e. need not form a codec. The encoder may discard some information in the original video sequence in order to represent the video in a more compact form (that is, at lower bitrate).
[0180] Video encoders may encode the video information in two phases. At first, pixel values in a certain picture area (or “block”) are predicted for example by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Then, the prediction error, i.e. the difference between the predicted block of pixels and the original block of pixels, is coded. This may be done by transforming the difference in pixel values using a specified transform (e.g. Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate).
[0181] In inter prediction, the sources of prediction are previously decoded pictures (a.k.a. reference pictures). Temporal inter prediction (a.k.a. temporal prediction) may refer to inter prediction where the reference picture is a previous picture in decoding order within the same layer as the current picture being encoded or decoded.
[0182] In intra block copy (IBC; a.k.a. intra-block-copy prediction or current picture referencing), prediction may be applied similarly to inter prediction, but the reference picture is the current picture and only previously decoded samples can be referred in the prediction process.
[0183] Inter-layer or inter-view prediction may be applied similarly to temporal inter prediction, but the reference picture is a decoded picture from another scalable layer or from another view, respectively.
[0184] In some cases, inter prediction may refer to temporal inter prediction only, while in other cases inter prediction may refer collectively to temporal inter prediction and any of intra block copy, inter-layer prediction, and inter-view prediction provided that they are performed with the same or similar process than temporal inter prediction.
[0185] Inter prediction or temporal prediction may sometimes be referred to as motion compensation or motion-compensated prediction.
[0186] Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in spatial or transform domain, i.e., either sample valuesor transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.
[0187] One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy-coded more efficiently when they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors and only the difference relative to the motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.
[0188] A video decoder may reconstruct the output video by applying prediction means similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in spatial pixel domain). After applying prediction and prediction error decoding means the decoder sums up the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) can also apply additional filtering means to improve the quality of the output video before passing it for display and / or storing it as prediction reference for the forthcoming frames in the video sequence.
[0189] A partitioning may be defined as a division of a set into subsets such that each element of the set is in exactly one of the subsets.
[0190] In some video codecs, a picture may be divided into coding units (CU) covering the area of the picture. A CU includes of one or more prediction units (PU) defining the prediction process for the samples within the CU and one or more transform units (TU) defining the prediction error coding process for the samples in the said CU. The CU may include of a rectangular block of samples with a size selectable from a predefined set of possible CU sizes. A CU with the maximum allowed size may be named as LCU (largest coding unit) or coding tree unit (CTU) and the video picture is divided into non-overlapping LCUs. An LCU can be further split into a combination of smaller CUs, e.g. by recursively splitting the LCU and resultant CUs. Each resulting CU may have at least one PU and at least one TU associated with it. Each PU and TU can be further split into smaller PUs and TUs in order to increase granularity of the prediction and prediction error coding processes, respectively. Each PU has prediction information associated with it defining what kind of a prediction is to be applied for the pixels within that PU (e.g. motion vector information for inter predicted PUs and intra prediction directionality information for intra predicted PUs).
[0191] Each TU may be associated with information describing the prediction error decoding process for the samples within the said TU (including e.g. DCT coefficient information). It may be signalled at CU level whether prediction error coding is applied or not for each CU. In the case there is no prediction error residual associated with the CU, it can be considered there are no TUs for the said CU. The division of the image into CUs, and division of CUs into PUs and TUs may be signalled in the bitstream allowing the decoder to reproduce the intended structure of these units.
[0192] In some coding formats, such as VVC, the samples are processed in units of coding tree blocks (CTB). In some coding formats, an encoder may select the size of a CTB. The array size for each luma CTB in both width and height may be denoted CtbSizeY in units of samples. A VVC encoder may select CtbSizeY on a sequence basis from values supported in the VVC standard (32, 64, 128), or the VVC encoder may be configured to use a certain CtbSizeY value. The width and height of the array for each chroma CTB may be denoted CtbWidthC and CtbHeightC, respectively, in units of samples.
[0193] In some coding formats, each CTB may be assigned a partition signalling to identify the block sizes for intra or inter prediction and for transform coding. The partitioning is a recursive quadtree partitioning. The root of the quadtree is associated with the CTB. The quadtree is split until a leaf is reached, which is referred to as the quadtree leaf. When the component width is not an integer number of the CTB size, the CTBs at the right component boundary are incomplete. When the component height is not an integer multiple of the CTB size, the CTBs at the bottom component boundary are incomplete.
[0194] In some coding formats, the coding block is the root node of two trees, the prediction tree and the transform tree. The prediction tree specifies the position and size of prediction blocks. The transform tree specifies the position and size of transform blocks. The splitting information for luma and chroma is identical for the prediction tree and may or may not be identical for the transform tree.
[0195] In some coding formats, the blocks and associated syntax structures may be grouped into "unit" structures as follows:- One transform block (monochrome picture) or three transform blocks (luma and chroma components of a picture in 4:2:0, 4:2:2 or 4:4:4 colour format) and the associated transform syntax structures units are associated with a transform unit.- One coding block (monochrome picture) or three coding blocks (luma and chroma), the associated coding syntax structures and the associated transform units are associated with a coding unit.- One CTB (monochrome picture) or three CTBs (luma and chroma), the associated coding tree syntax structures and the associated coding units are associated with a CTU.
[0196] A superblock in AVI is similar to a CTU in VVC. A superblock may be regarded as the largest coding block that the AV 1 specification supports. The size of the superblock is signalled in the sequence header to be 128 x 128 or 64 x 64 luma samples. A superblock may be partitioned into smaller coding blocks recursively. A coding block may have its own prediction and transform modes, independent of those of the other coding blocks.
[0197] In some coding modes of some video coding formats, the reference picture for inter prediction is indicated with an index to a reference picture list. The index may be coded with variable length coding, which usually causes a smaller index to have a shorter value for the corresponding syntax element. In some coding formats, two reference picture lists (reference picture list 0 and reference picture list 1) are generated for each bi-predictive (B) slice, and one reference picture list (reference picture list 0) is formed for each uni-predicted (P) slice.
[0198] In some coding formats, a set of reference picture lists may be present in a parameter set, such as a sequence parameter set, and the reference picture list(s) in use may be indicated e.g. in a slice header or a picture header. When none of the reference picture lists in a parameter set are considered suitable by an encoder for a present picture or slice, it may be possible to include a reference picture list in a respective syntax structure, such as a picture header or a slice header, directly.
[0199] An out-of-band transmission, signaling, or storage may refer to the capability of transmitting, signaling, or storing information in a manner that associates the information with a video bitstream. The out-of-band transmission may use a more reliable transmission mechanism compared to the protocols used for carrying coded video data, such as slices. The out-of-band transmission, signaling or storage may additionally or alternatively be used, e.g., for ease of access or session negotiation. For example, a sample entry of a track in a file conforming to the ISO Base Media File Format may comprise parameter sets, while the coded data in the bitstream is stored elsewhere in the file or in another file. Another example of out-of-band transmission, signaling, or storage comprises including information, such as NN and / or NN updates in a file format track that is separate from track(s) including coded video data.
[0200] The phrase along the bitstream (e.g., indicating along the bitstream) or along a coded unit of a bitstream (e.g., indicating along a coded tile) may be used in claims and described embodiments to refer to transmission, signaling, or storage in a manner that the ‘out-of-band’ data is associated with, but not included within, the bitstream or the coded unit, respectively. The phrase decoding along the bitstream or along a coded unit of a bitstream or alike may refer to decoding the referred out-of-band data (which may be obtained from out-of-band transmission, signaling, or storage) that is associated with the bitstream or the coded unit, respectively. For example, the phrase along the bitstream may beused when the bitstream is included in a container file, such as a file conforming to the ISO Base Media File Format, and certain file metadata is stored in the file in a manner that associates the metadata to the bitstream, such as boxes in the sample entry for a track including the bitstream, a sample group for the track including the bitstream, or a timed metadata track associated with the track including the bitstream. In another example, the phrase along the bitstream may be used when the bitstream is made available as a stream over a communication protocol and a media description, such as a streaming manifest, is provided to describe the stream.
[0201] A bitstream may be defined as a sequence of bits or a sequence of syntax structures. A bitstream format may constrain the order of syntax structures in the bitstream.
[0202] A syntax element may be defined as an element of data represented in a bitstream. A syntax structure may be defined as zero or more syntax elements present together in a bitstream in a specified order.
[0203] Syntax structures may be specified, for example, using arithmetic, logical, relational, bitwise, and assignment operators similar to those available in many programming languages. For example, & may indicate a bit-wise ‘AND’ operation. Furthermore, syntax structures may be specified with reference to mathematical functions.
[0204] Syntax structures and semantics may use the values of variables derived from the values of syntax elements. Naming conventions may be defined for variables. For example, variables may be named by a mixture of lower case and upper case letter and without any underscore characters. V ariables starting with an upper case letter may be derived for the decoding of the current syntax structure and all depending syntax structures. Variables starting with an upper case letter may, in some cases, be used in the decoding process for later syntax structures without mentioning the originating syntax structure of the variable. Variables starting with a lower case letter may only be used in relation to the syntax structure or function they have been defined for.
[0205] An elementary unit for the output of a video encoder and the input of a video decoder, respectively, may be a network abstraction layer (NAL) unit. For transport over packet-oriented networks or storage into structured files, NAL units may be encapsulated into packets or similar structures. A bytestream format encapsulating NAL units may be used for transmission or storage environments that do not provide framing structures. The bytestream format may separate NAL units from each other by attaching a start code in front of each NAL unit. To avoid false detection of NAL unit boundaries, encoders may run a byte-oriented start code emulation prevention algorithm, which may add an emulation prevention byte to the NAL unit payload, when a start code would have occurredotherwise. In order to enable straightforward gateway operation between packet and stream-oriented systems, start code emulation prevention may be performed regardless of whether the bytestream format is in use or not. A NAL unit may be defined as a syntax structure including an indication of the type of data to follow and bytes including that data in the form of a raw byte sequence payload interspersed as necessary with emulation prevention bytes. A raw byte sequence payload (RBSP) may be defined as a syntax structure including an integer number of bytes that is encapsulated in a NAL unit. An RBSP is either empty or has the form of a string of data bits including syntax elements followed by an RBSP stop bit and followed by zero or more subsequent bits equal to 0.
[0206] A bitstream may be defined to logically include a syntax structure, such as a NAL unit, when the syntax structure is transmitted along the bitstream but may be included in the bitstream according to the bitstream format. A bitstream may be defined to natively comprise a syntax structure, when the bitstream includes the syntax structure.
[0207] In some coding formats or standards, a bitstream may be in the form of a network abstraction layer (NAL) unit stream or a byte stream, that forms the representation of coded pictures and associated data forming one or more coded video sequences.
[0208] In some formats or standards, a first bitstream may be followed by a second bitstream in the same logical channel, such as in the same file or in the same connection of a communication protocol. An elementary stream (in the context of video coding) may be defined as a sequence of one or more bitstreams.
[0209] In some coding formats, such as AVI, a bitstream may comprise a sequence of open bitstream units (OBUs). An OBU comprises a header and a payload, wherein the header identifies a type of the OBU. Furthermore, an OBU may include the OBU size, which may follow the OBU header and precede the OBU payload within the OBU, wherein the OBU size may be indicative of a size of the pay load in bytes.
[0210] A bitstream of OBUs may be formatted as a so-called low-overhead bitstream format, which comprises a sequence of OBUs, wherein the OBU header indicates the OBU size (without the size field) in bytes. Alternatively, a bitstream of OBUs may be formatted as a so-called length delimited bitstream format, in which each temporal unit starts with a size field (temporal_unit_size) indicating the size of the temporal unit payload in bytes, and each frame unit starts with a size field (frame_unit_size) indicating the size of the frame unit payload in bytes, and each OBU is preceded by a size field indicating the size of the OBU in bytes.
[0211] The OBU types may include the following.- Sequence Header includes information that applies to the entire sequence and whether to enable certain coding tools.- Temporal Delimiter indicates the frame presentation time stamp. All displayable frames following a temporal delimiter OBU will use this time stamp, until the next temporal delimiter OBU arrives. A temporal delimiter and its subsequent OBUs of the same time stamp are referred to as a temporal unit. In the context of scalable coding, the compression data associated with all representations of a frame at various spatial and fidelity resolutions will be in the same temporal unit.- Frame Header sets up the coding information for a given frame, including signaling inter or intraframe type, indicating the reference frames and signaling probability model update method.- Tile Group includes the tile data associated with a frame. Each tile can be independently decoded. The collective reconstructions form the reconstructed frame after potential loop filtering.- Frame includes the frame header and tile data. The frame OBU is largely equivalent to a frame header OBU and a tile group OBU but allows less overhead cost.- Metadata carries information, such as high dynamic range, scalability, and timecode.- Tile List includes tile data similar to a tile group OBU. However, each tile here has an additional header that indicates its reference frame index and position in the current frame. This allows the decoder to process a subset of tiles and display the corresponding part of the frame, without the need to fully decode all the tiles in the frame.
[0012] In some coding standards, NAL units include a header and payload. The NAL unit header indicates the type of the NAL unit. In some coding standards, the NAL unit header indicates a scalability layer identifier (e.g., called nuh_layer_id in H.265 / HEVC and H.266 / VVC), which may be used, e.g., for indicating spatial or quality layers, views of a multiview video, or auxiliary layers (such as depth maps or alpha planes). In some coding standards, the NAL unit header includes a temporal sublayer identifier, which may be used for indicating temporal subsets of the bitstream, such as a 30-frames-per- second subset of a 60-frames-per-second bitstream.
[0213] Bitstreams or coded video sequences may be encoded to be temporally scalable as follows. Each picture may be assigned to a particular temporal sub-layer. A temporal sub-layer may be equivalently called a sub-layer, temporal sublayer, sublayer, or temporal level. Temporal sub-layers may be enumerated, e.g., from 0 upwards. The lowest temporal sub-layer, sub-layer 0, may be decoded independently. Pictures at temporal sub-layer 1 may be predicted from reconstructed pictures at temporal sub-layers 0 and 1. Pictures at temporal sub-layer 2 may be predicted from reconstructed pictures at temporal sub-layers 0, 1, and 2, and so on. In other words, a picture at temporal sub-layer N does not use any picture at temporal sub-layer greater than N as a reference for inter prediction. Thebitstream created by excluding all pictures greater than or equal to a selected sub-layer value and including pictures remains conforming.
[0214] Each picture of a temporally scalable bitstream may be assigned with a temporal identifier (also known as TID, temporal layer identifier, temporal sublayer identifier, or temporal layer ID), which may be, for example, assigned to a variable Temporalld. The temporal identifier may, for example, be indicated in a NAL unit header or in an OBU extension header. Temporalld equal to 0 corresponds to the lowest temporal level. The bitstream created by excluding all coded pictures having a Temporalld greater than or equal to a selected value and including all other coded pictures remains conforming. Consequently, a picture having Temporalld equal to tid_value does not use any picture having a Temporalld greater than tid_value as a prediction reference.
[0215] NAL units may be categorized into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units are typically coded slice NAL units.
[0216] A non-VCL NAL unit may be, for example, one of the following types: a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a supplemental enhancement information (SEI) NAL unit, an access unit delimiter, an end of sequence NAL unit, an end of bitstream NAL unit, or a filler data NAL unit. Parameter sets may be needed for the reconstruction of decoded pictures, whereas many of the other non-VCL NAL units are not necessary for the reconstruction of decoded sample values.
[0217] Video coding specifications may enable the use of supplemental enhancement information (SEI) messages or alike. Some video coding specifications include SEI NAL units, and some video coding specifications include both prefix SEI NAL units and suffix SEI NAL units, where the former type may start a picture unit or alike and the latter type may end a picture unit or alike. An SEI NAL unit may include one or more SEI messages, which are not required for the decoding of output pictures but may assist in related processes, such as picture output timing, post-processing of decoded pictures, rendering, error detection, error concealment, and resource reservation. Several SEI messages are specified in H.264 / AVC, H.265 / HEVC, H.266 / VVC, and H.274 / VSEI standards, and the user data SEI messages enable organizations and companies to specify SEI messages for their own use. The standards may include the syntax and semantics for the specified SEI messages but a process for handling the messages in the recipient may not be defined. Consequently, encoders may be required to follow the standard specifying a SEI message when they create SEI message(s), and decoders may not be required to process SEI messages for output order conformance. One of the example reasons to include the syntax and semantics of SEI messages in standards is to allow different system specifications to interpret the supplemental information identically and hence interoperate. It is intended that system specificationsmay require the use of particular SEI messages both in the encoding end and in the decoding end, and additionally the process for handling particular SEI messages in the recipient may be specified.
[0218] Some video coding specifications enable metadata OBUs. A metadata OBU comprises a type field, which specifies the type of metadata.
[0019] A coded video sequence (C V S) may be defined as a sequence of coded pictures in decoding order that is independently decodable and is followed by another coded video sequence or the end of the bitstream.
[0220] A coded layer video sequence (CLVS) may be defined as a sequence of pictures and associated other data within the same scalable layer (e.g., with the same value of nuh_layer_id in VVC) that is decodable independently of other pictures in the same layer.
[0221] Some codecs use a concept of picture order count (POC). A value of POC is derived for each picture and is non-decreasing with increasing picture position in output order. POC therefore indicates the output order of pictures. POC may be used in the decoding process for example for implicit scaling of motion vectors and for reference picture list initialization. Furthermore, POC may be used in the verification of output order conformance. The variable including a POC value of a picture may be referred to as PicOrderCntVal.
[0222] A Decoded Picture Buffer (DPB) may be used in the encoder and / or in the decoder. There may be two reasons to buffer decoded pictures, for references in inter prediction and for reordering decoded pictures into output order. Some coding formats, such as HEVC, provide a great deal of flexibility for both reference picture marking and output reordering, separate buffers for reference picture buffering and output picture buffering may waste memory resources. Hence, the DPB may include a unified decoded picture buffering process for reference pictures and output reordering. A decoded picture may be removed from the DPB when it is no longer used as a reference and is not needed for output.
[0223] Output order may be defined as the order in which the decoded pictures are output from the decoded picture buffer (for the decoded pictures that are to be output from the decoded picture buffer).
[0224] Output time may be defined as a time when a decoded picture is to be output from a decoder or from the DPB of a decoder (for the decoded pictures that are to be output from the DPB), for example as specified by a hypothetical reference decoder (HRD) specification according to theoutput timing DPB operation.
[0225] Pictures having the same output order may be defined to mean the same as pictures having the same output time.
[0226] Decoding order may be defined as the order in which syntax elements are processed by the decoding process. It may be required that syntax elements are ordered in a bitstream in their decoding order.
[0227] A decoder and / or an HRD may comprise a picture output process. The output process may be considered to be a process in which the decoder provides decoded and cropped pictures as the output of the decoding process. The output process is typically a part of video coding standards, typically as a part of the hypothetical reference decoder specification. In output cropping, lines and / or columns of samples may be removed from decoded pictures according to a cropping rectangle to form output pictures. A cropped decoded picture may be defined as the result of cropping a decoded picture based on the conformance cropping window specified e.g. in the sequence parameter set that is referred to by the corresponding coded picture.
[0228] In some video coding specifications, encoders may control whether a decoded (and cropped) picture is output by the picture output process or a similar decoder-side process. Encoders may include in or along a bitstream one or more syntax elements for controlling picture output. For example, bitstream syntax may include a flag (e.g., called pic_output_flag or ph_pic_output_flag) e.g. in a picture header and / or an image segment header (e.g. a slice header). The semantics of pic_output_flag may be specified in a manner that when pic_output_flag is equal to 0, the respective decoded picture is not output (by the picture output process or alike), and when pic_output_flag is equal to 1, the respective decoded picture is output unless otherwise concluded in the decoding process.
[0229] An identifier may be defined as a syntax element that identifies a syntax structure. A value of the identifier may for example differ in different instances of the same syntax structure, such as a parameter set. A particular instance of the syntax structure may be referenced through its identifier value. For example, a parameter set that is referenced by the (de)coding of a coded video slice may be identified by providing the identifier value of the parameter set in a header of the coded video slice.
[0230] An indicator (ide) may be defined as a syntax element whose value indicates a selection among more than two values (for which semantics have been specified). An indicator syntax element may have _idc postfix in its name.
[0231] Images may be split into independently codable and decodable image segments (e.g., slices or tiles or tile groups). Such image segments may enable parallel processing. Image segments may be coded as separate units in the bitstream, such as VCL NAL units in H.264 / AVC, HEVC, and VVC. Coded image segments may comprise a header and a payload, wherein the header includes parameter values needed for decoding the payload.
[0232] In some video coding formats, such as HEVC and VVC, a picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of coding tree units (CTU) that covers a rectangular region of a picture. The partitioning to tiles forms a grid that may be characterized by a list of tile column widths (in CTUs) and a list of tile row heights (in CTUs). For encoding and / or decoding, the CTUs in a tile are scanned in raster scan order within that tile. In HEVC, tiles are ordered in the bitstream consecutively in the raster scan order of the tile grid.
[0233] In some video coding formats, such as AVI, a picture may be partitioned into tiles, and a tile includes an integer number of complete superblocks that collectively form a complete rectangular region of a picture. In-picture prediction across tile boundaries may be disabled. The minimum tile size may be one superblock, and the maximum tile size in the presently specified levels in AV 1 is 4096 x 2304 in terms of luma sample count. The picture is partitioned into a tile grid of one or more tile rows and one or more tile columns. The tile grid may be signaled in the picture header to have a uniform tile size or nonuniform tile size, where in the latter case the tile row heights and tile column widths are signaled. The superblocks in a tile are scanned in raster scan order within that tile.
[0234] In some video coding formats, such as VVC, a slice includes an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture. Consequently, each vertical slice boundary is always also a vertical tile boundary. It is possible that a horizontal boundary of a slice is not a tile boundary but includes horizontal CTU boundaries within a tile; this occurs when a tile is split into multiple rectangular slices, each of which includes an integer number of consecutive complete CTU rows within the tile.
[0235] In some video coding formats, such as VVC, two modes of slices are supported, namely the raster-scan slice mode and the rectangular slice mode. In the raster-scan slice mode, a slice includes a sequence of complete tiles in a tile raster scan of a picture. In the rectangular slice mode, a slice includes either a number of complete tiles that collectively form a rectangular region of the picture or a number of consecutive complete CTU rows of one tile that collectively form a rder within the rectangular region corresponding to that slice.
[0236] In HEVC, a slice includes an integer number of CTUs. The CTUs are scanned in the raster scan order of CTUs within tiles or within a picture when tiles are not in use. A slice may include an integer number of tiles, or a slice can be included in a tile.
[0237] In some video coding formats, such as AVI, a tile group OBU carries one or more complete tiles. The first and last tiles of in the tile group OBU may be indicated in the tile group OBU before the coded tile data. Tiles within a tile group OBU may appear in a tile raster scan of a picture.
[0238] A coded picture may be defined as a coded representation of a picture. In some video coding formats, a coded picture may comprise VCL NAL units. In some coding formats, a coded picture may be defined as a coded representation of a picture including all coding tree units of the picture.
[0239] In some coding formats, an access unit (AU) may be defined as a set of NAL units that are associated with each other according to a specified classification rule, are consecutive in decoding order, and include at most one coded picture at any scalability layer (e.g., with any specific value of nuh_layer_id in some coding formats, such as HEVC or VVC). In some coding formats, an access unit comprises one or more complete picture units. In some coding formats, in addition to including the VCL NAL units of a coded picture, an access unit may also include non-VCL NAL units associated with the coded picture. Said specified classification rule may, for example, associate pictures with the same output time or picture order count value into the same access unit.
[0240] A picture unit may be defined as a set of NAL units that are associated with each other according to a specified classification rule, are consecutive in decoding order, and include exactly one coded picture. Some non-VCL NAL units, such as prefix SEI NAL units, are allowed to precede all VCL NAL units of the same picture unit in decoding order and may therefore start a picture unit but are not allowed to follow all VCL NAL units of the same picture unit in decoding order. Some non-VCL NAL units, such as suffix SEI NAL units, are allowed to follow all VCL NAL units of the same picture unit in decoding order but are not allowed to precede all VCL NAL units of the same picture unit in decoding order. A simplified structure of a picture unit is presented below, where decoding order of NAL units is from the top towards the bottom, dashed boxes indicate optional syntax structures or concepts, and solid boxes indicate mandatory syntax structures or concepts. FIG. 5 illustrates a picture unit 500. The picture unit 500 is required to include exactly one coded picture 502, and a coded picture is required to have at least one coded slice, e.g., 504-1 and 504-2.
[0241] In some coding formats, such as AVI, a coded video sequence comprises one or more temporal units. A temporal unit includes of a series of OBUs starting from a temporal delimiter, optional sequence headers, optional metadata OBUs, a sequence of one or more frame headers, each followedby zero or more tile group OBUs as well as optional padding OBUs. A temporal unit may be defined to comprise all the OBUs that are associated with a specific, distinct time instant. A temporal unit may comprise a temporal delimiter OBU, and all the OBUs that follow, up to but not including the next temporal delimiter. A temporal delimiter OBU may be defined as an indication that the following OBUs will have a different presentation / decoding time stamp from the one of the last frame prior to the temporal delimiter.
[0242] Video coding standards may specify profiles and levels. A profile may be regarded as a subset of algorithmic features of the standard. Alternatively, a profile may be defined as a specified subset of the syntax of a coding standard. A level may be defined as a set of limits to the coding parameters that impose a set of constraints in decoder resource consumption. Alternatively, a level may be defined as a defined set of constraints on the values that may be taken by the syntax elements and variables of a coding standard. The same set of levels may be defined for all profiles, with most aspects of the definition of each level being in common across different profiles, although aspects may also differ between profiles. The profile and level can be used to signal properties of a media stream, as well as to signal the capability of a media decoder. Each pair of profile and level may be considered to form an "interoperability point."
[0243] Through the combination of a profile and a level, a decoder can declare, without actually attempting the decoding process, whether it is capable of decoding a stream. When the decoder is not capable of decoding the stream, it may cause the decoder to crash, operate slower than real-time, and / or discard data due to buffer overflows.
[0244] A concept of tier has been specified and used in HEVC and may be similarly specified and used in other codecs. A tier may be defined as a specified category of level constraints imposed on values of the syntax elements in the bitstream, where the level constraints are nested within a tier and a decoder conforming to a certain tier and level would be capable of decoding all bitstreams that conform to the same tier or the lower tier of that level or any level below it.
[0245] Scalable video coding
[0246] Scalable video coding may refer to coding structure where one bitstream may include multiple representations of the content, for example, at different bitrates, resolutions or frame rates. In these cases the receiver can extract the desired representation depending on its characteristics (e.g., resolution that matches best the display device). Alternatively, a server or a network element may extract the portions of the bitstream to be transmitted to the receiver depending on, e.g., the network characteristics or processing capabilities of the receiver. A meaningful decoded representation may beproduced by decoding only certain parts of a scalable bitstream. A scalable bitstream typically include of a ‘base layer’ providing the lowest quality video available and one or more enhancement layers that enhance the video quality when received and decoded together with the lower layers. In order to improve coding efficiency for the enhancement layers, the coded representation of that layer typically depends on the lower layers. For example, the motion and mode information of the enhancement layer can be predicted from lower layers. Similarly, the pixel data of the lower layers can be used to create prediction for the enhancement layer.
[0247] A scalable bitstream may include a ‘base layer’ providing the lowest quality video available and one or more enhancement layers that enhance the video quality when received and decoded together with the lower layers. In order to improve coding efficiency for the enhancement layers, the coded representation of that layer may depend on the lower layers. E.g., the motion and mode information of the enhancement layer may be predicted from lower layers. Similarly, the pixel data of the lower layers can be used to create prediction for the enhancement layer.
[0248] A scalable video codec for quality scalability (also known as signal-to-noise or SNR) and / or spatial scalability may be implemented as follows. For a base layer, a conventional non-scalable video encoder and decoder is used. The reconstructed / decoded pictures of the base layer are included in the reference picture buffer for an enhancement layer. In H.264 / AVC, HEVC, and similar codecs using reference picture list(s) for inter prediction, the base layer decoded pictures may be inserted into a reference picture list(s) for coding / decoding of an enhancement layer picture similarly to the decoded reference pictures of the enhancement layer. Consequently, the encoder may choose a base-layer reference picture as inter prediction reference and indicate its use, e.g., with a reference picture index in the coded bitstream. The decoder decodes from the bitstream, for example from a reference picture index, that a base-layer picture is used as inter prediction reference for the enhancement layer. When a decoded base-layer picture is used as prediction reference for an enhancement layer, it is referred to as an inter-layer reference picture.
[0249] It needs to be understood that the description of scalable video coding may be generalized to any scalability hierarchy with more than two layers. In this case, a second enhancement layer may depend on a first enhancement layer in encoding and / or decoding processes, and the first enhancement layer may therefore be regarded as the base layer for the encoding and / or decoding of the second enhancement layer. Furthermore, it needs to be understood that there may be inter-layer reference pictures from more than one layer in a reference picture buffer or reference picture lists of an enhancement layer, and each of these inter-layer reference pictures may be considered to reside in a base layer or a reference layer for the enhancement layer being encoded and / or decoded. Furthermore, it needs to be understood that other types of inter-layer processing than reference-layer pictureupsampling may take place instead or additionally. For example, the bit-depth of the samples of the reference-layer picture may be converted to the bit-depth of the enhancement layer and / or the sample values may undergo a mapping from the color space of the reference layer to the color space of the enhancement layer.
[0250] A scalable video coding and / or decoding scheme may use multi-loop coding and / or decoding, which may be characterized as follows. In the encoding / decoding, a base layer picture may be reconstructed / decoded to be used as a motion-compensation reference picture for subsequent pictures, in coding / decoding order, within the same layer or as a reference for inter-layer (or inter-view or inter-component) prediction. The reconstructed / decoded base layer picture may be stored in the decoded picture buffer (DPB). An enhancement layer picture may likewise be reconstructed / decoded to be used as a motion-compensation reference picture for subsequent pictures, in coding / decoding order, within the same layer or as reference for inter-layer (or inter-view or inter-component) prediction for higher enhancement layers, when any. In addition to reconstructed / decoded sample values, syntax element values of the base / reference layer or variables derived from the syntax element values of the base / reference layer may be used in the inter-lay er / inter-component / inter- view prediction.
[0251] Inter-layer prediction may be defined as prediction in a manner that is dependent on data elements (e.g., sample values or motion vectors) of reference pictures from a different layer than the layer of the current picture (being encoded or decoded). Many types of inter-layer prediction exist and may be applied in a scalable video encoder / decoder.
[0252] The types of inter-layer prediction may comprise, but are not limited to, one or more of the following: inter-layer sample prediction, inter-layer motion prediction, inter-layer residual prediction. In inter-layer sample prediction, at least a subset of the reconstructed sample values of a source picture for inter-layer prediction are used as a reference for predicting sample values of the current picture. In inter-layer motion prediction, at least a subset of the motion vectors of a source picture for inter-layer prediction are used as a reference for predicting motion vectors of the current picture. Typically, predicting information on which reference pictures are associated with the motion vectors is also included in inter-layer motion prediction. For example, the reference indices of reference pictures for the motion vectors may be inter-layer predicted and / or the picture order count or any other identification of a reference picture may be inter-layer predicted. In some cases, inter-layer motion prediction may also comprise prediction of block coding mode, header information, block partitioning, and / or other similar parameters. In some cases, coding parameter prediction, such as inter-layer prediction of block partitioning, may be regarded as another type of inter-layer prediction. In inter-layer residual prediction, the prediction error or residual of selected blocks of a source picture for inter-layer prediction is used for predicting the current picture.
[0253] A direct reference layer may be defined as a layer that may be used for inter-layer prediction of another layer for which the layer is the direct reference layer. A direct predicted layer may be defined as a layer for which another layer is a direct reference layer. An indirect reference layer may be defined as a layer that is not a direct reference layer of a second layer but is a direct reference layer of a third layer that is a direct reference layer or indirect reference layer of a direct reference layer of the second layer for which the layer is the indirect reference layer. An indirect predicted layer may be defined as a layer for which another layer is an indirect reference layer. A dependent layer may be a directed predicted layer or an indirect predicted layer. An independent layer may be defined as a layer that does not have direct reference layers. In other words, an independent layer is not predicted using inter-layer prediction. A non-base layer may be defined as any other layer than the base layer, and the base layer may be defined as the lowest layer in the bitstream. An independent non-base layer may be defined as a layer that is both an independent layer and a non-base layer.
[0254] A multi-layer bitstream is a bitstream comprising multiple layers, which may be, but are not limited to, base and enhancement layers as discussed above for scalable video coding. A multi-layer bitstream may additionally or alternatively comprise independent layers that do not have inter-layer prediction relationship between each other and may even represent different types of content.
[0255] Scalability modes or scalability dimensions may include, but are not limited, to one or more of the following:- Quality scalability: Base layer pictures are coded at a lower quality than enhancement layer pictures, which may be achieved, for example, by using a greater quantization parameter value (e.g., a greater quantization step size for transform coefficient quantization) in the base layer than in the enhancement layer.- Spatial scalability: Base layer pictures are coded at a lower resolution (e.g., have fewer samples) than enhancement layer pictures. Spatial scalability and quality scalability may sometimes be considered the same type of scalability.- Bit-depth scalability: Base layer pictures are coded at lower bit-depth (e.g., 8 bits) than enhancement layer pictures (e.g., 10 or 12 bits).- Dynamic range scalability: Scalable layers represent a different dynamic range and / or images obtained using a different tone mapping function and / or a different optical transfer function.- Chroma format scalability: Base layer pictures provide lower spatial resolution in chroma sample arrays (e.g., coded in 4:2:0 chroma format) than enhancement layer pictures (e.g., 4:4:4 format).- Color gamut scalability: enhancement layer pictures have a richer / broader color representation range than that of the base layer pictures - for example the enhancement layer may have UHDTV (ITU-R BT.2020) color gamut and the base layer may have the ITU-R BT.709 color gamut.- Region-of-interest (ROI) scalability: An enhancement layer represents a spatial subset of the base layer. ROI scalability may be used together with other types of scalabilities, e.g., quality or spatial scalability so that the enhancement layer provides higher subjective quality for the spatial subset.- View scalability, which may also be referred to as multiview coding. In an example, the base layer represents a first view or a first camera, whereas an enhancement layer represents a second view or a second camera. In another example, the base layer represents a first set of views, which may be for example frame-packed, whereas an enhancement layer represents a second set of views.- Depth scalability, which may also be referred to as depth-enhanced coding. A layer or some layers of a bitstream may represent texture view(s), while other layer or layers may represent depth view(s).
[0256] In the above scalability examples, base layer information may be used to code enhancement layer to minimize the additional bitrate overhead.
[0257] Scalability may be enabled in two example ways. Either by introducing new coding modes for performing prediction of pixel values or syntax from lower layers of the scalable representation; or by placing the lower layer pictures to the reference picture buffer (decoded picture buffer, (DPB)) of the higher layer. The first approach (coding modes) is more flexible and thus may provide better coding efficiency in most cases. The reference frame -based scalability approach may be implemented more efficiently with minimal changes to single layer codecs, as compared to the first approach, while still achieving majority of the coding efficiency gains available. For example, a reference frame-based scalability codec may be implemented by utilizing the same hardware or software implementation for all the layers, by simply managing of the DPB management by external means.
[0258] An output layer set (OLS) may be defined as a set of layers where one or more layers in the set of layers are indicated to be output layers. Pictures of an output layer may be determined to be output by the decoder similarly to determining whether a picture is output or not in single-layer bitstreams, as described earlier. Pictures that are among the layers of an OLS but not among output layers are not output by the decoder. When multiple OLSs are indicated for a bitstream, the decoder may be instructed, e.g., thorough an interface which OLS is used in decoding and output. OLSs may be indicated, e.g., in a VPS.
[0259] In some video coding formats, an encoder may indicate a profile and a level for an output layer set. e.g., in a VPS. The capability associated with the indicated profile and level may be required in order to decode the output layer set. The output layer set may be constrained by the constraints defined for the indicated profile and level.
[0260] It has been proposed, e.g., in JVET-O1150, available from [https: / / www.jvet- experts.org / doc_end_user / documents / 15_Gothenburg / wgl l / JVET-O1150-v2.zip (last accessed on January 3, 2024)], that temporal sublayers may be used for any type of scalability. A mapping of scalability dimensions to sublayer identifiers could be provided, e.g., in a VPS or in an SEI message.
[0261] Information on neural-network post-filter characteristics (NNPFC) and neural-network post-filter activation (NNPFA) supplemental enhancement information (SEI) messages
[0262] The NNPFC SEI message and the NNPFA SEI message have been described in version 3 of the versatile supplemental enhancement information (VSEI) standard.
[0263] The syntax structure specifying the NNPFC SEI message may be called nn_post_filter_characteristics. The syntax structure specifying the NNPFA SEI message may be called nn_po st_filter_acti vation .
[0264] The NNPFC SEI message comprises the nnpfc_id syntax element, which includes an identifying number that may be used to identify a post-processing filter. A base post-processing filter is the filter that is included in or identified by the first NNPFC SEI message, in decoding order, that has a particular nnpfc_id value within a coded layer video sequence (CLVS). When there is a second NNPFC SEI message that has the same nnpfc_id value that defines the base post-processing filter, an update relative to the base post-processing filter is applied to obtain a post-processing filter associated with the nnpfc_id value. The update may be obtained by decoding the coded neural network bitstream in the second NNPFC SEI message. Otherwise, the post-processing filter associated with the nnpfc_id value is assigned to be the same as the base post-processing filter.
[0265] The NNPFC SEI message comprises the nnpfc_mode_idc syntax element, the semantics of which may be defined as follows:
[0266] nnpfc_mode_idc equal to 1 specifies that the base post-processing filter or the update relative to the base post-processing filter associated with the nnpfc_id value is a neural network identified by the Uniform Resource Identifier (URI) nnpfc_uri with the format identified by the tag URI nnpfc_tag_uri.
[0267] nnpfc_mode_idc equal to 0 indicates that this SEI message includes an ISO / IEC 15938-17 bitstream that specifies the base post-processing filter or updates relative to the base post-processing filter with the same nnpfc_id value.
[0268] The NNPFC SEI message may also comprise:- Purpose of the post-processing filter, which may comprise, but may not be limited to, one or more of the following:■ Visual quality improvement;■ Chroma upsampling from the 4:2:0 chroma format to the 4:2:2 or 4:4:4 chroma format, or from the 4:2:2 chroma format to the 4:4:4 chroma format;■ Increasing the width or height of the input picture;■ Frame rate upsampling;■ Bit depth upsampling; or■ Colorization.- Formatting of the input tensors that are given as input to the neural network inference- Formatting of the output tensors that are resulting from the neural network inference; and- Characterization of the complexity of the neural network.
[0269] The NNPFC SEI message syntax comprises nnpfc_base_flag. nnpfc_base_flag equal to 1 specifies that the SEI message specifies the base NNPF. nnpfc_base_flag equal to 0 specifies that the SEI message specifies an update relative to the base NNPF.
[0270] When nnpfc_base_flag is equal to 0, the following applies:- This SEI message defines an update relative to the preceding base NNPF in decoding order with the same nnpfc_id value. Updates are not cumulative but rather each update is applied on the base NNPF, which is the NNPF specified by the first NNPFC SEI message, in decoding order, that has a particular nnpfc_id value within the current CLVS. The NNPF defined by this SEI message is obtained by applying the update defined by this SEI message relative to the base NNPF with the same nnpfc_id value.- This SEI message pertains to the current decoded picture and all subsequent decoded pictures of the current layer, in output order, until the end of the current CLVS or up to but excluding the decoded picture that follows the current decoded picture in output order within the current CLVS and is associated with a subsequent NNPFC SEI message, in decoding order, having nnpfc_base_flag equal to 0 and that particular nnpfc_id value within the current CLVS, whichever is earlier.
[0271] The NNPFC SEI message syntax includes the nnpfc_num_input_pics_minusl syntax element. nnpfc_num_input_pics_minusl plus 1 specifies the number of pictures used as input for the NNPF. The variable numlnputPics may be set equal to nnpfc_num_input_pics_minusl + 1.
[0272] A frame rate upsampling filter may interchangeably be called a picture rate upsampling filter. Such a filter generates or interpolates one or more pictures between a pair of pictures given asinput to the filter. It is also possible to have a frame rate upsampling filter where the number of input pictures may be greater than 2. Such a frame rate upsampling filter may generate pictures between more than one pair of input pictures. A frame rate upsampling filter may comprise a neural network, in which case the generation of the interpolated pictures between a pair of input pictures is performed by the inference of the neural network. It is possible to have a frame rate upsampling filter that extrapolates a picture before input picture(s) or after input picture(s), instead of or in addition to between input pictures.
[0273] When the filtering purpose comprises frame rate upsampling, the NNPFC SEI message includes nnpfc_interpolated_pics[ i ] syntax elements for the values of i in the range of 0, inclusive, to nnpfc_num_input_pics_minusl, exclusive. nnpfc_interpolated_pics[ i ] specifies the number of interpolated pictures generated by the NNPF between the i-th and the ( i + 1 )-th picture used as input for the NNPF.
[0274] The NNPFC SEI message syntax may comprise an indication, which may be called nnpfc_absent_input_pic_zero_flag, that indicates how pictures that would not originate from the current bitstream are expected to be replaced in the input tensor. nnpfc_absent_input_pic_zero_flag equal to 1 indicates that the NNPF expects an input picture that is not present in the current bitstream to be represented sample arrays with sample values equal to 0. nnpfc_absent_input_pic_flag equal to 0 indicates that the NNPF expects an input picture that is not present in the current bitstream to be represented by the closest input picture in output order within the current bitstream.
[0275] The NNPFC SEI message syntax may comprise an indication, which may be called nnpfc_auxiliary_inp_idc, that indicates when auxiliary input data in addition to sample array(s) of input picture(s) is present in the input tensor of the NNPF. nnpfc_auxiliary_inp_idc greater than 0 indicates that auxiliary input data is present in the input tensor of the NNPF. Specific semantics may be specified for specific non-zero values of nnpfc_auxiliary_inp_idc. nnpfc_auxiliary_inp_idc equal to 0 indicates that auxiliary input data is not present in the input tensor.
[0276] The NNPFA SEI message specifies the neural-network post-processing filter (NNPF) that may be used for post-processing filtering for the current picture, or for post-processing filtering for the current picture and one or more other pictures. The NNPFA SEI message comprises the nnpfa_target_id syntax element, which indicates that the neural-network post-processing filter with nnpfc_id equal to nnpfa_target_id may be used for post-processing filtering for the indicated persistence. The indicated persistence may be the current picture only (indicated by nnpfa_persistence_flag equal to 0). Alternatively, the NNPF activation may be indicated to be persistent by nnpfa_persistence_flag equal to 1, in which case the persistence of the NNPF activation may last until the end of the current CLVSor the next picture, in output order, in the current layer associated with a NNPFA SEI message with the same nnpfa_target_id as the current SEI message.
[0277] The NNPFA SEI message syntax may comprise a syntax element indicative when the base post-processing filter or the latest post-processing filter is activated, where the latest post-processing filter is defined by the base post-processing filter relative to which the latest filter update, when any, has been applied. The syntax element may be called nnpfa_target_base_flag. nnpfa_target_base_flag equal to 1 specifies that the target NNPF is the base NNPF with nnpfc_id equal to nnpfa_target_id. nnpfa_target_base_flag equal to 0 specifies that the target NNPF is the NNPF specified by the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id that precedes the first VCL NAL unit of the current picture in decoding order and is not a repetition of the NNPFC SEI message that includes the base NNPF.
[0278] The NNPFA SEI message syntax may comprise indications which ones of the filtered pictures corresponding to the input pictures are output by the NNPF process. For the i-th input picture that is filtered by the NNPF, the NNPFA SEI message syntax may comprise nnpfa_output_flag[ i ] syntax element, which when equal to 0, specifies that the filtered picture is not output by the NNPF process, and when equal to 1, specifies that the filtered picture is output by the NNPF process.
[0279] In relation to an NNPFA SEI message, two sets of pictures may be defined, namely nnpfcTargetPictures and nnpfaTargetPictures. nnpfcTargetPictures may be defined to be the set of pictures to which the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id that precedes the current NNPFA SEI message in decoding order pertains. nnpfaTargetPictures may be defined to be the set of pictures for which the target NNPF is activated by the current NNPFA SEI message. It may be required for a conforming bitstream that any picture included in nnpfaTargetPictures shall also be included in nnpfcTargetPictures.
[0280] An NNPF process comprises performing the NNPF inference for given input pictures. The NNPF inference may be performed in a patch-wise manner so that the entire picture area gets filtered. The NNPF inference may be followed by outputting NNPF-generated pictures in their increasing index order, where all NNPF-generated pictures that were interpolated by the NNPF are output and those NNPF-generated pictures that correspond to any input pictures to the NNPF are output as specified in the semantics of the NNPFA SEI message.
[0281] A general post-processing filtering process using NNPFs may be described as follows. Input to this process is a bitstream BitstreamToFilter. Output of this process is a list of NNPF output pictures ListNnpfOutputPics. First, BitstreamToFilter is decoded, and the list CroppedDecodedPicturesis set to be the list of the cropped decoded pictures in output order resulted from decoding BitstreamToFilter. Second, the filtering process for one picture, as described below, is repeatedly invoked, in output order, for each cropped decoded picture that is in CroppedDecodedPictures and for which one or more NNPFs are activated. The order of the pictures in ListNnpfOutputPics is in output order. It may be required that within ListNnpfOutputPics there shall be no more than one picture pertaining to any particular output time instance. When for any particular picture in CroppedDecodedPictures there are multiple NNPFs activated and only one the NNPFs is allowed to be chosen to be applied although any of the NNPFs may be chosen, the above constraint shall apply regardless of which NNPF is chosen to be applied to the particular picture.
[0282] A filtering process for one picture using an NNPF may be described as follows. The filtering process for one picture using an NNPF may be applied to each cropped decoded picture, referred to as the current picture, that is in CroppedDecodedPictures and for which one or more NNPFs are activated. When applying an NNPF to the current picture, the filtered and / or interpolated pictures are generated by the NNPF by applying the NNPF process to the current picture. When applying an NNPF to the current picture, the order of the pictures generated by the NNPF by applying the NNPF process being stored into the output tensor of the NNPF is in output order. When the applied NNPF is the last NNPF that is applied to the current picture, the pictures generated by the NNPF and output by the NNPF process are included into ListNnpfOutputPics, in the same order as when the pictures are stored into the output tensor of the NNPF.
[0283] The use of NNPFC and NNPFA SEI messages for VVC has been described in version 3 of the versatile video coding (VVC) standard. It is to be understood that NNPFC and NNPFA SEI message may be similarly used for any other video coding specification.
[0284] When NNPFC and NNPFA SEI messages are used for VVC, a decoder selects input pictures for the NNPF. The input pictures may be selected in reverse output order starting from a picture for which the NNPF is activated through an NNPFA SEI message. The input pictures may be indexed, starting from index 0 that is assigned for the picture for which the NNPF is activated through an NNPFA SEI message. In an example, the decoder selects the input picture with index i, where i is greater than 0, to be the latest cropped decoded output picture, in output order, that precedes the input picture with index i-1 in output order. When there is no cropped decoded output picture, in output order, that precedes the input picture with index i- 1 in output order as a result of decoding the bitstream, it may be considered that the input picture with index i is not present in the current bitstream (e.g., missing) and the subsequent input pictures, when any, with index i+1 to numlnputPics-l, inclusive, are likewise missing. A missing input picture may be treated like described above in relation to nnpfc_absent_input_pic_zero_flag syntax element.
[0285] When NNPFC and NNPFA SEI messages are used for VVC and a picture rate upsampling NNPF that interpolates pictures between a single pair of input pictures is activated persistently until the end of the bitstream, the NNPF is applied repeatedly at the end of the bitstream for different sets of input pictures up to but excluding a set of input pictures that would cause creation of any interpolated picture after the last picture of the bitstream in output order. In these sets of input pictures, some of the pictures may be missing and may be, for example, replaced by the last picture within the bitstream in output order.
[0286] Visual temporal extrapolation
[0287] It is to be understood that, in various embodiments, the terms visual temporal extrapolation, temporal extrapolation, and video prediction may be used interchangeably. Visual temporal extrapolation may be defined as a method, algorithm, or process that generates one or more pictures in the future given one or more past pictures as input. Visual temporal extrapolation may be realized by, but is not necessarily based on or limited to, neural network inference.
[0288] Use cases for visual temporal extrapolation
[0289] Use cases for visual temporal extrapolation include:- Very low delay computer vision for domains like robotics and autonomous driving, where extrapolated future pictures facilitate anticipatory decision making.- Increase of the rendered picture rate in very low-latency applications, such as cloud gaming, relative to the decoded picture rate.- Reduction of the end-to-end delay in low-latency applications through extrapolating and displaying future pictures before they are received.- Generative face video for very low bitrate video coding.
[0290] Information on generative Al, including temporal extrapolation and generative face video
[0291] The term generative Artificial Intelligence (Al), or generative modeling, or generative machine learning (and other similar terms), are commonly used to indicate a class of models learned from data that are capable of generating new data. Some example generative models are based on neural networks. The basic components or layers of a generative NN are usually not different from the components or layers of a non-generative NN. Example of such components are non-linear layers, fully- connected layers, normalization layers, attention layers, and the like.
[0292] One typical example of neural network architecture that allows for generating text data is a Trasformer-based “decoder”, where “decoder” may not refer to a decoder that is part of a codec performing compression of input data into a small bitstream. Instead, the decoder is a neural network that gets a set of input words or parts of words or tokens extracted from input words, and outputs a set of output words or parts of words or tokens. At inference time, such a NN is run in auto-regressive mode, where the generated word(s) or token(s) is provided as part of the input word(s) or token(s). In order for such a NN to generate data, it is trained to predict the next word(s) (or an estimate of a probability distribution over the next words) given a set of input words. The NN would be based on the Transformer architecture, which comprises the use of the self-attention mechanism, where an attention score is assigned to each input token or word based on all other input tokens or words, including the previously generated words or tokens. During training of a decoder-style Transformer architecture, the future data items (words or tokens) are masked so not to leak information from the future. In some cases, decoder-style Transformer architectures are referred to as “uni-directional” (because they use or process information from left-to-right), as opposed to some encoder-stype Transformer architectures that are referred to as “bi-directional” (because they use or process information from left-to-right and from right-to-left).
[0293] Another example of generative modeling is visual temporal extrapolation, where a picture is generated by a NN based on one or more previously decoded or generated pictures and on one or more other data items. The one or more previously decoded or generated pictures may be pictures decoded by a process that does not involve generative modeling, such as a traditional codec, e.g., a V VC -compliant codec. The one or more data items may include parameters or features that describe the differences between the one or more previously decoded pictures and the current picture to be temporally extrapolated. Examples of such parameters are facial parameters (such as facial landmarks and their positions or differential positions with respect to the facial landmarks of a previous picture), or parameters of other objects. The one or more data items may be signaled from encoder to decoder.
[0294] FIG. 6 illustrates a generic block diagram of a generative face video codec (GFVC). The input base picture 602 is encoded with any video encoder 604, such as a VVC encoder, into a coded base picture of a bitstream 606. The input subsequent pictures, e.g., subsequent pictures 608-1, 608-2 are fed into the analysis model 610 for extracting feature parameters. The feature parameters are encoded 612 to generate a feature bitstream 614. Feature parameters may be encoded independently of feature parameters of any other picture, e.g., for the first subsequent picture following the base picture in decoding order. Alternatively, feature parameters may be encoded 612 in a predictive manner with reference to earlier encoded feature parameters in decoding order, wherein feature residual may be encoded relative to predicted feature parameters. Feature residual may undergo quantization and entropy coding as part of feature encoding. Encoded feature parameters are included in or along thefeature bitstream 614, e.g., in a generative face video SEI message. For a subsequent picture, the video encoder may encode a dummy picture or a drive picture, as described in the subsequent section. Encoded feature parameters of a subsequent picture may be included in or along the respective coded dummy picture or drive picture, e.g., in the same picture unit.
[0295] The decoder 616 decodes a coded base picture from the bitstream 606. For a subsequent coded dummy or drive picture, the feature decoder 618 decodes encoded feature parameters. The decoded feature parameters and the decoded base picture and optionally a decoded drive picture are input to a generative neural network model inference 620, which generates an output picture sequence 622.
[0296] Information on GFV SEI message
[0297] In JVET, a Generative Face Video (GFV) SEI message has been proposed for a new version of the Versatile Supplemental Enhancement Information (VSEI) standard, in order to support visual temporal extrapolation of faces in videos. One of the main functions of the GFV SEI message is to signal facial parameters, which represent an auxiliary input to the NN that performs the temporal extrapolation. The latest document describing the GFV SEI message is JVET-AF0234, available from(last accessed on January 3, 2024)]. An example summary is provided as follows.
[0298] The GFV SEI message defines an interface between a video decoder and a generator NN that performs visual temporal extrapolation of faces in videos.
[0299] The following are the types of pictures considered in the GFV SEI message:- Base picture: a decoded output picture that may be used by the generative network to generate a novel face picture. It is coded by a traditional codec such as VVC-compliant codecs.- Dummy picture: picture unit of the minimum allowed resolution that includes only SEI messages.- (optional) Primary picture (also called drive picture or driving picture) that can be optionally input to a generative NN to improve background texture and / or facial details. This drive picture may also be coded by a traditional codec such as a VVC-compliant codec.
[0300] A second neural network, referred to as Translator NN, is used to convert the facial parameters (signaled within a GFV SEI message) into the following converted parameters:- 15 3D-keypoints- One 3x3 matrix- One 1x3 matrix
[0301] The input to the generator NN may be one or more of the following:- The (previously decoded) base pic.- The converted parameters (output of Translator NN)- (optionally) The primary / drive picture.
[0302] The generative face video (GFV) SEI message proposed in JVET-AF0234 makes use of dummy pictures, e.g., pictures of the minimum allowed resolution, that are only used to include GFV SEI messages. It is asserted that dummy pictures seem to be redundant and cause unnecessary bitrate overhead.
[0303] Legacy decoders would display drive pictures and dummy pictures as proposed in JVET- AF0234, which would be annoying for human observers, since they are not intended for displaying and may not represent image content that is similar to base pictures.
[0304] It has been argued that a generative video SEI message, such as GFV SEI message, does not achieve interoperability between different implementations, since support of SEI messages in decoders may be optional and the generative neural network associated with the generative video SEI message may not have been specified in an exact manner.
[0305] Some embodiments are described in relation to a generative video SEI message. The generative face video (GFV) SEI message described above is an example of a generative video SEI message. A generative video SEI message may be defined as an SEI message that may comprise syntax elements from which input(s) to the generative neural network inference are derived and / or may cause invocation of the generative neural network inference. It is to be understood that embodiments are not limited to the GFV SEI message but apply to any generative video SEI message.
[0306] Some of the embodiments described herein avoid coding of dummy pictures and hence making the decoded output of legacy decoders (e.g., not capable of running a generative neural network) more reasonable for displaying.
[0307] In some embodiments, the generative neural network is integrated in scalable video coding.
[0308] Enabling multiple generative face video SEI message in a picture unit, each invoking inference of generative NN
[0309] In this section, embodiments to avoid coding of dummy pictures are presented. Instead of a dummy picture comprising a GFV SEI message, an encoder encodes a suffix SEI NAL unit comprising a GFV SEI message in the picture unit that includes the base or drive picture referenced by the GFV SEI message. Multiple GFV SEI messages with different payload are allowed in the same picture unit. Each new instance of a GFV SEI message is intended to invoke the generative neural network (NN) inference.
[0310] FIG. 7 illustrates a picture unit generated by an encoder according to embodiments described in this section. The picture unit 700 includes a coded picture 702, and a coded picture includes at least one coded slice NAL unit, e.g., 704-1 and 704-2. The picture unit 700 also includes at least one suffix SEI NAL unit, e.g., 706-1 and 706-2. Each suffix SEI NAL unit includes a generative video SEI message.
[0311] Encoding a generative video SEI message in a suffix SEI NAL unit of a picture unit comprising a coded base or drive picture
[0312] In an embodiment, an encoder:- receives a first input picture;- encodes the first input picture into a first coded picture, wherein the first coded picture may be a coded base picture or a coded drive picture and a first picture unit comprises the first coded picture;- receives a second input picture;- extracts feature parameters from the second input picture;- encodes the feature parameters into encoded feature parameters; and- includes the encoded feature parameters in the first picture unit.
[0313] In an additional embodiment, the encoder:- includes the encoded feature parameters in a generative video supplemental enhancement information (SEI) message that follows, in decoding order, the first coded picture.
[0314] In an additional embodiment, the encoder:- includes the generative video SEI message in a suffix SEI network abstraction layer (NAL) unit within the first picture unit.
[0315] Encoding a generative video SEI message in a prefix SEI NAL unit of a picture unit succeeding a coded base or drive picture in decoding order
[0316] In an embodiment, an encoder:- receives a first input picture;- encodes the first input picture into a first coded picture, wherein the first coded picture may be a coded base picture or a coded drive picture;- receives a second input picture;- extracts feature parameters from the second input picture;- encodes the feature parameters into encoded feature parameters; and- includes the encoded feature parameters in a subsequent picture unit, wherein the subsequent picture unit is intended to comprise a subsequent coded picture, which may be a second coded drive picture or a second coded base picture.
[0317] In an additional embodiment, the encoder:- includes the encoded feature parameters in a generative video supplemental enhancement information (SEI) message that precedes, in decoding order, the subsequent coded picture.
[0318] In an additional embodiment, the encoder:- includes the generative video SEI message in a prefix SEI network abstraction layer (NAL) unit within the subsequent picture unit.
[0319] Invoking generative neural network inference
[0320] In an embodiment, when the encoded feature parameters are decoded, they are intended to invoke the inference of a generative neural network.
[0321] In an embodiment, the inference of a generative neural network is performed in the decoder side as signaled with a generative video SEI message, such as a GFV SEI message. When there are multiple generative video SEI messages, such as multiple GFV SEI messages, with different content in the same picture unit, the inference of the generative neural network may be performed separately for each of these generative video SEI messages.
[0322] In an embodiment, an encoder additionally includes an activation signal to invoke a generative neural network inference along the encoded feature parameters. For example, an encoder may include a neural-network post-filter activation (NNPFA) SEI message next to the generative video SEI message in decoding order to invoke the inference of a generative neural network.
[0323] In an embodiment, an encoder additionally includes an activation signal to identify a generative neural network to be used for inference, which is invoked by other means, such as the generative video SEI message. For example, an encoder may encode a neural-network post-filteractivation (NNPFA) SEI message preceding the generative video SEI message in decoding order, to invoke the inference of a generative neural network. The NNPFA SEI message identifies the NNPF (e.g., whether the base NNPF or an updated NNPF for the given nnpfa_target_id is in use) and does not imply inference. The inference of the identified NNPF is performed in the decoder side as signaled with the generative video SEI message(s).
[0324] Identifying an instance of a generative video SEI message
[0325] In an embodiment, an encoder includes an identifier or a counter value for each generative video SEI message, wherein the identifier or the counter value identifies the generative video SEI message from other generative video SEI messages in the same picture units. When the identifier or the counter value is the same as for a previous generative video SEI message in the same picture unit, the current generative video SEI message may be concluded to be a copy of the previous generative video SEI message.
[0326] Example embodiments related to generative face video (GFV) SEI message
[0327] In an example embodiment, an encoder generates multiple GFV SEI messages with different payload in the same picture unit. Each new instance of a GFV SEI message invokes the generative neural network (NN) inference. Instead of a dummy picture including a GFV SEI message, an encoder uses a suffix SEI NAL unit comprising a GFV SEI message in the picture unit that includes the base or drive picture referenced by the GFV SEI message.
[0328] In an additional embodiment, the encoder adds a counter syntax element (gfv_cnt) in the GFV SEI message. With the counter syntax elements, it is possible to differentiate between repetitions and new instances of GFV SEI messages within a picture unit.
[0329] In a first example embodiment, the counter syntax element may use the following syntax and semantics:
[0330] gfv_cnt specifies a GFV SEI message instance count value for this gfv_id value within a picture unit.
[0331] In an example, the values of gfv_cnt may be constrained as follows: The gfv_cnt of the first GFV SEI message, in decoding order, with a particular value of gfv_id within picture unit shall be equal to 0. When gfv_cnt assigned to currGfvCnt is greater than 0, a GFV SEI message with the same gfv_id value and gfv_cnt equal to currGfvCnt - 1 shall precede the current GFV SEI message in decoding order in the same picture unit. The value of gfv_cnt shall be in the range of 0 to 65 535, inclusive.
[0332] It may be required that when there are multiple GFV SEI messages with the same values of gfv_id and gfv_cnt present in the same picture unit, the SEI messages shall have the same SEI payload content.
[0333] The generative neural network inference may be invoked in increasing order of gfv_cnt to generate a video picture per each GFV SEI message that has gfv_base_pic_flag equal to 0 and a unique value of gfv_cnt within a picture unit.
[0334] In an embodiment, gfv_base_pic_flag equal to 1 indicates the current decoded output picture corresponds to a base picture. gfv_base_pic_flag equal to 0 indicates the current decoded output picture does not correspond to a base picture or this SEI message does not specify syntax elements for a base picture. When gfv_cnt is greater than 0, it may be required that gfv_base_pic_flag shall be equal to 0.
[0335] In an alternative embodiment, gfv_base_pic_flag is present only when gfv_cnt is equal to 0, as indicated in the following syntax. When gfv_base_pic_flag is not present, it is inferred to be equal to 0.
[0336] In a second example embodiment, the counter syntax element may use the following syntax and semantics:
[0337] gfv_cnt specifies a GFV SEI message instance count value for this gfv_id value within a picture unit.
[0338] In an example, the values of gfv_cnt may be constrained as follows: The gfv_cnt of the first GFV SEI message, in decoding order, with a particular value of gfv_id within picture unit and with gfv_base_pic_flag equal to 0 shall be equal to 0. When gfv_cnt assigned to currGfvCnt is greater than 0, a GFV SEI message with the same gfv_id value and gfv_cnt equal to currGfvCnt - 1 shall precede the current GFV SEI message in decoding order in the same picture unit. The value of gfv_cnt shall be in the range of 0 to 65 535, inclusive.
[0339] It may be required that when there are multiple GFV SEI messages with the same values of gfv_id and gfv_cnt present in the same picture unit, the SEI messages shall have the same SEI payload content.
[0340] The generative neural network inference may be invoked in increasing order of gfv_cnt to generate a video picture per each GFV SEI message that has gfv_base_pic_flag equal to 0 and a unique value of gfv_cnt within a picture unit.
[0341] Decoding a generative video SEI message from a suffix SEI NAL unit of a picture unit comprising a coded base or drive picture
[0342] In an embodiment, a decoder:- receives a first coded picture from a first picture unit, wherein the first coded picture may be a coded base picture or a coded drive picture;- decodes first coded picture to a first decoded picture;- receives encoded feature parameters from the first picture unit;- creates a feature input tensor from the encoded feature parameters; and- invokes a generative neural network inference with the first decoded picture and the feature input tensor as inputs to generate a picture.
[0343] In an embodiment, a decoder is characterized by an interface that accepts less than a picture unit to be passed for decoding, and decoding of the first coded picture is invoked prior to receiving or creating the feature input tensor.
[0344] Decoding a generative video SEI message from a prefix SEI NAL unit of a picture unit succeeding a coded base or drive picture in decoding order
[0345] In an embodiment, a decoder:- receives a first coded picture, wherein the first coded picture may be a coded base picture or a coded drive picture;- decodes the first coded picture to a first decoded picture;- receives encoded feature parameters from a subsequent picture unit;- creates a feature input tensor from the encoded feature parameters; and- invokes a generative neural network inference with the first decoded picture and the feature input tensor as inputs to generate a picture.
[0346] In an embodiment, a decoder is characterized by an interface that accepts less than a picture unit to be passed for decoding, and creation of the feature input tensor is performed prior to receiving or decoding a subsequent coded picture from the subsequent picture unit.
[0347] Approach based on multiple layers and generative video SEI message! s')
[0348] In this section, embodiments where drive and dummy pictures reside in a dependent layer are presented. Consequently, legacy decoders may decode only the independent layer and avoid accidental displaying of drive and dummy pictures.
[0349] Base pictures are at an independent layer. When an access unit has a base picture in an independent layer, it also may also have a picture in the dependent layer predicting only from the independent layer and having no residual. This picture in the dependent layer may be referred to as a replica picture since the decoded base picture and the decoded replica picture are typically identical. The replica picture may be used as a reference for temporal inter prediction within the dependent layer.
[0350] Drive pictures may be marked as output pictures (e.g., have pic_output_flag equal to 1 in VVC), since they are needed as input in the generative NN performed as post-processing.
[0351] Dummy pictures may be predicted from a preceding picture in decoding order within the dependent layer, may be coded without prediction residual, and may be marked as output pictures (e.g., have ph_pic_output_flag equal to 1 in VVC). Alternatively, dummy pictures may be encoded in any manner that causes small bit rate overhead without considering the content of decoded dummy pictures, and may be marked not to be output (e.g., have ph_pic_output_flag equal to 0 in VVC).
[0352] Any dependent layer picture unit including a dummy or drive picture includes a generative video SEI message (such as a GFV SEI message).
[0353] A legacy decoder, such as a VVC Main 10 profile decoder, would decode only the independent layer, e.g., display only the base pictures.
[0354] FIG. 8 illustrates of a bitstream generated by an encoder according to embodiments described in this section is presented. NAL units (e.g., coded slice NAL units or non-VCL NAL units) are belonging to a layer are enclosed in a box with rounded corners 804, 806. FIG.8 includes an independent layer 804 and a dependent layer 806. The decoding order of access units, e.g., a first access 802 and a second access unit 803, is from left to right. Within an access unit, the independent layer 804 is decoded before the dependent layer 806. Within a picture unit, e.g., second picture unit 808 and a coded picture, e.g., a coded picture 810, a first coded picture 812, and a second coded picture 814, NAL units are decoded from top to bottom.
[0355] In an embodiment, an encoder:- receives a first input picture;- encodes the first input picture into a first coded picture, wherein the first coded picture is a coded base picture, the first coded picture resides in an independent layer, and a first access unit comprises the first coded picture;- receives a second input picture;- encodes a second coded picture, wherein the second coded picture may be a coded drive picture based on the second input picture or a coded dummy picture, the second coded picture resides in a dependent layer, a second picture unit comprises the second coded picture, and a second access unit comprises the second coded picture;- extracts feature parameters from the second input picture;- encodes the feature parameters into encoded feature parameters; and- includes the encoded feature parameters in the second picture unit.
[0356] In an embodiment, the encoder encodes a coded replica picture in the first access unit, wherein the coded replica picture resides in the dependent layer and is characterized to being predicted only from the independent layer and having no prediction residual.
[0357] In an embodiment, the includes the encoded feature parameters in a generative video supplemental enhancement information (SEI) message in the second picture unit.
[0358] In an example, a decoder may be capable of decoding only an independent layer and hence decode only the coded base pictures.
[0359] In an embodiment, the encoder predicts a coded dummy picture from the previous picture, in decoding order, within the dependent layer.
[0360] Approach based on multiple layers without generative video SEI message! s)
[0361] In this section, embodiments where the decoding of coded pictures of a dependent layer requires generative neural network inference are presented. The decoding of a dependent layer is indicated to be possible under a profile that includes the capability of the generative neural network, thus legacy decoders will not attempt to decode the dependent layer.
[0362] In the embodiments of this section, reference pictures are indicated conventionally. Thus, inter-layer prediction takes place between pictures of the same access unit.
[0363] An independent layer has base pictures, dummy pictures and drive pictures.
[0364] When the independent layer has a base picture, the same access unit has a picture in the dependent layer that is inter-layer predicted from the base picture and may have zero residual, and both the independent-layer and the dependent-layer pictures are marked to be output (e.g., have ph_pic_output_flag equal to 1 according to VVC).
[0365] Drive pictures at the independent layer are marked not to be output (e.g., have ph_pic_output_flag equal to 0 in VVC), since they are meant only to carry content used as input to the generative NN and not to be displayed directly.
[0366] Dummy pictures at the independent layer are marked not to be output (e.g., have ph_pic_output_flag equal to 0 according to VVC).
[0367] When the independent layer has a dummy or drive picture, the same access unit includes a picture in the dependent layer that uses exactly two pictures as reference, namely the inter-layer reference picture, which is either a dummy picture or a drive picture, and the dependent-layer base picture, which may be indicated as a short-term or long-term reference picture.
[0368] In some embodiments, the decoding process of the dependent layer picture comprises running the inference of the generative NN
[0369] In some embodiments, the output layer set including both the dependent layer and the independent layer is indicated to conform to a new profile, which may for example be referred to as Scalable Generative Main 10. A legacy decoder does not know the new profile and hence would not try to decode both layers. Only the enhancement layer may be an output layer in this output layer set.
[0370] In some embodiments, the output layer set including only the independent layer is indicated to conform to an existing profile, e.g., Main 10. A legacy decoder may be able to decode the independent layer and would display the base pictures but not the dummy or drive pictures, which is a desired behavior.
[0371] FIG. 9 illustrates a bitstream generated by an encoder according to embodiments described in this section. Arrows 902, 904, and 906 indicate inter prediction, wherein horizontal arrow(s) 906 indicate temporal inter prediction and vertical arrows 902 and 904 indicate inter-layer prediction. The arrow head points to the picture being predicted, e.g., a first coded DL picture 908 and a second coded DL picture 910 and the arrow source indicates the picture used as reference for prediction, e.g., a first coded BL picture (base) 912 and a first coded BL picture (base) 914.
[0372] In an embodiment, an encoder:- receives a first input picture;- encodes the first input picture into a first coded base-layer picture, wherein the first coded baselayer picture is a coded base picture, the first coded base-layer picture resides in an independent layer, and a first access unit comprises the first coded base-layer picture;- encodes a first coded dependent-layer picture, wherein the first coded dependent-layer picture resides in a dependent layer, and the first access unit comprises the first coded dependent-layer picture;- receives a second input picture;- encodes a second coded base-layer picture, wherein the second coded base-layer picture may be a coded drive picture based on the second input picture or a coded dummy picture, the second coded base-layer picture resides in the independent layer, and a second access unit comprises the second coded base-layer picture;- encodes a second coded dependent-layer picture, wherein the second coded dependent-layer picture resides in a dependent layer, and the second access unit comprises the second coded dependentlayer picture;- indicates in or along the second coded dependent-layer picture that the second coded dependentlayer picture may use the first coded dependent-layer picture and the second coded base-layer picture as reference for prediction;- extracts feature parameters from the second input picture;- encodes the feature parameters into encoded feature parameters; and- indicates that the encoded feature parameters are in use for the decoding of the second coded dependent-layer picture.
[0373] In an embodiment, the encoder encodes the first dependent-layer picture as a coded replica picture, wherein the coded replica picture is characterized to being predicted only from the independent layer and having no prediction residual.
[0374] Indicating output status
[0375] In an embodiment, the encoder encodes an indication in or along the second coded baselayer picture to omit output.
[0376] In an embodiment, the encoder indicates that the first coded base-layer picture is output.
[0377] In an embodiment, the encoder indicates that the first coded dependent-layer picture is output.
[0378] In an embodiment, the encoder indicates that the first coded dependent-layer picture and the second coded dependent-layer picture are output.
[0379] In an embodiment, the encoder indicates that the independent layer and the dependent layer form an output layer set and the dependent layer is an output layer of the output layer set.
[0380] In an embodiment, the encoder indicates a profile for the output layer set, wherein the profile includes the generative neural network capability.
[0381] Deriving feature input tensor
[0382] In an embodiment, the encoder creates a feature input tensor from the encoded feature parameters.
[0383] Alternatives for indicating the encoded feature parameters or the feature input tensor
[0384] In an embodiment, the encoder includes the encoded feature parameters or the feature input tensor in an adaptation parameter set with a type indicating generative video, such as generative face video.
[0385] In an embodiment, the encoder includes the encoded feature parameters or the feature input tensor in a picture header of the second coded dependent-layer picture. For example, the picture header extension of the VVC picture header syntax may comprise the encoded feature parameters.
[0386] Generating a decoded dependent-layer picture using the generative neural network and excluding syntax elements not relevant for reconstructing the generated picture
[0387] In an embodiment, the encoder:- reconstructs a first decoded dependent-layer picture from the first coded dependent-layer picture; and- reconstructs a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependent-layer picture and the feature input tensor as inputs to generate a picture.
[0388] In an embodiment, the syntax of the second coded dependent-layer picture excludes syntax elements that are not relevant to reconstructing the second decoded dependent-layer picture with the generative neural network.
[0389] In an embodiment, the syntax of one or more syntax structures of the second dependentlayer excludes syntax elements that are not relevant for reconstructing the second decoded dependentlayer picture with the generative neural network. For example, the syntax of parameter sets, picture header, slice header, and / or slice data syntax structures for the dependent layer may exclude syntax elements that are not relevant to reconstructing the second decoded dependent-layer picture with the generative neural network.
[0390] Using the generative neural network to reconstruct prediction units
[0391] In an embodiment, the encoder:- reconstructs a first decoded dependent-layer picture from the first coded dependent-layer picture;- reconstructs one or more prediction units of a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependent-layer picture and the feature input tensor as inputs; and- encodes difference between the one or more prediction units and collocated units of the second input picture as prediction residual of the second coded dependent-layer picture.
[0392] Using the generative neural network to generate an inter-layer reference picture
[0393] In an embodiment, the encoder:- reconstructs a first decoded dependent-layer picture from the first coded dependent-layer picture;- invokes a generative neural network inference with the first decoded dependent-layer picture and the feature input tensor as inputs to generate an inter-layer reference picture; and- encodes the second coded dependent-layer picture with reference to the inter-layer reference picture.
[0394] In an embodiment, the encoder:- encodes difference between the second input picture and the inter-layer reference picture as prediction residual of the second coded dependent-layer picture.
[0395] Decoding
[0396] In an embodiment, a decoder:- receives a first coded base-layer picture, wherein the first coded base-layer picture is a coded base picture, the first coded base-layer picture resides in an independent layer, and a first access unit comprises the first coded base-layer picture;- decodes the first coded base-layer picture into a first decoded base-layer picture;- receives a first coded dependent-layer picture, wherein the first coded dependent-layer picture resides in a dependent layer, and the first access unit comprises the first coded dependent-layer picture;- decodes the first coded dependent-layer picture into a first decoded dependent-layer picture;- receives a second coded base-layer picture, wherein the second coded base-layer picture may be a coded drive picture based on the second input picture or a coded dummy picture, the second coded base-layer picture resides in the independent layer, and a second access unit comprises the second coded base-layer picture;- in response to the second coded base-layer picture being a coded drive picture, decodes the second coded base-layer picture into a second decoded base-layer picture;- receives a second coded dependent-layer picture, wherein the second coded dependent-layer picture resides in a dependent layer, and the second access unit comprises the second coded dependentlayer picture;- decodes from or along the second coded dependent-layer picture that the second coded dependentlayer picture may use the first coded dependent-layer picture and the second coded base-layer picture as reference for prediction;- obtains a feature input tensor that is in use for the decoding of the second coded dependent-layer picture; and- reconstructs a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependent-layer picture and the feature input tensor as inputs to generate a picture.
[0397] It is to be noted that respective encoder embodiments are applicable for most of the decoder embodiments.
[0398] Approach based on multiple layers and diagonal inter-layer prediction without generative video SEI message! s')
[0399] In this section, embodiments where the decoding of coded pictures of a dependent layer requires generative neural network inference are presented. The decoding of a dependent layer is indicated to be possible under a profile that includes the capability of the generative neural network, thus legacy decoders will not attempt to decode the dependent layer.
[0400] In the embodiments of this section, inter-layer prediction may take place "diagonally", e.g., from a base or drive picture present in a first access that precedes a second access unit including the current dependent-layer picture in decoding order.
[0401] Base pictures and drive pictures are present at an independent layer.
[0402] Base pictures are marked to be output (e.g., have ph_pic_output_flag equal to 1 according to VVC).
[0403] Drive pictures are marked not to be output (e.g., have ph_pic_output_flag equal to 0 according to VVC), since they are meant only to carry content used as input to the generative NN and not to be displayed directly.
[0404] A dependent layer includes a picture at every temporal position where the decoder is running the generative NN. In some embodiments, the dependent layer picture is a dummy picture, e.g., includes no residual information relative to an inter-layer reference picture. In some embodiments, the dependent layer picture may include residual relative to an inter-layer reference picture.
[0405] In some embodiments, rather than using a conventional inter-layer reference picture (e.g., upsampled independent layer picture), the generative NN generates the inter-layer reference picture, i.e., the generative NN inference is regarded as an inter-layer process.
[0406] Various embodiments are presented to indicate the base picture and / or the drive picture used as input for the generative NN.
[0407] FIG. 10 illustrates of an example bitstream generated by an encoder according to some embodiments described in this section. The first coded picture 1002 may be a base picture, and optionally there may be a third coded picture 1004 that is a drive picture. The base picture 1002, or both the base picture 1002 and the drive picture 1004, are used as input to the generative NN to generate an inter-layer reference picture in some embodiments. In the illustration, the drive picture 1004 is present in the second access unit 1006, but it may generally be present in any access unit as long as its decoding order precedes the second coded picture 1008.
[0408] In an embodiment, an encoder:- receives a first input picture;- encodes the first input picture into a first coded picture, wherein the first coded picture is a coded base picture or a coded drive picture, the first coded picture resides in an independent layer, and a first access unit comprises the first coded picture;- receives a second input picture;- extracts feature parameters from the second input picture;- encodes the feature parameters into encoded feature parameters;- encodes a second coded picture, wherein the second coded picture resides in a dependent layer, and the second access unit comprises the second coded picture; and- indicates that the encoded feature parameters are in use for the decoding of the second coded picture.
[0409] In an embodiment, the encoder:- reconstructs a first decoded picture from the first coded picture;- creates a feature input tensor from the encoded feature parameters;- invokes a generative neural network inference with the first decoded picture and the feature input tensor as inputs to generate a first inter-layer reference picture; and- encodes the second coded picture with reference to the first inter-layer reference picture.
[0410] In an embodiment, the encoder:- omits encoding prediction residual for the second coded picture.
[0411] In an embodiment, the encoder:- encodes difference between the second input picture and the first inter-layer reference picture as prediction residual of the second coded picture.
[0412] Indicating output status
[0413] In an embodiment, in response to the first coded picture being a coded base picture, the encoder indicates that the first coded picture is output.
[0414] In an embodiment, in response to the first coded picture being a coded drive picture, the encoder indicates that the first coded picture is not output.
[0415] In an embodiment, the encoder indicates that the independent layer and the dependent layer form an output layer set and the dependent layer is an output layer of the output layer set.
[0416] In an embodiment, the encoder indicates that the second coded picture is output.
[0417] Indicating when an independent-layer picture is a base picture or a drive picture for decoding of the dependent layer
[0418] In an embodiment, the encoder:- indicates in or along the first coded picture an indication indicating that the first coded picture is a coded base picture or that the first coded picture is a coded drive picture.
[0419] In an embodiment, the encoder indicates with a picture header extra bit that the first coded picture is a coded base picture or that the first coded picture is a coded drive picture. The picture header syntax may, for example, comply with VVC.
[0420] Indicating inter-layer reference pictures
[0421] In an embodiment, the encoder:- indicates that the second coded picture uses inter-layer prediction from the independent layer.
[0422] In an embodiment, the encoder:- infers that the inter-layer prediction from the independent layer refers to the first inter-layer reference picture.
[0423] In an embodiment, the encoder:- receives a third input picture; and- encodes the third input picture into a third coded picture, wherein the third coded picture is a coded drive picture, the third coded picture resides in an independent layer, and the third coded picture follows the first coded picture in decoding order and precedes the second coded picture in decoding order.
[0424] In an embodiment, in response to the third coded picture residing in the second access unit, the encoder omits indicating that the third coded picture is coded drive picture and infers that the third coded picture is a coded drive picture.
[0425] In an embodiment, the encoder:- indicates in or along the third coded picture an indication indicating that the third coded picture is a coded drive picture.
[0426] In an embodiment, the encoder indicates with a picture header extra bit that the third coded picture is a coded drive picture. The picture header syntax may, for example, comply with VVC.
[0427] In an embodiment:- the encoder reconstructs the third decoded picture from the third coded picture; and- the third decoded picture is used in the invocation of a generative neural network inference as an additional input.
[0428] In an embodiment, the encoder:- infers when the inter-layer prediction from the independent layer refers to the first inter-layer reference picture or the third decoded picture.
[0429] In an example embodiment using VVC syntax, the first appearance of a particular ilrp_idx value in ref_pic_list_struct may be inferred to refer to the inter-layer reference picture (i.e., the interlayer reference picture generated with the previous decoded base picture as input), and when there is a second appearance of the same ilrp_idx value in the same ref_pic_list_struct, it refers to the third decoded picture (i.e., the previous decoded drive picture). In some embodiments, the third coded picture may be required to reside in the second access unit.
[0430] Embodiments related to syntax comprising encoded feature parameters and / or feature input tensor under the section "Approach based on multiple layers without generative video SEI message(s)" also apply as embodiments in this section.
[0431] In an embodiment, the encoder:- indicates in a reference picture list syntax structure information indicating that the first coded picture is used as input for generating the inter-layer reference picture.
[0432] In an embodiment, the information indicating that the first coded picture is used as input for generating the inter-layer reference picture is a picture order count difference between a current picture and the first coded picture.
[0433] In an embodiment, the reference picture list syntax structure of VVC is extended as follows, wherein sps_diagonal_inter_layer_pred_enabled_flag may be present in an SPS extension and is inferred to be equal to 0 when the SPS extension is not present:i delta_poc_ilrp_sign_flag[ listldx ][ rplsldx ][ i ] s u(l) s
[0434] abs_delta_poc_ilrp[ listldx ][ rplsldx ][ i ] specifies the absolute picture order count (POC) difference between the POC of the picture where this reference picture list is applied and the POC of the indicated inter-layer reference picture.
[0435] delta_poc_ilrp_sign_flag[ listldx ][ rplsldx ][ i ] equal to 1 specifies the POC of the indicated inter-layer reference picture is greater than the POC of the picture where this reference picture list is applied. delta_poc_ilrp_sign_flag[ listldx ][ rplsldx ][ i ] equal to 0 specifies the POC of the indicated inter-layer reference picture is less than the POC of the picture where this reference picture list is applied.
[0436] FIG. 11 is an example apparatus, which may be implemented in hardware, caused to perform integrating generative neural network in multimedia coding. The apparatus 1100 comprises atleast one processor 1102 (e.g., an FPGA and / or CPU), one or more memories 1104 including computer program code 1105, the computer program code 1105 having instructions to carry out the methods described herein, wherein the at least one memory 1104 and the computer program code 1105 are configured to, with the at least one processor 1102, cause the apparatus 1100 to implement circuitry, a process, component, module, or function (implemented with control module 1106) to implement the examples described herein, including integrating generative neural network in multimedia coding. Optionally included encoder 1180 of the control module 1106 performs encoding, and optionally included decoder 1190 implements decoding. The memory 1104 may be a non -transitory memory, a transitory memory, a volatile memory (e.g., RAM), or a non-volatile memory (e.g., ROM).
[0437] The apparatus 1100 includes a display and / or I / O interface 1108, which includes user interface (UI) circuitry and elements, that may be used to display features or a status of the methods described herein (e.g., as one of the methods is being performed or at a subsequent time), or to receive input from a user such as with using a keypad, camera, touchscreen, touch area, microphone, biometric recognition, one or more sensors, and the like. The apparatus 1100 includes one or more communication, e.g., network (N / W) interfaces (I / F(s)) 1110. The communication I / F(s) 1110 may be wired and / or wireless and communicate over the Internet / other network(s) via any communication technique including via one or more links 1124. The communication I / F(s) 1110 may comprise one or more transmitters or one or more receivers.
[0438] The transceiver 1116 comprises one or more transmitters 1118 and one or more receivers 1120. The transceiver 1116 and / or communication I / F(s) 1110 may comprise standard well-known components such as an amplifier, filter, frequency-converter, (de)modulator, and encoder / decoder circuitries and one or more antennas, such as antennas 1114 used for communication over wireless link 1122.
[0439] The control module 1106 of the apparatus 1100 comprises one of or both parts 1106-1 and / or 1106-2, which may be implemented in a number of ways. The control module 1106 may be implemented in hardware as control module 1106-1, such as being implemented as part of the one or more processors 1102. The control module 1106-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the control module 1106 may be implemented as control module 1106-2, which is implemented as computer program code (having corresponding instructions) 1105 and is executed by the one or more processors 1102. For instance, the one or more memories 1104 store instructions that, when executed by the one or more processors 1102, cause the apparatus 1100 to perform one or more of the operations as described herein. Furthermore, the one or more processors 1102, one or more memories 1104, and example algorithms (e.g., as flowcharts and / or signaling diagrams), encoded as instructions, programs, or code, are meansfor causing performance of the operations described herein.
[0440] The apparatus 1100 to implement the functionality of control 1106 may correspond to any of the apparatuses depicted herein. Alternatively, apparatus 1100 and its elements may not correspond to any of the other apparatuses depicted herein, as apparatus 1100 may be part of a self- organizing / optimizing network (SON) node or other node, such as a node in a cloud.
[0441] The apparatus 1100 may also be distributed throughout the network (e.g., internet 28) including within and between apparatus 1100 and any network element (such as a base station 24 and / or apparatus 90).
[0442] Interface 1112 enables data communication and signaling between the various items of apparatus 1100, as shown in FIG. 11. For example, the interface 1112 may be one or more buses such as address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. Computer program code (e.g., instructions) 1105, including control 1106 may comprise object- oriented software configured to pass data or messages between objects within computer program code 1105. The apparatus 1100 need not comprise each of the features mentioned, or may comprise other features as well. The various components of apparatus 1100 may at least partially reside in a common housing 1128, or a subset of the various components of apparatus 1100 may at least partially be located in different housings, which different housings may include housing 1128.
[0443] FIG. 12 shows a schematic representation of non-volatile memory media 1200a (e.g. computer / compact disc (CD) or digital versatile disc (DVD)) and 1200b (e.g. universal serial bus (USB) memory stick) and 1200c (e.g. cloud storage for downloading instructions and / or parameters 1202 or receiving emailed instructions and / or parameters 1202) storing instructions and / or parameters 1202 which when executed by a processor allows the processor to perform one or more of the operations of the methods described herein.
[0444] FIG. 13 is an example method 1300 to implement the embodiments described herein, in accordance with an embodiment. At 1302 the method 1300 includes encoding, a suffix enhancement information unit comprising a face video enhancement information message in the picture unit that comprises a base or drive picture referenced by the face video enhancement information, to generate an coded picture. At 1304 the method 1300 includes signaling the encoded picture.
[0445] The method 1300 may be performed with an apparatus described herein, for example, the any apparatus of FIG. 1 to FIG. 4, any apparatus of FIG. 11, or any other apparatus described herein.
[0446] FIG. 14 is another example method 1400 to implement the embodiments described herein, in accordance with an embodiment. At 1402 the method 1400 includes receiving a first input picture. At 1404 the method 1400 includes encoding the first input picture into a first coded picture, wherein the first coded picture comprises a coded base picture or a coded drive picture, and wherein a first picture unit comprises the first coded picture. At 1406 the method 1400 includes receiving a second input picture. At 1408 the method 1400 includes extracting feature parameters from the second input picture. At 1410 the method 1400 includes encoding the feature parameters into encoded feature parameters. At 1412 the method 1400 includes including the encoded feature parameters in the first picture unit or in a subsequent picture unit, wherein the subsequent picture unit comprises a subsequent coded picture, wherein the subsequent coded picture comprises a second coded drive picture or a second coded base picture.
[0447] The method 1400 may be performed with an apparatus described herein, for example, the any apparatus of FIG. 1 to FIG. 4, any apparatus of FIG. 11, or any other apparatus described herein.
[0448] FIG. 15 is yet another example method 1500 to implement the embodiments described herein, in accordance with an embodiment. At 1502 the method 1500 includes receiving a first coded picture from a first picture unit, wherein the first coded picture comprises a coded base picture or a coded drive picture. At 1504 the method 1500 includes decoding the first coded picture to generate a first decoded picture. At 1506 the method 1500 includes receiving encoded feature parameters from the first picture unit or from a subsequent picture unit. At 1508 the method 1500 includes creating a feature input tensor from the encoded feature parameters. At 1510 the method 1500 includes invoking a generative neural network inference with the first decoded picture and the feature input tensor as inputs to generate a picture.
[0449] The method 1500 may be performed with an apparatus described herein, for example, the any apparatus of FIG. 1 to FIG. 4, any apparatus of FIG. 11, or any other apparatus described herein.
[0450] FIG. 16 is still another example method 1600 to implement the embodiments described herein, in accordance with an embodiment. At 1602 the method 1600 includes receiving a first input picture. At 1604 the method 1600 includes encoding the first input picture into a first coded picture, wherein the first coded picture is a coded base picture, and wherein the first coded picture resides in an independent layer, and wherein a first access unit comprises the first coded picture. At 1606 the method 1600 includes receiving a second input picture. At 1608 the method 1600 includes encoding a second coded picture, wherein the second coded picture comprises a coded drive picture based on the second input picture or a coded dummy picture, and wherein the second coded picture resides in a dependentlayer, and wherein a second picture unit comprises the second coded picture, and wherein a second access unit comprises the second coded picture. At 1610 the method 1600 includes extracting feature parameters from the second input picture. At 1612 the method 1600 includes encoding the feature parameters into encoded feature parameters. At 1614 the method 1600 includes including the encoded feature parameters in the second picture unit.
[0451] The method 1600 may be performed with an apparatus described herein, for example, the any apparatus of FIG. 1 to FIG. 4, any apparatus of FIG. 11, or any other apparatus described herein.
[0452] FIG. 17 is still another example method 1700 to implement the embodiments described herein, in accordance with an embodiment. At 1702 the method 1700 includes receiving a first input picture. At 1704 the method 1700 includes encoding the first input picture into a first coded base-layer picture, wherein the first coded base-layer picture is a coded base picture, and wherein the first coded base-layer picture resides in an independent layer, and wherein a first access unit comprises the first coded base-layer picture. At 1706 the method 1700 includes encoding a first coded dependent-layer picture, wherein the first coded dependent-layer picture resides in a dependent layer, and wherein the first access unit comprises the first coded dependent-layer picture. At 1708 the method 1700 includes receiving a second input picture. At 1710 the method 1700 includes encoding a second coded baselayer picture, wherein the second coded base-layer picture comprises a coded drive picture based on the second input picture or a coded dummy picture, and wherein the second coded base-layer picture resides in the independent layer, and wherein a second access unit comprises the second coded base-layer picture. At 1712 the method 1700 includes encoding a second coded dependent-layer picture, wherein the second coded dependent-layer picture resides in a dependent layer, and the second access unit comprises the second coded dependent-layer picture. At 1714 the method 1700 includes indicating in or along the second coded dependent-layer picture that the second coded dependent-layer picture uses the first coded dependent-layer picture and the second coded base-layer picture as reference for prediction. At 1716 the method 1700 includes extracting feature parameters from the second input picture. At 1718 the method 1700 includes encoding the feature parameters into encoded feature parameters. At 1720 the method 1700 includes indicating that the encoded feature parameters are to be used for decoding of the second coded dependent-layer picture.
[0453] The method 1700 may be performed with an apparatus described herein, for example, the any apparatus of FIG. 1 to FIG. 4, any apparatus of FIG. 11, or any other apparatus described herein.
[0454] FIG. 18 is still another example method 1800 to implement the embodiments described herein, in accordance with an embodiment. At 1802 the method 1800 includes reconstructing a first decoded dependent-layer picture from a first coded dependent-layer picture. At 1804 the method 1800includes reconstructing a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependent-layer picture and a feature input tensor as inputs to generate a picture.
[0455] The method 1800 may be performed with an apparatus described herein, for example, the any apparatus of FIG. 1 to FIG. 4, any apparatus of FIG. 11, or any other apparatus described herein.
[0456] FIG. 19 is still another example method 1900 to implement the embodiments described herein, in accordance with an embodiment. At 1902 the method 1900 includes reconstructing a first decoded dependent-layer picture from a first coded dependent-layer picture. At 1904 the method 1900 includes reconstructing one or more prediction units of a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependent-layer picture and a feature input tensor as inputs. At 1906 the method 1900 includes encoding a difference between the one or more prediction units and collocated units of the second input picture as prediction residual of the second coded dependent-layer picture.
[0457] The method 1900 may be performed with an apparatus described herein, for example, the any apparatus of FIG. 1 to FIG. 4, any apparatus of FIG. 11, or any other apparatus described herein.
[0458] FIG. 20 is still another example method 2000 to implement the embodiments described herein, in accordance with an embodiment. At 2002 the method 2000 includes reconstructing a first decoded dependent-layer picture from a first coded dependent-layer picture. At 2004 the method 2000 includes invoking a generative neural network inference with the first decoded dependent-layer picture and a feature input tensor as inputs to generate an inter-layer reference picture. At 2006 the method 2000 includes encoding a second coded dependent-layer picture with reference to the inter-layer reference picture.
[0459] The method 2000 may be performed with an apparatus described herein, for example, the any apparatus of FIG. 1 to FIG. 4, any apparatus of FIG. 11, or any other apparatus described herein.
[0460] FIG. 21 is still another example method 2100 to implement the embodiments described herein, in accordance with an embodiment. At 2102 the method 2100 includes receiving a first coded base-layer picture, wherein the first coded base-layer picture is a coded base picture, the first coded base-layer picture resides in an independent layer, and a first access unit comprises the first coded baselayer picture. At 2104 the method 2100 includes decoding the first coded base-layer picture into a first decoded base-layer picture. At 2106 the method 2100 includes receiving a first coded dependent-layer picture, wherein the first coded dependent-layer picture resides in a dependent layer, and the first accessunit comprises the first coded dependent-layer picture. At 2108 the method 2100 includes decoding the first coded dependent-layer picture into a first decoded dependent-layer picture. At 2110 the method 2100 includes receiving a second coded base-layer picture, wherein the second coded base-layer picture may be a coded drive picture based on the second input picture or a coded dummy picture, the second coded base-layer picture resides in the independent layer, and a second access unit comprises the second coded base-layer picture. At 2112 the method 2100 includes, in response to the second coded base-layer picture being a coded drive picture, decoding the second coded base-layer picture into a second decoded base-layer picture. At 2114 the method 2100 includes receiving a second coded dependent-layer picture, wherein the second coded dependent-layer picture resides in a dependent layer, and the second access unit comprises the second coded dependent-layer picture. At 2116 the method 2100 includes decoding from or along the second coded dependent-layer picture that the second coded dependent-layer picture may use the first coded dependent-layer picture and the second coded base-layer picture as reference for prediction. At 2118 the method 2100 includes obtaining a feature input tensor that is in use for the decoding of the second coded dependent-layer picture. At 2120 the method 2100 includes reconstructing a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependent-layer picture and the feature input tensor as inputs to generate a picture.
[0461] The method 2100 may be performed with an apparatus described herein, for example, the any apparatus of FIG. 1 to FIG. 4, any apparatus of FIG. 11, or any other apparatus described herein.
[0462] FIG. 22 is still another example method 2200 to implement the embodiments described herein, in accordance with an embodiment. At 2202 the method 2200 includes receiving a first input picture. At 2204 the method 2200 includes encoding the first input picture into a first coded picture, wherein the first coded picture is a coded base picture or a coded drive picture, and wherein the first coded picture resides in an independent layer, and wherein a first access unit comprises the first coded picture. At 2206 the method 2200 includes receiving a second input picture. At 2208 the method 2200 includes extracting feature parameters from the second input picture. At 2210 the method 2200 includes encoding the feature parameters into encoded feature parameters, encoding a second coded picture, wherein the second coded picture resides in a dependent layer, and wherein the second access unit comprises the second coded picture. At 2212 the method 2200 includes indicating that the encoded feature parameters are used for decoding the second coded dependent-layer picture.
[0463] The method 2200 may be performed with an apparatus described herein, for example, the any apparatus of FIG. 1 to FIG. 4, any apparatus of FIG. 11, or any other apparatus described herein.
[0464] As described above, FIGs. 13 to 22 include flowcharts of an apparatus (e.g. 50, 1100, or any other apparatuses described herein), method, and computer program product according to certain example embodiments. It will be understood that each block of the flowcharts, and combinations of blocks in the flowcharts, may be implemented by various means, such as hardware, firmware, processor, circuitry, and / or other devices associated with execution of software including one or more computer program instructions. For example, one or more of the procedures described above may be embodied by computer program instructions. In this regard, the computer program instructions which embody the procedures described above may be stored by a memory (e.g. 58 or 1104) of an apparatus employing an embodiment of the present invention and executed by processing circuitry (e.g. 56 or 1102) of the apparatus. As will be appreciated, any such computer program instructions may be loaded onto a computer or other programmable apparatus (e.g., hardware) to produce a machine, such that the resulting computer or other programmable apparatus implements the functions specified in the flowchart blocks. These computer program instructions may also be stored in a computer-readable memory that may direct a computer or other programmable apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture, the execution of which implements the function specified in the flowchart blocks. The computer program instructions may also be loaded onto a computer or other programmable apparatus to cause a series of operations to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide operations for implementing the functions specified in the flowchart blocks.
[0465] A computer program product is therefore defined in those instances in which the computer program instructions, such as computer-readable program code portions, are stored by at least one non- transitory computer -readable storage medium with the computer program instructions, such as the computer-readable program code portions, being configured, upon execution, to perform the functions described above, such as in conjunction with the flowchart(s) of FIGs. 13 to 22. In other embodiments, the computer program instructions, such as the computer-readable program code portions, need not be stored or otherwise embodied by a non-transitory computer-readable storage medium, but may, instead, be embodied by a transitory medium with the computer program instructions, such as the computer- readable program code portions, still being configured, upon execution, to perform the functions described above.
[0466] Accordingly, blocks of the flowcharts support combinations of means for performing the specified functions and combinations of operations for performing the specified functions for performing the specified functions. It will also be understood that one or more blocks of the flowcharts, and combinations of blocks in the flowcharts, may be implemented by special purpose hardware-basedcomputer systems which perform the specified functions, or combinations of special purpose hardware and computer instructions.
[0467] In some embodiments, certain ones of the operations above may be modified or further amplified. Furthermore, in some embodiments, additional optional operations may be included. Modifications, additions, or amplifications to the operations above may be performed in any order and in any combination.
[0468] Some embodiments have been described in relation to base and drive pictures. It is to be understood that embodiments may be similarly realized with only base picture(s).
[0469] Some embodiments have been described in relation to a single base picture. It is to be understood embodiments may be similarly realized with multiple base pictures.
[0470] Some embodiments have been described in relation to a single drive picture. It is to be understood embodiments may be similarly realized with multiple drive pictures.
[0471] Some embodiments have been described in relation to one or more generative neural networks. It is to be understood that embodiments can be likewise realized with any neural networks, such as convolutional neural networks. For example, in some embodiments, the inference of a convolutional neural network may be regarded as an inter-layer process to generate an inter-layer reference picture.
[0472] Some embodiments have been described in relation to specific SEI messages, such as the GFV SEI message and / or NNPFA SEI message. It is to be understood that embodiments are not limited to these specific SEI messages but can be realized with any similar SEI messages.
[0473] Some embodiments have been described in relation to SEI messages. It is to be understood that embodiments are not limited to SEI messages but can be realized with any similar syntax structures, such as metadata OBUs.
[0474] Some embodiments have been described in relation to certain terms, such as picture unit, applicable to some video coding formats, such as VVC. It is to be understood that embodiments are not limited to video coding formats where the terms are applicable but can be realized with any video coding formats with their respective terms. Some examples of respective terms in different video coding formats have been described above. For example, an access unit in one video coding format may bereferred to as a temporal unit in another video coding format. In another example, a picture unit may be regarded to comprise a coded frame and associated metadata.
[0475] In the above, some example embodiments have been described with the help of syntax of the bitstream. It needs to be understood, however, that the corresponding structure and / or computer program may reside at the encoder for generating the bitstream and / or at the decoder for decoding the bitstream.
[0476] In the above, where example embodiments have been described with reference to an encoder, it needs to be understood that the resulting bitstream and the decoder have corresponding elements in them. Likewise, where example embodiments have been described with reference to a decoder, it needs to be understood that the encoder has structure and / or computer program for generating the bitstream to be decoded by the decoder.
[0477] Many modifications and other embodiments of the inventions set forth herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the inventions are not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims. Moreover, although the foregoing descriptions and the associated drawings describe example embodiments in the context of certain example combinations of elements and / or functions, it should be appreciated that different combinations of elements and / or functions may be provided by alternative embodiments without departing from the scope of the appended claims. In this regard, for example, different combinations of elements and / or functions than those explicitly described above are also contemplated as may be set forth in some of the appended claims. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
[0478] It should be understood that the foregoing description is only illustrative. Various alternatives and modifications may be devised by those skilled in the art. For example, features recited in the various dependent claims could be combined with each other in any suitable combination(s). In addition, features from different embodiments described above could be selectively combined into a new embodiment. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims.
[0479] References to a ‘computer’, ‘processor’, etc. should be understood to encompass not only computers having different architectures such as single / multi-processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device such as instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device, and the like.
[0480] As used herein, the term ‘circuitry’ may refer to any of the following: (a) hardware circuit implementations, such as implementations in analog and / or digital circuitry, and (b) combinations of circuits and software (and / or firmware), such as (as applicable): (i) a combination of processor(s) or (ii) portions of processor(s) / software including digital signal processor(s), software, and memory(ies) that work together to cause an apparatus to perform various functions, and (c) circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation, even when the software or firmware is not physically present. This description of ‘circuitry’ applies to uses of this term in this application. As a further example, as used herein, the term ‘circuitry’ would also cover an implementation of merely a processor (or multiple processors) or a portion of a processor and its (or their) accompanying software and / or firmware. The term ‘circuitry’ would also cover, for example and when applicable to the particular element, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, or another network device.
[0481] Circuitry or Circuit: As used in this application, the term ‘circuitry’ or ‘circuit’ may refer to one or more or all of the following:(a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); and(b) combinations of hardware circuits and software, such as (as applicable):(i) a combination of analog and / or digital hardware circuit(s) with software / firmware; and(ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and(c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.
[0482] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example, and when applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
Claims
CLAIMSWhat is claimed is:
1. An apparatus comprising at least one processor; and at least one non -transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: encoding, a suffix enhancement information unit comprising a face video enhancement information message in the picture unit that comprises a base or drive picture referenced by the face video enhancement information message, to generate an encoded picture; and signaling the encoded picture.
2. The apparatus of claim 1, wherein the face video enhancement information message comprises a generative video supplemental enhancement information message.
3. The apparatus of any of the claims 1 or 2, wherein the suffix enhancement information unit suffix supplemental enhancement information network abstraction layer unit.
4. An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving a first input picture; encoding the first input picture into a first coded picture, wherein the first coded picture comprises a coded base picture or a coded drive picture, and wherein a first picture unit comprises the first coded picture; receiving a second input picture; extracting feature parameters from the second input picture; encoding the feature parameters into encoded feature parameters; and including the encoded feature parameters in the first picture unit or in a subsequent picture unit, wherein the subsequent picture unit comprises a subsequent coded picture, wherein the subsequent coded picture comprises a second coded drive picture or a second coded base picture.
5. The apparatus of claim 4, wherein the apparatus is further caused to perform: including the encoded feature parameters in a generative video (GV) supplemental enhancement information (SEI) message that follows the first coded picture in decoding order.
6. The apparatus of claim 5, wherein the apparatus is further caused to perform: including includes the generative video SEI message in a suffix SEI network abstraction layer (NAL) unit within the first picture unit.
7. The apparatus of any of the claims 4 to 6, wherein when the encoded feature parameters are decoded, the encoded feature parameters invoke a generative neural network inference.
8. The apparatus of any of the claims 4 to 7, wherein the apparatus is further caused to perform: including an activation signal to invoke the generative neural network inference along the encoded feature parameters.
9. The apparatus of any of the claims 4 to 8, wherein the apparatus is further caused to perform: including an identifier or a counter value in the generative video SEI message, wherein the identifier or the counter value identifies the generative video SEI message from other generative video SEI messages in the first picture unit or the subsequent picture unit.
10. The apparatus of claim 4, wherein the apparatus is further caused to perform: generating multiple generative face video (GFV) supplemental enhancement information (SEI) messages with different payload in the first picture unit or the subsequent picture unit, wherein each instance of a GFV SEI message invokes a generative neural network (NN) inference.
11. The apparatus of claim 10, wherein the apparatus is caused to perform: using a suffix SEI NAL unit comprising a GFV SEI message in the first picture unit that comprises a base picture or a drive picture referenced by the GFV SEI message.
12. The apparatus of claim 11, wherein the apparatus is caused to perform: adding a counter syntax element in the GFV SEI message.
13. An apparatus comprising at least one processor; and at least one non -transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving a first coded picture from a first picture unit, wherein the first coded picture comprises a coded base picture or a coded drive picture; decoding the first coded picture to generate a first decoded picture; receiving encoded feature parameters from the first picture unit or from a subsequent picture unit; creating a feature input tensor from the encoded feature parameters; andinvoking a generative neural network inference with the first decoded picture and the feature input tensor as inputs to generate a picture.
14. An apparatus of claim 13, wherein the apparatus is caused to perform accepting, by an interface, less than a picture unit to be passed for decoding, and wherein decoding of the first coded picture is invoked prior to receiving or creating the feature input tensor.
15. An apparatus of claim 13, wherein the apparatus is caused to perform: accepting, by an interface , less than a picture unit to be passed for decoding, and wherein creating the feature input tensor is performed prior to receiving or decoding a subsequent coded picture from the subsequent picture unit.
16. The apparatus of any of the claims 13 to 15, wherein the apparatus is further caused to perform: decoding the encoded feature parameters; and invoking the generative neural network inference based on decoding the encoded feature parameters.
17. The apparatus of any of the claims 13 to 16, wherein the apparatus is further caused to perform: decoding an identifier or a counter value from the generative video SEI message, wherein the identifier or the counter value identifies the generative video SEI message from other generative video SEI messages in the first picture unit or the subsequent picture unit.
18. The apparatus of claim 17, wherein when the identifier or the counter value is the same as for a previous generative video SEI message in the first picture unit or the subsequent picture unit, the current generative video SEI message is concluded to be a copy of the previous generative video SEI message.
19. The apparatus of claim 13, wherein the apparatus is further caused to perform: decoding multiple generative face video (GFV) supplemental enhancement information (SEI) messages with different payload from the first picture unit or the subsequent picture unit, wherein each instance of a GFV SEI message invokes a generative neural network (NN) inference.
20. The apparatus of claim 19, wherein the apparatus is caused to perform: decoding a suffix SEI NAL unit comprising a GFV SEI message from the first picture unit that comprises a base picture or a drive picture referenced by the GFV SEI message.
21. The apparatus of claim 20, wherein the apparatus is caused to perform: decoding a counter syntax element from the GFV SEI message.
22. An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving a first input picture; encoding the first input picture into a first coded picture, wherein the first coded picture is a coded base picture, and wherein the first coded picture resides in an independent layer, and wherein a first access unit comprises the first coded picture; receiving a second input picture; encoding a second coded picture, wherein the second coded picture comprises a coded drive picture based on the second input picture or a coded dummy picture, and wherein the second coded picture resides in a dependent layer, and wherein a second picture unit comprises the second coded picture, and wherein a second access unit comprises the second coded picture; extracting feature parameters from the second input picture; encoding the feature parameters into encoded feature parameters; and including the encoded feature parameters in the second picture unit.
23. The apparatus of claim 22, wherein the apparatus is further caused to perform: encoding a coded replica picture in the first access unit, wherein the coded replica picture resides in the dependent layer and is predicted from the independent layer, and wherein the coded replica picture comprises no prediction residual.
24. The apparatus of any of the claims 22 or 23, wherein the apparatus is further caused to perform: including the encoded feature parameters in a generative video supplemental enhancement information (SEI) message in the second picture unit.
25. The apparatus of any of the claims 22 to 24, wherein the apparatus is further caused to perform: predicting the coded dummy picture from a previous picture, in decoding order, within the dependent layer.
26. An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving a first input picture;encoding the first input picture into a first coded base-layer picture, wherein the first coded base-layer picture is a coded base picture, and wherein the first coded base-layer picture resides in an independent layer, and wherein a first access unit comprises the first coded baselayer picture; encoding a first coded dependent-layer picture, wherein the first coded dependent-layer picture resides in a dependent layer, and wherein the first access unit comprises the first coded dependent-layer picture; receiving a second input picture; encoding a second coded base-layer picture, wherein the second coded base-layer picture comprises a coded drive picture based on the second input picture or a coded dummy picture, and wherein the second coded base-layer picture resides in the independent layer, and wherein a second access unit comprises the second coded base-layer picture; encoding a second coded dependent-layer picture, wherein the second coded dependent-layer picture resides in the dependent layer, and the second access unit comprises the second coded dependent-layer picture; indicating in or along the second coded dependent-layer picture that the second coded dependent-layer picture uses the first coded dependent-layer picture and the second coded baselayer picture as reference for prediction; extracting feature parameters from the second input picture; encoding the feature parameters into encoded feature parameters; and indicating that the encoded feature parameters are to be used for decoding of the second coded dependent-layer picture.
27. The apparatus of claim 26, wherein the apparatus is further caused to perform encoding the first dependent-layer picture as a coded replica picture, wherein the coded replica picture is predicted from the independent layer, and wherein the coded replica picture no prediction residual.
28. The apparatus of any of the claims 26 or 27, wherein the apparatus is further caused to perform: encoding an indication in or along the second coded base-layer picture to omit output of a reconstructed picture decoded from the second coded base-layer picture.
29. The apparatus of any of the claims 26 or 27, wherein the apparatus is further caused to perform: indicating that a first reconstructed base-layer picture decoded from the first coded base-layer picture is output;indicating that a first reconstructed dependent-layer picture decoded from the first coded dependent-layer picture is output; indicating that a second reconstructed dependent-layer picture decoded from the second coded dependent-layer picture is output; indicating that an independent layer and a dependent layer form an output layer set and the dependent layer is an output layer of the output layer set; or indicating a profile to the output layer set, wherein the profile includes the generative neural network capability.
30. The apparatus of any of the claims 26 to 28, wherein the apparatus is further caused to perform: creating a feature input tensor from the encoded feature parameters.
31. The apparatus of any of the claims 26 to 30, wherein the apparatus is further caused to perform: including the encoded feature parameters or the feature input tensor in an adaptation parameter set with a type indicating generative video; or including the encoded feature parameters or the feature input tensor in a picture header of the second coded dependent-layer picture.
32. An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: reconstructing a first decoded dependent-layer picture from a first coded dependentlayer picture; and reconstructing a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependent-layer picture and a feature input tensor as inputs to generate a picture.
33. The apparatus of claim 32, wherein a syntax of the second coded dependent-layer picture excludes syntax elements that are not relevant to reconstructing the second decoded dependent-layer picture with the generative neural network.
34. The apparatus of claim 32, a syntax of one or more syntax structures of the second dependent-layer excludes syntax elements that are not relevant for reconstructing the second decoded dependent-layer picture with the generative neural network.
35. An apparatus comprising at least one processor; and at least one non -transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: reconstructing a first decoded dependent-layer picture from a first coded dependentlayer picture; reconstructing one or more prediction units of a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependentlayer picture and a feature input tensor as inputs; and encoding a difference between the one or more prediction units and collocated units of the second input picture as prediction residual of the second coded dependent-layer picture.
36. An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: reconstructing a first decoded dependent-layer picture from a first coded dependentlayer picture; invoking a generative neural network inference with the first decoded dependentlayer picture and a feature input tensor as inputs to generate an inter-layer reference picture; and encoding a second coded dependent-layer picture with reference to the inter-layer reference picture.
37. The apparatus of claim 36, wherein the apparatus is further caused to perform: encoding difference between the second input picture and the inter-layer reference picture as prediction residual of the second coded dependent-layer picture.
38. An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving a first coded base-layer picture, wherein the first coded base-layer picture is a coded base picture, the first coded base-layer picture resides in an independent layer, and a first access unit comprises the first coded base-layer picture; decoding the first coded base-layer picture into a first decoded base-layer picture; receiving a first coded dependent-layer picture, wherein the first coded dependent-layer picture resides in a dependent layer, and the first access unit comprises the first coded dependent-layer picture;decoding the first coded dependent-layer picture into a first decoded dependent-layer picture; receiving a second coded base-layer picture, wherein the second coded base-layer picture may be a coded drive picture based on the second input picture or a coded dummy picture, the second coded base-layer picture resides in the independent layer, and a second access unit comprises the second coded base-layer picture; in response to the second coded base-layer picture being the coded drive picture, decoding the second coded base-layer picture into a second decoded base-layer picture; receiving a second coded dependent-layer picture, wherein the second coded dependent-layer picture resides in the dependent layer, and the second access unit comprises the second coded dependent-layer picture; decoding from or along the second coded dependent-layer picture that the second coded dependent-layer picture may use the first coded dependent-layer picture and the second coded base-layer picture as reference for prediction; obtaining a feature input tensor that is in use for the decoding of the second coded dependent-layer picture; and reconstructing a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependent-layer picture and the feature input tensor as inputs to generate a picture.
39. An apparatus comprising at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving a first input picture; encoding the first input picture into a first coded picture, wherein the first coded picture is a coded base picture or a coded drive picture, and wherein the first coded picture resides in an independent layer, and wherein a first access unit comprises the first coded picture; receiving a second input picture; extracting feature parameters from the second input picture; encoding the feature parameters into encoded feature parameters; encoding a second coded picture, wherein the second coded picture resides in a dependent layer, and wherein the second access unit comprises the second coded picture; and indicating that the encoded feature parameters are used for decoding the second coded dependent-layer picture.
40. The apparatus of claim 39, wherein the apparatus is further caused to perform: reconstructing a first decoded picture from the first coded picture;creating a feature input tensor from the encoded feature parameters; invoking a generative neural network inference with the first decoded picture and the feature input tensor as inputs to generate a first inter-layer reference picture; and encoding the second coded picture with reference to the first inter-layer reference picture.
41. The apparatus of any of the claims 39 or 40, wherein the apparatus is further caused to perform: omitting encoding prediction residual for the second coded picture.
42. The apparatus of any of the claims 39 to 41, wherein the apparatus is further caused to perform: encoding difference between the second input picture and the first inter-layer reference picture as prediction residual of the second coded picture.
43. The apparatus of any of the claims 39 to 41, wherein the apparatus is further caused to perform: in response to the first coded picture being the coded base picture, indicating that the first coded picture is an output; in response to the first coded picture being the coded drive picture, indicating that the first coded picture is not the output; indicating that the independent layer and the dependent layer form an output layer set and the dependent layer is an output layer of the output layer set; or indicating that the second coded picture is output.
44. The apparatus of any of the claims 39 to 43, wherein the apparatus is further caused to perform: indicating in or along the first coded picture an indication indicating that the first coded picture is the coded base picture or that the first coded picture is the coded drive picture; or indicating with a picture header extra bit that the first coded picture is the coded base picture or that the first coded picture is the coded drive picture.
45. The apparatus of any of the claims 39 to 44, wherein the apparatus is further caused to perform: indicating that the second coded picture uses inter-layer prediction from the independent layer.
46. The apparatus of any of the claims 39 to 44, wherein the apparatus is further caused to perform: inferring that the inter-layer prediction from the independent layer refers to the first inter-layer reference picture.
47. The apparatus of any of the claims 39 to 44, wherein the apparatus is further caused to perform: receiving a third input picture; and encoding the third input picture into a third coded picture, wherein the third coded picture is the coded drive picture, and wherein the third coded picture resides in the independent layer, and the third coded picture follows the first coded picture in decoding order and precedes the second coded picture in the decoding order.
48. The apparatus of claim 47, wherein the apparatus is further caused to perform: in response to the third coded picture residing in the second access unit, omitting indicating that the third coded picture is the coded drive picture and infers that the third coded picture is the coded drive picture.
49. The apparatus of claim 47, wherein the apparatus is further caused to perform: indicating in or along the third coded picture an indication indicating that the third coded picture is the coded drive picture.
50. The apparatus of claim 47, wherein the apparatus is further caused to perform: indicating using a picture header extra bit that the third coded picture is the coded drive picture.
51. The apparatus of any of the claims 47 to 50, wherein the apparatus is further caused to perform: reconstructing the third decoded picture from the third coded picture; and using the third decoded picture in the invocation of the generative neural network inference as an additional input.
52. The apparatus of any of the claims 47 to 51, wherein the apparatus is further caused to perform: inferring when the inter-layer prediction from the independent layer refers to the first inter-layer reference picture or the third decoded picture.
53. The apparatus of any of the claims 47 to 51, wherein the apparatus is further caused to perform: indicating in a reference picture list syntax structure information indicating that the first coded picture is used as input for generating the inter-layer reference picture.
54. The apparatus of claim 53, wherein the information indicating that the first coded picture is used as input for generating the inter-layer reference picture comprises a picture order count difference between a current picture and the first coded picture.
55. A method comprising: encoding, a suffix enhancement information unit comprising a face video enhancement information message in the picture unit that comprises a base or drive picture referenced by the face video enhancement information message, to generate an encoded picture; and signaling the encoded picture.
56. The method of claim 55, wherein the face video enhancement information message comprises a generative video supplemental enhancement information message.
57. The method of any of the claims 55 or 56, wherein the suffix enhancement information unit suffix supplemental enhancement information network abstraction layer unit.
58. A method comprising: receiving a first input picture; encoding the first input picture into a first coded picture, wherein the first coded picture comprises a coded base picture or a coded drive picture, and wherein a first picture unit comprises the first coded picture; receiving a second input picture; extracting feature parameters from the second input picture; encoding the feature parameters into encoded feature parameters; and including the encoded feature parameters in the first picture unit or in a subsequent picture unit, wherein the subsequent picture unit comprises a subsequent coded picture, wherein the subsequent coded picture comprises a second coded drive picture or a second coded base picture.
59. The method of claim 58 further comprising: including the encoded feature parameters in a generative video (GV) supplemental enhancement information (SEI) message that follows the first coded picture in decoding order.
60. The method of claim 59 further comprising: including includes the generative video SEI message in a suffix SEI network abstraction layer (NAL) unit within the first picture unit.
61. The method of any of the claims 58 to 60, wherein when the encoded feature parameters are decoded, the encoded feature parameters invoke a generative neural network inference.
62. The method of any of the claims 58 to 61 further comprising: including an activation signal to invoke the generative neural network inference along the encoded feature parameters.
63. The method of any of the claims 58 to 62 further comprising: including an identifier or a counter value in the generative video SEI message, wherein the identifier or the counter value identifies the generative video SEI message from other generative video SEI messages in the first picture unit or the subsequent picture unit.
64. The method of 58 further comprising: generating multiple generative face video (GFV) supplemental enhancement information (SEI) messages with different payload in the first picture unit or the subsequent picture unit, wherein each instance of a GFV SEI message invokes a generative neural network (NN) inference.
65. The method of claim 64 further comprising: using a suffix SEI NAL unit comprising a GFV SEI message in the first picture unit that comprises a base picture or a drive picture referenced by the GFV SEI message.
66. The method of claim 65 further comprising: adding a counter syntax element in the GFV SEI message.
67. A method comprising: receiving a first coded picture from a first picture unit, wherein the first coded picture comprises a coded base picture or a coded drive picture; decoding the first coded picture to generate a first decoded picture; receiving encoded feature parameters from the first picture unit or from a subsequent picture unit; creating a feature input tensor from the encoded feature parameters; and invoking a generative neural network inference with the first decoded picture and the feature input tensor as inputs to generate a picture.
68. A method of claim 67 further comprising: accepting, by an interface, less than a picture unit to be passed for decoding, and wherein decoding of the first coded picture is invoked prior to receiving or creating the feature input tensor.
69. A method of claim 67 accepting, by an interface, less than a picture unit to be passed for decoding, and wherein creating the feature input tensor is performed prior to receiving or decoding a subsequent coded picture from the subsequent picture unit.
70. The method of any of the claims 67 to 69 further comprising: decoding the encoded feature parameters; and invoking the generative neural network inference based on decoding the encoded feature parameters.
71. The method of any of the claims 67 to 70 further comprising: decoding an identifier or a counter value from the generative video SEI message, wherein the identifier or the counter value identifies the generative video SEI message from other generative video SEI messages in the first picture unit or the subsequent picture unit.
72. The method of claim 71, wherein when the identifier or the counter value is the same as for a previous generative video SEI message in the first picture unit or the subsequent picture unit, the current generative video SEI message is concluded to be a copy of the previous generative video SEI message.
73. The method of 67 further comprising: decoding multiple generative face video (GFV) supplemental enhancement information (SEI) messages with different payload from the first picture unit or the subsequent picture unit, wherein each instance of a GFV SEI message invokes a generative neural network (NN) inference.
74. The method of claim 73 further comprising: decoding a suffix SEI NAL unit comprising a GFV SEI message from the first picture unit that comprises a base picture or a drive picture referenced by the GFV SEI message.
75. The method of claim 74 further comprising: decoding a counter syntax element from the GFV SEI message.
76. A method comprising: receiving a first input picture; encoding the first input picture into a first coded picture, wherein the first coded picture is a coded base picture, and wherein the first coded picture resides in an independent layer, and wherein a first access unit comprises the first coded picture; receiving a second input picture;encoding a second coded picture, wherein the second coded picture comprises a coded drive picture based on the second input picture or a coded dummy picture, and wherein the second coded picture resides in a dependent layer, and wherein a second picture unit comprises the second coded picture, and wherein a second access unit comprises the second coded picture; extracting feature parameters from the second input picture; encoding the feature parameters into encoded feature parameters; and including the encoded feature parameters in the second picture unit.
77. The method of claim 76 further comprising: encoding a coded replica picture in the first access unit, wherein the coded replica picture resides in the dependent layer and is predicted from the independent layer, and wherein the coded replica picture comprises no prediction residual.
78. The method of any of the claims 76 or 77 further comprising: including the encoded feature parameters in a generative video supplemental enhancement information (SEI) message in the second picture unit.
79. The method of any of the claims 76 to 78 further comprising: predicting the coded dummy picture from a previous picture, in decoding order, within the dependent layer.
80. A method comprising: receiving a first input picture; encoding the first input picture into a first coded base-layer picture, wherein the first coded base-layer picture is a coded base picture, and wherein the first coded base-layer picture resides in an independent layer, and wherein a first access unit comprises the first coded baselayer picture; encoding a first coded dependent-layer picture, wherein the first coded dependent-layer picture resides in a dependent layer, and wherein the first access unit comprises the first coded dependent-layer picture; receiving a second input picture; encoding a second coded base-layer picture, wherein the second coded base-layer picture comprises a coded drive picture based on the second input picture or a coded dummy picture, and wherein the second coded base-layer picture resides in the independent layer, and wherein a second access unit comprises the second coded base-layer picture;encoding a second coded dependent-layer picture, wherein the second coded dependent-layer picture resides in the dependent layer, and the second access unit comprises the second coded dependent-layer picture; indicating in or along the second coded dependent-layer picture that the second coded dependent-layer picture uses the first coded dependent-layer picture and the second coded baselayer picture as reference for prediction; extracting feature parameters from the second input picture; encoding the feature parameters into encoded feature parameters; and indicating that the encoded feature parameters are to be used for decoding of the second coded dependent-layer picture.
81. The method of claim 80 further comprising: encoding the first dependent-layer picture as a coded replica picture, wherein the coded replica picture is predicted from the independent layer, and wherein the coded replica picture no prediction residual.
82. The method of any of the claims 80 or 81 further comprising: encoding an indication in or along the second coded base-layer picture to omit output of a reconstructed picture decoded from the second coded base-layer picture.
83. The method of any of the claims 80 or 81 further comprising: indicating that a first reconstructed base-layer picture decoded from the first coded base-layer picture is output; indicating that a first reconstructed dependent-layer picture decoded from the first coded dependent-layer picture is output; indicating that a second reconstructed dependent-layer picture decoded from the second coded dependent-layer picture is output; indicating that an independent layer and a dependent layer form an output layer set and the dependent layer is an output layer of the output layer set; or indicating a profile to the output layer set, wherein the profile includes the generative neural network capability.
84. The method of any of the claims 80 to 82 further comprising: creating a feature input tensor from the encoded feature parameters.
85. The method of any of the claims 80 to 84 further comprising: including the encoded feature parameters or the feature input tensor in an adaptation parameter set with a type indicating generative video; orincluding the encoded feature parameters or the feature input tensor in a picture header of the second coded dependent-layer picture.
86. A method comprising: reconstructing a first decoded dependent-layer picture from a first coded dependentlayer picture; and reconstructing a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependent-layer picture and a feature input tensor as inputs to generate a picture.
87. The method of claim 86, wherein a syntax of the second coded dependent-layer picture excludes syntax elements that are not relevant to reconstructing the second decoded dependentlayer picture with the generative neural network.
88. The method of claim 86, a syntax of one or more syntax structures of the second dependent-layer excludes syntax elements that are not relevant for reconstructing the second decoded dependent-layer picture with the generative neural network.
89. A method comprising: reconstructing a first decoded dependent-layer picture from a first coded dependentlayer picture; reconstructing one or more prediction units of a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependentlayer picture and a feature input tensor as inputs; and encoding a difference between the one or more prediction units and collocated units of the second input picture as prediction residual of the second coded dependent-layer picture.
90. A method comprising: reconstructing a first decoded dependent-layer picture from a first coded dependentlayer picture; invoking a generative neural network inference with the first decoded dependentlayer picture and a feature input tensor as inputs to generate an inter-layer reference picture; and encoding a second coded dependent-layer picture with reference to the inter-layer reference picture.
91. The method of claim 90 further comprising: encoding difference between the second input picture and the inter-layer reference picture as prediction residual of the second coded dependent-layer picture.
92. A method comprising: receiving a first coded base-layer picture, wherein the first coded base-layer picture is a coded base picture, the first coded base-layer picture resides in an independent layer, and a first access unit comprises the first coded base-layer picture; decoding the first coded base-layer picture into a first decoded base-layer picture; receiving a first coded dependent-layer picture, wherein the first coded dependent-layer picture resides in a dependent layer, and the first access unit comprises the first coded dependent-layer picture; decoding the first coded dependent-layer picture into a first decoded dependent-layer picture; receiving a second coded base-layer picture, wherein the second coded base-layer picture may be a coded drive picture based on the second input picture or a coded dummy picture, the second coded base-layer picture resides in the independent layer, and a second access unit comprises the second coded base-layer picture; in response to the second coded base-layer picture being the coded drive picture, decoding the second coded base-layer picture into a second decoded base-layer picture; receiving a second coded dependent-layer picture, wherein the second coded dependent-layer picture resides in the dependent layer, and the second access unit comprises the second coded dependent-layer picture; decoding from or along the second coded dependent-layer picture that the second coded dependent-layer picture may use the first coded dependent-layer picture and the second coded base-layer picture as reference for prediction; obtaining a feature input tensor that is in use for the decoding of the second coded dependent-layer picture; and reconstructing a second decoded dependent-layer picture by invoking a generative neural network inference with the first decoded dependent-layer picture and the feature input tensor as inputs to generate a picture.
93. A method comprising: receiving a first input picture; encoding the first input picture into a first coded picture, wherein the first coded picture is a coded base picture or a coded drive picture, and wherein the first coded picture resides in an independent layer, and wherein a first access unit comprises the first coded picture;receiving a second input picture; extracting feature parameters from the second input picture; encoding the feature parameters into encoded feature parameters, encoding a second coded picture, wherein the second coded picture resides in a dependent layer, and wherein the second access unit comprises the second coded picture; and indicating that the encoded feature parameters are used for decoding the second coded dependent-layer picture.
94. The method of claim 93 further comprising: reconstructing a first decoded picture from the first coded picture; creating a feature input tensor from the encoded feature parameters; invoking a generative neural network inference with the first decoded picture and the feature input tensor as inputs to generate a first inter-layer reference picture; and encoding the second coded picture with reference to the first inter-layer reference picture.
95. The method of any of the claims 93 or 94 further comprising: omitting encoding prediction residual for the second coded picture.
96. The method of any of the claims 93 to 95 further comprising: encoding difference between the second input picture and the first inter-layer reference picture as prediction residual of the second coded picture.
97. The method of any of the claims 93 to 95 further comprising: in response to the first coded picture being the coded base picture, indicating that the first coded picture is an output; in response to the first coded picture being the coded drive picture, indicating that the first coded picture is not the output; indicating that the independent layer and the dependent layer form an output layer set and the dependent layer is an output layer of the output layer set; or indicating that the second coded picture is output.
98. The method of any of the claims 93 to 97 further comprising: indicating in or along the first coded picture an indication indicating that the first coded picture is the coded base picture or that the first coded picture is the coded drive picture; orindicating with a picture header extra bit that the first coded picture is the coded base picture or that the first coded picture is the coded drive picture.
99. The method of any of the claims 93 to 98 further comprising: indicating that the second coded picture uses inter-layer prediction from the independent layer.
100. The method of any of the claims 93 to 98 further comprising: inferring that the interlayer prediction from the independent layer refers to the first inter-layer reference picture.
101. The method of any of the claims 93 to 98 further comprising: receiving a third input picture; and encoding the third input picture into a third coded picture, wherein the third coded picture is the coded drive picture, and wherein the third coded picture resides in the independent layer, and the third coded picture follows the first coded picture in decoding order and precedes the second coded picture in the decoding order.
102. The method of claim 101 further comprising: in response to the third coded picture residing in the second access unit, omitting indicating that the third coded picture is the coded drive picture and infers that the third coded picture is the coded drive picture.
103. The method of claim 101 further comprising: indicating in or along the third coded picture an indication indicating that the third coded picture is the coded drive picture.
104. The method of claim 101 further comprising: indicating using a picture header extra bit that the third coded picture is the coded drive picture.
105. The method of any of the claims 101 to 104 further comprising: reconstructing the third decoded picture from the third coded picture; and using the third decoded picture in the invocation of the generative neural network inference as an additional input.
106. The method of any of the claims 101 to 105 further comprising: inferring when the inter-layer prediction from the independent layer refers to the first inter-layer reference picture or the third decoded picture.
107. The method of any of the claims 101 to 105 further comprising: indicating in areference picture list syntax structure information indicating that the first coded picture is used as input for generating the inter-layer reference picture.
108. The method of claim 107, wherein the information indicating that the first coded picture is used as input for generating the inter-layer reference picture comprises a picture order count difference between a current picture and the first coded picture.
109. An apparatus comprising means for performing the methods as claimed in any of the claims 55 to 108.
110. A computer readable medium comprising program instructions which, when executed by an apparatus, cause the apparatus to perform the methods as claimed in any of the claims 55 to 108.
111. The computer readable medium of claim 110, wherein the computer readable medium comprises a non-transitory computer readable medium.