Adaptive input picture selection in post filter groups

EP4740481A1Pending Publication Date: 2026-05-13NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
NOKIA TECHNOLOGIES OY
Filing Date
2024-06-13
Publication Date
2026-05-13

AI Technical Summary

Technical Problem

Existing multimedia systems face challenges in adaptively selecting and applying post-processing filters, particularly when previous filters output a different number of pictures than input for current filters, or when applying filters in a hierarchical manner, leading to inefficiencies in data processing and quality enhancement.

Method used

An apparatus and method that receive signaling information about post-processing filters, using this information to infer how to apply filters, including neural network post-filters, for adaptive input picture selection and hierarchical processing, ensuring appropriate filter usage based on input and output conditions.

Benefits of technology

Enables efficient and adaptive application of post-processing filters, improving data processing and quality enhancement by accurately selecting input pictures and applying filters in a hierarchical manner, even when previous filters output differently than expected, thereby optimizing multimedia data handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024055806_09012025_PF_FP_ABST
    Figure IB2024055806_09012025_PF_FP_ABST
Patent Text Reader

Abstract

An apparatus includes circuitry configured to: receive an encoding of at least one picture; receive signaling comprising information related to a group of at least one post-processing filter; and use the information related to the group of the at least one post-processing filter to infer whether and how to use at least one post-processing filter in the group of the at least one post-processing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one post-processing filter in the group when a previous post-processing filter in a cascade outputs a different number of pictures than what is used as input for a current post-processing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.
Need to check novelty before this filing date? Find Prior Art

Description

ADAPTIVE INPUT PICTURE SELECTION IN POST FILTER GROUPSTECHNICAL FIELD

[0001] The examples and non-limiting embodiments relate generally to multimedia transport and, more particularly, to signaling information about multiple post processing filters.BACKGROUND

[0002] It is known to perform data compression and decoding in a multimedia system.SUMMARY

[0003] Example 1: An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, causes the apparatus at least to: receive an encoding of at least one picture; receive signaling comprising information related to a group of at least one post-processing filter; and use the information related to the group of the at least one post-processing filter to infer whether and how to use at least one post-processing filter in the group of the at least one post-processing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one post-processing filter in the group when a previous post-processing filter in a cascade outputs a different number of pictures than what is used as input for a current post-processing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.

[0004] Example 2: The apparatus of example 1, wherein the at least one post-processing filter comprises a neural network post-filter.

[0005] Example 3: The apparatus of any of examples 1 to 2, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: apply a picture rate upsampling post-processing filter to obtain a second frequency from an input having a first frequency; and apply the picture rate upsampling post-processing filter to obtain a third frequency from an input having the second frequency; wherein the signaled information indicates that the picture rate upsampling post-processing filter is to be applied to obtain the second frequency from the input having the first frequency, and that the picture rate upsampling post-processing filter is to be applied to obtain the third frequency from the input having the second frequency.

[0006] Example 4: The apparatus of any of examples 1 to 3, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: decode one or more indications identifying input pictures used as an input for a post-processing filter in the group of the at least onepost-processing filter.

[0007] Example 5: The apparatus of any of examples 1 to 4, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: decode one or more indications identifying input pictures used as input for each post-processing filter in the group of the at least one post-processing filter.

[0008] Example 6: The apparatus of any of examples 1 to 5, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: infer one or more input pictures for an initial post-processing filter in the group of the at least one post-processing filter; and decode one or more indications identifying input pictures for each subsequent post-processing filter in the group; wherein each subsequent post-processing filter in the group follows the initial postprocessing filter.

[0009] Example 7: The apparatus of any of examples 1 to 6, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: decode one or more indications identifying input pictures for a post-processing filter in the group of the at least one post-processing filter from a neural-network post-filter activation supplemental enhancement information message that is comprised in a processing order nesting supplemental enhancement information message.

[0010] Example 8: The apparatus of any of examples 1 to 7, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: decode one or more indications identifying input pictures for a post-processing filter in the group of the at least one post-processing filter from a neural-network post-filter extended activation supplemental enhancement information message.

[0011] Example 9: The apparatus of any of examples 1 to 8, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: derive a list of candidate input pictures for a second or later post-processing filter in the group of the at least one post-processing filter from filtered pictures that are output by previous post-processing filters in the the group of the at least one post-processing filter, when any; interpolated pictures that are output by a postprocessing filter process of previous post-processing filters in the group of the at least one postprocessing filter, when any; and candidate input pictures for a first post-processing filter in the the group of the at least one post-processing filter.

[0012] Example 10: The apparatus of example 9, wherein an order of pictures in the list of candidate input pictures is pre-defined.

[0013] Example 11: The apparatus of any of examples 9 to 10, wherein an order of pictures inthe list of candidate input pictures is an inverse output order.

[0014] Example 12: The apparatus of any of examples 9 to 11, wherein the candidate input pictures within the list of candidate input pictures are non-overlapping in output time, the list comprising up to one picture per an output time.

[0015] Example 13: The apparatus of any of examples 9 to 12, wherein a picture resulting from a subsequent post-processing filter in the group of the at least one post-processing filter precedes a picture resulting from a previous post-processing filter in the group of the at least one postprocessing filter in the list of candidate input pictures, when the picture resulting from the subsequent post-processing filter in the the group of the at least one post-processing filter has a same output order as the picture resulting from the previous post-processing filter in the group of the at least one post-processing filter.

[0016] Example 14: The apparatus of any of examples 1 to 13, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: decode, given a list of candidate input pictures for a post-processing filter in the group of the at least one post-processing filter, one or more indications that indicate which of the candidate input pictures in the list are selected as input pictures for the post-processing filter in the group of the at least one post-processing filter.

[0017] Example 15: The apparatus of example 14, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: decode, given the list of candidate input pictures for the post-processing filter in the group of the at least one post-processing filter, one or both of: an indication for the post-processing filter when all the pictures in the list of candidate input pictures are input pictures to the post-processing filter, or one or more skip counts corresponding to how many pictures in the list of candidate input pictures are skipped when selecting pictures from the list of candidate input pictures to be used as input pictures.

[0018] Example 16: The apparatus of any of examples 14 to 15, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: decode, given the list of candidate input pictures for the post-processing filter in the group of the at least one post-processing filter, a bit mask, where a bit position in the bit mask corresponds to a picture in the list of candidate input pictures, and a value of a bit indicates whether a picture in a respective position within the list of candidate input pictures is selected as an input picture to the post-processing filter in the group of the at least one post-processing filter.

[0019] Example 17: The apparatus of any of examples 1 to 16, wherein the instructions, whenexecuted by the at least one processor, cause the apparatus at least to: decode a description of the group of the at least one post-processing filter that comprises more than one occurrence of the same post-processing filter.

[0020] Example 18: An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, causes the apparatus at least to: include, in a bitstream, an encoding of at least one picture; determine information related to a group of at least one post-processing filter; and indicate, in the bitstream, the information related to the group of the at least one post-processing filter; wherein the information related to the group of the at least one post-processing filter is configured to be used to infer whether and how to use at least one post-processing filter in the group of the at least one post-processing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one post-processing filter in the group when a previous post-processing filter in a cascade outputs a different number of pictures than what is used as input for a current post-processing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.

[0021] Example 19: The apparatus of example 18, wherein the at least one post-processing filter comprises a neural network post-filter.

[0022] Example 20: The apparatus of any of examples 18 to 19, wherein the signaled information indicates that a picture rate upsampling post-processing filter is to be applied to obtain a second frequency from an input having a first frequency, and that the picture rate upsampling postprocessing filter is to be applied to obtain a third frequency from an input having the second frequency.

[0023] Example 21: The apparatus of any of examples 18 to 20, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: encode one or more indications identifying input pictures used as input for a post-processing filter in the group of the at least one post-processing filter.

[0024] Example 22: The apparatus of any of examples 18 to 21, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: encode one or more indications identifying input pictures used as input for each post-processing filter in the group of the at least one post-processing filter.

[0025] Example 23: The apparatus of any of examples 18 to 22, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: infer one or more input pictures for an initial post-processing filter in the group of the at least one post-processing filter; and encodeone or more indications identifying input pictures for each subsequent post-processing filter in the group; wherein each subsequent post-processing filter in the group follows the initial postprocessing filter.

[0026] Example 24: The apparatus of any of examples 18 to 23, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: encode one or more indications identifying input pictures for a post-processing filter in the group of the at least one post-processing filter from a neural-network post-filter activation supplemental enhancement information message that is comprised in a processing order nesting supplemental enhancement information message.

[0027] Example 25: The apparatus of any of examples 18 to 24, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: encode one or more indications identifying input pictures for a post-processing filter in the group of the at least one post-processing filter from a neural-network post-filter extended activation supplemental enhancement information message.

[0028] Example 26: The apparatus of any of examples 18 to 25, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: derive a list of candidate input pictures for a second or later post-processing filter in the group of the at least one post-processing filter from filtered pictures that are output by previous post-processing filters in the group of the at least one post-processing filter, when any, interpolated pictures that are output by a post-processing filter process of previous post-processing filters in the group of the at least one post-processing filter, when any, and candidate input pictures for a first post-processing filter in the group of the at least one post-processing filter.

[0029] Example 27: The apparatus of example 26, wherein an order of pictures in the list of candidate input pictures is pre-defined.

[0030] Example 28: The apparatus of any of examples 26 to 27, wherein an order of pictures in the list of candidate input pictures is an inverse output order.

[0031] Example 29: The apparatus of any of examples 26 to 28, wherein the candidate input pictures within the list of candidate input pictures are non-overlapping in output time, the list comprising up to one picture per an output time.

[0032] Example 30: The apparatus of any of examples 26 to 29, wherein a picture resulting from a subsequent post-processing filter in the group of the at least one post-processing filter precedes a picture resulting from a previous post-processing filter in the group of the at least one postprocessing filter in the list of candidate input pictures, when the picture resulting from the subsequentpost-processing filter in the group of the at least one post-processing filter has a same output order as the picture resulting from the previous post -processing filter in the group of the at least one postprocessing filter.

[0033] Example 31: The apparatus of any of examples 18 to 30, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: encode, given a list of candidate input pictures for a post-processing filter in the group of the at least one post-processing filter, one or more indications that indicate which of the candidate input pictures in the list are selected as input pictures for the post-processing filter in the group of the at least one post-processing filter.

[0034] Example 32: The apparatus of example 31, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: encode, given the list of candidate input pictures for the post-processing filter in the group of the at least one post-processing filter, one or both of: an indication for the post-processing filter when all the pictures in the list of candidate input pictures are input pictures to the post-processing filter, or one or more skip counts corresponding to how many pictures in the list of candidate input pictures are skipped when selecting pictures from the list of candidate input pictures to be used as input pictures.

[0035] Example 33: The apparatus of any of examples 31 to 32, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: encode, given the list of candidate input pictures for the post-processing filter in the group of the at least one post-processing filter, a bit mask, where a bit position in the bit mask corresponds to a picture in the list of candidate input pictures, and a value of a bit indicates whether a picture in a respective position within the list of candidate input pictures is selected as an input picture to the post-processing filter in the group of the at least one post-processing filter.

[0036] Example 34: The apparatus of any of examples 18 to 33, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: encode a description of the group of the at least one post-processing filter that comprises more than one occurrence of the same post-processing filter.

[0037] Example 35: A method comprising: receiving an encoding of at least one picture; receiving signaling comprising information related to a group of at least one post-processing filter; and using the information related to the group of the at least one post-processing filter to infer whether and how to use at least one post-processing filter in the group of the at least one postprocessing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one post-processing filter in the group when a previous post-processingfilter in a cascade outputs a different number of pictures than what is used as input for a current postprocessing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.

[0038] Example 36: A method comprising: including, in a bitstream, an encoding of at least one picture; determining information related to a group of at least one post-processing filter; and indicating, in the bitstream, the information related to the group of the at least one post-processing filter; wherein the information related to the group of the at least one post-processing filter is configured to be used to infer whether and how to use at least one post-processing filter in the group of the at least one post-processing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one post-processing filter in the group when a previous post-processing filter in a cascade outputs a different number of pictures than what is used as input for a current post-processing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.

[0039] Example 37: An apparatus comprising: means for receiving an encoding of at least one picture; means for receiving signaling comprising information related to a group of at least one postprocessing filter; and means for using the information related to the group of the at least one postprocessing filter to infer whether and how to use at least one post-processing filter in the group of the at least one post-processing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one post-processing filter in the group when a previous post-processing filter in a cascade outputs a different number of pictures than what is used as input for a current post-processing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.

[0040] Example 38: An apparatus comprising: means for including, in a bitstream, an encoding of at least one picture; means for determining information related to a group of at least one postprocessing filter; and means for indicating, in the bitstream, the information related to the group of the at least one post-processing filter; wherein the information related to the group of the at least one post-processing filter is configured to be used to infer whether and how to use at least one postprocessing filter in the group of the at least one post-processing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one postprocessing filter in the group when a previous post-processing filter in a cascade outputs a different number of pictures than what is used as input for a current post-processing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.

[0041] Example 39: A non-transitory program storage device readable by a machine, tangiblyembodying a program of instructions executable by the machine for performing operations, the operations comprising: receiving an encoding of at least one picture; receiving signaling comprising information related to a group of at least one post-processing filter; and using the information related to the group of the at least one post-processing filter to infer whether and how to use at least one post-processing filter in the group of the at least one post-processing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one postprocessing filter in the group when a previous post-processing filter in a cascade outputs a different number of pictures than what is used as input for a current post-processing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.

[0042] Example 40: A non-transitory program storage device readable by a machine, tangibly embodying a program of instructions executable by the machine for performing operations, the operations comprising: including, in a bitstream, an encoding of at least one picture; determining information related to a group of at least one post-processing filter; and indicating, in the bitstream, the information related to the group of the at least one post-processing filter; wherein the information related to the group of the at least one post-processing filter is configured to be used to infer whether and how to use at least one post-processing filter in the group of the at least one post-processing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one post-processing filter in the group when a previous post-processing filter in a cascade outputs a different number of pictures than what is used as input for a current postprocessing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The foregoing embodiments and other features are explained in the following description, taken in connection with the accompanying drawings, wherein:

[0044] FIG. 1 shows schematically an electronic device employing embodiments of the examples described herein.

[0045] FIG. 2 shows schematically a user equipment suitable for employing embodiments of the examples described herein.

[0046] FIG. 3 further shows schematically electronic devices employing embodiments of the examples described herein connected using wireless and wired network connections.

[0047] FIG. 4 shows schematically a block chart of an encoder used for data compression on a general level.

[0048] FIG. 5 shows a system pipeline for VCM.

[0049] FIG. 6 shows an example where two postfilters are to be used in cascade.

[0050] FIG. 7 shows an example implementing filters of different complexity.

[0051] FIG. 8 shows an example implementing two filters in parallel, one filter for visual enhancement and another filter for machine enhancement.

[0052] FIG. 9 shows an example where one or more coefficients are used to weight the contribution of the output of each of the two postfilters.

[0053] FIG. 10 shows an example where the output of a first filter is used for displaying and the output of a second filter is used as input to one or more machine analysis tasks.

[0054] FIG. 11 shows an example using two filters of low complexity for visual enhancement and spatial upsampling, and using two other filters of high complexity for visual enhancement and spatial upsampling.

[0055] FIG. 12 shows an example using a combination of two filters for visual enhancement, a filter for machine enhancement and a filter for frame upsampling.

[0056] FIG. 13 is a block diagram illustrating a system in accordance with an example.

[0057] FIG. 14 is an example apparatus configured to implement the examples described herein.

[0058] FIG. 15 shows a representation of an example of non-volatile memory media used to store instructions that implement the examples described herein.

[0059] FIG. 16 is an example method performed with a decoder, based on the examples described herein.

[0060] FIG. 17 is an example method performed with an encoder, based on the examples described herein.

[0061] FIG. 18 shows an example of hierarchical activation of a picture rate upsampling filter NNPF, where NNPF[x] indicates the x-th activation of the NNPF.

[0062] FIG. 19 shows an example of visual quality enhancement followed by picture rate upsampling, where NNPF[x] are applied in ascending order of x.

[0063] FIG. 20 is an example method, based on the examples described herein.

[0064] FIG. 21 is an example method, based on the examples described herein.DETAILED DESCRIPTION OF EXAMPLE EMBODIMENTS

[0065] Described herein is a method and apparatus for signaling information about multiple postprocessing filters.

[0066] The following describes in detail a suitable apparatus and possible mechanisms for a video / image encoding process according to embodiments. In this regard reference is first made to FIG. 1 and FIG. 2, where FIG. 1 shows an example block diagram of an apparatus 50. The apparatus may be an Internet of Things (loT) apparatus configured to perform various functions, such as for example, gathering information by one or more sensors, receiving or transmitting information, analyzing information gathered or received by the apparatus, or the like. The apparatus may comprise a video coding system, which may incorporate a codec. FIG. 2 shows a layout of an apparatus according to an example embodiment. The elements of FIG. 1 and FIG. 2 are explained next.

[0067] The electronic device 50 may for example be a mobile terminal or user equipment of a wireless communication system, a sensor device, a tag, or other lower power device. However, it would be appreciated that embodiments of the examples described herein may be implemented within any electronic device or apparatus which may process data by neural networks.

[0068] The apparatus 50 may comprise a housing 30 for incorporating and protecting the device. The apparatus 50 further may comprise a display 32 in the form of a liquid crystal display. In other embodiments of the examples described herein the display may be any suitable display technology suitable to display an image or video. The apparatus 50 may further comprise a keypad 34. In other embodiments of the examples described herein any suitable data or user interface mechanism may be employed. For example the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.

[0069] The apparatus may comprise a microphone 36 or any suitable audio input which may be a digital or analog signal input. The apparatus 50 may further comprise an audio output device which in embodiments of the examples described herein may be any one of: an earpiece 38, speaker, or an analog audio or digital audio output connection. The apparatus 50 may also comprise a battery (or in other embodiments of the examples described herein the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator). The apparatus 50 mayfurther comprise a camera 42 capable of recording or capturing images and / or video. The apparatus 50 may further comprise an infrared port for short range line of sight communication to other devices. In other embodiments the apparatus 50 may further comprise any suitable short range communication solution such as for example a Bluetooth wireless connection or a USB / firewire wired connection.

[0070] The apparatus 50 may comprise a controller 56, processor or processor circuitry for controlling the apparatus 50. The controller 56 may be connected to memory 58 which in embodiments of the examples described herein may store both data in the form of image and audio data and / or may also store instructions for implementation on the controller 56. The controller 56 may further be connected to codec circuitry 54 suitable for carrying out coding and / or decoding of audio and / or video data or assisting in coding and / or decoding carried out by the controller.

[0071] The apparatus 50 may further comprise a card reader 48 and a smart card 46, for example a UICC and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network.

[0072] The apparatus 50 may comprise radio interface circuitry 52 connected to the controller and suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system or a wireless local area network. The apparatus 50 may further comprise an antenna 44 connected to the radio interface circuitry 52 for transmitting radio frequency signals generated at the radio interface circuitry 52 to other apparatus(es) and / or for receiving radio frequency signals from other apparatus(es).

[0073] The apparatus 50 may comprise a camera capable of recording or detecting individual frames which are then passed to the codec 54 or the controller for processing. The apparatus may receive the video image data for processing from another device prior to transmission and / or storage. The apparatus 50 may also receive either wirelessly or by a wired connection the image for coding / decoding. The structural elements of apparatus 50 described above represent examples of means for performing a corresponding function.

[0074] With respect to FIG. 3, an example of a system within which embodiments of the examples described herein can be utilized is shown. The system 10 comprises multiple communication devices which can communicate through one or more networks. The system 10 may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network (such as a GSM, UMTS, CDMA, LTE, 4G, 5G network etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a Bluetooth personal area network, an Ethernet local area network, a token ring local area network, a wide areanetwork, and the Internet.

[0075] The system 10 may include both wired and wireless communication devices and / or apparatus 50 suitable for implementing embodiments of the examples described herein.

[0076] For example, the system shown in FIG. 3 shows a mobile telephone network 11 and a representation of the internet 28. Connectivity to the internet 28 may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.

[0077] The example communication devices shown in the system 10 may include, but are not limited to, an electronic device or apparatus 50, a combination of a personal digital assistant (PDA) and a mobile telephone 14, a PDA 16, an integrated messaging device (IMD) 18, a desktop computer 20, a notebook computer 22, or a head -mounted apparatus 21. The head -mounted apparatus 21 may be a head-mounted display (HMD), or glasses having a device such as a camera configured to encode and / or decode images and / or video. The apparatus 50 may be stationary or mobile when carried by an individual who is moving. The apparatus 50 may also be located in a mode of transport including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle or any similar suitable mode of transport.

[0078] The embodiments may also be implemented in a set-top box; i.e. a digital TV receiver, which may / may not have a display or wireless capabilities, in tablets or (laptop) personal computers (PC), which have hardware and / or software to process neural network data, in various operating systems, and in chipsets, processors, DSPs and / or embedded systems offering hardware / software based coding.

[0079] Some or further apparatus may send and receive calls and messages and communicate with service providers through a wireless connection 25 to a base station 24. The base station 24 may be connected to a network server 26 that allows communication between the mobile telephone network 11 and the internet 28. The system may include additional communication devices and communication devices of various types.

[0080] The communication devices may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocol-internet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service(MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11, 3GPP Narrowband loT and any similar wireless communication technology. A communications device involved in implementing various embodiments of the examples described herein may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.

[0081] In telecommunications and data networks, a channel may refer either to a physical channel or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire, whereas a logical channel may refer to a logical connection over a multiplexed medium, capable of conveying several logical channels. A channel may be used for conveying an information signal, for example a bitstream, from one or several senders (or transmitters) to one or several receivers.

[0082] The embodiments may also be implemented in so-called loT devices. The Internet of Things (loT) may be defined, for example, as an interconnection of uniquely identifiable embedded computing devices within the existing Internet infrastructure. The convergence of various technologies has and may enable many fields of embedded systems, such as wireless sensor networks, control systems, home / building automation, etc. to be included in the Internet of Things (loT). In order to utilize the Internet loT devices are provided with an IP address as a unique identifier. loT devices may be provided with a radio transmitter, such as a WLAN or Bluetooth transmitter or a RFID tag. Alternatively, loT devices may have access to an IP-based network via a wired network, such as an Ethernet-based network or a power-line connection (PLC).

[0083] An MPEG-2 transport stream (TS), specified in ISO / IEC 13818-1 or equivalently in ITU- T Recommendation H.222.0, is a format for carrying audio, video, and other media as well as program metadata or other metadata, in a multiplexed stream. A packet identifier (PID) is used to identify an elementary stream (a.k.a. packetized elementary stream) within the TS. Hence, a logical channel within an MPEG-2 TS may be considered to correspond to a specific PID value.

[0084] Available media file format standards include ISO base media file format (ISO / IEC 14496-12, which may be abbreviated ISOBMFF) and file format for NAL unit structured video (ISO / IEC 14496-15), which derives from the ISOBMFF.

[0085] Recently, Hypertext Transfer Protocol (HTTP) has been widely used for the delivery of real-time multimedia content over the Internet, such as in video streaming applications. Several commercial solutions for adaptive streaming over HTTP, such as Microsoft® Smooth Streaming, Apple® Adaptive HTTP Live Streaming and Adobe® Dynamic Streaming, have been launched as well as standardization projects have been carried out. Adaptive HTTP streaming (AHS) was firststandardized in Release 9 of 3rd Generation Partnership Project (3GPP) packet- switched streaming (PSS) service (3GPP TS 26.234 Release 9: “Transparent end-to-end packet-switched streaming service (PSS); protocols and codecs”). MPEG took 3GPP AHS Release 9 as a starting point for the MPEG DASH standard (ISO / IEC 23009-1: “Dynamic adaptive streaming over HTTP (DASH)-Part 1: Media presentation description and segment formats,” International Standard, 2nd Edition, 2014). 3GPP continued to work on adaptive HTTP streaming in communication with MPEG and published 3GP-DASH (Dynamic Adaptive Streaming over HTTP; 3GPP TS 26.247: “Transparent end-to-end packet-switched streaming Service (PSS); Progressive download and dynamic adaptive Streaming over HTTP (3GP-DASH)”. MPEG DASH and 3GP-DASH are technically close to each other and may therefore be collectively referred to as DASH. Some concepts, formats, and operations of DASH are described below as an example of a video streaming system, wherein the embodiments may be implemented. The embodiments of the invention are not limited to DASH, but rather the description is given for one possible basis on top of which the invention may be partly or fully realized.

[0086] In DASH, the multimedia content may be stored on an HTTP server and may be delivered using HTTP. The content may be stored on the server in two parts: Media Presentation Description (MPD), which describes a manifest of the available content, its various alternatives, their URL addresses, and other characteristics; and segments, which include the actual multimedia bitstreams in the form of chunks, in a single file or multiple files. The MDP provides the necessary information for clients to establish a dynamic adaptive streaming over HTTP. The MPD includes information describing media presentation, such as an HTTP- uniform resource locator (URL) of each Segment to make GET Segment request. To play the content, the DASH client may obtain the MPD e.g. by using HTTP, email, thumb drive, broadcast, or other transport methods. By parsing the MPD, the DASH client may become aware of the program timing, media-content availability, media types, resolutions, minimum and maximum bandwidths, and the existence of various encoded alternatives of multimedia components, accessibility features and required digital rights management (DRM), media-component locations on the network, and other content characteristics. Using this information, the DASH client may select the appropriate encoded alternative and start streaming the content by fetching the segments using e.g. HTTP GET requests. After appropriate buffering to allow for network throughput variations, the client may continue fetching the subsequent segments and also monitor the network bandwidth fluctuations. The client may decide how to adapt to the available bandwidth by fetching segments of different alternatives (with lower or higher bitrates) to maintain an adequate buffer.

[0087] In DASH, hierarchical data model is used to structure media presentation as follows. A media presentation consists of a sequence of one or more Periods, each Period includes one or moreGroups, each Group includes one or more Adaptation Sets, each Adaptation Sets includes one or more Representations, each Representation consists of one or more Segments. A Representation is one of the alternative choices of the media content or a subset thereof typically differing by the encoding choice, e.g. by bitrate, resolution, language, codec, etc. The Segment includes certain duration of media data, and metadata to decode and present the included media content. A Segment is identified by a URI and can typically be requested by a HTTP GET request. A Segment may be defined as a unit of data associated with an HTTP-URL and optionally a byte range that are specified by an MPD.

[0088] The DASH MPD complies with Extensible Markup Language (XML) and is therefore specified through elements and attributes as defined in XML.

[0089] Real-time Transport Protocol (RTP) is widely used for real-time transport of timed media such as audio and video. RTP may operate on top of the User Datagram Protocol (UDP), which in turn may operate on top of the Internet Protocol (IP). RTP is specified in Internet Engineering Task Force (IETF) Request for Comments (RFC) 3550, available from [www.ietf.org / rfc / rfc3550.txt (last accessed on September 29, 2021)]. In RTP transport, media data is encapsulated into RTP packets. Typically, each media type or media coding format has a dedicated RTP pay load format.

[0090] RTP is designed to carry a multitude of multimedia formats, which permits the development of new formats without revising the RTP standard. To this end, the information required by a specific application of the protocol is not included in the generic RTP header. For a class of applications (e.g., audio, video), an RTP profile may be defined. For a media format (e.g., a specific video coding format), an associated RTP payload format may be defined. Every instantiation of RTP in a particular application may require a profile and payload format specifications. For example, an RTP profile for audio and video conferences with minimal control is defined in RFC 3551, and an Audio-Visual Profile with Feedback (AVPF) is specified in RFC 4585. The profile may define a set of static payload type assignments and / or may use a dynamic mechanism for mapping between a payload format and a payload type (PT) value using Session Description Protocol (SDP). The latter mechanism is used for newer video codec such as RTP payload format for H.264 defined in RFC 6184 or RTP Payload Format for HEVC defined in RFC 7798.

[0091] An RTP session is an association among a group of participants communicating with RTP. It is a group communications channel which can potentially carry a number of RTP streams. An RTP stream is a stream of RTP packets comprising media data. An RTP stream is identified by an SSRC belonging to a particular RTP session. SSRC refers to either a synchronization source or a synchronization source identifier that is the 32-bit SSRC field in the RTP packet header. Asynchronization source is characterized in that all packets from the synchronization source form part of the same timing and sequence number space, so a receiver device may group packets by synchronization source for playback. Examples of synchronization sources include the sender of a stream of packets derived from a signal source such as a microphone or a camera, or an RTP mixer. Each RTP stream is identified by a SSRC that is unique within the RTP session.

[0092] A point-to-point RTP session includes two endpoints, communicating using unicast. Both RTP and RTCP traffic are conveyed endpoint to endpoint.

[0093] Many multipoint audio-visual conferences operate utilizing a centralized unit, which may be called Multipoint Control Unit (MCU). An MCU may implement the functionality of an RTP translator or an RTP mixer. An RTP translator may be a media translator that may modify the media inside the RTP stream. A media translator may for example decode and re-encode the media content (e.g., transcode the media content). An RTP mixer is a middlebox that aggregates multiple RTP streams that are part of a session by generating one or more new RTP streams. An RTP mixer may manipulate the media data. One common application for a mixer is to allow a participant to receive a session with a reduced number of resources compared to receiving individual RTP streams from all endpoints. A mixer can be viewed as a device terminating the RTP streams received from other endpoints in the same RTP session. Using the media data carried in the received RTP streams, a mixer generates derived RTP streams that are sent to the receiving endpoints.

[0094] The Session Description Protocol (SDP) may be used to convey media details, transport addresses, and other session description metadata, when initiating multimedia teleconferences, voice-over-IP calls, or other multimedia delivery sessions. SDP is a format for describing multimedia communication sessions for the purposes of announcement and invitation. SDP does not deliver any media streams itself but may be used between endpoints e.g., for negotiation of network metrics, media types, and / or other associated properties. SDP is extensible for the support of new media types and formats.

[0095] SDP uses attributes to extend the core protocol. Attributes can appear within the Session or Media sections and are scoped accordingly as session-level or media-level. New attributes can be added to the standard through registration with IANA. A media description may include any number of "a=" lines (attribute-fields) that are media description specific. Session-level attributes convey additional information that applies to the session as a whole rather than to individual media descriptions.

[0096] The "fmtp" attribute of SDP allows parameters that are specific to a particular format to be conveyed in a way that SDP does not have to understand them. The format must be one of theformats specified for the media. Format-specific parameters, semicolon separated, may be any set of parameters required to be conveyed by SDP and given unchanged to the media tool that will use this format. At most one instance of this attribute is allowed for each format.

[0097] The SDP offer / answer model specifies a mechanism in which endpoints achieve a common operating point of media details and other session description metadata when initiating the multimedia delivery session. One endpoint, the offerer sends a session description (the offer) to the other endpoint, the answerer. The offer includes all the media parameters needed to exchange media with the offerer, including codecs, transport addresses, and protocols to transfer media. When the answerer receives an offer, it elaborates an answer and sends it back to the offerer. The answer includes the media parameters that the answerer is willing to use for that particular session. SDP may be used as the format for the offer and the answer.

[0098] An initial SDP offer includes zero or more media streams, wherein each media stream is described by an "m=" line and its associated attributes. Zero media streams implies that the offerer wishes to communicate, but that the streams for the session will be added at a later time through a modified offer.

[0099] A direction attribute may be used in the SDP offer / answer model as follows. If the offerer wishes to only send media on a stream to its peer, it marks the stream as sendonly with the "a=sendonly" attribute. If the offerer wishes to only receive media from its peer, it marks the stream as recvonly. If the offerer wishes to both send and receive media with its peer, it may include an "a=sendrecv" attribute in the offer, or it may omit it, since sendrecv is the default.

[0100] In the SDP offer / answer model, the list of media formats for each media stream comprises the set of formats (codecs and any parameters associated with the codec, in the case of RTP) that the offerer is capable of sending and / or receiving (depending on the direction attributes). If multiple formats are listed, it means that the offerer is capable of making use of any of those formats during the session and thus the answerer may change formats in the middle of the session, making use of any of the formats listed, without sending a new offer. For a sendonly stream, the offer indicates those formats the offerer is willing to send for this stream. For a recvonly stream, the offer indicates those formats the offerer is willing to receive for this stream. For a sendrecv stream, the offer indicates those codecs or formats that the offerer is willing to send and receive with. The list of media formats in the "m=" line is listed in the order preference, the first entry in the list being the most preferred.

[0101] SDP may be used for declarative purposes, e.g., for describing a stream available to be received over a streaming session. For example, SDP may be included in Real Time StreamingProtocol (RTSP).

[0102] A Multipurpose Internet Mail Extension (MIME) is an extension to an email protocol which makes it possible to transmit and receive different kinds of data files on the Internet, for example video, audio, images, and software. An internet media type is an identifier used on the Internet to indicate the type of data that a file includes. Such internet media types may also be called as content types. Several MIME type / subtype combinations exist that can include different media formats. Content type information may be included by a transmitting entity in a MIME header at the beginning of a media transmission. A receiving entity thus may need to examine the details of such media content to determine if the specific elements can be rendered given an available set of codecs. Especially when the end system has limited resources, or the connection to the end system has limited bandwidth, it may be helpful to know from the content type alone if the content can be rendered.

[0103] One of the original motivations for MIME is the ability to identify the specific media type of a message part. However, due to various factors, it is not always possible from looking at the MIME type and subtype to know which specific media formats are included in the body part or which codecs are indicated in order to render the content. Optional media parameters may be provided in addition to the MIME type and subtype to provide further details of the media content.

[0104] Optional media parameters may be conveyed in SDP, e.g., using the "a=fmtp" line of SDP. Optional media parameters may be specified to apply for certain direction attribute(s) with an SDP offer / answer and / or for declarative purposes. Optional media parameters may be specified not to apply for certain direction attribute(s) with an SDP offer / answer and / or for declarative purposes. Semantics of optional media parameters may depend on and may differ based on which direction attribute(s) of an SDP offer / answer they are used with and / or whether they are used for declarative purposes.

[0105] An example of an optional media parameter specified in the VVC RTP pay load format is sprop-sei. When present, sprop-sei conveys one or more SEI messages that describe bitstream characteristics. A decoder can rely on the bitstream characteristics that are described in the SEI messages carried within sprop-sei for the entire duration of the session, independently of the persistence scopes of the SEI messages specified in H.274 / VSEI or VVC. The value of sprop-sei may be defined as a comma-separated list, where each list element is a base64 representation (as defined in RFC 4648) of an SEI NAL unit.

[0106] In an example, empty or truncated SEI message payloads are allowed in sprop-sei to indicate the capability of the offerer to encode these SEI messages with any SEI message payloadthat starts with the bits included in sprop-sei (if any).

[0107] In an example, an optional recv-sei MIME parameter in an SDP answer indicate which SEI messages must be present in the bitstream from the offerer to the answerer. The SEI messages in recv-sei may be allowed or required to be of the following types: the SEI message may be required to be the same as directly included in the sprop-sei of the offer; the SEI message may be required to have a type indicated in a SEI manifest SEI message in the sprop-sei of the offer; the SEI message may be required to have a type and start with the respective content indicated by a SEI prefix indication SEI message included within the srop-sei. When an SEI message with a particular SEI message type is present in sprop-sei of an offer and is absent in recv-sei in an answer, it may indicate that the answerer does not support the processing the SEI message included in sprop-sei.

[0108] In an example, an optional recv-sei MIME parameter is introduced along the following principles: The SEI messages included in recv-sei in an answer indicate which SEI messages must be present in the bitstream from the offerer to the answerer. When recv-sei is present in an answer, the SEI message types in recv-sei may be required to be the same as or a subset of those included in the sprop-sei of the offer. When recv-sei is present in an answer, the SEI message payload for a particular SEI message type in recv-sei may be required to start with the SEI message payload bits present (if any) in recv-sei for the same SEI message type. It is suggested to allow empty or truncated SEI message payloads in recv-sei to indicate that the answerer has the liberty to encode any values for remaining of the SEI message payload (as long as they conform to the specification where the SEI message is specified). When an SEI message with a particular SEI message type is present in sprop-sei of an offer and is absent in recv-sei in an answer, it may indicate that the answerer does not support the processing the SEI message included in sprop-sei.

[0109] In an example, an optional MIME parameter in an SDP offer from the offerer to the answerer includes an SEI NAL unit including an SEI manifest SEI message and one or more SEI prefix indication SEI messages to indicate a capability of an encoder in the offerer to encode SEI messages as indicated by the semantics of the SEI manifest SEI message and the one or more SEI prefix indication SEI messages and encode a bitstream obeying the constraints implied by the encoded SEI messages.

[0110] In an example, an optional MIME parameter in an SDP answer from the answerer to the offerer includes an SEI NAL unit including an SEI manifest SEI message and one or more SEI prefix indication SEI messages to indicate a requirement or preference of the offerer for encoded SEI messages included in the bitstream from the offerer to the answerer, as indicated by the semantics of the SEI manifest SEI message and the one or more SEI prefix indication SEI messages.

[0111] FIG. 4 shows a block diagram of a general structure of a video encoder. FIG. 4 presents an encoder for two layers, but it would be appreciated that presented encoder could be similarly extended to encode more than two layers. FIG. 4 illustrates a video encoder comprising a first encoder section 500 for a base layer and a second encoder section 502 for an enhancement layer. Each of the first encoder section 500 and the second encoder section 502 may comprise similar elements for encoding incoming pictures. The encoder sections 500, 502 may comprise a pixel predictor 302, 402, prediction error encoder 303, 403 and prediction error decoder 304, 404. FIG. 4 also shows an embodiment of the pixel predictor 302, 402 as comprising an inter-predictor 306, 406 (Pinter), an intra-predictor 308, 408 (Pintra), a mode selector 310, 410, a filter 316, 416 (F), and a reference frame memory 318, 418 (RFM). The pixel predictor 302 of the first encoder section 500 receives 300 base layer images (Io,n) of a video stream to be encoded at both the inter-predictor 306 (which determines the difference between the image and a motion compensated reference frame 318) and the intra-predictor 308 (which determines a prediction for an image block based only on the already processed parts of the current frame or picture). The output of both the inter-predictor and the intra-predictor are passed to the mode selector 310. The intra-predictor 308 may have more than one intra-prediction modes. Hence, each mode may perform the intra-prediction and provide the predicted signal to the mode selector 310. The mode selector 310 also receives a copy of the base layer picture 300. Correspondingly, the pixel predictor 402 of the second encoder section 502 receives 400 enhancement layer images (Ii,n) of a video stream to be encoded at both the interpredictor 406 (which determines the difference between the image and a motion compensated reference frame 418) and the intra-predictor 408 (which determines a prediction for an image block based only on the already processed parts of the current frame or picture). The output of both the inter-predictor and the intra-predictor are passed to the mode selector 410. The intra-predictor 408 may have more than one intra-prediction modes. Hence, each mode may perform the intra-prediction and provide the predicted signal to the mode selector 410. The mode selector 410 also receives a copy of the enhancement layer picture 400.

[0112] Depending on which encoding mode is selected to encode the current block, the output of the inter-predictor 306, 406 or the output of one of the optional intra-predictor modes or the output of a surface encoder within the mode selector is passed to the output of the mode selector 310, 410. The output of the mode selector is passed to a first summing device 321, 421. The first summing device may subtract the output of the pixel predictor 302, 402 from the base layer picture 300 / enhancement layer picture 400 to produce a first prediction error signal 320, 420 (Dn) which is input to the prediction error encoder 303, 403.

[0113] The pixel predictor 302, 402 further receives from a preliminary reconstructor 339, 439 the combination of the prediction representation of the image block 312, 412 (P’n) and the output338, 438 (D’n) of the prediction error decoder 304, 404. The preliminary reconstructed image 314, 414 (I’n) may be passed to the intra-predictor 308, 408 and to the filter 316, 416. The filter 316, 416 receiving the preliminary representation may filter the preliminary representation and output a final reconstructed image 340, 440 (R’n) which may be saved in a reference frame memory 318, 418. The reference frame memory 318 may be connected to the inter-predictor 306 to be used as the reference image against which a future base layer picture 300 is compared in inter-prediction operations. Subject to the base layer being selected and indicated to be the source for inter-layer sample prediction and / or inter-layer motion information prediction of the enhancement layer according to some embodiments, the reference frame memory 318 may also be connected to the inter-predictor 406 to be used as the reference image against which a future enhancement layer picture 400 is compared in inter-prediction operations. Moreover, the reference frame memory 418 may be connected to the inter-predictor 406 to be used as the reference image against which a future enhancement layer picture 400 is compared in inter-prediction operations.

[0114] Filtering parameters from the filter 316 of the first encoder section 500 may be provided to the second encoder section 502 subject to the base layer being selected and indicated to be the source for predicting the filtering parameters of the enhancement layer according to some embodiments.

[0115] The prediction error encoder 303, 403 comprises a transform unit 342, 442 (T) and a quantizer 344, 444 (Q). The transform unit 342, 442 transforms the first prediction error signal 320, 420 to a transform domain. The transform is, for example, the DCT transform. The quantizer 344, 444 quantizes the transform domain signal, e.g. the DCT coefficients, to form quantized coefficients.

[0116] The prediction error decoder 304, 404 receives the output from the prediction error encoder 303, 403 and performs the opposite processes of the prediction error encoder 303, 403 to produce a decoded prediction error signal 338, 438 which, when combined with the prediction representation of the image block 312, 412 at the second summing device 339, 439, produces the preliminary reconstructed image 314, 414. The prediction error decoder 304, 404 may be considered to comprise a dequantizer 346, 446 (Q1), which dequantizes the quantized coefficient values, e.g. DCT coefficients, to reconstruct the transform signal and an inverse transformation unit 348, 448 (T1), which performs the inverse transformation to the reconstructed transform signal wherein the output of the inverse transformation unit 348, 448 includes reconstructed block(s). The prediction error decoder may also comprise a block filter which may filter the reconstructed block(s) according to further decoded information and filter parameters.

[0117] The entropy encoder 330, 430 (E) receives the output of the prediction error encoder 303,403 and may perform a suitable entropy encoding / variable length encoding on the signal to provide error detection and correction capability. The outputs of the entropy encoders 330, 430 may be inserted into a bitstream e.g. by a multiplexer 508 (M).

[0118] Fundamentals of neural networks

[0119] A neural network (NN) is a computation graph consisting of several layers of computation. Each layer consists of one or more units, where each unit performs an elementary computation. A unit is connected to one or more other units, and the connection may have associated with a weight. The weight may be used for scaling the signal passing through the associated connection. Weights are learnable parameters, i.e., values which can be learned from training data. There may be other learnable parameters, such as those of batch-normalization layers.

[0120] Two of the most widely used architectures for neural networks are feed-forward and recurrent architectures. Feed-forward neural networks are such that there is no feedback loop: each layer takes input from one or more of the layers before and provides its output as the input for one or more of the subsequent layers. Also, units inside a certain layer take input from units in one or more of preceding layers, and provide output to one or more of following layers.

[0121] Initial layers (those close to the input data) extract semantically low-level features such as edges and textures in images, and intermediate and final layers extract more high-level features. After the feature extraction layers there may be one or more layers performing a certain task, such as classification, semantic segmentation, object detection, denoising, style transfer, super-resolution, etc. In recurrent neural nets, there is a feedback loop, so that the network becomes stateful, i.e., it is able to memorize information or a state.

[0122] Neural networks are being utilized in an ever-increasing number of applications for many different types of device, such as mobile phones. Examples include image and video analysis and processing, social media data analysis, device usage data analysis, etc.

[0123] An important property of neural nets (and other machine learning tools) is that they are able to learn properties from input data, either in supervised way or in unsupervised way. Such learning is a result of a training algorithm, or of a meta-level neural network providing the training signal.

[0124] In general, the training algorithm consists of changing some properties of the neural network so that its output is as close as possible to a desired output. For example, in the case of classification of objects in images, the output of the neural network can be used to derive a class or category index which indicates the class or category that the object in the input image belongs to.Training usually happens by minimizing or decreasing the output’s error, also referred to as the loss. Examples of losses are mean squared error, cross-entropy, etc. In recent deep learning techniques, training is an iterative process, where at each iteration the algorithm modifies the weights of the neural net to make a gradual improvement of the network’s output, i.e., to gradually decrease the loss.

[0125] As used herein, the terms “model”, “neural network”, “neural net” and “network” interchangeably, and also the weights of neural networks are sometimes referred to as learnable parameters or simply as parameters.

[0126] Training a neural network is an optimization process, but the final goal is different from the typical goal of optimization. In optimization, the only goal is to minimize a function. In machine learning, the goal of the optimization or training process is to make the model learn the properties of the data distribution from a limited training dataset. In other words, the goal is to learn to use a limited training dataset in order to learn to generalize to previously unseen data, i.e., data which was not used for training the model. This is usually referred to as generalization. In practice, data is usually split into at least two sets, the training set and the validation set. The training set is used for training the network, i.e., to modify its learnable parameters in order to minimize the loss. The validation set is used for checking the performance of the network on data which was not used to minimize the loss, as an indication of the final performance of the model. In particular, the errors on the training set and on the validation set are monitored during the training process to understand the following things:

[0127] If the network is learning at all - in this case, the training set error should decrease, otherwise the model is in the regime of underfitting.

[0128] If the network is learning to generalize - in this case, also the validation set error needs to decrease and to be not too much higher than the training set error. If the training set error is low, but the validation set error is much higher than the training set error, or it does not decrease, or it even increases, the model is in the regime of overfitting. This means that the model has just memorized the training set’ s properties and performs well only on that set, but performs poorly on a set not used for tuning its parameters.

[0129] While the above background information on neural networks and related training algorithms may be valid at the time when this document was written, the field of neural networks and machine learning in general is developing at a fast pace. Thus, it is to be understood that at least some of the embodiments described herein are not limited to the definition of a neural network, or a machine learning model, or a training algorithm that was given in the background information above.

[0130] Lately, neural networks have been used for compressing and de-compressing data such as images, i.e., in an image codec. The most widely used architecture for realizing one component of an image codec is the auto-encoder, which is a neural network consisting of two parts: a neural encoder and a neural decoder (we refer to these simply as encoder and decoder, even though we may refer to algorithms which are learned from data instead of being tuned by hand). The encoder takes as input an image and produces a code which requires less bits than the input image. This code may be obtained by applying a binarization or quantization process to the output of the encoder. The decoder takes in this code and reconstructs the image which was input to the encoder.

[0131] Such encoder and decoder are usually trained to minimize a combination of bitrate and distortion, where the distortion may be based on one or more of the following metrics: Mean Squared Error (MSE), Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), or similar. These metrics are meant to be correlated to the human visual perception quality, so that minimizing or maximizing one or more of these metrics results into improving the visual quality of the decoded image as perceived by humans.

[0132] Some video coding and video metadata specifications

[0133] The Advanced Video Coding standard (which may be abbreviated H.264, AVC or H.264 / AVC) was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunications Standardization Sector of International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). The H.264 / AVC standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, also known as MPEG -4 Part 10 Advanced Video Coding (AVC). There have been multiple versions of the H.264 / AVC standard, each integrating new extensions or features to the specification. These extensions include Scalable Video Coding (SVC) and Multiview Video Coding (MVC).

[0134] The High Efficiency Video Coding standard (which may be abbreviated H.265, HEVC or H.265 / HEVC) was developed by the Joint Collaborative Team - Video Coding (JCT-VC) of VCEG and MPEG. The standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2, also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC). Extensions to H.265 / HEVC include scalable, multiview, three-dimensional, and fidelity range extensions, which may be referred to as SHVC, MV-HEVC, 3D-HEVC, and REXT, respectively. The references in this description to H.265 / HEVC, SHVC, MV-HEVC, 3D-HEVC and REXT that have been made for the purpose ofunderstanding definitions, structures or concepts of these standard specifications are to be understood to be references to the latest versions of these standards that were available before the date of this application, unless otherwise indicated.

[0135] Versatile Video Coding (which may be abbreviated VVC, H.266, or H.266 / VVC) is a video compression standard developed as the successor to HEVC. VVC is specified in ITU-T Recommendation H.266 and equivalently in ISO / IEC 23090-3, which is also referred to as MPEG- I Part 3.

[0136] A specification of the AV 1 bitstream format and decoding process were developed by the Alliance of Open Media (AOM). The AVI specification was published in 2018. AOM is reportedly working on the AV2 specification.

[0137] ITU-T Recommendation H.274, which is equivalent to ISO / IEC 23002-7, may be called "versatile supplemental enhancement information messages for coded video bitstreams" and be referred to as "versatile supplemental enhancement information" or VSEI. The VSEI standard specifies the syntax and semantics of video usability information (VUI) parameters and supplemental enhancement information (SEI) messages. The VUI parameters and SEI messages defined in the VSEI standard are designed to be conveyed within coded video bitstreams in a manner specified in a video coding specification or to be conveyed by other means determined by the specifications for systems that make use of such coded video bitstreams. The VSEI standard is intended for use with VVC coded video bitstreams, although it is drafted in a manner intended to be sufficiently generic that it may also be used with other types of coded video bitstreams. VUI parameters and SEI messages may, for example, assist in processes related to decoding, display or other purposes.

[0138] Fundamentals of video / image coding

[0139] An elementary unit for the input to an encoder and the output of a decoder, respectively, in most cases is a picture. A picture given as an input to an encoder may also be referred to as a source picture, and a picture decoded by a decoded may be referred to as a decoded picture or a reconstructed picture.

[0140] The source and decoded pictures are each comprised of one or more sample arrays, such as one of the following sets of sample arrays:- Luma (Y) only (monochrome).- Luma and two chroma (YCbCr or YCgCo).- Green, Blue and Red (GBR, also known as RGB).- Arrays representing other unspecified monochrome or tri-stimulus color samplings (for example, YZX, also known as XYZ).

[0141] In the following, these arrays may be referred to as luma (or L or Y) and chroma, where the two chroma arrays may be referred to as Cb and Cr or Cg and Co; regardless of the actual color representation method in use. The actual color representation method in use may be indicated, e.g., in a coded bitstream e.g., using the Video Usability Information (VUI) syntax. A component may be defined as an array or single sample from one of the three sample arrays (luma and two chroma) or the array or a single sample of the array that compose a picture in monochrome format.

[0142] A picture may be defined to be either a frame or a field. A frame comprises a matrix of luma samples and possibly the corresponding chroma samples. A field is a set of alternate sample rows of a frame and may be used as encoder input, when the source signal is interlaced. Chroma sample arrays may be absent (and hence monochrome sampling may be in use) or chroma sample arrays may be subsampled when compared to luma sample arrays.

[0143] Some chroma formats may be summarized as follows:- In monochrome sampling there is only one sample array, which may be nominally considered the luma array.- In 4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array.- In 4:2:2 sampling, each of the two chroma arrays has the same height and half the width of the luma array.- In 4:4:4 sampling when no separate color planes are in use, each of the two chroma arrays has the same height and width as the luma array.

[0144] Coding formats or standards may allow to code sample arrays as separate color planes into the bitstream and respectively decode separately coded color planes from the bitstream. When separate color planes are in use, each one of them is separately processed (by the encoder and / or the decoder) as a picture with monochrome sampling.

[0145] Video codec consists of an encoder that transforms the input video into a compressed representation suited for storage / transmission and a decoder that can decompress the compressed video representation back into a viewable form. Typically encoder discards some information in the original video sequence in order to represent the video in a more compact form (that is, at lower bitrate).

[0146] Typical hybrid video codecs, for example ITU-T H.263 and H.264, encode the video information in two phases. Firstly pixel values in a certain picture area (or “block”) are predicted for example by motion compensation means (finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (using the pixel values around the block to be coded in a specified manner). Secondly the prediction error, i.e. the difference between the predicted block of pixels and the original block of pixels, is coded. This is typically done by transforming the difference in pixel values using a specified transform (e.g. Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size or transmission bitrate).

[0147] Inter prediction, which may also be referred to as temporal prediction, motion compensation, or motion-compensated prediction, exploits temporal redundancy. In inter prediction the sources of prediction are previously decoded pictures (a.k.a. reference pictures).

[0148] In temporal inter prediction, the sources of prediction are previously decoded pictures in the same scalable layer. In intra block copy (IBC; a.k.a. intra-block-copy prediction), prediction may be applied similarly to temporal inter prediction but the reference picture is the current picture and only previously decoded samples can be referred in the prediction process. Inter-layer or inter-view prediction may be applied similarly to temporal inter prediction, but the reference picture is a decoded picture from another scalable layer or from another view, respectively. In some cases, inter prediction may refer to temporal inter prediction only, while in other cases inter prediction may refer collectively to temporal inter prediction and any of intra block copy, inter-layer prediction, and interview prediction provided that they are performed with the same or similar process than temporal prediction. Inter prediction, temporal inter prediction, or temporal prediction may sometimes be referred to as motion compensation or motion-compensated prediction.

[0149] Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in spatial or transform domain, i.e., either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.

[0150] One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy-coded more efficiently if they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors and only the difference relative tothe motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.

[0151] The decoder reconstructs the output video by applying prediction means similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in spatial pixel domain). After applying prediction and prediction error decoding means the decoder sums up the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) can also apply additional filtering means to improve the quality of the output video before passing it for display and / or storing it as prediction reference for the forthcoming frames in the video sequence.

[0152] Image and video codecs may use a set of filters, which may enhance the visual quality of the predicted visual content. Filters may be applied either in-loop or out-of-loop, or both. In-loop filters (which may be also called loop filters) are used in reconstructing prediction reference that may be used for predicting forthcoming video signal. In other words, in the case of in-loop filters, the filter applied on one block in the currently encoded frame may affect the encoding of another block in the same frame and / or in another frame which is predicted from the current frame. An inloop filter may affect the bitrate and / or the visual quality. In fact, an enhanced block may cause a smaller residual (difference between original block and predicted-and-filtered block), thus requiring less bits to be encoded. An out-of-the loop filter (which may also be called a post-processing filter or a post-filter) may be applied on a frame or part of a frame after it has been reconstructed, the filtered visual content may not be used as a source for prediction, and thus it may only impact the visual quality of the frames that are output by the decoder.

[0153] In typical video codecs the motion information is indicated with motion vectors associated with each motion compensated image block. Each of these motion vectors represents the displacement of the image block in the picture to be coded (in the encoder side) or decoded (in the decoder side) and the prediction source block in one of the previously coded or decoded pictures.

[0154] In order to represent motion vectors efficiently those are typically coded differentially with respect to block specific predicted motion vectors. In typical video codecs the predicted motion vectors are created in a predefined way, for example calculating the median of the encoded or decoded motion vectors of the adjacent blocks.

[0155] Another way to create motion vector predictions is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures andsignalling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index is typically predicted from adjacent blocks and / or or co-located blocks in temporal reference picture.

[0156] Moreover, typical high efficiency video codecs employ an additional motion information coding / decoding mechanism, often called merging / merge mode, where all the motion field information, which includes motion vector and corresponding reference picture index for each available reference picture list, is predicted and used without any modification / correction. Similarly, predicting the motion field information is carried out using the motion field information of adjacent blocks and / or co-located blocks in temporal reference pictures and the used motion field information is signalled among a list of motion field candidate list filled with motion field information of available adjacent / co-located blocks.

[0157] In typical video codecs the prediction residual after motion compensation is first transformed with a transform kernel (like DCT) and then coded. The reason for this is that often there still exists some correlation among the residual and transform can in many cases help reduce this correlation and provide more efficient coding.

[0158] Typical video encoders utilize Lagrangian cost functions to find optimal coding modes, e.g. the desired Macroblock mode and associated motion vectors. This kind of cost function uses a weighting factor X to tie together the (exact or estimated) image distortion due to lossy coding methods and the (exact or estimated) amount of information that is required to represent the pixel values in an image area:C = D + R where C is the Lagrangian cost to be minimized, D is the image distortion (e.g. Mean Squared Error) with the mode and motion vectors considered, and R the number of bits needed to represent the required data to reconstruct the image block in the decoder (including the amount of data to represent the candidate motion vectors).

[0159] An out-of-band transmission, signaling, or storage may refer to the capability of transmitting, signaling, or storing information in a manner that associates the information with a video bitstream. The out-of-band transmission may use a more reliable transmission mechanism compared to the protocols used for carrying coded video data, such as slices. The out-of-band transmission, signaling or storage may additionally or alternatively be used, e.g., for ease of access or session negotiation. For example, a sample entry of a track in a file conforming to the ISO BaseMedia File Format may comprise parameter sets, while the coded data in the bitstream is stored elsewhere in the file or in another file. Another example of out-of-band transmission, signaling, or storage comprises including information, such as NN and / or NN updates in a file format track that is separate from track(s) including coded video data.

[0160] The phrase along the bitstream (e.g., indicating along the bitstream) or along a coded unit of a bitstream (e.g., indicating along a coded tile) may be used in claims and described embodiments to refer to transmission, signaling, or storage in a manner that the ‘out-of-band’ data is associated with, but not included within, the bitstream or the coded unit, respectively. The phrase decoding along the bitstream or along a coded unit of a bitstream or alike may refer to decoding the referred out-of-band data (which may be obtained from out-of-band transmission, signaling, or storage) that is associated with the bitstream or the coded unit, respectively. For example, the phrase along the bitstream may be used when the bitstream is included in a container file, such as a file conforming to the ISO Base Media File Format, and certain file metadata is stored in the file in a manner that associates the metadata to the bitstream, such as boxes in the sample entry for a track including the bitstream, a sample group for the track including the bitstream, or a timed metadata track associated with the track including the bitstream. In another example, the phrase along the bitstream may be used when the bitstream is made available as a stream over a communication protocol and a media description, such as a streaming manifest, is provided to describe the stream.

[0161] A bitstream may be defined as a sequence of bits or a sequence of syntax structures. A bitstream format may constrain the order of syntax structures in the bitstream.

[0162] A syntax element may be defined as an element of data represented in a bitstream. A syntax structure may be defined as zero or more syntax elements present together in a bitstream in a specified order.

[0163] Syntax structures may be specified, for example, using arithmetic, logical, relational, bitwise, and assignment operators similar to those available in many programming languages. For example, & may indicate a bit-wise ‘AND’ operation. Furthermore, syntax structures may be specified with reference to mathematical functions.

[0164] Syntax structures and semantics may use the values of variables derived from the values of syntax elements. Naming conventions may be defined for variables. For example, variables may be named by a mixture of lower case and upper case letter and without any underscore characters. Variables starting with an upper case letter may be derived for the decoding of the current syntax structure and all depending syntax structures. Variables starting with an upper case letter may, in some cases, be used in the decoding process for later syntax structures without mentioning theoriginating syntax structure of the variable. Variables starting with a lower case letter may only be used in relation to the syntax structure or function they have been defined for.

[0165] An elementary unit for the output of a video encoder and the input of a video decoder, respectively, may be a network abstraction layer (NAL) unit. For transport over packet-oriented networks or storage into structured files, NAL units may be encapsulated into packets or similar structures. A bytestream format encapsulating NAL units may be used for transmission or storage environments that do not provide framing structures. The bytestream format may separate NAL units from each other by attaching a start code in front of each NAL unit. To avoid false detection of NAL unit boundaries, encoders may run a byte-oriented start code emulation prevention algorithm, which may add an emulation prevention byte to the NAL unit payload, when a start code would have occurred otherwise. In order to enable straightforward gateway operation between packet and stream-oriented systems, start code emulation prevention may be performed regardless of whether the bytestream format is in use or not. A NAL unit may be defined as a syntax structure including an indication of the type of data to follow and bytes including that data in the form of a raw byte sequence payload interspersed as necessary with emulation prevention bytes. A raw byte sequence pay load (RBSP) may be defined as a syntax structure including an integer number of bytes that is encapsulated in a NAL unit. An RBSP is either empty or has the form of a string of data bits including syntax elements followed by an RBSP stop bit and followed by zero or more subsequent bits equal to 0.

[0166] A bitstream may be defined to logically include a syntax structure, such as a NAL unit, when the syntax structure is transmitted along the bitstream but may be included in the bitstream according to the bitstream format. A bitstream may be defined to natively comprise a syntax structure, when the bitstream includes the syntax structure.

[0167] In some coding formats or standards, a bitstream may be in the form of a network abstraction layer (NAL) unit stream or a byte stream, that forms the representation of coded pictures and associated data forming one or more coded video sequences.

[0168] In some formats or standards, a first bitstream may be followed by a second bitstream in the same logical channel, such as in the same file or in the same connection of a communication protocol. An elementary stream (in the context of video coding) may be defined as a sequence of one or more bitstreams.

[0169] In some coding formats, such as AVI, a bitstream may comprise a sequence of open bitstream units (OBUs). An OBU comprises a header and a payload, wherein the header identifies a type of the OBU. Furthermore, the header may comprise a size of the payload in bytes.

[0170] In some coding standards, NAL units include a header and payload. The NAL unit header indicates the type of the NAL unit. In some coding standards, the NAL unit header indicates a scalability layer identifier (e.g., called nuh_layer_id in H.265 / HEVC and H.266 / VVC), which may be used, e.g., for indicating spatial or quality layers, views of a multiview video, or auxiliary layers (such as depth maps or alpha planes). In some coding standards, the NAL unit header includes a temporal sublayer identifier, which may be used for indicating temporal subsets of the bitstream, such as a 30-frames-per-second subset of a 60-frames-per-second bitstream.

[0171] Bitstreams or coded video sequences may be encoded to be temporally scalable as follows. Each picture may be assigned to a particular temporal sub-layer. A temporal sub-layer may be equivalently called a sub-layer, temporal sublayer, sublayer, or temporal level. Temporal sub-layers may be enumerated, e.g., from 0 upwards. The lowest temporal sub-layer, sub-layer 0, may be decoded independently. Pictures at temporal sub-layer 1 may be predicted from reconstructed pictures at temporal sub-layers 0 and 1. Pictures at temporal sub-layer 2 may be predicted from reconstructed pictures at temporal sub-layers 0, 1, and 2, and so on. In other words, a picture at temporal sub-layer N does not use any picture at temporal sub-layer greater than N as a reference for inter prediction. The bitstream created by excluding all pictures greater than or equal to a selected sub-layer value and including pictures remains conforming.

[0172] Each picture of a temporally scalable bitstream may be assigned with a temporal identifier (also known as temporal layer identifier, temporal sublayer identifier, or temporal layer ID), which may be, for example, assigned to a variable Temporalld. The temporal identifier may, for example, be indicated in a NAL unit header or in an OBU extension header. Temporalld equal to 0 corresponds to the lowest temporal level. The bitstream created by excluding all coded pictures having a Temporalld greater than or equal to a selected value and including all other coded pictures remains conforming. Consequently, a picture having Temporalld equal to tid_value does not use any picture having a Temporalld greater than tid_value as a prediction reference.

[0173] NAL units may be categorized into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units are typically coded slice NAL units.

[0174] A non-VCL NAL unit may be, for example, one of the following types: a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a supplemental enhancement information (SEI) NAL unit, an access unit delimiter, an end of sequence NAL unit, an end of bitstream NAL unit, or a filler data NAL unit. Parameter sets may be needed for the reconstruction of decoded pictures, whereas many of the other non-VCL NAL units are not necessary for the reconstruction of decoded sample values.

[0175] Video coding specifications may enable the use of supplemental enhancement information (SEI) messages or alike. Some video coding specifications include SEI NAL units, and some video coding specifications include both prefix SEI NAL units and suffix SEI NAL units, where the former type can start a picture unit or alike and the latter type can end a picture unit or alike. An SEI NAL unit includes one or more SEI messages, which are not required for the decoding of output pictures but may assist in related processes, such as picture output timing, post-processing of decoded pictures, rendering, error detection, error concealment, and resource reservation. Several SEI messages are specified in H.264 / AVC, H.265 / HEVC, H.266 / VVC, and H.274 / VSEI standards, and the user data SEI messages enable organizations and companies to specify SEI messages for their own use. The standards may include the syntax and semantics for the specified SEI messages but a process for handling the messages in the recipient might not be defined. Consequently, encoders may be required to follow the standard specifying a SEI message when they create SEI message(s), and decoders might not be required to process SEI messages for output order conformance. One of the reasons to include the syntax and semantics of SEI messages in standards is to allow different system specifications to interpret the supplemental information identically and hence interoperate. It is intended that system specifications can require the use of particular SEI messages both in the encoding end and in the decoding end, and additionally the process for handling particular SEI messages in the recipient can be specified.

[0176] The syntax of an SEI message may use the function more_data_in_payload( ). The return value of more_data_in_payload( ) may be specified to be TRUE when the SEI message payload or the VUI payload includes more syntax element(s) compared to the current position for parsing and FALSE otherwise. more_data_in_payload( ) may be specified as follows:If byte_aligned( ) is equal to TRUE and the current position in an SEI message syntax structure or vui_parameters( ) syntax structure is 8 * payloadSize bits from the beginning of the syntax structure, the return value of more_data_in_payload( ) is equal to FALSE.Otherwise, the return value of more_data_in_payload( ) is equal to TRUE.

[0177] The function byte_aligned( ) may be specified as follows:If the current position in the bitstream is a byte-aligned position, i.e., the current position is an integer multiple of 8 bits from the position of the first bit in the bitstream, the return value of byte_aligned( ) is equal to TRUE.Otherwise, the return value of byte_aligned( ) is equal to FALSE.

[0178] Some video coding specifications enable metadata OBUs. A metadata OBU comprises atype field, which specifies the type of metadata.

[0179] A coded video sequence (CVS) may be defined as a sequence of coded pictures in decoding order that is independently decodable and is followed by another coded video sequence or the end of the bitstream.

[0180] A coded layer video sequence (CLVS) may be defined as a sequence of pictures and associated other data within the same scalable layer (e.g., with the same value of nuh_layer_id in VVC) that is decodable independently of other pictures in the same layer.

[0181] Some codecs use a concept of picture order count (POC). A value of POC is derived for each picture and is non-decreasing with increasing picture position in output order. POC therefore indicates the output order of pictures. POC may be used in the decoding process for example for implicit scaling of motion vectors and for reference picture list initialization. Furthermore, POC may be used in the verification of output order conformance. The variable including a POC value of a picture may be referred to as PicOrderCntVal.

[0182] A Decoded Picture Buffer (DPB) may be used in the encoder and / or in the decoder. There may be two reasons to buffer decoded pictures, for references in inter prediction and for reordering decoded pictures into output order. Some coding formats, such as HEVC, provide a great deal of flexibility for both reference picture marking and output reordering, separate buffers for reference picture buffering and output picture buffering may waste memory resources. Hence, the DPB may include a unified decoded picture buffering process for reference pictures and output reordering. A decoded picture may be removed from the DPB when it is no longer used as a reference and is not needed for output.

[0183] Output order may be defined as the order in which the decoded pictures are output from the decoded picture buffer (for the decoded pictures that are to be output from the decoded picture buffer).

[0184] Output time may be defined as a time when a decoded picture is to be output from a decoder or from the DPB of a decoder (for the decoded pictures that are to be output from the DPB), for example as specified by a hypothetical reference decoder specification according to the output timing DPB operation.

[0185] Pictures having the same output order may be defined to mean the same as pictures having the same output time.

[0186] Decoding order may be defined as the order in which syntax elements are processed bythe decoding process. It may be required that syntax elements are ordered in a bitstream in their decoding order.

[0187] An identifier may be defined as a syntax element that identifies a syntax structure. A value of the identifier may for example differ in different instances of the same syntax structure, such as a parameter set. A particular instance of the syntax structure may be referenced through its identifier value. For example, a parameter set that is referenced by the (de)coding of a coded video slice may be identified by providing the identifier value of the parameter set in a header of the coded video slice.

[0188] An indicator (ide) may be defined as a syntax element whose value indicates a selection among more than two values (for which semantics have been specified). An indicator syntax element may have _idc postfix in its name.

[0189] A uniform resource identifier (URI) may be defined as a string of characters used to identify a name of a resource. Such identification enables interaction with representations of the resource over a network, using specific protocols. A URI is defined through a scheme specifying a concrete syntax and associated protocol for the URI. The uniform resource locator (URL) and the uniform resource name (URN) are forms of URI. A URL may be defined as a URI that identifies a web resource and specifies the means of acting upon or obtaining the representation of the resource, specifying both its primary access mechanism and network location. A URN may be defined as a URI that identifies a resource by name in a particular namespace. A URN may be used for identifying a resource without implying its location or how to access it.

[0190] Scalable video coding

[0191] A scalable bitstream may include a "base layer" providing the lowest quality video available and one or more enhancement layers that enhance the video quality when received and decoded together with the lower layers. In order to improve coding efficiency for the enhancement layers, the coded representation of that layer may depend on the lower layers. E.g., the motion and mode information of the enhancement layer can be predicted from lower layers. Similarly, the pixel data of the lower layers can be used to create prediction for the enhancement layer.

[0192] A scalable video codec for quality scalability (also known as SignaLto-Noise or SNR) and / or spatial scalability may be implemented as follows. For a base layer, a conventional non- scalable video encoder and decoder is used. The reconstructed / decoded pictures of the base layer are included in the reference picture buffer for an enhancement layer. In H.264 / AVC, HEVC, and similar codecs using reference picture list(s) for inter prediction, the base layer decoded picturesmay be inserted into a reference picture list(s) for coding / decoding of an enhancement layer picture similarly to the decoded reference pictures of the enhancement layer. Consequently, the encoder may choose a base-layer reference picture as inter prediction reference and indicate its use e.g., with a reference picture index in the coded bitstream. The decoder decodes from the bitstream, for example from a reference picture index, that a base-layer picture is used as inter prediction reference for the enhancement layer. When a decoded base-layer picture is used as prediction reference for an enhancement layer, it is referred to as an inter-layer reference picture.

[0193] Scalability modes or scalability dimensions may include but are not limited to the following:

[0194] Quality scalability: Base layer pictures are coded at a lower quality than enhancement layer pictures, which may be achieved for example using a greater quantization parameter value (i.e., a greater quantization step size for transform coefficient quantization) in the base layer than in the enhancement layer.

[0195] Spatial scalability: Base layer pictures are coded at a lower resolution (i.e., have fewer samples) than enhancement layer pictures. Spatial scalability and quality scalability may sometimes be considered the same type of scalability.

[0196] Bit-depth scalability: Base layer pictures are coded at lower bit-depth (e.g., 8 bits) than enhancement layer pictures (e.g., 10 or 12 bits).

[0197] Dynamic range scalability: Scalable layers represent a different dynamic range and / or images obtained using a different tone mapping function and / or a different optical transfer function.

[0198] Chroma format scalability: Base layer pictures provide lower spatial resolution in chroma sample arrays (e.g., coded in 4:2:0 chroma format) than enhancement layer pictures (e.g., 4:4:4 format).

[0199] Color gamut scalability: enhancement layer pictures have a richer / broader color representation range than that of the base layer pictures - for example the enhancement layer may have UHDTV (ITU-R BT.2020) color gamut and the base layer may have the ITU-R BT.709 color gamut.

[0200] Region-of-interest (ROI) scalability: An enhancement layer represents a spatial subset of the base layer. ROI scalability may be used together with other types of scalabilities, e.g., quality or spatial scalability so that the enhancement layer provides higher subjective quality for the spatial subset.

[0201] View scalability, which may also be referred to as multiview coding. The base layer represents a first set of views, whereas an enhancement layer represents a second set of views.

[0202] Depth scalability, which may also be referred to as depth-enhanced coding. A layer or some layers of a bitstream may represent texture view(s), while other layer or layers may represent depth view(s).

[0203] In all of the above scalability cases, base layer information could be used to code enhancement layer to minimize the additional bitrate overhead.

[0204] Scalability can be enabled in two basic ways. Either by introducing new coding modes for performing prediction of pixel values or syntax from lower layers of the scalable representation or by placing the lower layer pictures to the reference picture buffer (decoded picture buffer, DPB) of the higher layer. The first approach is more flexible and thus can provide better coding efficiency in most cases. However, the second, reference frame -based scalability, approach can be implemented very efficiently with minimal changes to single layer codecs while still achieving majority of the coding efficiency gains available. Essentially a reference frame -based scalability codec can be implemented by utilizing the same hardware or software implementation for all the layers, just taking care of the DPB management by external means.

[0205] In ROI scalability, spatial correspondence of an ROI enhancement layer in relation to its reference layer(s) is indicated. In VVC, scaling windows can be used to indicate this spatial correspondence.

[0206] It has been proposed, e.g. in JVET-O1150 (https: / / www.jvet- experts.org / doc_end_user / documents / 15_Gothenburg / wgl 1 / JVET-O1150-v2.zip), that temporal sublayers could be used for any type of scalability. A mapping of scalability dimensions to sublayer identifiers could be provided e.g. in a VPS or in an SEI message.

[0207] Neural-network post-filter characteristics (NNPFC) and neural-network post-filter activation (NNPFA) SEI messages

[0208] The neural-network post-filter characteristics (NNPFC) SEI message and the neural- network post-filter activation (NNPFA) SEI message have been described in document JVET- AD2006, which specifies a draft amendment to the versatile supplemental enhancement information (VSEI) standard.

[0209] The syntax structure specifying the NNPFC SEI message may be called nn_post_filter_characteristics. The syntax structure specifying the NNPFA SEI message may becalled nn_post_filter_activation.

[0210] The NNPFC SEI message comprises the nnpfc_id syntax element, which includes an identifying number that may be used to identify a post-processing filter. A base post-processing filter is the filter that is included in or identified by the first NNPFC SEI message, in decoding order, that has a particular nnpfc_id value within a coded layer video sequence (CLVS). If an NNPFC SEI message is neither the first NNPFC SEI message, in decoding order, in the current CLVS nor a repetition of the first NNPFC SEI message, in decoding order, that has a particular nnpfc_id value within the current CLVS, the NNPFC SEI message defines an update relative to the base postprocessing filter, and the update relative to the base post-processing filter is applied to obtain a postprocessing filter associated with the nnpfc_id value. The update may be obtained by decoding the coded neural network bitstream in the second NNPFC SEI message (when nnpfc_mode_idc is equal to 0) or through the Uniform Resource Identifier defining the update (when nnpfc_mode_idc is equal to 1). Otherwise (i.e., when there is no update defined by an NNPFC SEI message), the postprocessing filter associated with the nnpfc_id value is assigned to be the same as the base postprocessing filter.

[0211] The NNPFC SEI message comprises nnpfc_mode_idc syntax element, the semantics of which may be defined as follows:

[0212] nnpfc_mode_idc equal to 1 specifies that the base post-processing filter or the update relative to the base post-processing filter associated with the nnpfc_id value is a neural network identified by the Uniform Resource Identifier (URI) nnpfc_uri with the format identified by the tag URI nnpfc_tag_uri.

[0213] nnpfc_mode_idc equal to 0 indicates that this SEI message includes an ISO / IEC 15938- 17 bitstream that specifies the base post-processing filter or updates relative to the base postprocessing filter with the same nnpfc_id value.

[0214] The NNPFC SEI message may also comprise:

[0215] 1. Purpose of the post-processing filter, which may comprise one or more of the following: visual quality improvement, chroma upsampling from the 4:2:0 chroma format to the 4:2:2 or 4:4:4 chroma format, or from the 4:2:2 chroma format to the 4:4:4 chroma format, increasing the width or height, frame rate upsampling, bit depth upsampling, colorization.

[0216] 2. Formatting of the input tensors that are given as input to the neural network inference

[0217] 3. Formatting of the output tensors that are resulting from the neural network inference

[0218] 4. Characterization of the complexity of the neural network

[0219] The NNPFC SEI message syntax includes the nnpfc_num_input_pics_minus 1 syntax element. nnpfc_num_input_pics_minusl plus 1 specifies the number of pictures used as input for the NNPF. The variable numlnputPics may be set equal to nnpfc_num_input_pics_minusl + 1.

[0220] A frame rate upsampling filter may interchangeably be called a picture rate upsampling filter. Such a filter generates or interpolates one or more pictures between a pair of pictures given as input to the filter. It is also possible to have a frame rate upsampling filter where the number of input pictures may be greater than 2. Such a frame rate upsampling filter may generate pictures between more than one pair of input pictures. A frame rate upsampling filter may comprise a neural network, in which case the generation of the interpolated pictures between a pair of input pictures is performed by the inference of the neural network. It is possible to have a frame rate upsampling filter that extrapolates a picture before input picture(s) or after input picture(s), instead of or in addition to between input pictures.

[0221] When the filtering purpose comprises frame rate upsampling, the NNPFC SEI message includes nnpfc_interpolated_pics[ i ] syntax elements for the values of i in the range of 0, inclusive, to nnpfc_num_input_pics_minusl, exclusive. nnpfc_interpolated_pics[ i ] specifies the number of interpolated pictures generated by the NNPF between the i-th and the ( i + 1 )-th picture used as input for the NNPF.

[0222] The NNPFC SEI message syntax may comprise an indication, which may be called nnpfc_absent_input_pic_zero_flag, that indicates how pictures that would not originate from the current bitstream are expected to be replaced in the input tensor. nnpfc_absent_input_pic_zero_flag equal to 1 indicates that the NNPF expects an input picture that is not present in the current bitstream to be represented sample arrays with sample values equal to 0. nnpfc_absent_input_pic_flag equal to 0 indicates that the NNPF expects an input picture that is not present in the current bitstream to be represented by the closest input picture in output order within the current bitstream.

[0223] The NNPFA SEI message specifies the neural-network post-processing filter that may be used for post-processing filtering for the current picture, or for post-processing filtering for the current picture and one or more other pictures. The NNPFA SEI message comprises the nnpfa_target_id syntax element, which indicates that the neural-network post-processing filter with nnpfc_id equal to nnfpa_target_id may be used for post-processing filtering for the indicated persistence. The indicated persistence may be the current picture only (indicated by nnpfa_persistence_flag equal to 0). Alternatively, the NNPF activation may be indicated to be persistent by nnpfa_persistence_flag equal to 1 , in which case the persistence of the NNPF activationmay last until the end of the current CLVS or the next picture, in output order, in the current layer associated with a NNPFA SEI message with the same nnpfa_target_id as the current SEI.

[0224] The NNPFA SEI message syntax may comprise a syntax element indicative if the base post-processing filter or the latest post-processing filter is activated, where the latest post-processing filter is defined by the base post-processing filter relative to which the latest filter update, if any, has been applied. The syntax element may be called nnpfa_target_base_flag. nnpfa_target_base_flag equal to 1 specifies that the target NNPF is the base NNPF with nnpfc_id equal to nnpfa_target_id. nnpfa_target_base_flag equal to 0 specifies that the target NNPF is the NNPF specified by the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id that precedes the first VCE NAE unit of the current picture in decoding order and is not a repetition of the NNPFC SEI message that includes the base NNPF.

[0225] In relation to an NNPFA SEI message, two sets of pictures may be defined, namely nnpfcTargetPictures and nnpfaTargetPictures. nnpfcTargetPictures may be defined to be the set of pictures to which the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id that precedes the current NNPFA SEI message in decoding order pertains. nnpfaTargetPictures may be defined to be the set of pictures for which the target NNPF is activated by the current NNPFA SEI message. It may be required for a conforming bitstream that any picture included in nnpfaTargetPictures shall also be included in nnpfcTargetPictures.

[0226] An NNPF process comprises performing the NNPF inference for given input pictures. The NNPF inference may be performed in a patch-wise manner so that the entire picture area gets filtered. The NNPF inference may be followed by outputting NNPF-gen erated pictures in their increasing index order, where all NNPF-generated pictures that were interpolated by the NNPF are output and those NNPF-generated pictures that correspond to any input pictures to the NNPF are output as specified in the semantics of the NNPFA SEI message.

[0227] A general post-processing filtering process using NNPFs may be described as follows. Input to this process is a bitstream BitstreamToFilter. Output of this process is a list of NNPF output pictures EistNnpfOutputPics. First, BitstreamToFilter is decoded, and the list CroppedDecodedPictures is set to be the list of the cropped decoded pictures in output order resulted from decoding BitstreamToFilter. Second, the filtering process for one picture, as described below, is repeatedly invoked, in output order, for each cropped decoded picture that is in CroppedDecodedPictures and for which one or more NNPFs are activated. The order of the pictures in EistNnpfOutputPics is in output order. It may be required that within EistNnpfOutputPics there shall be no more than one picture pertaining to any particular output time instance. When for anyparticular picture in CroppedDecodedPictures there are multiple NNPFs activated and only one the NNPFs is allowed to be chosen to be applied although any of the NNPFs may be chosen, the above constraint shall apply regardless of which NNPF is chosen to be applied to the particular picture.

[0228] A filtering process for one picture using an NNPF may be described as follows. The filtering process for one picture using an NNPF may be applied to each cropped decoded picture, referred to as the current picture, that is in CroppedDecodedPictures and for which one or more NNPFs are activated. When applying an NNPF to the current picture, the filtered and / or interpolated pictures are generated by the NNPF by applying the NNPF process to the current picture. When applying an NNPF to the current picture, the order of the pictures generated by the NNPF by applying the NNPF process being stored into the output tensor of the NNPF is in output order. When the applied NNPF is the last NNPF that is applied to the current picture, the pictures generated by the NNPF and output by the NNPF process are included into ListNnpfOutputPics, in the same order as when the pictures are stored into the output tensor of the NNPF.

[0229] The use of NNPFC and NNPFA SEI messages for VVC has been described in document JVET-AD2005, which specifies a draft amendment to the versatile video coding (VVC) standard. It is to be understood that NNPFC and NNPFA SEI message may be similarly used for any other video coding specification.

[0230] When NNPFC and NNPFA SEI messages are used for VVC, a decoder selects input pictures for the NNPF. The input pictures may be selected in reverse output order starting from a picture for which the NNPF is activated through an NNPFA SEI message. The input pictures may be indexed, starting from index 0 that is assigned for the picture for which the NNPF is activated through an NNPFA SEI message. In an example, the decoder selects the input picture with index i, where i is greater than 0, to be the latest cropped decoded output picture, in output order, that precedes the input picture with index i-1 in output order. If there is no cropped decoded output picture, in output order, that precedes the input picture with index i- 1 in output order as a result of decoding the bitstream, it may be considered that the input picture with index i is not present in the current bitstream (i.e., missing) and the subsequent input pictures, if any, with index i+1 to numlnputPics-l, inclusive, are likewise missing. A missing input picture may be treated like described above in relation to nnpfc_absent_input_pic_zero_flag syntax element.

[0231] When NNPFC and NNPFA SEI messages are used for VVC and a picture rate upsampling NNPF that interpolates pictures between a single pair of input pictures is activated persistently until the end of the bitstream, the NNPF is applied repeatedly at the end of the bitstream for different sets of input pictures up to but excluding a set of input pictures that would cause creation of anyinterpolated picture after the last picture of the bitstream in output order. In these sets of input pictures, some of the pictures may be missing and may be, for example, replaced by the last picture within the bitstream in output order.

[0232] SEI manifest SEI message

[0233] An SEI manifest SEI message has been specified for example in the H.265 / HEVC standard and the H.266 / VVC standard. An SEI manifest SEI message conveys information on SEI messages that are indicated as expected (i.e., likely) to be present or not present in a coded video sequence (CVS) or a bitstream. Such information may include the following:The indication that certain types of SEI messages are expected (i.e., likely) to be present (although not guaranteed to be present) in the CVS.For each type of SEI message that is indicated as expected (e.g., likely) to be present in the CVS, the degree of expressed necessity of interpretation of the SEI messages of this type, as follows: o The degree of necessity of interpretation of an SEI message type may be indicated as "necessary", "unnecessary", or "undetermined". o An SEI message is indicated by the encoder (i.e., the content producer) as being "necessary" when the information conveyed by the SEI message is considered as necessary for interpretation by the decoder or receiving system in order to properly process the content and enable an adequate user experience; it does not mean that the bitstream is required to include the SEI message in order to be a conforming bitstream. It is at the discretion of the encoder to determine which SEI messages are to be considered as necessary in a particular CVS.The indication that certain types of SEI messages are expected (e.g., likely) not to be present (although not guaranteed not to be present) in the CVS.

[0234] The content of an SEI manifest SEI message may, for example, be used by transport-layer or systems-layer processing elements to determine whether the CVS is suitable for delivery to a receiving and decoding system, based on whether the receiving system can properly process the CVS to enable an adequate user experience or whether the CVS satisfies the application needs.

[0235] It may be required that an SEI NAL unit including an SEI manifest SEI message does not include any other SEI messages other than SEI prefix indication SEI messages. When present in anSEI NAL unit, the SEI manifest SEI message may be required to be the first SEI message in the SEI NAL unit.

[0236] SEI prefix indication SEI message

[0237] An SEI prefix indication SEI message has been specified in the H.265 / HEVC standard and the H.266 / VVC standard. The SEI prefix indication SEI message carries one or more SEI prefix indications for SEI messages of a particular value of SEI payload type (payloadType). Each SEI prefix indication is a bit string that follows the SEI payload syntax of that value of payloadType and includes a number of complete syntax elements starting from the first syntax element in the SEI payload.

[0238] Each SEI prefix indication for an SEI message of a particular value of payloadType indicates that one or more SEI messages of this value of payloadType are expected or likely to be present in the coded video sequence (CVS) and to start with the provided bit string. A starting bit string would typically include only a true subset of an SEI payload of the type of SEI message indicated by the payloadType, may include a complete SEI pay load, and shall not include more than a complete SEI payload. It is not prohibited for SEI messages of the indicated value of payloadType to be present that do not start with any of the indicated bit strings.

[0239] SEI prefix indications should provide sufficient information for indicating what type of processing is needed or what type of content is included. The former (type of processing) indicates decoder-side processing capability, e.g., whether some type of frame unpacking is needed. The latter (type of content) indicates, for example, whether the bitstream includes subtitle captions in a particular language.

[0240] The content of an SEI prefix indication SEI message may, for example, be used by transport-layer or systems-layer processing elements to determine whether the CVS is suitable for delivery to a receiving and decoding system, based on whether the receiving system can properly process the CVS to enable an adequate user experience or whether the CVS satisfies the application needs.

[0241] SEI processing order

[0242] The SEI processing order SEI message has been described in document JVET-AA2027. The SEI processing order SEI message carries information indicating the preferred processing order, as determined by the encoder (i.e., the content producer), for different types of SEI messages that may be present in the bitstream. When an SEI processing order SEI message is present, it is present in the first access unit of the coded video sequence (CVS). The SEI processing order SEI messagepersists in decoding order from the current access unit until the end of the CVS. The SEI processing order SEI message comprises a list of pairs, each pair comprising a SEI payload type value po_sei_payload_type[ i ] and a processing order value po_sei_processing_order[ i ]. po_sei_payload_type[ i ] specifies the value of payloadType for the i-th SEI message for which information is provided in the SEI processing order SEI message. po_sei_processing_order[ i ] indicates the preferred order of processing any SEI message with payloadType equal to po_sei_payload_type[ i ]. po_sei_processing_order[ m ] greater than 0 and less than po_sei_processing_order[ n ] indicates any SEI message with payloadType equal to po_sei_payload_type[ m ], when present, should be processed before any SEI message with payloadType equal to po_sei_payload_type[ n ], when present. po_sei_processing_order[ m ] greater than 0 and equal to po_sei_processing_order[ n ] indicates that the preferred order of processing of SEI messages with payloadTypes equal to po_sei_payload_type[ m ] and po_sei_payload_type[ n ] is unknown or unspecified or determined by external means not specified in this Specification. po_sei_processing_order[ i ] equal to 0 specifies that the preferred order of processing SEI messages with payloadType equal to po_sei_payload_type[ i ] is unknown or unspecified or determined by external means.

[0243] In a later draft of the SEI processing order (SPO) SEI message is described in the subsequent paragraphs.

[0244] Different types of SEI messages may have the same payloadType value and are differentiated by values of syntax elements in the SEI pay load. Such differentiation by values of syntax elements in the SEI payload is to be performed by comparing values sent using po_sei_prefix_data_bit[ i ][ j ] syntax elements (when present) or values sent as SEI messages within a processing order nesting SEI message (when present). For example, neural-network post-filter characteristics (NNPFC) SEI messages can be differentiated by having different nnpfc_id values.

[0245] The SEI processing order SEI message includes the po_id syntax element. po_id contains an identifying number to identify the SPO SEI message. When an SPO SEI message with a particular value of po_id is present in any access unit of a CVS, an SPO SEI message with that particular value of po_id shall be present in the first access unit of the CVS in decoding order. The number of SEI messages and the payloadType codes of the SEI messages indicated within each SPO SEI message with the same value of po_id persist in decoding order from the current access unit until the end of the CVS in output order.

[0246] When an SPO SEI message with a particular value of po_id is present in any access unit of a CVS, an SPO SEI message with that particular value of po_id shall be present in the first accessunit of the CVS in decoding order. The number of SEI messages and the payloadType codes of the SEI messages indicated within each SPO SEI message with the same value of po_id persist in decoding order from the current access unit until the end of the CVS in output order.

[0247] The SPO SEI message can carry one or more SEI prefix indications of a particular payloadType. When present, each SEI prefix indication is a bit string that follows the SEI payload syntax of that value of payloadType and contains a number of complete syntax elements starting from the first syntax element in the SEI payload. These SEI prefix indications should provide sufficient information to determine the specific processing order for types of SEI messages having the same value of payloadType but a different preferred processing order.

[0248] A processing chain may be defined to consist of a list of types of SEI messages identified by an SPO SEI message in the preferred processing order indicated in the SPO SEI message. A processing chain may be additionally or alternatively defined to include one or more processing steps. A processing step may be interchangeably called a process or a processing stage. In some cases, a processing chain may comprise alternative or parallel processing steps. The list of types of SEI message of a processing chain may also comprise types of SEI messages that define properties, rather than processing steps, wherein the properties may, for example, describe the video content at the respective processing step.

[0249] Each type of SEI message in the processing chain indicated by an SPO SEI message is identified by the syntax elements po_sei_payload_type[ i ], po_sei_wrapping_flag[ i ], po_sei_processing_order[ i ] and, when present, po_num_bits_in_prefix_indication_minusl[ i ] and po_prefix_data_bit[ i ] [ j ] .

[0250] An SEI message type is not required to belong to any processing chain and may belong to any number of processing chains identified by SPO SEI messages with different po_id values.

[0251] Each SEI message of an SEI message type identified within the SPO SEI message has the same persistence scope as if the SEI message was carried outside of the SPO SEI message and not identified within an SPO SEI message.

[0252] Processing chains may be alternatives to each other, i.e., such that at most processing chain is chosen to be applied, or they can be complementary, i.e., such that more than one processing chain is chosen and applied separately, with each processing chain generating one output. At most one processing chain can be chosen to be applied by a decoding system at one time.

[0253] The processing order nesting (PON) SEI message includes one or more SEI messages that should be applied only as parts of the processing chain identified by an associated SEI processingorder SEI message and should not be applied in a manner that would contradict with the processing chain identified by the associated SEI processing order SEI message.

[0254] The SEI messages contained in a PON SEI message are referred to as PON-nested SEI messages.

[0255] An encoder may include multiple PON SEI messages in the same access unit. For example, a first PON SEI message in an access unit can contain a PON-nested SEI message that applies to multiple processing chains and one or more other PON SEI messages in the same access unit that apply to a single processing chain only.

[0256] A decoding system may apply a processing chain in a breadth first manner as follows. It is to be understood that other realizations are also possible. For example, a decoding system may choose not to decode the entire bitstream befoe applying the processing chain.- First, the bitstream is decoded, and the list PoPicList is set to be the list of the cropped decoded pictures in output order that resulted from decoding the bitstream, and a processing chain is chosen.- For each of the SEI message types of the chosen processing chain, the following applies in a non-decreasing order of the corresponding po_sei_processing_order[ i ] values:- The following applies for each picture picA in PoPicList in output order, when an SEI message associated with the i-th SEI message type persists for picA or a picture for which an NNPF that generated picA was activated by a preceding process in the processing chain:- When picA is not a cropped decoded picture, the following exceptions apply for the interpretation of the SEI message:- The interface variables for purposes of interpretation of the SEI message are derived from picA instead of the syntax elements indicating properties for the respective cropped decoded picture.- The semantics of the SEI message, or the semantics of the SEI message and, when the SEI message is an NNPFA SEI message, the associated NNPFC SEI message, apply to pictures in PoPicList instead of cropped decoded pictures.When the i-th SEI message type is present in SpoProcessingList, the process implied by the SEI message is performed and PoPicList is updated by replacing pictures with the corresponding processed pictures, if any, resulting from the process and insertingthe other pictures, if any, resulting from the process into PoPicList so that the output order is obeyed.

[0257] A decoding system may apply a processing chain in a depth first manner as follows. It is to be understood that other realizations are also possible. For example, a decoding system may choose to apply a processing stage to a particular picture picA (to obtain a resulting processed picture procPicA) and one or more more subsequent pictures before processing procPicA with the next processing stage.- First, the bitstream is decoded, and the list PoPicList is set to be the list of the cropped decoded pictures in output order that resulted from decoding the bitstream, and a processing chain is chosen.- The following is repeatedly applied, in output order, for each picture picA in PoPicList for which a set of SEI messages associated with SEI message types in SpoProcessingList of the chosen processing chain persist for picA, the following applies:- The following applies for each of the set of SEI messages in a non-decreasing order of the corresponding po_sei_processing_order[ i ] values:- When the current SEI message is not the first in the set of the SEI messages, the following exceptions apply for the interpretation of the SEI message:- The interface variables for purposes of interpretation of the SEI message are derived from the pictures in the updated PoPicList instead of the syntax elements indicating properties for the respective cropped decoded pictures.- The semantics of the SEI message, or of the SEI message and, when the SEI message is an NNPFA SEI message, the associated NNPFC SEI message, apply to pictures in PoPicList instead of cropped decoded pictures.- The process implied by the SEI message is invoked repeatedly, in output order, for picA and each of the pictures in PoPicList that are, or correspond to, interpolated or extrapolated pictures generated by the application of the process implied by any preceding SEI message, if any, to picA. After each invocation of the process, PoPicList is updated by replacing pictures with the corresponding processed pictures, if any, resulting from the process and inserting the other pictures, if any, resulting from the process into PoPicList so that the output order is obeyed.

[0258] Background information on Video Coding for Machines (VCM)

[0259] Reducing the distortion in image and video compression is often intended to increase human perceptual quality, as humans are considered to be the end users, i.e. consuming / watching the decoded image. Recently, with the advent of machine learning, especially deep learning, there is a rising number of machines (i.e., autonomous agents) that analyze data independently from humans and that may even take decisions based on the analysis results without human intervention. Examples of such analysis are object detection, scene classification, semantic segmentation, video event detection, anomaly detection, pedestrian tracking, etc. Example use cases and applications are selfdriving cars, video surveillance cameras and public safety, smart sensor networks, smart TV and smart advertisement, person re-identification, smart traffic monitoring, drones, etc. This may raise the following question: when decoded data is consumed by machines, shouldn’t we aim at a different quality metric -other than human perceptual quality- when considering media compression in intermachine communications? Also, dedicated algorithms for compressing and decompressing data for machine consumption are likely to be different than those for compressing and decompressing data for human consumption. The set of tools and concepts for compressing and decompressing data for machine consumption is referred to here as Video Coding for Machines.

[0260] It is likely that the receiver-side device has multiple “machines” or neural networks (NNs). These multiple machines may be used in a certain combination which is for example determined by an orchestrator sub-system. The multiple machines may be used for example in succession, based on the output of the previously used machine, and / or in parallel. For example, a video which was compressed and then decompressed may be analyzed by one machine (NN) for detecting pedestrians, by another machine (another NN) for detecting cars, and by another machine (another NN) for estimating the depth of all the pixels in the frames.

[0261] Also, please notice that we use the term “receiver-side” or “decoder-side” to refer to the physical or abstract entity or device which includes one or more machines, and runs these one or more machines on some encoded and eventually decoded video representation which is encoded by another physical or abstract entity or device, the “encoder-side device”.

[0262] The encoded video data may be stored into a memory device, for example as a file. The stored file may later be provided to another device.

[0263] Alternatively, the encoded video data may be streamed from one device to another.

[0264] FIG. 5 is a general illustration of the pipeline 500 of Video Coding for Machines. A VCM encoder 504 encodes the input video 502 into a bitstream 506. A bitrate 510 may be computed 508 from the bitstream 506 in order to evaluate the size of the bitstream 506. A VCM decoder 512 decodes the bitstream output 506 by the VCM encoder 504. The output 514 of the VCM decoder512 is referred in FIG. 5 as “Decoded data for machines”. This data 514 may be considered as the decoded or reconstructed video. However, in some implementations of this pipeline 500, this data 514 may not have the same or similar characteristics as the original video 502 which was input to the VCM encoder 504. For example, this data 514 may not be easily understandable by a human by simply rendering the data onto a screen. The output 514 of VCM decoder 512 is then input to one or more task neural networks (516, 518, 520, 522). In FIG. 5, for the sake of illustrating that there may be any number of task-NNs, there are three example task-NNs, namely a task -NN 516 for object detection, a task-NN 518 for object segmentation, a task-NN 3 for object tracking, and a nonspecified one (Task-NN X 522). The goal of VCM is to obtain a low bitrate while guaranteeing that the task-NNs (516, 518, 520, 522) still perform well in terms of the evaluation metric associated to each task.

[0265] As shown in FIG. 5, a performance (532) of the first task (e.g. object detection) is evaluated (524) and, a performance (534) of the second task (e.g. object segmentation) is evaluated (526), a performance (536) of the third task (e.g. object tracking) is evaluated (528), and a performance (538) of the unspecified task is evaluated (530). The evaluated performances (532, 534, 536, 538) are collectively given as 540.

[0266] When a conventional video encoder, such as a H.266 / VVC encoder, is used as a VCM encoder, one or more of the following approaches may be used to adapt the encoding to be suitable to machine analysis tasks (1-4 as follows):

[0267] 1. One or more regions of interest (ROIs) may be detected. An ROI detection method may be used. For example, ROI detection may be performed using a task NN, such as an object detection NN. In some cases, ROI boundaries of a group of pictures or an intra period may be spatially overlaid and rectangular areas may be formed to cover the ROI boundaries. The detected ROIs (or rectangular areas, likewise) may be used in one or more of the following ways:

[0268] The quantization parameter (QP) may be adjusted spatially in a manner that ROIs are encoded using finer quantization step size(s) than other regions. For example, QP may be adjusted CTU-wise.

[0269] The video is preprocessed to include only the ROIs, while the other areas are replaced by one or more constant values or removed.

[0270] A grid is formed in a manner that a single grid cell covers a ROI. Grid rows or grid columns that include no ROIs are downsampled as preprocessing to encoding.

[0271] 2. Quantization parameter of the highest temporal sublayer(s) is increased (i.e. coarserquantization is used) when compared to practices for human watchable video.

[0272] 3. The original video is temporally downsampled as preprocessing prior to encoding. A frame rate upsampling method may be used as postprocessing subsequent to decoding, if machine analysis at the original frame rate is desired.

[0273] 4. A filter is used to preprocess the input to the conventional encoder. The filter may be a machine learning based filter, such as a convolutional neural network.

[0274] The examples described herein provide solutions to at least the following problems: 1. How can a receiver evaluate whether to use one or more post-processing filters, based on one or more criteria?, and 2. How can a receiver be informed about how to use two or more post-processing filters for the same picture or for the same picture unit?

[0275] The examples described herein provide mechanisms for signalling information, from an encoder to a decoder, about a group of one or more post-processing filters, such as neural network based post-processing filters, and / or other post-processing operations. In the following, even though the term post-processing filter is used in embodiments, it is to be understood that the embodiments generally apply to any post-processing operation. The signalled information may be signalled in- band or out-of-band, with respect to the encoded content (such as an encoded video).

[0276] The one or more post-processing filters comprised in the group of one or more postprocessing filters (to which the signalled information applies) may be indicated in several possible ways.

[0277] In one embodiment, the signalled information is part of a nesting SEI message, e.g., an SEI message that comprises one or more other SEI messages. The other SEI messages may comprise, among other, one or more NNPFC SEI messages (as specified in H.266 / VVC and VSEI standard specifications) which describe characteristics of post-processing filters comprised in the group of one or more post-processing filters.

[0278] In another embodiment, the signalled information is part of a (non-nesting) SEI message. The signalled information may comprise one or more filter identifiers (or filter IDs), that identify the post-processing filters comprised in the group of one or more post-processing filters.

[0279] In one embodiment, the signalled information may comprise indicating how two or more post-processing filters comprised in the group of one or more post-processing filters are to be used.

[0280] In one embodiment, the signalled information may comprise indicating that the two ormore post-processing filters are to be used in cascade, for at least one picture of the video sequence. Furthermore, in an additional embodiment, the signalled information may comprise an order for the two or more post-processing filters to be used in cascade.

[0281] In one embodiment, the signalled information may comprise indicating that two or more post-processing filters in a group are alternatives, for at least one picture of the video sequence. It is up to the receiver to choose among those, for example based on criteria such as complexity and availability of resources.

[0282] In one embodiment, the signalled information may comprise indicating that at least one of the inputs to two or more post-processing filters in the group comprises the same data or substantially the same data, for at least one picture of the video sequence.

[0283] In one embodiment, the signalled information may comprise indicating that the two or more post-processing filters take the same data as input, and that their outputs are to be combined based at least on a combination operation, for at least one picture of the video sequence. In one embodiment, the signalled information may comprise an indication of the combination operation. In one embodiment, the combination operation may be performed based at least on one or more coefficients. In one embodiment, the one or more coefficients are predetermined. In another embodiment, the signalled information may comprise the one or more coefficients that may be used for performing the combination operation.

[0284] In one embodiment, the signalled information may comprise indicating that the two or more post-processing filters take the same data as input, and that their outputs may be used separately for different purposes or goals, for at least one picture of the video sequence.

[0285] In one embodiment, the signalled information comprises one or more sets of expected gains associated to respective one or more post-processing filters in the group of post-processing filters, or comprises one set of expected gains associated to all the postfilters in the group of postfilters, for the case where a purpose of the one or more post-processing filters or all the postfilters, respectively, is visual enhancement, or enhancement of machine analysis task(s), or any other purpose where the goal is to obtain a gain in terms of one or more metrics. Such an expected gain may have been determined during or after a development stage of the postfilter, based at least on a performance of the postfilter on a validation dataset. In the set of expected gains associated to a certain postfilter or to the postfilters in a certain group of postfilters, there may be one or more expected gains associated to that postfilter or to the postfilters in that group of postfilters, where different expected gains may be expressed in terms of different metrics.

[0286] In one embodiment, the signalled information comprises one or more sets of actual gains associated to respective one or more post-processing filters in the group of post-processing filters, or comprises one set of actual gains associated to all the postfilters in the group of postfilters, for the case where a purpose of the one or more post-processing filters or all the postfilters, respectively, is visual enhancement, or enhancement of machine analysis task(s), or any other purpose where the goal is to obtain a gain in terms of one or more metrics. In the set of actual gains associated to a certain postfilter or to the postfilters in a certain group of postfilters, there may be one or more actual gains associated to that postfilter or to the postfilters in that group of postfilters, where different actual gains may be expressed in terms of different metrics.

[0287] In one embodiment, the signalled information may comprise one or more filter identifiers (filter IDs) that identify respective one or more post-processing filters with the same purpose, where only one of the identified one or more post-processing filters is used or activated for any input picture. The signalled information may indicate or may be considered to imply for a decoder that either all of the identified one or more post-processing filters are intended to be applied as activated or none of them are intended to be applied. In an additional embodiment, the signalled information may comprise an indication that a portion of the properties of the identified one or more postprocessing filters are shared (i.e., in common) for all of them.

[0288] The examples described herein include mechanisms for signalling information, from an encoder to a decoder, about a group of one or more post-processing filters, such as neural network based post-processing filters, and / or other post-processing operations. In the following, even though the term post-processing filter (or, for short, postfilter or filter) is used in embodiments, it is to be understood that the embodiments generally apply to any post-processing operation. The signalled information may be signalled in-band or out-of-band, with respect to the encoded content (such as an encoded video). Even though most embodiments and examples are described in terms of signalling the information in SEI messages (for example, in a “postfilter group SEI message” (PFG SEI message)) for convenience, other means to carry the signalled information may be possible.

[0289] There may be one or more groups, where each group may comprise one or more postprocessing filters. In the following, for the sake of simplicity, embodiments are described with reference to one group.

[0290] Two or more groups may comprise or refer to the same set of postfilters, or may comprise or refer to respective two or more disjoints sets of postfilters, or may comprise or refer to respective two or more overlapping (i.e., joint) sets of postfilters.

[0291] In at least some of the syntax tables in this document: any byte alignment related syntaxmay have been ignored for the sake of simplicity.

[0292] Indicating the postfilters belonging to a group

[0293] The one or more post-processing filters comprised in a group of one or more postprocessing filters (to which the signalled information applies) may be indicated in different possible ways. In the following, several embodiments describe possible ways for indicating which postfilters belong to a group.

[0294] In one embodiment, the signalled information is part of a nesting SEI message, e.g., an SEI message that comprises one or more other SEI messages. The one or more other SEI messages may comprise one or more SEI messages that describe or comprise information about respective one or more postfilters that belong to the group represented by this nesting SEI message. In one example, the one or more other SEI messages may comprise, among others, one or more NNPFC SEI messages (as specified in H.266 / VVC and VSEI standard specifications) which describe characteristics of respective one or more post-processing filters comprised in the group of one or more post-processing filters.

[0295] In one example, the nesting SEI message comprises the following: an identifier for the group of postfilters, a syntax element indicative of the count of nested NNPFC SEI messages, a first NNPFC SEI message, describing characteristics of a first post-processing filter, a second NNPFC SEI message, describing characteristics of a second post-processing filter, a gain SEI message (or, alternatively, just gain information, represented by one or more syntax elements), and other syntax elements, according to some of the embodiments, such as information about the order of filters, etc. This example may be represented, in terms of syntax, as follows:

[0296] Where pfg_id is an identifier for the group of filters defined by this SEI message, pfg_num_filters_minus2+2 indicates the number of NNPFC SEI messages present in the nesting SEI message postfilter_group(), nn_post_filter_characteristics( ) indicates an NNPFC SEI message,nnpf_gain_sei_message() indicates a gain SEI message, nnpf_order() indicates a syntax structure indicating information about the order of the filters whose characteristics are indicated in the nn_post_filter_characteristics( ) SEI messages when they are used in cascade. It is to be understood that instead of pfg_num_filters_minus2, any other syntax element indicative of the count of the nested NNPFC SEI messages could be used in the syntax structure. For example, pfg_num_filters_minus2 could be replaced by pfg_num_filters_minusl, which indicates the number of NNPFC SEI messages present in the nesting SEI message minus 1. It may be allowed nnpf_order() to be absent, in which case a default order of cascading filters, such as the order that they are listed in the for loop, may be used. In one example, the nnpf_order() syntax structure may comprise pfg_order[ i ] syntax elements for each value of i in the range of 0 to pfg_num_filters_minus2 + 1, inclusive, where examples of specifying the semantics of pfg_order[ i ] are provided subsequently.

[0297] In another example, the nesting SEI message comprises the following: a first NNPFA SEI message activating a first post-processing filter, a second NNPFA SEI message activating a second post-processing filter, a gain SEI message (or, alternatively, just gain information, represented by one or more syntax elements), and other syntax elements, according to some of the embodiments described herein, such as information about the order of filters, etc. This example may be represented, in terms of syntax, as follows:

[0298] Where nn_post_filter_activation( ) indicates an NNPFA SEI message. Other syntax elements and structures are like in the example above.

[0299] In an additional embodiment, where the nesting SEI message comprises one or more NNPFA SEI messages, the persistence scope of the group of postfilters is equal to the shortest persistence scope among the persistence scopes of all the one or more NNPFs activated by the respective one or more NNPFA SEI messages. In an alternative additional embodiment, where the nesting SEI message comprises one or more NNPFA SEI messages, the persistence scope of the group of postfilters is equal to the longest persistence scope among the persistence scopes of all the one or more NNPFs activated by the respective one or more NNPFA SEI messages. In a yetalternative additional embodiment, where the nesting SEI message comprises one or more NNPFA SEI messages, the persistence scopes of all the NNPFs in the group of postfilters are required to be the same or substantially the same.

[0300] In another embodiment, the signalled information is part of a (non-nesting) SEI message. The signalled information may comprise one or more filter identifiers (or filter IDs) that identify respective one or more post-processing filters comprised in the group of one or more post-processing filters, to which other information comprised in the signalled information applies.

[0301] In one example, the SEI message comprises the following: one or more filter IDs, other syntax elements, according to some of the embodiments described herein, such as an identifier for the filter group, information about the gain brought by the one or more postfilters indicated by the one or more filter IDs, information about the order of filters, etc. This example may be represented, in terms of syntax, as follows:

[0302] Where pfg_id is an identifier for the group of filters defined by this SEI message, pfg_num_filters_minus2+2 indicates the number of postfilters in the group of postfilters that this SEI message refers to, pfg_filter_id[ i ] indicates an identifier that identifies an i-th filter whose characteristics are indicated by an NNPFC SEI message (where the identifier is to be matched to nnpfc_id in the NNPFC SEI message), pfg_gain[ i ] indicates a gain that is brought by the i-th postfilter identified by pfg_filter_id[ i ], pfg_order[ i ] indicates an order for the i-th postfilter identified by pfg_filter_id[ i ].

[0303] In various embodiments, filter identifiers and filter group identifiers may have the same value space, i.e., there is no filter identifier value that would also be a filter group identifier value. In various embodiments, where a syntax element, such as pfg_filter_id[ i ], is used, the syntax element may be used to identify a post-filter when its value is equal to a post-filter identifier value or a filter group, when its value is equal to a filter group identifier value.

[0304] In another embodiment, a new mode indicator value is defined for the NNPFC SEI message, which is used to refer to a filter group through a filter group identifier. A filter group may be specified with any other embodiment, for example with a post-filter group SEI message. Syntax elements or groups of syntax elements that are present for the NNPFC SEI message may be used to indicate the properties for the filter group. For example, syntax elements related to complexity may be used to indicate the total complexity of the filter group. The filter group may be activated by an NNPFA SEI message with the nnpfa_target_id value equal to the nnpfc_id value of the NNPFC SEI message defining the filter group. In one example, the following syntax of the NNPFC SEI message may be used:

[0305] Where nnpfc_group_id indicates that this SEI message refers to the filter group defined in the post-filter group SEI message with pfg_id equal to nnpfc_group_id, and the syntax elements and derived variables defined in this SEI message apply to the referenced filter group.

[0306] In another embodiment, a new mode indicator value is defined for the NNPFC SEI message, which is used to define a filter group. Syntax elements or groups of syntax elements that are present for the NNPFC SEI message may be used to indicate the properties for the filter group. For example, syntax elements related to complexity may be used to indicate the total complexity of the filter group. The filter group may be activated by an NNPFA SEI message with the nnpfa_target_id value equal to the nnpfc_id value of the NNPFC SEI message defining the filter group. In one example, the following syntax of the NNPFC SEI message may be used:

[0307] Where nnpfc_pfg_num_filters_minus2+2 indicates the number of postfilters in the group of postfilters that this SEI message refers to, nnpfc_pfg_filter_id[ i ] indicates an identifier that identifies an i-th filter in this filter group whose nnpfc_id is equal to nnpfc_pfg_filter_id[ i ]. The postfilter with nnpfc_id equal to nnpfc_pfg_filter_id[ i ] is defined by one or more other NNPFC SEI messages.

[0308] Related to the above embodiments where the NNPFC SEI message is modified to have a new mode indicator value for defining or referring to a filter group, it is remarked that certain syntax elements, according to some of the embodiments, such as information about the order of filters and / or the gain that is brought by the filter group may be included in the NNPFC SEI message.

[0309] Related to the above embodiments where the NNPFC SEI message is modified to have a new mode indicator value for defining or referring to a filter group, it is remarked that certain syntax elements, such as those related to input tensor generation or output tensor interpretation, may be excluded from the NNPFC SEI message under the condition that the new mode is in use. The input tensor generation for the filter group is available in the NNPFC SEI message for the first filter in execution order, and the output tensor interpretation for the filter group is available in the NNPFC SEI message for the last filter in execution order. In one example, the following syntax of the NNPFC SEI message may be used where the input and output formatting related syntax elements are present when the nnpfc_mode_idc is equal to 0 or 1 but not present for the new mode indicator value:

[0310] In another embodiment, new mode indicator value(s) are defined for the NNPFC SEI message, which are used to indicate that the input picture(s) to the post-filter are output picture(s) resulting from an indicated post-filter. The filter group may be activated by an NNPFA SEI message with the nnpfa_target_id value equal to the nnpfc_id value of the NNPFC SEI message of the last post-filter in the processing order. The NNPFC SEI messages that are linked with each other form a filter group. In one example, the following syntax of the NNPFC SEI message may be used:

[0311] Where nnpfc_mode_idc equal to 2 and 3 have the semantics of nnpfc_mode_idc equal to 0 and 1, respectively, and additionally indicate that this SEI message specifies an NNPF for which the input picture results from the NNPFs indicated by nnpfc_source_filter_id[ i ] values. nnpfc_num_source_filters_minusl + 1 indicates the count of the NNPFs from which input picture(s) to the NNPF defined by this NNPFC SEI message are obtained. nnpfc_source_filter_id[ i ] indicates the nnpfc_id value of the i-th NNPF used to obtain input picture(s) to the NNPF defined by this NNPFC SEI message.

[0312] INDICATING HOW TO USE THE FILTERS IN A GROUP

[0313] In one embodiment, the signalled information may comprise indicating how two or more post-processing filters comprised in the group of one or more post-processing filters are to be used.

[0314] Cascaded filters

[0315] In one embodiment, the signalled information may comprise indicating that the two or more post-processing filters are to be used in cascade, for at least one picture of the video sequence. Furthermore, in an additional embodiment, the signalled information may comprise an order for thetwo or more post-processing filters to be used in cascade.

[0316] FIG. 6 shows an illustrative example, where two postfilters “Filter 1” (604) and “Filter 2” (606) are to be used in cascade, where “Filter 1” (604) performs visual enhancement and “Filter 2” (606) performs frame upsampling (for example, as specified by the respective NNPFC SEI messages that describe the characteristics of the two filters, including their purpose). The “Decoded picture” 602 refers to a cropped decoded picture that is output by a video decoder. The cropped decoded picture 602 represents one of the inputs to the “Filter 1” (604). At least one of the outputs (605) of “Filter 1” (604) represents an input (605) to “Filter 2” (606). At least one of the outputs 607 of “Filter 2” (606) represents the final output 607 of the postprocessing stage and may be used for displaying 608.

[0317] The following is an example syntax table for signalling information for this embodiment:

[0318] Where pfg_cascade_flag indicates whether the filters identified by pfg_filter_id are to be used in cascade for at least one picture of the video sequence, pfg_order[ i ] indicates the order of the i-th filter identified by pfg_filter_id[ i ]. For example, the order may be indicated as an integer number, where 0 represents the first position in the cascaded filter chain, 1 represents the second position, etc.

[0319] In another example syntax table, the postfilters that are to be used in cascade may be a subset with respect to the postfilters belonging to the group of postfilters that this SEI message refers to, as follows:

[0320] Where pfg_cascade_flag[ i ] indicates whether the filter identified by pfg_filter_id[ i ] is to be used in cascade for at least one picture of the video sequence.

[0321] This last example may be useful, for example, when information that is common to a group of postfilters is to be signalled, while only a subset of those postfilters will be used in cascade for at least one picture of the video sequence.

[0322] In an alternative embodiment, there may not be a flag indicating that the postfilters in the group of postfilters are to be used in cascade. Instead, the order indication for a certain postfilter (e.g., as indicating by pfg_order[ i ]) would implicitly comprise the information that the corresponding postfilter is to be used in cascade. For example, if the values of pfg_order[ m ] and pfg_order[ n ] are different and pfg_order[ m ] is greater than pfg_order[ n ], then the m-th filter is to be used before the n-th filter, and at least one of the inputs to the n-th filter is at least one of the outputs of the m-th filter.

[0323] In an embodiment, an activation SEI message may be used to indicate that, for a certain picture, two or more postfilters are to be used. The way that the multiple postfilters are to be used (e.g., in cascade) and the order of the postfilters in the cascade are as indicated by the postfilter group SEI message. The following is an example of an SEI message for activating multiple postfilters for a certain picture.

[0324] Where nnmpfa_num_filters_minusl plus 1 indicates the number of postfilters that areactivated for the current picture, nnmpfa_target_id[ i ] identifies the i-th postfilter that is part of a group of postfilters as indicated in a postfilter group SEI message (for example by pfg_filter_id[ i ]), nnmpfa_cancel_flag indicates whether this SEI message cancels the persistence of the multiple postfilters identified by nnmpfa_target_id, nnmpfa_persistence_flag indicates the persistence of the multiple postfilters identified by nnmpfa_target_id.

[0325] In an embodiment, an activation SEI message may be used to indicate that, for a certain picture, the NNPF activated by this SEI message follows in cascaded processing order the one or more NNPFs indicated in this SEI message.

[0326] The following is an example of an SEI message for activating a postfilter that follows in cascaded processing order the NNPF indicated in this SEI message.

[0327] Where nncpfa_target_id identifies that the postfilter with nnpfc_id equal to nncpfa_target_id is activated by this SEI message, nncpfa_source_id identifies the postfilter that precedes the postfilter activated by this SEI message in cascaded processing order, nncpfa_cancel_flag indicates whether this SEI message cancels the persistence of the postfilter identified by nncpfa_target_id, nncpfa_persistence_flag indicates the persistence of the postfilter identified by nncpfa_target_id. The output tensors of the postfilter identified by the nncpfa_source_id may be used to derive an input tensor for the postfilter identified by nncpfa_target_id.

[0328] The following is an example of an SEI message for activating a postfilter that follows in cascaded processing order the one or more NNPFs indicated in this SEI message.

[0329] Where nncpfa_target_id identifies that the postfilter with nnpfc_id equal to nncpfa_target_id is activated by this SEI message, nncpfa_num_filters_minusl plus 1 indicates the number of postfilters that precede the postfilter activated by this SEI message in cascaded processing order, nncpfa_source_id[ i ] identifies the i-th postfilter that precedes the postfilter activated by this SEI message in cascaded processing order, nncpfa_cancel_flag indicates whether this SEI message cancels the persistence of the postfilters identified by nncpfa_target_id, nncpfa_persistence_flag indicates the persistence of the postfilter identified by nncpfa_target_id. The output tensors of one or more of the postfilters identified by the nncpfa_source_id[ i ] values may be used to derive an input tensor for the postfilter identified by nncpfa_target_id.

[0330] Alternative filters

[0331] In one embodiment, the signalled information may comprise indicating that two or more post-processing filters in a group are alternatives, for at least one picture of the video sequence. It is up to the receiver to choose among those, for example based on criteria such as complexity and availability of resources.

[0332] In one embodiment, the signalled information indicating that two or more post-processing filters in a group are alternatives may be further pre-defined to indicative of, or further indicate, one or more of the following: i) alternatives are for initial selection, e.g., at the start of a bitstream or a CLVS; ii) alternatives can be switched dynamically, e.g., during a bitstream or a CLVS; iii) alternatives have the same number of input pictures and the same number of output pictures and / or can be activated with a single set of activation SEI messages that identify the NNPF group of alternatives or a single NNPF among the NNPF group of alternatives.

[0333] In one embodiment, when the signalled information indicates that two or more postprocessing filters in a group are alternatives and it is pre-defined or indicated that the alternatives can be switched dynamically, e.g., during a bitstream or a CLVS, a decoder or alike may determine picture units at which such switching from a first alternative (being activated before) to a second alternative (to be activated) may take place. In a first additional embodiment, a decoder determines a picture unit suitable for switching from a first alternative to a second alternative, when no input pictures of the second alternative are the same as those for the first alternative. In a second additionalembodiment, a decoder determines a picture unit suitable for switching from a first alternative to a second alternative, when no NNPF output pictures of the first alternative and the second alternative would overlap in output time.

[0334] In one embodiment, the signalled information indicating that two or more post-processing filters in a group are alternatives may comprise a specific identifier value or a specific syntax element indicating that applying no filter is an alternative to one or more NNPFs included in the same NNPF group of alternatives.

[0335] In one embodiment, the two or more postfilters are different with respect to at least one feature or characteristic, such as complexity, gain, spatial and / or temporal upsampling factor, use of one or more auxiliary inputs, etc.

[0336] A typical use case is where the two or more postfilters have the same purpose, e.g., they perform the same or substantially the same task, for example both filters perform visual enhancement, or both filters perform frame -rate upsampling, etc.

[0337] FIG. 7 an illustrative example, where two postfilters “Filter 1” (706) and “Filter 2” (708) are to be used as alternative filters, where both “Filter 1” (706) and “Filter 2” (708) perform visual enhancement (for example, as specified by the respective NNPFC SEI messages that describe the characteristics of the two filters, including their purpose). “Filter 1” (706) and “Filter 2” (708) have different complexity (for example, as specified by the respective NNPFC SEI messages that describe the characteristics of the two filters). “Decoded picture” (702) refers to a cropped decoded picture that is output by a video decoder. The cropped decoded picture 702 represents one of the inputs to the “Filter 1” (706) and “Filter 2” (708). “Switch” (704, 710) refers to an operation that may not be actually present in the post-processing stage, but that represents the choice between one of the two filters (706, 708) for a certain picture, i.e., only one filter is activated for a certain picture. In this example, for a certain picture for which one of these two filters (706, 708) may be activated, a receiver chooses (e.g. using switch 710) which of the two filters (706, 708) to actually use, based at least on the complexity of those two filters. The output 707 of filter 706 or the output 709 of filter 708 is used for display 712. In general, a receiver may choose based on one or more of the following features (1-10 as follows):

[0338] 1. Complexity in terms of number of parameters of the postfilters.

[0339] 2. Complexity in terms of type of parameters of the postfilters.

[0340] 3. Complexity in terms of number of Multiply- Accumulate (MAC) operations per sample of the postfilters.

[0341] 4. Complexity in terms of size required to store the uncompressed parameters of the postfilters.

[0342] 5. Complexity in terms of size required to run or execute the postfilters, which may account for storing both the parameters and any intermediate and final outputs of the postfilters, into temporary memory such as a GPU RAM memory or CPU RAM memory. The required size of the input to the postfilters may influence the size required to run or execute the postfilters.

[0343] 6. Available and / or predicted storage space (where predicted refers to the predicted storage space for the time when the postfilter will be used).

[0344] 7. Available and / or predicted computational capability (e.g., in terms of supported number of MAC operations per sample).

[0345] 8. Available and / or predicted temporary memory, such as GPU RAM memory and CPURAM memory.

[0346] 9. Available and / or predicted electrical power.

[0347] 10. Gain provided by the postfilters, with respect to one or more quality metrics.

[0348] The following is an example syntax table:

[0349] Where pfg_alternative_filters_flag indicates whether the filters identified by pfg_filter_id are alternative filters and thus a receiver can choose which filters to use.

[0350] In another example syntax table, the alternative postfilters may be a subset with respect to the postfilters belonging to the group of postfilters that this SEI message refers to, as follows:

[0351] This last example may be useful, for example, when information that is common to a group of postfilters is to be signalled, while only a subset of those postfilters are alternative filters for at least one picture of the video sequence.

[0352] In another example syntax table, the indication of the alternative filters is included in the NNPFC SEI message, for example as follows:

[0353] Where nnpfc_alternative_filters_flag equal to 1 indicates that the filters identified by nnpfc_source_filter_id[ i ] are alternative filters, and the output picture(s) of any one of these alternative filters may be used as the input picture(s) for the NNPF defined by this NNPFC SEI message. In an additional embodiment, nnpfc_alternative_filters_flag equal to 0 indicates that the filters identified by nnpfc_source_filter_id[ i ] are a group of filters, and the output pictures of all the filters in the group of filters are used as the input pictures for the NNPF defined by this NNPFC SEI message.

[0354] In an alternative embodiment, there may not be a flag indicating that the postfilters in the group of postfilters are to be used as alternative filters. Instead, an order indication may be signaled (e.g., by using a syntax element pfg_order[ i ] for each i-th filter), where the order indication for a certain postfilter (e.g., as indicating by pfg_order[ i ]) would implicitly comprise the information that the corresponding postfilter is to be used as alternative filter. For example, if the values ofpfg_order[ m ] and pfg_order[ n ] are same, then the m-th filter and the n-th filter are to be used as alternative filters.

[0355] In an embodiment, an activation SEI message may be used to indicate that, for a certain picture, two or more postfilters are available. The way that the multiple postfilters are to be used (e.g., as alternative filters) is as indicated by the postfilter group SEI message. The following is an example of an activation SEI message that considers multiple postfilters for a certain picture.

[0356] Where nnmpfa_num_filters_minusl plus 1 indicates the number of postfilters that are alternatives to be activated for the current picture, nnmpfa_target_id[ i ] identifies the i-th postfilter that is part of a group of postfilters as indicated in a postfilter group SEI message, nnmpfa_cancel_flag indicates whether this SEI message cancels the persistence of the multiple postfilters identified by nnmpfa_target_id, nnmpfa_persistence_flag indicates the persistence of the multiple postfilters identified by nnmpfa_target_id.

[0357] In an embodiment, when two or more postfilters are indicated to be used as alternative filters, the signalled information may comprise an indication of one or more main differences among the two or more postfilters. In one example, the main difference between two postfilters that are indicated to be used as alternative filters is the complexity of the two postfilters. In another example, the main difference between two postfilters that are indicated to be used as alternative filters is that one postfilter may improve subjective visual quality and another postfilter may improve objective visual quality. In yet another example, the main differences between two postfilters that are indicated to be used as alternative filters is their complexity and the expected or actual gain that they may provide. In yet another example, the main difference between two postfilters that are indicated to be used as alternative filters is that one postfilter takes data derived from a quantization parameter as an auxiliary input, and another postfilter does not take data derived from a quantization parameter as an auxiliary input.

[0358] Parallel filters

[0359] In one embodiment, the signalled information may comprise indicating that at least one of the inputs to two or more post-processing filters in the group comprises the same data or substantially the same data, for at least one picture of the video sequence.

[0360] Such two or more postfilters may be referred to as parallel filters in this embodiment and related embodiments, or as filters that are run or executed in parallel. However, it is to be understood that such two or more postfilters may be run or executed in any temporal order, either simultaneously (at same time) or sequentially (different times) or at overlapping times. Another possible term for filters for which at least part of their input is same or substantially same may be forking filters.

[0361] FIG. 8 is an illustrative example, where two filters “Filter 1” (804) and “Filter 2” (806) are used in parallel, i.e., they take in the same input data, which is the cropped decoded picture 802 that is output by a video decoder.

[0362] Parallel filters with combination

[0363] In one embodiment, the signalled information may comprise indicating that the two or more post-processing filters take the same data as input, and that their outputs are to be combined based at least on a combination operation, for at least one picture of the video sequence. In one embodiment, the signalled information may comprise an indication of the combination operation. In one embodiment, the combination operation may be performed based at least on one or more coefficients. In one embodiment, the one or more coefficients are predetermined. In another embodiment, the signalled information may comprise the one or more coefficients that may be used for performing the combination operation.

[0364] FIG. 9 is an illustrative example of this embodiment. “Filter 1” (904) and “Filter 2” (906) are two postfilters for the purpose of visual enhancement (for example, as specified by the respective NNPFC SEI messages that describe the characteristics of the two filters). One of the inputs to the two postfilters is a cropped decoded picture denoted as “Decoded picture” 902. The two outputs of the two postfilters, namely output 905 of filter 904 and output 907 of filter 906, are combined by the “Combination” block (908), based on one or more coefficients denoted as “Signalled coefficients” 910. The output 911 of the combination operation 908 is the final output 911 of the post-processing stage and may be used for displaying 912. The combination operation 908 may be a linear combination where the one or more coefficients 910 are used to weight the contribution of the output (905, 907) of each of the two postfilters (904, 906). The one or more coefficients (910) are signalled from an encoder to a decoder, for example as part of the postfilter group SEI message.

[0365] The following is an example syntax table:

[0366] Where pfg_parallel_filters_comb_flag indicates whether the filters identified by pfg_filter_id are to be run in parallel (i.e., same input) and by combining their output.

[0367] In this example, the combination and any combination coefficients may be predefined in a standard specification (e.g., VSEI), such as a weighted average with equal weights for all the postfilters to be run in parallel.

[0368] The following is another example syntax table:

[0369] Where pfg_parallel_filters_comb_flag indicates whether the filters identified by pfg_filter_id are to be run in parallel (i.e., same input) and by combining their output, pfg_comb_mode_idc indicates a combination operation (or combination mode), such as linear combination, pfg_comb_coeff[ i ] indicates combination coefficients to be used for combining the output of the i-th filter indicated by pfg_filter_id[ i ] by means of the combination operation indicated by pfg_comb_mode_idc.

[0370] In another example, an additional flag is used to indicate whether the combination coefficients are present in the signalled information, as follows:

[0371] Where pfg_parallel_filters_comb_flag indicates whether the filters identified by pfg_filter_id are to be run in parallel (i.e., same input) and by combining their output, pfg_comb_coeff_present_flag indicates whether the combination coefficients are present, pfg_comb_mode_idc indicates a combination operation (or combination mode), such as linear combination, pfg_comb_coeff[ i ] indicates combination coefficients to be used for combining the output of the i-th filter indicated by pfg_filter_id[ i ] by means of the combination operation indicated by pfg_comb_mode_idc.

[0372] In an alternative embodiment, there may not be a flag indicating that the postfilters in the group of postfilters are to be used in parallel. Instead, the order indication for a certain postfilter (e.g., as indicating by pfg_order[ i ]) would implicitly comprise the information that the corresponding postfilter is to be used in parallel. For example, if the values of pfg_order[ m ] and pfg_order[ n ] are same, then the m-th filter and the n-th filter are to be used in parallel.

[0373] In one embodiment, the one or more coefficients or an update to the one or more coefficients are signalled for one or more pictures to which the group of postfilters is to be applied, such as within an activation SEI message.

[0374] The following is an example syntax table for an activation SEI message for activating multiple postfilters for a certain picture, demonstrating this embodiment.

[0375] Where nnmpfa_num_filters_minusl plus 1 indicates the number of postfilters that are activated in parallel for the current picture, nnmpfa_target_id[ i ] identifies the i-th postfilter that is part of a group of postfilters as indicated in a postfilter group SEI message and that are to be run in parallel with combination of their output, nnmpfa_comb_coeff[ i ] indicates one or more combination coefficients to be used for combining the output of the i-th filter indicated by nnmpfa_target_id[ i ] by means of the combination operation indicated by pfg_comb_mode_idc in the postfilter group SEI message or by means of a default combination operation (e.g., as specified in a standard specification), nnmpfa_cancel_flag indicates whether this SEI message cancels the persistence of the multiple postfilters identified by nnmpfa_target_id, nnmpfa_persistence_flag indicates the persistence of the multiple postfilters identified by nnmpfa_target_id.

[0376] In another example embodiment for an activation SEI message syntax, a cascaded postfilter activation SEI message comprises the one or more coefficients for the combination as follows:

[0377] Where nncpfa_comb_mode_idc indicates a combination operation (or combination mode), such as linear combination. nncpfa_comb_coeff[ i ] indicates combination coefficients to beused for combining the output of the i-th filter indicated by nncpfa_source_id[ i ] by means of the combination operation indicated by nncpfa_comb_mode_idc. The presence of nncpfa_comb_coeff[ i ] may be conditional on the combination mode, i.e., nncpfa_comb_mode_idc value.

[0378] Parallel filters without combination

[0379] In one embodiment, the signalled information may comprise indicating that the two or more post-processing filters take the same data as input, and that their outputs may be used separately for different purposes or goals, for at least one picture of the video sequence.

[0380] FIG. 10 illustrates an example of this embodiment, where a group of postfilters comprises two postfilters denoted as “Filter 1” (1004) and “Filter 2” (1006). “Filter 1” (1004) performs visual enhancement and “Filter 2” (1006) performs machine enhancement (i.e., enhancement of one or more machine analysis tasks), for example as specified by the respective NNPFC SEI messages that describe the characteristics of the two filters. The two postfilters (1004, 1006) take a cropped decoded picture denoted by “Decoded picture” 1002 as input. The output 1005 of “Filter 1” 1004 is used for displaying 1008 whereas the output 1007 of “Filter 2” 1006 is used as input to one or more machine analysis tasks 1010.

[0381] The following is an example syntax table for this embodiment:

[0382] Where pfg_parallel_filters_nocomb_flag indicates whether the filters identified by pfg_filter_id are to be run in parallel (i.e., same input) without combining their output.

[0383] Alternating filters

[0384] In one embodiment, the signalled information may comprise indicating that the two or more post-processing filters in a group are activated in an alternating manner so that at most one postfilter of the group is activated for any picture. The two or more post-processing filters may, for example, have the same purpose but be trained with a different data set. An encoder may selectwhich of the two or more post-processing filters in the group is to be applied per each picture and indicate the applied filter with the NNPFA SEI message, for instance. The signalled information may be beneficial to conclude, for example, the complexity of the applied post-processing filters to be limited by the highest complexity of any of the filters in the group, instead of, for example, the cumulative complexity of all the filters in the group.

[0385] Indicating an identifier of a group of postfilters

[0386] In at least some of the previous embodiments, the signalled information may comprise indicating an identifier of a group of post-processing filters. The identifier of a group of postprocessing filters may be used to identify which group and associated information should be applied to one or more pictures.

[0387] The following is an example syntax table for the postfilter group SEI message:

[0388] Where pfg_id indicates an identifier for the group of postfilters that this SEI message refers to.

[0389] This identifier may be used or referred to by an activation SEI message, for identifying the group of postfilters to be activated or considered for activation for one or more pictures. The following is an example syntax table for the activation SEI message.

[0390] Where nnmpfa_group_id indicates an identifier for the group of postfilters that this SEImessage refers to, which is to be matched to the value of the syntax element pfg_id of a postfilter group SEI message.

[0391] In one example, two postfilters are part of two groups (a first group and a second group), where the first group is represented by a first PFG SEI message and the second group is represented by a second PFG SEI message. The first PFG SEI message indicates that the two postfilters in the first group are to be used in cascade, whereas the second PFG SEI message indicates that the two postfilters in the second group are to be used as alternative filters.

[0392] In an embodiment, an encoder selects an identifier of a group of post-processing filters (e.g., pfg_id) in a manner that it does not overlap with any of NNPF identifiers (e.g., nnpfc_id).

[0393] In an embodiment, an activation SEI message, such as an NNPFA SEI message, activates a group of post-processing filters when its identifier (e.g., nnpfa_target_id) is equal to an identifier of a group of post-processing filters and activates a single postfilter when its identifier (e.g., nnpfa_target_id) is equal to an NNPF identifier (e.g., nnpfc_id).

[0394] Pre-defined post-processing

[0395] In an embodiment for encoding or decoding, one or more filter identifier values, such as certain nnpfc_id, nnpfa_target_id, pfg_filter_id[ i ] and / or nnmpfa_target_id[ i ] values, indicate operations from a pre-defined set of operations. The pre-defined set of operations may comprise, but might not be limited to, one or more of the following: sample-wise combination with averaging, sample-wise weighted combination, sample-wise multiplicative weighting, horizontal flipping (a.k.a. horizontal mirroring), vertical flipping (a.k.a. vertical mirroring), color space transformation, or resampling (wherein the target spatial resolution may be inferred or indicated). This embodiment may be used together with other embodiments to include pre-defined post-processing with neural- network post-filter(s) in an indicated processing order.

[0396] Combining different embodiments on how to use filters in a group - Separate usages

[0397] One embodiment may comprise features of several of the previous embodiments on how to use filters in a group, where only one type of usage of multiple filters is allowed. The signalled information may comprise indicating how the two or more postfilters are to be used, at least for one picture of the video sequence.

[0398] The following is an example syntax table for this embodiment:

[0399] Where pfg_usage_idc indicates how the filters identified by pfg_filter_id are to be used. Different values of pfg_usage_idc may indicate that the filters are to be used in cascade, or as alternative, or in parallel with combination, or in parallel without combination. For example, the meaning of different values for pfg_usage_idc may be as in the following table:

[0400] In another example embodiment, the usage indicator is added conditioned on a new mode indicator value defined for the NNPFC SEI message, which is used to define a filter group. In one example, the following syntax of the NNPFC SEI message may be used:

[0401] Where nnpfc_pfg_usage_idc is defined like pfc_usage_idc above.

[0402] In another embodiment, an indicator indicates how two or more post-filters or other postprocessing operations are combined to form a set of input pictures for an NNPF. The indicator may be indicative of, but might not be limited to, one or more of the following (where the value in parenthesis is assumed in the syntax example below): (0) alternative filters, (1) sample-wise combination of the output picture(s) of the filters with averaging, (2) sample-wise weighted combination of the output picture(s) of the filters, (3) sample-wise multiplicative weighting of the output picture(s) of the filters, (4) concatenating the output picture(s) of the filters. In one example, the following syntax of the NNPFC SEI message may be used:

[0403] Where nnpfc_filter_usage_idc indicates the method to combine the output pictures of two or more post-filters or other post-processing operations to form a set of input pictures to the NNPF defined by this NNPFC SEI message. nnpfc_filter_usage_idc equal to 0 indicates that the output picture(s) of the NNPF with nnpfc_id equal to nnpfc_source_filter_id[ i ] with any value of imay be used as the input picture(s) to the NNPF defined by this NNPFC SEI message. nnpfc_filter_usage_idc equal to 1 indicates that each sample in location (x,y) in the cldx-th sample array of the picldx-th input picture inputPic[ picldx ][ cldx ][ y ][ x ] is an average of the respective sample of the output picture outputPic[ i ][ picldx ][ cldx ][ y ][ x ] of all the NNPFs identified by nnpfc_id equal to nnpfc_source_filter_id[ i ] for all values of i. nnpfc_filter_usage_idc equal to 2 indicates that each sample in location (x,y) in the cldx-th sample array of the picldx-th input picture inputPic[ picldx ][ cldx ] [ y ] [ x ] is equal toE(outputPic[ i ][ picldx ][ cldx ][ y ][ x ] * (nnpfc_comb_weight_minusl[ i ] + 1)) -?( nnpfc_comb_divisor_minus 1 + 1 ). nnpfc_filter_usage_idc equal to 3 indicates that each sample in location (x,y) in the cldx-th sample array of the picldx-th input picture inputPic[ picldx ][ cldx ][ y ][ x ] is equal to the product of outputPic[ i ][ picldx ][ cldx ][ y ][ x ] for all values of i. nnpfc_filter_usage_idc equal to 3 indicates that the number of input pictures to the NNPF defined by this NNPFC SEI message is equal to the sum of the number of output pictures in the NNPFs identified by nnpfc_id equal to nnpfc_source_filter_id[ i ] for all values of i, and the output pictures are ordered in increasing order of the filter index i to form the input pictures to the NNPF defined by this NNPFC SEI message.

[0404] In one embodiment, the signalled information may comprise indicating that, for at least one picture of the video sequence, some of two or more post-processing filters are to be used in cascade, some other of the two or more postfilters are alternatives, some other of the two or more postfilters are to be used in parallel with combination of their outputs, some other of the two or more postfilters are to be used in parallel without combining their outputs.

[0405] Number of post-filter inferences

[0406] In one embodiment, an encoder may encode, or a decoder may decode, an indication indicative of the number of inferences of an NNPF when the NNPF is applied as one filter among a single invocation or execution of a group of filters. As a response to the indication, the encoder or the decoder may perform the indicated number of inferences of the NNPF for consecutive sets of input pictures. In an additional embodiment, the output pictures resulting from these inferences of the NNPF are provided in output order as input pictures to the subsequent filter(s) in the processing order of the group of filters. In an alternative additional embodiment, the output pictures and the unfiltered input pictures of the NNPF are provided in output order as input pictures to the subsequent filter(s) in the processing order of the group of filters. In one example, the following syntax may be used:

[0407] Where pfg_num_inferences_minusl[ i ] + 1 indicates the number of inferences for the i-th filter in this post-filter group for a single execution of the post-filter group.

[0408] In an alternative embodiment, an encoder or a decoder may infer the number of inferences of an NNPF per a single inference of another NNPF of the same processing chain.

[0409] In an additional embodiment, when a first NNPF precedes a second NNPF in a cascaded processing order and the number of input pictures for the second NNPF (denoted num!nputPicsNnpf2) is greater than the number of output pictures of the first NNPF (denoted numOutputPicsNnpfl), it may be required that num!nputPicsNnpf2 is an integer multiple of numOutputPicsNnpf 1 and the number of inferences for the first NNPF may be derived to be equal to num!nputPicsNnpf2 / numOutputPicsNnpfl per each inference of the second NNPF. In an additional embodiment, when a first NNPF precedes a second NNPF in a cascaded processing order and the number of input pictures for the second NNPF (denoted num!nputPicsNnpf2) is less than the number of output pictures of the first NNPF (denoted numOutputPicsNnpfl), it may be required that numOutputPicsNnpfl is an integer multiple of num!nputPicsNnpf2 and the number of inferences for the second NNPF may be derived to be equal to numOutputPicsNnpfl / num!nputPicsNnpf2 per each inference of the first NNPF.

[0410] In an additional embodiment, a first NNPF takes multiple input pictures for a single inference, the first NNPF carries out picture rate upsampling (i.e., creates intermediate pictures in between at least one pair of consecutive input pictures in output order) and may additionally filter zero or more of the input pictures (e.g., by enhancing visual quality). Let processedPicsNnpfl be the sequence of pictures, in output or display order, comprising the input pictures of the first NNPF that are not filtered by the first NNPF, the pictures filtered by the first NNPF (if any), and the intermediate pictures created by the first NNPF by a single inference of the first NNPF for a single set of input pictures. A second NNPF follows the first NNPF in a cascaded processing order. The pictures of processedPicsNnpfl are used as the input pictures to the second NNPF. Let numProcessedPicsNnpf 1be the number of pictures processedPicsNnpf 1. In an additional embodiment, the number of input pictures to the second NNPF (denoted num!nputPicsNnpf2) is equal to numProcessedPicsNnpfl, and the number of inferences for the second NNPF is derived to be equal to 1 per each inference of the first NNPF. In an alternative additional embodiment, it may be required that the number of input pictures to the second NNPF (denoted numInputPicsNnpf2) is an integer multiple of numProcessedPicsNnpfl, and the number of inferences for the first NNPF is derived to be equal to numInputPicsNnpf2 / numProcessedPicsNnpfl per each inference of the second NNPF. In an alternative additional embodiment, it may be required that numProcessedPicsNnpfl is an integer multiple of the number of input pictures to the second NNPF (denoted numInputPicsNnpf2), and the number of inferences for the second NNPF is derived to be equal to numProcessedPicsNnpfl / numInputPicsNnpf2 per each inference of the first NNPF.

[0411] An additional technical problem to be solved by the examples described herein is encoding signaling and decoding signaling for one or more of the following cases: 1. Multiple activated NNPFs for a picture in a cascading manner, 2. Multiple activated NNPFs (with the same or different purposes) for a picture in an alternative manner (choose to apply none, filter a only, or filter b only), 3. Multiple activated NNPFs (with the same or different purposes) for a picture in a cascading and alternative manner (choose to apply none, filter a only, filter b only, or both filter a and b in a cascaded manner in the indicated order), 4. Multiple activated NNPFs for a picture while the receiver may process the bitstream multiple times to generate multiple different results (for each time, to choose to apply none, filter a only, or filter b only), and 5. Multiple activated NNPF cascades in an alternative manner (choose to apply none, both filter a and b in a cascaded manner in the indicated order, or both filter c and d in a cascaded manner in the indicated order). Specifically, there is a need to identify the input pictures used for an NNPF in a filter cascade.

[0412] Indicating the input pictures for an NNPF in an NNPF group

[0413] Several embodiments are presented below that enable specifying filter groups, such as filter cascades, with selected input pictures to each NNPF in the filter group. These embodiments are advantageous, since they enable, for example, the following cases: 1) The NNPF process of the (i - l)-th NNPF in a cascade outputs a different number of pictures than what is used as input for the i-th NNPF in the cascade. 2) NNPFs are applied in a hierarchical fashion. For example, the same picture rate upsampling NNPF is first applied to obtain 60 Hz from 30 Hz input and subsequently applied to obtain 120 Hz from the 60 Hz signal.

[0414] In one embodiment, an encoder encodes and / or a decoder decodes one or more indications identifying the input pictures used as input for an NNPF in an NNPF group.

[0415] In an embodiment, indication(s) identifying the input pictures for each NNPF in an NNPF group are encoded and / or decoded. In an embodiment, input pictures for the initial NNPF in an NNPF group are inferred and indication(s) identifying the input pictures for each subsequent NNPF (following the initial NNPF in processing order) in an NNPF group are encoded and / or decoded.

[0416] In an embodiment, indication(s) identifying the input pictures for an NNPF in an NNPF group are encoded into and / or decoded from a post-filter group activation SEI message or alike. In an embodiment, indication(s) identifying the input pictures for an NNPF in an NNPF group are encoded into and / or decoded from a postfilter group SEI message or alike.

[0417] In an embodiment, a list of candidate input pictures for the i-th NNPF in the NNPF group is derived according to a pre-defined method. The order of pictures in the list of candidate input pictures is pre-defined and may be, e.g., the inverse output order.

[0418] In an embodiment, a list of candidate input pictures for the i-th NNPF in the NNPF group is derived from input pictures to the O-th NNPF in the NNPF group (i.e., the initial NNPF in the NNPF group, such as the first NNPF of an NNPF group specifiying a filter cascade) or from filtered or interpolated pictures that are output by the NNPF process of the previous NNPFs in the NNPF group. The order of pictures in the list of candidate input pictures is pre-defined and may be, e.g., the inverse output order.

[0419] In an embodiment, a list of candidate input pictures for the i-th NNPF in the NNPF group is derived from candidate input pictures to the O-th NNPF in the NNPF group or from filtered or interpolated pictures that are output by the NNPF process of the previous NNPFs in the NNPF group. The order of pictures in the list of candidate input pictures is pre-defined and may be, e.g., the inverse output order.

[0420] In an embodiment, a list of candidate input pictures for the i-th NNPF in the NNPF group is derived from filtered or interpolated pictures that are output by the NNPF process of the previous NNPF in the NNPF group, i.e., the (i-l)-th NNPF in the NNPF group. The order of pictures in the list of candidate input pictures is pre-defined and may be, e.g., the inverse output order.

[0421] In an embodiment, the candidate input pictures within the list of candidate input pictures may be non-overlapping in output time, i.e., the list comprises only up to one picture per an output time. If the union of the candidate input pictures to the O-th NNPF in the NNPF group and the filtered and interpolated pictures that are output by the NNPF process of the previous NNPFs in the NNPF group includes multiple pictures of the same output time, a pre-defined method may be used to select which picture is kept in the list of candidate pictures. For example, the candidate input pictures tothe O-th NNPF in the NNPF group may be added to the list of candidate input pictures first, followed by adding the pictures that are output by the NNPF process of the i-th NNPF in the NNPF group in increasing order of i. If a picture with a certain output time already exists in the list of candidate input pictures when adding the pictures that are output by the NNPF process, that picture is replaced by the corresponding picture that is having the same output time and is being added to the list of candidate input pictures.

[0422] In an embodiment, a list of candidate input pictures for the i-th NNPF in the NNPF group is derived from all filtered or interpolated pictures that are output by the NNPF process of the previous NNPFs in the NNPF group or are candidate inputs to the O-th NNPF in the NNPF group. The order of pictures in the list of candidate input pictures is pre-defined and may be, e.g., the inverse output order. If there are two pictures having the same output order, their respective order in the list of candidate input pictures may be pre-defined, e.g., a picture resulting from i-th NNPF in the NNPF group may precede the picture having the same output time or the same output order from the j -th NNPF in the NNPF group, when j is less than i.

[0423] In an embodiment, if the list of candidate input pictures for the i-th NNPF in the NNPF group consists of one and only one picture, an encoder and / or a decoder concludes that the picture is selected as an input picture for the i-th NNPF in the NNPF group.

[0424] In an embodiment, given a list of candidate input pictures for the i-th NNPF in the NNPF group, an encoder encodes and / or a decoder decodes one or more indications indicative which of the candidate input pictures in the list are selected as input pictures for the i-th NNPF in the NNPF group.

[0425] In an embodiment, given a list of candidate input pictures for the i-th NNPF in the NNPF group, an encoder encodes and / or a decoder decodes one or both of: an indication for the i-th NNPF if all the pictures in the list of candidate input pictures are input pictures to the NNPF; one or more skip counts corresponding to how many pictures in the list of candidate input pictures are skipped when selecting pictures from the list of candidate input pictures to be used as input pictures. The skip counts may be represented by variable-length codewords, such as ue(v). In an example, the list of candidate input pictures for the i-th NNPF in an NNPF group consists of pictures 5, 4, 3, 2, and 1 (which may, for example, be POC values), the i-th NNPF inputs two input pictures, and the skip counts are 1 and 2. The first skip count, i.e., skip count 1, would cause the first picture in the list of candidate input pictures to be skipped and thus picture 4 to be selected as an input picture. The second skip count would cause two subsequent candidate input pictures, i.e., pictures 3 and 2 to be skipped ad thus picture 1 to be selected as an input picture. In conclusion, in this example pictures4 and 1 are given as input to the i-th NNPF in the NNPF group.

[0426] In an embodiment, given a list of candidate input pictures for the i-th NNPF in the NNPF group, an encoder encodes and / or a decoder decodes indices of the selected input pictures among the list of candidate input pictures. The indices may be represented by fixed-length codewords, such as u(v), where the number of bits v (which may interchangeably be referred to as the length of the syntax element) may be selected as described below.

[0427] In an embodiment, the indices are relative to the list of candidate input pictures, in which case the indices may have a length determined by the number of candidate input pictures, e.g., 1 bit for 2 candidate input pictures, 2 bits for up to 4 candidate input pictures, 3 bits for up to 8 candidate pictures, and so on. In an example, the list of candidate input pictures for the i-th NNPF in an NNPF group consists of pictures 5, 4, 3, 2, and 1 (which may, for example, be POC values), the i-th NNPF inputs two input pictures, and the indices of the selected input pictures are 1 and 4. The first index (1) selects picture 4 from the list and the second index (4) selects picture 1 from the list. In conclusion, in this example pictures 4 and 1 are given as input to the i-th NNPF in the NNPF group.

[0428] In another embodiment, the first index of a selected input picture is relative to the list of the candidate input pictures and each subsequent index is relative to the candidate pictures that follow the previous selected input picture in the list of the candidate input pictures. The indices may have a length that is determined by the number of candidate input pictures remaining in the list, e.g., 1 bit for 2 remaining candidate input pictures, 2 bits for up to 4 remaining candidate input pictures, 3 bits for up to 8 remaining candidate pictures, and so on. In an example, the list of candidate input pictures for the i-th NNPF in an NNPF group consists of pictures 5, 4, 3, 2, and 1 (which may, for example, be POC values), the i-th NNPF inputs two input pictures, and the indices of the selected input pictures are 1 and 2. The first index (1) selects picture 4 from the list and the second index (2) selects picture 1 from the remaining list of candidate pictures, i.e., pictures 3, 2, 1. In conclusion, in this example pictures 4 and 1 are given as input to the i-th NNPF in the NNPF group.

[0429] In any of the above-described embodiments, if the number of pictures following the latest selected input picture in the list of candidate input pictures is equal to the number of input pictures yet to be selected, an encoder and / or a decoder concludes that all the pictures following the latest selected input picture in the list of candidate input pictures are selected as input pictures to the i-th NNPF in the NNPF group.

[0430] In an embodiment, given a list of candidate input pictures for the i-th NNPF in the NNPF group, an encoder encodes and / or a decoder decodes a bit mask, where a bit position corresponds to a picture in the list of candidate input pictures and the value of a bit indicates whether the picture inthe respective position within the list of candidate input pictures is selected as an input picture to the i-th NNPF in the NNPF group. The bit mask may have a length that is up to the number of pictures in the list of candidate input pictures or until it is indicative of the last input picture being selected, whichever is smaller. In an example, the list of candidate input pictures for the i-th NNPF in an NNPF group consists of pictures 5, 4, 3, 2, and 1 (which may, for example, be POC values), the i-th NNPF inputs two input pictures, and the bit mask is 01001. The bit mask indicates that in this example pictures 4 and 1 are given as input to the i-th NNPF in the NNPF group. In another example, the list of candidate input pictures for the i-th NNPF in an NNPF group consists of pictures 5, 4, 3, 2, and 1 (which may, for example, be POC values), the i-th NNPF inputs two input pictures, and the bit mask is 011. The bit mask indicates that in this example pictures 4 and 3 are given as input to the i-th NNPF in the NNPF group.

[0431] In an embodiment, an encoder indicates and / or a decoder decodes an indication of a method used for selecting input pictures for the i-th NNPF in the NNPF group from the list of candidate input pictures. For example, an identifier or index may pre-defined, e.g., in a coding standard, to two or more methods described above, and an encoder may indicate the identifier or index value and / or a decoder may decode the identifier or index value indicative of the method.

[0432] Combining different embodiments on how to use filters in a group - Combined usages

[0433] One embodiment may comprise features of several of the previous embodiments on how to use filters in a group, where two or more types of usage of multiple filters are supported. The signalled information may comprise indicating how the two or more postfilters are to be used, at least for one picture of the video sequence. For example, a first group of postfilters is to be used in cascade, a second group of postfilters is to be used in cascade, where the first group and second group are to be used as alternatives.

[0434] In one example, the signalling information comprises indicating whether a certain element of the group specified in a postfilter group SEI message is a postfilter (e.g., as specified by an NNPFC SEI message) or another group of postfilters. The following is an example syntax table:

[0435] Where pfg_num_elements_minus2+2 indicates the number of elements in the group specified in this postfilter group SEI message, pfg_group_element_type[ i ] indicates a type of an i-th element in the group specified in this postfilter group SEI message, pfg_filter_id[ i ] indicates an identifier for a postfilter (for example, an identifier of a NNPFC SEI message), pfg_group_id[ i ] indicates an identifier of a postfilter group SEI message (to be matched with pfg_id of another PFG SEI message). For example, the meaning of different values for pfg_group_element_type may be as in the following table:

[0436] Another example of syntax table that indicates the elements based on SEI messages (e.g., NNPFC SEI messages and PFG SEI messages) is as follows:

[0437] Where nn_post_filter_characteristics() indicates an NNPFC SEI message, post_filter_group() indicates a postfilter group SEI message.

[0438] FIG. 11 illustrates an example. In this example, “Filter 1” (1106) performs visual enhancement and has low complexity, “Filter 2” (1108) performs spatial upsampling and has low complexity, “Filter 3” (1110) performs visual enhancement and has high complexity, “Filter 4” (1112) performs spatial upsampling and has high complexity . A first group 1121 comprises a second group 1122 and a third group 1123 and indicates that the second group 1122 is to be used as alternative with respect to the third group 1123. The second group 1122 comprises the filters “Filter 1” (1106) and “Filter 2” (1108) and indicates that those filters (1106, 1108) are to be used in cascade. The third group 1123 comprises the filters “Filter 3” (1110) and “Filter 4” (1112) and indicates that those filters (1110, 1112) are to be used in cascade. Referring to the syntax table of the PFG SEI message above, a first PFG SEI message may comprise the following content (1-5 as follows):

[0439] 1. npfg_num_elements_minus2+2 would be equal to 2.

[0440] 2. pfg_usage_idc would be equal to 1 (indicating alternative filters).

[0441] 3. pfg_group_element_type

[0000] would be equal to 1.

[0442] 4. pfg_group_element_type

[0001] would be equal to 1.

[0443] 5. A second and a third PFG SEI messages would be present, where the second PFG SEI message comprises signalling information for the second group 1122 and a third PFG SEI message comprises signalling information for the third group 1123.

[0444] The second PFG SEI message may comprise the following content (1-5 as follows):

[0445] 1. npfg_num_elements_minus2+2 would be equal to 2.

[0446] 2. pfg_usage_idc would be equal to 0 (indicating cascaded filters).

[0447] 3. pfg_group_element_type

[0000] would be equal to 0.

[0448] 4. pfg_group_element_type

[0001] would be equal to 0.

[0449] 5. Two NNPFC SEI messages would be present, indicating characteristics for “Filter 1” (1106) and “Filter 2” (1108).

[0450] The third PFG SEI message may comprise the following content (1-5 as follows):

[0451] 1. npfg_num_elements_minus2+2 would be equal to 2.

[0452] 2. pfg_usage_idc would be equal to 0 (indicating cascaded filters).

[0453] 3. pfg_group_element_type

[0000] would be equal to 0.

[0454] 4. pfg_group_element_type

[0001] would be equal to 0.

[0455] 5. Two NNPFC SEI messages would be present, indicating characteristics for “Filter 3” (1110) and “Filter 4” (1112).

[0456] In FIG. 11, “Decoded picture” (1102) refers to a decoded picture that is output by a video decoder or video codec. The cropped decoded picture 1102 represents one of the inputs to filter 1106 and filter 1110. The output 1107 of filter 1106 is used as an input to filter 1108. The output 1111 of filter 1110 is used as an input to filter 1112. “Switch” (1104, 1114) refers to an operation that may not be actually present in the post-processing stage, but that represents the choice between the second group 1122 comprising filter 1106 and filter 1108 and the third group 1123 comprising filter 1110 and filter 1112 for a certain picture, e.g., only one group of filters is activated for a certain picture. In this example, for a certain picture for which one of these two groups may be activated, a receiver chooses (e.g. using switch 1114) which of the second group 1122 or the third group 1123 to actually use, for example based on complexity. The output 1109 of the second group 1122 comprising filter 1106 and filter 1108 or the output 1113 of the third group 1123 comprising filter 1110 and filter 1112 is used for display 1116.

[0457] FIG. 12 illustrates another example. In this example, using as input decoded picture 1202, “Filter 1” (1202) and “Filter 2” (1204) perform visual enhancement and are to be used in parallel with combination of their outputs. “Combination” 1210 performs a combination operation of the output 1203 of “Filter 1” (1202) and the output 1205 of “Filter 2” (1204). “Filter 3” (1214) performs frame-rate upsampling based on the output 1211 of the combination operation 1210. The output 1215 of “Filter 3” (1214) is used for displaying 1216. “Filter 4” (1206) performs enhancement for one or more machine analysis tasks (1212). The output 1207 of “Filter 4” (1206) is used as input to one or more machine analysis tasks (1212).

[0458] The two outputs of the two postfilters, namely output 1203 of filter 1202 and output 1205 of filter 1204, are combined by the “Combination” block (1210), based on one or more coefficients denoted as “Signalled coefficients” 1208. The combination operation 1210 may be a linear combination where the one or more coefficients 1208 are used to weight the contribution of the output (1203, 1205) of each of the two postfilters (1202, 1204). The one or more coefficients (1208) are signalled from an encoder to a decoder, for example as part of a postfilter group SEI message.

[0459] A first group 1221 comprises a second group 1222 and the postfilter “Filter 4” (1206), and indicates that they are to be used in parallel without combination. The second group 1222 comprises a third group 1223 and the postfilter “Filter 3” (1214), and indicates that they are to be used in cascade. The third group 1223 comprises the two postfilters “Filter 1” (1202) and “Filter 2” (1204), and indicates that they are to be used in parallel with combination 1210 of their outputs.

[0460] Referring to the syntax table of the PFG SEI message above, a first PFG SEI message may comprise the following content (1-6 as follows):

[0461] 1. npfg_num_elements_minus2+2 would be equal to 2.

[0462] 2. pfg_usage_idc would be equal to 3 (indicating parallel filters without combination).

[0463] 3. pfg_group_element_type

[0000] would be equal to 1 (indicating a group)

[0464] 4. pfg_group_element_type

[0001] would be equal to 0 (indicating a postfilter)

[0465] 5. A second PFG SEI message would be present, that comprises signalling information for the second group 1222.

[0466] 6. An NNPFC SEI message would be present, indicating characteristics for “Filter 4” (1206).

[0467] The second PFG SEI message may comprise the following content (1-6 as follows):

[0468] 1. npfg_num_elements_minus2+2 would be equal to 2.

[0469] 2. pfg_usage_idc would be equal to 0 (indicating cascaded filters or groups).

[0470] 3. pfg_group_element_type

[0000] would be equal to 1 (indicating a group).

[0471] 4. pfg_group_element_type

[0001] would be equal to 0 (indicating a postfilter).

[0472] 5. A third PFG SEI message would be present, that comprises signalling information forthe third group 1223.

[0473] 6. An NNPFC SEI message would be present, indicating characteristics for “Filter 3” (1214).

[0474] The third PFG SEI message may comprise the following content (1-5 as follows):

[0475] 1. npfg_num_elements_minus2+2 would be equal to 2.

[0476] 2. pfg_usage_idc would be equal to 2 (indicating parallel filters with combination of their outputs).

[0477] 3. pfg_group_element_type

[0000] would be equal to 0 (indicating a postfilter).

[0478] 4. pfg_group_element_type

[0001] would be equal to 0 (indicating a postfilter).

[0479] 5. Two NNPFC SEI messages would be present, indicating characteristics for “Filter 1” (1202) and “Filter 2” (1204).

[0480] In an embodiment, all filter group identifiers and filter identifiers are unique. A target identifier, such as nnpfa_target_id or nnpfc_pfg_filter_id[ i ], may refer to a post-filter or to a filter group. Consequently, an NNPFA SEI message may activate a filter group, an NNPFC SEI message extended with mode indicating a filter group may define a filter group that comprises another filter group.

[0481] In an embodiment, a postfilter group SEI message identifies a non-cyclic graph representation of filters. Any representation format for a non-cyclic graph may be used.

[0482] Overview of Adaptive Input Picture Selection in Post-Filter Groups

[0483] Described herein is the following (1-3):

[0484] 1. NNPF group characteristics (NNPFGC) SEI message, which includes: a. An identifier of the NNPF group, b. A grouping type of the NNPF group, where 0 indicates a cascade of NNPFs and 1 indicates alternative NNPFs, c. The purpose of the NNPF group, the presence of which is conditioned by the grouping type indicating a cascade of NNPFs. Note: When the grouping type indicates alternative filters, their purpose is available in the NNPFC SEI messages defining the filters, d. The number of members in the NNPF group, where each member can be an NNPF or an NNPF group, e. The identifiers of the NNPFs or NNPF groups belonging to this NNPF group, f. A flag indicating if the complexity information for the NNPF group is present, g. Complexityinformation specified like in the NNPFC SEI message.

[0485] 2. NNPF group activation (NNPFGA) SEI message, which is used for activating an NNPF group of a cascade of NNPFs. Note: NNPFs that are defined as alternatives in an NNPF group are activated through NNPFA SEI messages, which is asserted to provide better compatibility with VSEI v3 and allow defining alternative filters that have different number of input or output pictures (and hence could require being activated at different picture units or having NNPFA SEI messages with different content). For each filter in the NNPF group, a list of candidate input pictures for the i-th NNPF in the NNPF group is derived from all non-overlapping filtered or interpolated pictures that are output by the NNPF process of the previous NNPFs in the NNPF group or are candidate inputs to the O-th NNPF in the NNPF group.

[0486] The proposed NNPFGA SEI message includes (a-d):

[0487] a. nnpfga_target_id, nnpfga_cancel_flag, and nnpfga_persistence_flag specified like nnpfa_target_id, nnpfa_cancel_flag, and nnpfa_persistence_flag, respectively, but applying to an NNPF group rather than to a single NNPF,

[0488] b. The number of filters in the NNPF group

[0489] c. For each filter in the NNPF group, the NNPFGA SEI message includes: i. nnpfga_target_base_flag[ i ] specified for the i-th NNPF in the NNPF group like nnpfa_target_base_flag, ii. An indication if the input pictures for the i-th NNPF in the NNPF group are formed from the list of candidate input pictures without skipping. If not, the following is indicated: 1. The number of input pictures for the i-th NNPF in the NNPF group, 2. The picture count that is skipped in the list of candidate input pictures when selecting input pictures for the i-th NNPF in the NNPF group

[0490] d. nnpfga_num_output_entries[ i ] and nnpfga_output_flag[ i ][ j ] specified for the i-th NNPF in the NNPF group like nnpfa_num_output_entries and nnpfa_output_flag[ j ]

[0491] 3. The general post-processing filtering process using NNPFs is proposed to be updated as follows: a. It is clarified that the bitstream may be filtered multiple times to produce different results, b. Selecting between alternatives indicated in an NNPFGC SEI message is added, c. Applying an NNPF cascade as indicated in NNPFGC and NNPFGA SEI messages is added.

[0492] Description how the design goals are supported

[0493] An embodiment for encoding may be formed according to each of the following numberedparagraphs. An embodiment for decoding may be formed to decode the SEI messages described in each of the following numbered paragraphs and perform the neural-network post-filtering according to the decoded SEI messages.

[0494] 1. Multiple activated NNPFs for a picture in a cascading manner. An NNPFGC SEI message with nnpfgc_grouping_type equal to 0 (cascading) is included in a CLVS and the NNPF group is activated through an NNPFGA SEI message.

[0495] 2. Multiple activated NNPFs (with the same or different purposes) for a picture in an alternative manner (choose to apply none, filter a only, or filter b only). NNPFC SEI messages defining filters a and b are included in a CLVS. An NNPFGC SEI message with nnpfgc_grouping_type equal to 1 (alternatives) is included in the CLVS, including the nnpfc_id values of filters a and b. The alternative NNPFs are activated with NNPFA SEI messages with nnpfa_target_id indicating nnpfc_id values of filters a and b.

[0496] 3. Multiple activated NNPFs (with the same or different purposes) for a picture in a cascading+alternative manner (choose to apply none, filter a only, filter b only, or both filter a and b in a cascaded manner in the indicated order). NNPFC SEI messages defining filters a and b are included in a CLVS. An NNPFGC SEI message with nnpfgc_grouping_type equal to 0 (cascading) is included in the CLVS with filters a and b identified. Another NNPFGC SEI message with nnpfgc_grouping_type equal to 1 (alternatives) is included in the CLVS with nnpfc_id values of filters a and b and the nnpfgc_id value of the filter cascade. Filters a and b are activated with NNPFA SEI messages and the filter cascade a+b is activated with NNPFGA SEI message(s).

[0497] 4. Multiple activated NNPFs for a picture while the receiver may process the bitstream multiple times to generate multiple different results (for each time, to choose to apply none, filter a only, or filter b only). NNPFC SEI messages defining filters a and b are included in a CLVS. Filters a and b are activated with NNPFA SEI messages. Since there is no NNPFGC SEI message indicating that filters and b form a cascade or are alternatives, the receiver may choose to process the bitstream multiple times applying different filters.

[0498] 5. Multiple activated NNPF cascades in an alternative manner (choose to apply none, both filter a and b in a cascaded manner in the indicated order, or both filter c and d in a cascaded manner in the indicated order). NNPFC SEI messages defining filters a, b, c and d are included in a CLVS. A first NNPFGC SEI message with nnpfgc_grouping_type equal to 0 (cascading) is included in the CLVS with filters a and b identified. A second NNPFGC SEI message with nnpfgc_grouping_type equal to 0 (cascading) is included in the CLVS with filters c and d identified. A third NNPFGC SEI message with nnpfgc_grouping_type equal to 1 (alternatives) is included in the CLVS withnnpfgc_id values of the filter cascades a+b and c+d. The filter cascade a+b is activated with NNPFGA SEI message(s) and the filter cascade c+d is activated with other NNPFGA SEI message(s).

[0499] Example syntax and semantics of an NNPFGC SEI message are presented below. Embodiments may be realized using selected features of the presented syntax and semantics.

[0500] Syntax of NNPFGC SEI message

[0501] Semantics of NNPFGC SEI message

[0502] The neural-network post-filter group characteristics (NNPFGC) SEI message specifies a neural network post-filter (NNPF) group. It is indicated by the SEI message if the NNPF group defines an NNPF cascade or defines NNPFs or NNPF groups of NNPF cascades that are alternatives to each other. The use of NNPF groups of NNPF cascades for specific pictures is indicated with neural -network post-filter group activation (NNPFGA) SEI messages.

[0503] nnpfgc_id includes an identifying number that may be used to identify an NNPF group. The value of nnpfgc_id shall be in the range of 0 to 232- 2, inclusive. Values of nnpfgc_id from 256 to 511, inclusive, and from 231to 232- 2, inclusive, are reserved for future use by ITU-T I ISO / IEC. Decoders conforming to this edition of this document encountering an NNPFGC SEI message withnnpfgc_id in the range of 256 to 511, inclusive, or in the range of 231to 232- 2, inclusive, shall ignore the SEI message. The value of nnpfgc_id shall not be equal to any nnpfc_id value of any NNPFC SEI message present in the same CLVS. When the value of nnpfgc_id of an NNPFGC SEI message nnpfgcSeiA is equal to the value of nnpfgc_id of another NNPFGC SEI message nnpfgcSeiB present in the same CLVS, nnpfgcSeiA and nnpfgcSeiB shall be identical.

[0504] nnpfgc_grouping_type equal to 0 indicates that this SEI message specifies a group of cascaded neural-network post-filters. nnpfgc_grouping_type equal to 1 indicates that the NNPFs or NNPF groups identified by the nnpfgc_member_id[ i ] are alternatives to each other out of which only one should be applied.

[0505] nnpfgc_purpose has the semantics of nnpfc_purpose but with the exception that the semantics are specified for the NNPF group defined by this SEI message rather than the NNPF defined by an NNPFC SEI message.

[0506] nnpfgc_num_members_minus2 plus 2 indicates the number of NNPFs or NNPF groups in the NNPF group that this SEI message defines.

[0507] nnpfgc_member_id[ i ] indicates that if there is an NNPF with nnpfc_id equal to nnpfgc_member_id[ i ] defined in the CLVS, the i-th member in the NNPF group defined by this SEI message as follows. If there is an NNPF with nnpfc_id equal to nnpfgc_member_id[ i ] defined in the CLVS, the i-th member in the NNPF group defined by this SEI message is an NNPF that has nnpfc_id equal to nnpfgc_member_id[ i ]. Otherwise (there is no NNPF with nnpfc_id equal to nnpfgc_member_id[ i ] defined in the CLVS), the i-th member in the NNPF group defined by this SEI message is an NNPF group with nnpfgc_id equal to nnpfgc_member_id[ i ].

[0508] One or more of the following constraints in this paragraph may be imposed. When an nnpfgc_member_id[ i ] value references an nnpfgc_id value of an NNPFGC SEI message nnpfgcSei, it is a requirement of bitstream conformance that the NNPFGC SEI message nnpfgcSei shall have nnpfgc_grouping_type equal to 0. When nnpfgc_grouping_type is equal to 0, it is a requirement of bitstream conformance that there is an NNPF with nnpfc_id value equal to nnpfgc_member_id[ i ] defined in the CLVS. When nnpfgc_grouping_type is equal to 1, it is a requirement of bitstream conformance that there is an NNPF with nnpfc_id value equal to nnpfgc_member_id[ i ] or an NNPF group with nnpfgc_id value equal to nnpfgc_member_id[ i ] defined in the CLVS.

[0509] When nnpfgc_grouping_type is equal to 0, the NNPFs with nnpfc_id equal to nnpfgc_member_id[ i ] are performed in cascade in increasing order of i, as activated by an NNPFGA SEI message with nnpfga_target_id equal to nnpfgc_id.

[0510] nnpfgc_complexity_info_present_flag, nnpfgc_parameter_type_idc, nnpfgc_log2_parameter_bit_length_minus3, nnpfgc_num_parameters_idc, nnpfgc_num_kmac_operations_idc, and nnpfgc_total_kilobyte_size have the semantics of nnpfc_complexity_info_present_flag, nnpfc_parameter_type_idc, nnpfc_log2_parameter_bit_length_minus3, nnpfc_num_parameters_idc, nnpfc_num_kmac_operations_idc, and nnpfc_total_kilobyte_size, respectively, but with the exception that the semantics are specified for the NNPF group defined by this SEI message rather than the NNPF defined by an NNPFC SEI message. When nnpfgc_grouping_type is equal to 1, nnpfgc_complexity_info_present_flag shall be equal to 0.

[0511] It is to be understood that variations of the presented NNPFGC SEI message syntax and semantics are possible. For example, nnpfgc_complexity_info_present_flag equal to 1 may be allowed for nnpfgc_grouping_type equal to 0, in which case each syntax element defining a complexity may, for example, specify the maximum value for any alternative defined in the NNPF group. In another example, it may be indicated in the NNPFGC SEI message or pre-defined if the alternatives have the same purpose, and if so, the nnpfgc_purpose syntax element may be present also when nnpfgc_grouping_type is equal to 1. In yet another example, nnpfcg_member_id[ i ] of an NNPF cascade may be allowed to refer to an NNPF group of any grouping type.

[0512] Example syntax and semantics of an NNPFGA SEI message are presented below. Embodiments may be realized using selected features of the presented syntax and semantics.

[0513] Syntax of NNPFGA SEI message

[0514] Semantics of NNPFGA SEI message

[0515] The neural-network post-filter group activation (NNPFGA) SEI message activates or deactivates the possible use of the target neural-network post-processing filter group (NNPFG), identified by nnpfga_target_id, for post-processing filtering of a set of pictures. For a particular picture for which the NNPFG is activated, the target NNPFG is the NNPFG specified by the last NNPFGC SEI message with nnpfgc_id equal to nnpfga_target_id, that precedes the first VCE NAE unit of the current picture in decoding order and the NNPFs of the target NNPFG are defined by the NNPFC SEI messages that have nnpfc_id equal to any nnpfgc_member_id[ i ] value of the target NNPFG and are present in the current picture unit or precede the current picture in decoding order.

[0516] Use of this SEI message requires the definition of the following variables:

[0517] - Input picture width and height in units of luma samples, denoted herein byInitCropped Width [ idx ] and InitCroppedHeight[ idx ], respectively, of the candidate input pictures with index idx in the range of 0 to numCandlnputPics - 1, inclusive, that may be used as input for the NNPF group.

[0518] - Euma sample array InitCroppedYPic[ idx ] and chroma sample arrays InitCropped Cb Pic [ idx ] and InitCroppedCrPic[ idx ], when present, of the candidate input pictures with index idx in the range of 0 to numCandlnputPics - 1, inclusive, that may be used as input for the NNPF group.

[0519] - Bit depth BitDepthy for the luma sample array of the candidate input pictures.

[0520] - Bit depth BitDepthc for the chroma sample arrays, if any, of the candidate input pictures.

[0521] - A chroma format indicator, denoted herein by ChromaFormatldc.

[0522] - When nnpfc_auxiliary_inp_idc is equal to 1, a filtering strength control value arrayStrengthControlVal[ idx ] that shall include real numbers in the range of 0 to 1, inclusive, of the candidate input pictures with index idx in the range of 0 to numCandlnputPics - 1, inclusive.

[0523] Candidate input picture with index 0 corresponds to the picture for which the NNPF group is activated by this NNPFGA SEI message. Candidate input picture with index i in the range of 1 to numCandlnputPics - 1, inclusive, precedes the candidate input picture with index i - 1 in output order. Let cand!nputPicList

[0000] be the list of candidate input pictures in inverse output order.

[0524] nnpfga_target_id indicates the target NNPFG, which is specified by one or more NNPFGC SEI messages that pertain to the current picture and have nnpfgc_id equal to nnpfga_target_id.

[0525] The value of nnpfga_target_id shall be in the range of 0 to 232- 2, inclusive.

[0526] An NNPFGA SEI message with a particular value of nnpfga_target_id shall not be present in a current PU unless one or both of the following conditions are true:

[0527] - Within the current CLVS there is an NNPFGC SEI message with nnpfgc_id equal to the particular value of nnpfga_target_id and nnpfgc_grouping_type equal to 0 present in a PU preceding the current PU in decoding order.

[0528] - There is an NNPFGC SEI message with nnpfgc_id equal to the particular value of nnpfga_target_id and nnpfgc_grouping_type equal to 0 in the current PU.

[0529] When a PU includes both an NNPFGC SEI message with a particular value of nnpfgc_id and an NNPFGA SEI message with nnpfga_target_id equal to the particular value of nnpfgc_id, the NNPFGC SEI message shall precede the NNPFGA SEI message in decoding order.

[0530] nnpfga_cancel_flag equal to 1 indicates that the persistence of the target NNPFG established by any previous NNPFGA SEI message with the same nnpfga_target_id as the current SEI message is cancelled, i.e., the target NNPFG is no longer used unless it is activated by another NNPFGA SEI message with the same nnpfga_target_id as the current SEI message and nnpfga_cancel_flag equal to 0. nnpfga_cancel_flag equal to 0 indicates that the target NNPFG is activated for use.

[0531] nnpfga_persistence_flag specifies the persistence of the target NNPFG for the current layer.

[0532] nnpfga_persistence_flag equal to 0 specifies that the target NNPFG may be used for postprocessing filtering for the current picture only.

[0533] nnpfga_persistence_flag equal to 1 specifies that the target NNPFG may be used for postprocessing filtering for the current picture and all subsequent pictures of the current layer in outputorder until one or more of the following conditions are true:

[0534] - A new CL VS of the current layer begins.

[0535] - The bitstream ends.

[0536] - A picture in the current layer associated with a NNPFGA SEI message with the same nnpfga_target_id as the current SEI message and nnpfga_cancel_flag equal to 1 is output that follows the current picture in output order.

[0537] NOTE - The target NNPFG is not applied for this subsequent picture in the current layer associated with a NNPFGA SEI message with the same nnpfga_target_id as the current SEI message and nnpfga_cancel_flag equal to 1.

[0538] Let the nnpfgcTargetPictures be the set of pictures to which the last NNPFGC SEI message with nnpfgc_id equal to nnpfga_target_id that precedes the current NNPFGA SEI message in decoding order pertains. Let nnpfgaTargetPictures be the set of pictures for which the target NNPFG is activated by the current NNPFGA SEI message. It is a requirement of bitstream conformance that any picture included in nnpfgaTargetPictures shall also be included in nnpfgcTargetPictures.

[0539] nnpfga_num_filters_minus2 plus 2 indicates the number of NNPFs in the NNPF group that this SEI message activates. The value of nnpfga_num_filters_minus2 shall be equal to the value of nnpfgc_num_members_minus2 in the NNPFGC SEI message with nnpfgc_id equal to nnpfga_target_id.

[0540] nnpfga_target_base_flag[ i ] equal to 1 specifies that the i-th NNPF in the target NNPFG is the base NNPF with nnpfc_id equal to nnpfgc_member_id[ i ]. nnpfga_target_base_flag[ i ] equal to 0 specifies that the i-th NNPF in the target NNPFG is the NNPF specified by the last NNPFC SEI message with nnpfc_id equal to nnpfgc_member_id[ i ] that precedes the first VCL NAL unit of the current picture in decoding order and is not a repetition of the NNPFC SEI message that includes the base NNPF.

[0541] nnpfga_input_all_pics_flag[ i ] equal to 1 specifies that the input pictures to the i-th NNPF are selected from the list of candidate input pictures cand!nputPicList[ i ] without skipping. nnpfga_input_all_pics_flag[ i ] equal to 0 specifies that the input pictures to the i-th NNPF are selected from the list of candidate input pictures cand!nputPicList[ i ] in a manner that some candidate input pictures are skipped.

[0542] nnpfga_num_input_pics_minusl[ i ] specifies the number of input pictures for the i-th NNPF in the target NNPFG, i.e., the NNPF activated by the i-th loop entry. When present, nnpfga_num_input_pics_minusl[ i ] shall be equal to nnpfc_num_input_pics_minusl for an NNPF with nnpfc_id equal to nnpfgc_member_id[ i ] of an NNPFGC SEI message with nnpfgc_id equal to nnpfga_target_id. When not present, nnpfga_num_input_pics_minusl[ i ] is inferred to be equal to nnpfc_num_input_pics_minusl for an NNPF with nnpfc_id equal to nnpfgc_member_id[ i ] of an NNPFGC SEI message with nnpfgc_id equal to nnpfga_target_id.

[0543] nnpfga_input_pic_skip_count[ i ][ j ] specifies a j-th picture count that is skipped in the list of candidate input pictures cand!nputPicList[ i ] when selecting input pictures for the NNPF activated by the i-th loop entry. When nnpfga_input_pic_skip_count[ i ] [ j ] is not present, it is inferred to be equal to 0 for all values of j in the range of 0 to nnpfga_num_input_pics_minusl[ i ], inclusive. The variable numCandlnputPics, which indicates the number of candidate input pictures to the NNPF group, is derived as follows: numCandlnputPics = 0 for( j = 0; j <= nnpfga_num_input_pics_minusl

[0000] ; j++ ) numCandlnputPics += 1 + nnpfga_input_pic_skip_count

[0000] [ j ] (1)

[0544] Let cand!nputPicList[ m ] for m in the range of 1 to nnpfga_num_filters_minus2 + 1, inclusive, be a list of pictures in inverse output order that is initially empty and formed in decreasing order of n in the range of 0 to m - 1, inclusive, by including each picture that is output by the NNPF process of the n-th loop entry that has no corresponding picture already present in cand!nputPicList[ m ], and lastly including each picture present in cand!nputPicList

[0000] that has no corresponding picture already present in cand!nputPicList[ m ].

[0545] When a candidate input picture cand!nputPicList[ m ][ idx ] for any value of m in the range of 1 to nnpfga_num_filters_minus2 + 1, inclusive, is an NNPF output picture of the n-th NNPF process with the value of n being less than the value of m, the width and height of the candidate input picture are respectively equal to nnpfcOutputPicWidth and nnpfcOutputPicHeight of the NNPF output picture.

[0546] The list of input pictures inputPicList[ m ] to the NNPF of the m-th loop entry is derived as follows: for( k = 0, candldx = 0; k <= nnpfga_num_input_pics_minusl[ m ]; k++, candldx++ ) { candldx += nnpfga_input_pic_skip_count[ m ] [ k ]inputPicList[ m ] [ k ] = candInputPicList[ m ] [ candldx ] (2)}

[0547] It is a requirement of bitstream conformance that candldx shall not exceed the number of pictures in candInputPicList[ m ].

[0548] It is a requirement of bitstream conformance that the pictures present in inputPicList[ m ], for any value of m in the range of 1 to nnpfga_num_filters_minus2 + 1, inclusive, shall have the same width, height, bit depth, and chroma format.

[0549] For purposes of interpretation of the NNPFC SEI message with nnpfc_id equal to nnpfgc_member_id[ i ] in an NNPFGC SEI message with nnpfgc_id equal to nnpfga_target_id, the following variables are specified for the i-th loop entry:

[0550] - The variables BitDepthy, BitDepthc, and ChromaFormatldc are used as provided for the interpretation of this SEI message.

[0551] - CroppedWidth and CroppedHeight are set equal to the width and height of the pictures in inputPicList[ i ], respectively, in units of luma samples.

[0552] - For each input picture k in the range of 0 to nnpfga_num_input_pics_minusl[ i ], inclusive, the following applies: i) Cropped YPic[ k ], CroppedCbPic[ k ], and CroppedCrPic[ k ], when present, are set equal to respective sample array of inputPicList[ i ][ k ], ii) When nnpfc_auxiliary_inp_idc is equal to 1 for the NNPF with nnpfc_id equal to nnpfgc_member_id[ i ] in an NNPFGC SEI message with nnpfgc_id equal to nnpfga_target_id, the following applies: a) It is a requirement of bitstream conformance that inputPicList[ i ] [ k ] is the same as candInputPicList

[0000] [ idx ] for any value of idx in the range of 0 to numCandlnputPics - 1, inclusive, b) StrengthControlVal[ k ] is set equal to InitStrengthControlVal[ idx ].

[0553] nnpfga_num_output_entries[ i ] specifies the number of nnpfga_output_flag[ i ][ j ] syntax elements present in the NNPFGA SEI message. The value of nnpfga_num_output_entries[ i ] shall be in the range of 0 to NumlnpPicsInOutputTensor, inclusive, for an NNPF with nnpfc_id equal to nnpfgc_member_id[ i ] of an NNPFGC SEI message with nnpfgc_id equal to nnpfga_target_id.

[0554] nnpfga_output_flag[ i ][ j ] equal to 1 specifies that the NNPF-generated picture that corresponds to the input picture having index Inpldx[j ] derived for the i-th NNPF of the target NNPFG is output by the NNPF process activated by this loop entry, where the NNPF process is specified in the semantics of the NNPFC SEI message. nnpfga_output_flag[ i ][ j ] equal to 0specifies that the NNPF-gen erated picture that corresponds to the input picture having index Inpldx[ j ] derived for the i-th NNPF of the target NNPFG is not output by the NNPF process activated by this loop entry. When nnpfga_num_output_entries[ i ] is less than NumlnpPicsInOutputTensor derived for the i-th NNPF of the target NNPFG, nnpfga_output_flag[ i ][ j ] is inferred to be equal to 1 for each value of i in the range of nnpfga_num_output_entries[ i ] to NumlnpPicsInOutputTensor - 1, inclusive. Inpldx[ idx ] specifies the input picture index of the idx-th picture in inverse output order that is present in the output tensor of the NNPF and has a corresponding input picture. Consequently, Inpldx

[0000] is the input picture index of the latest input picture in output order that is present in the output tensor of the NNPF and has a corresponding input picture, Inpldx

[0001] is the input picture index of the second latest input picture in output order that is present in the output tensor of the NNPF and has a corresponding input picture, and so on.

[0555] Let NnpfgaOutputPicList, which is the list of pictures output by NNPF process of the NNPF group in output order, be initially empty and formed in decreasing order of n in the range of 0 to nnpfga_num_filters_minus2 + 1 , inclusive, by including each picture that is output by the NNPF process of the n-th loop entry that has no corresponding picture already present in NnpfgaOutputPicList.

[0556] Changes to the general post-processing filtering process using NNPFs

[0557] In an embodiment, a decoder performs the general post-processing filtering process using NNPFs as described below. In an embodiment, an encoder performs the general post-processing filtering process using NNPFs as described below with a change that CroppedDecodedPictures are obtained by reconstructing pictures as part of the encoding rather than by decoding the bitstream BitstreamToFilter.

[0558] General post-processing filtering process using NNPFs

[0559] Input to this process is a bitstream BitstreamToFilter. Output of this process is a list of NNPF output pictures ListNnpfOutputPics.

[0560] First, BitstreamToFilter is decoded, and the list CroppedDecodedPictures is set to be the list of the cropped decoded pictures in output order resulted from decoding BitstreamToFilter.

[0561] Second, NnpfCand is set to include any single NNPF or any single NNPF group.

[0562] Third, the filtering process for one picture, as described below, is repeatedly invoked, in output order, for each cropped decoded picture that is in CroppedDecodedPictures and for which thesingle NNPF included in NnpfCand or the single NNPF group included in NnpfCand, or one or more NNPFs or NNPF groups defined as alternatives in the NNPF group included in NnpfCand are activated.

[0563] The order of the pictures in ListNnpfOutputPics is in output order.

[0564] Within ListNnpfOutputPics there shall be no more than one picture pertaining to any particular output time instance. When for any particular picture in CroppedDecodedPictures there are multiple NNPFs activated and only one the NNPFs is allowed to be chosen to be applied although any of the NNPFs may be chosen, the above constraint shall apply regardless of which NNPF is chosen to be applied to the particular picture.

[0565] BitstreamToFilter may be processed multiple times to generate multiple different ListNnpfOutputPics through the second and third steps above. For each processing, NnpfCand may be selected to include a different NNPF or NNPF group from any of those selected previously for NnpfCand.

[0566] Filtering process for one picture

[0567] The filtering process specified in this subclause applies to each cropped decoded picture, referred to as the current picture, that is in CroppedDecodedPictures and for which one or more NNPFs or NNPF groups in NnpfCand are activated.

[0568] An NNPF or an NNPF group to be applied to the current picture is selected as follows:

[0569] - If NnpfCand includes a single NNPF and that NNPF is activated for the current picture according to an NNPFA SEI message, that NNPF is selected to be applied to the current picture.

[0570] - Otherwise, if NnpfCand includes an NNPF group with nnpfgc_grouping_type equal to 0 and that NNPF group is activated for the current picture according to an NNPFGA SEI message, that NNPF group is selected to be applied to the current picture.

[0571] - Otherwise (NnpfCand includes an NNPF group with nnpfgc_grouping_type equal to 1), the following applies (1-3):

[0572] 1. A set of candidate NNPFs or NNPF groups candSet is initially empty and then set to include the following:

[0573] The NNPFs that are activated for the current picture according to NNPFA SEImessages and are included in the NNPF group included in NnpfCand.

[0574] - The NNPF groups that are activated for the current picture according to NNPFGA SEI messages and are included in the NNPF group included in NnpfCand.

[0575] 2. For each candidate NNPF or NNPF group candFilter in candSet, the following applies:

[0576] - When one or more of the input pictures of candFilter are input pictures to the NNPF or NNPF group prevFilter that was used in any previous invocation of the filtering process for one picture for the same NnpfCand, candFilter is excluded from candSet.

[0577] 3. Any NNPF or NNPF group remaining in candSet is selected to be applied to the current picture.

[0578] When applying an NNPF to the current picture, the following applies:

[0579] - The filtered and / or interpolated pictures are generated by the NNPF by applying theNNPF process specified in the semantics of the NNPFC SEI message, in a patch-wise manner, to the current picture.

[0580] - The order of the pictures generated by the NNPF by applying the NNPF process being stored into the output tensor of the NNPF is in output order.

[0581] - The pictures generated by the NNPF and output by the NNPF process are included into ListNnpfOutputPics, in the same order as when the pictures are stored into the output tensor of the NNPF.

[0582] When applying an NNPF group to the current picture, the following applies:

[0583] - The filtered and / or interpolated pictures are generated by applying the NNPF process specified in the semantics of the NNPFC SEI message, in a patch-wise manner, as specified in the semantics of the NNPFGA SEI message activating the NNPF group.

[0584] - The pictures in NnpfgaOutputPicList are included into ListNnpfOutputPics, in the same order as the pictures are stored in NnpfgaOutputPicList.

[0585] Example Embodiments

[0586] These example embodiments may be helpful for encoding operation, especially for hierarchical picture rate upsampling.

[0587] In an embodiment, an encoder encodes and / or a decoder decodes a description of an NNPF group that comprises more than one occurrence of the same NNPF. With the more than one occurrence of the same NNPF in the NNPF group, an encoder may indicate hierarchical application of the NNPF. Likewise, a decoder may conclude hierarchical application of the NNPF from the description of the NNPF group that includes more than one occurrence of the same NNPF in the NNPF group. Any embodiment related to indicating input pictures for a particular NNPF in an NNPF group may be used for an NNPF group that comprises more than occurrence of the same NNPF. Each occurrence may use a different set of input pictures, where a latter occurrence may use at least some of the output pictures from a previous occurrence, thereby achieving a hierarchical application of the same NNPF.

[0588] Hierarchical picture rate upsampling

[0589] This section presents an example embodiment where picture rate upsampling NNPFs are applied in a hierarchical fashion. The same picture rate upsampling NNPF is first applied to double the decoded picture rate (e.g., to obtain 60 Hz from 30 Hz) and subsequently applied to double the frame rate further (e.g., to obtain 120 Hz from 60 Hz). The picture rate is thus quadrupled by applying the picture rate upsampling filter hierarchically.

[0590] FIG. 18 presents a concrete example where the filter cascade is activated for picture 4. Decoded pictures 0 and 4 are first input to the NNPF to obtain an interpolated picture 2. Then, decoded picture 0 and interpolated picture 2 are provided to the NNPF to obtain interpolated picture 1. Finally, interpolated picture 2 and decoded picture 4 are provided to the NNPF to obtain interpolated picture 3.

[0591] Accordingly, FIG. 18 shows an example of hierarchical activation of a picture rate upsampling filter NNPF, where NNPF[x] indicates the x-th activation of the NNPF. In this example, the picture rate upsampling NNPF is defined by an NNPFC SEI message with nnpfc_id equal to 0. The NNPF takes two pictures as input and interpolates one picture as output.

[0592] The content of the NNPFGC SEI message defining a filter cascade for hierarchical picture rate upsampling is described as follows, where an integer value implies a particular value of the syntax element, and a descriptor type indicates that encoder has flexibility in selecting the value.

[0593] The content of the NNPFGA SEI message applying the filter cascade for hierarchical picture rate upsampling is described as follows:

[0594] Visual quality enhancement followed by picture rate upsampling

[0595] In this example embodiment, the following two NNPFs are defined:

[0596] - NNPFA, which takes one picture as input and enhances visual quality (nnpfc_id equal to0)

[0597] - NNPFB, which takes two pictures as input and interpolates an intermediate picture(nnpfc_id equal to 1)

[0598] An NNPF group is defined in NNPFGC SEI message indicating that the filters can be applied in cascade so that visual quality enhancement is performed first, followed by picture rate upsampling that uses the enhanced pictures as input.

[0599] FIG. 19 presents a concrete example where the filter cascade is activated for picture 2. Decoded picture 0 is first input to NNPFA to obtain a filtered picture 0, indicated by italics (0). Then, decoded picture 2 is filtered by NNPFA to obtain filtered picture 2, indicated by italics (2). Finally, filtered pictures 0 and 2 are provided to NNPFB to obtain interpolated picture 1. Thus, FIG. 19 shows an example of visual quality enhancment followed by picture rate upsampling, where NNPF[x] are applied in ascending order of x.

[0600] The content of the NNPFGC SEI message is described as follows, where an integer value implies a particular value of the syntax element, and a descriptor type indicates that encoder has flexibility in selecting the value.

[0601] The content of the NNPFGA SEI message applying the filter cascade is described as follows:

[0602] INDICATING PROCESSING ORDER OF A FILTER GROUP WITH RESPECT TO OTHER SEI MESSAGES

[0603] In an embodiment, an encoder includes, into or along a bitstream (e.g., in a SEI processing order SEI message), information indicative of the order of executing a filter group with respect to processing other SEI messages. In an embodiment, a decoder decodes, from or along a bitstream (e.g., from a SEI processing order SEI message), information indicative of the order of executing a filter group with respect to processing other SEI messages.

[0604] In an embodiment, an encoder includes, into or along a bitstream (e.g., in an SEI processing order SEI message), an indication, such as a flag, to indicate if a prefix of an SEI messageis indicated with the SEI message type in relation to their processing order. The prefix of an SEI message may be defined as a selected number of initial bytes of an SEI message. In an embodiment, an encoder includes, within the prefix of an SEI message, a prefix of an SEI message described in any other embodiment, such as postfilter_group( ), nn_multi_post_filter_activation( ), or NNPFC SEI message with a new nnpfc_mode_idc value. The prefix of the another SEI message may comprise an identifier of a NNPF group, a purpose of the NNPF group, complexity of the NNPF group and / or gain of the NNPF group, as described in other embodiments. In an embodiment, a decoder decodes, from or along a bitstream (e.g., from an SEI processing order SEI message), an indication, such as a flag, indicating if a prefix of an SEI message is indicated with the SEI message type in relation to their processing order. In an embodiment, a decoder decodes, from the prefix of an SEI message, a prefix of an SEI message described in any other embodiment, such as postfilter_group( ), nn_multi_post_filter_activation( ), or NNPFC SEI message with a new nnpfc_mode_idc value.

[0605] For example, the following syntax and semantics or alike may be used, where po_prefix_included[i] equal to 0 indicates that no prefix of the SEI message is included in the SEI processing order SEI message, and equal to 1 indicates that a prefix of the SEI message is included in the SEI processing order SEI message. When an SEI message defines or activates a filter group, po_prefix_included[i] may be set equal to 1 and prefix of the SEI message may include the filter group ID.

[0606] Wherein poPayloadType[ i ] is set equal to po_sei_payload_type[ i ]. po_sei_processing_order[ m ] greater than 0 and less than po_sei_processing_order[ n ] indicates any SEI message with payloadType equal to poPayloadType[ m ], when present, should be processed before any SEI message with payloadType equal to poPayloadType[ n ], when present. po_sei_processing_order[ i ] equal to 0 specifies that the preferred order of processing SEI messages with payloadType equal to poPayloadType[ i ] is unknown or unspecified or determined by external means. po_sei_processing_order[ m ] greater than 0 and equal to po_sei_processing_order[ n ] indicates that the m-th SEI message with payloadType equal to poPayloadType[ m ], when present, has the same input data as the n-th SEI message with payloadType equal to poPayloadType[ n ]. po_num_bits_in_prefix_indication_minusl[ i ] plus 1 specifies the number of bits in the i-th SEI prefix indication. po_sei_prefix_data_bit[ i ][ j ] specifies the j-th bit of the i-th SEI prefix indication. The bits po_sei_prefix_data_bit[ i ] [ j ] for j ranging from 0 to po_num_bits_in_prefix_indication_minus 1 [ i ] , inclusive, follow the syntax of the SEI payload with payloadType equal to po_prefix_sei_payload_type[ i ], and include a number of complete syntax elements starting from the first syntax element in the SEI payload syntax, and may or may not include all the syntax elements in the SEI payload syntax. The last bit of these bits (i.e., the bit sei_prefix_data_bit[ i ] [ num_bits_in_prefix_indication_minus 1 [ i ] ] ) may be required to be the last bit of a syntax element in the SEI payload syntax. byte_alignment_bit_equal_to_one may be required to be equal to 1. Other syntax elements are like described above.

[0607] ACTIVATING BASE POST-PROCESSING FILTER(S)

[0608] As described earlier, the NNPFA SEI message as specified in JVET-AC2032 activates the latest updated NNPF with nnpfc_id equal to nnpa_target_id. In an embodiment, the NNPFA SEI message is amended with an indication whether it activates the base NNPF or the latest updated NNPF having nnpfc_id equal to nnpfa_target_id. The activation of the base NNPF could be advantageous, for example, when an update is derived from a first group of pictures of a segment of pictures, where the segment is longer than the first group of pictures, and the picture content of the segment changes later considerably. It is therefore beneficial to enable activation of the base NNPF in the NNPFA SEI message.

[0609] In an embodiment, an encoder detects whether the base NNPF and an updated NNPF is beneficial for one or more consecutive pictures. For example, the encoder may filter the one or more consecutive pictures with the base NNPF and separately with updated NNPF and select the base orupdated NNPF based on which one performs better with one or more quality metrics, such as those discussed in the gain-related embodiments below. The encoder activates the base NNPF or the updated NNPF for the group of one or more consecutive pictures with an NNPFA SEI message.

[0610] In an embodiment, a decoder decodes, from an NNPFA SEI message, whether the base NNPF or an updated NNPF is to be activated, and accordingly activates the base NNPF or the updated NNPF. It is noted that the base NNPF is available in the decoder side, since it is kept as the basis for potential filter updates.

[0611] In one example, the following syntax may be used in the above-described embodiments:

[0612] Where nnpfa_base_flag equal to 1 specifies that the target NNPF is the base NNPF with nnpfc_id equal to nnpfa_target_id. nnpfa_base_flag equal to 0 specifies that the target NNPF is the NNPF specified by the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id, that precedes the first VCL NAL unit of the current picture in decoding order that is not a repetition of the NNPFC SEI message that includes the base NNPF. The other syntax elements have been described earlier or in JVET-AC2032.

[0613] In this example, the neural-network post-filter activation (NNPFA) SEI message activates or de-activates the possible use of the target neural-network post-processing filter (NNPF), identified by nnpfa_target_id, for post-processing filtering of a set of pictures. For a particular picture for which the NNPF is activated, the target NNPF is derived as follows: If nnpfa_base_flag is equal to 1, the target NNPF is the base NNPF with nnpfc_id equal to nnpfa_target_id. Otherwise (nnpfa_base_flag is equal to 0), the target NNPF is the NNPF specified by the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id, that precedes the first VCL NAL unit of the current picture in decoding order that is not a repetition of the NNPFC SEI message that includes the base NNPF.

[0614] In an embodiment, an activation SEI message that activates a group of postfilters isamended with an indication whether it activates the base NNPFs or the latest updated NNPFs of the group of postfilters. The above-described embodiments similarly apply for a group of postfilters. In one example, the following syntax may be used:

[0615] Where nnmpfa_base_flag equal to 1 specifies that the target NNPF group comprises the base NNPFs in the NNPF group with identifier nnmpfa_group_id. nnmpfa_base_flag equal to 0 specifies that the target NNPF group is the NNPF group where each postfilter is specified by the last NNPFC SEI message with nnpfc_id equal to identifiers belonging to the NNPF group, that precedes the first VCL NAL unit of the current picture in decoding order that is not a repetition of the NNPFC SEI message that includes the base NNPF.

[0616] In an embodiment, an activation SEI message that activates a group of postfilters is amended, for each NNPF in the group of postfilters, with an indication whether the SEI activation message activates the base NNPF or the latest updated NNPF. The above-described embodiments similarly apply for a group of postfilters.

[0617] INDICATING ACTIVATION OF AN NNPF FOR AN SEI PROCESSING ORDER

[0618] Extension of an NNPFA SEI message

[0619] In an embodiment, the NNPFA SEI message is extended by adding new syntax elements for controlled selection of input pictures to the NNPF inference at the end of the syntax structure.

[0620] In an embodiment, when an encoder uses the extended NNPFA SEI message syntax that includes new syntax elements for controlled selection of input pictures, the NNPFA SEI message is included in a processing order nesting SEI message. Consequently, decoders or parsers that comply with the earlier NNPFA SEI message syntax do not accidentally decode or parse the NNPFA SEI message with the extended syntax.

[0621] In an embodiment, when a decoder parses an NNPFA SEI message included in aprocessing order nesting SEI message, the decoder also decodes the new syntax elements for controlled selected of input pictures to the NNPF inferece, when present.

[0622] In an embodiment, the NNPFA SEI message is extended with the following syntax. The additional syntax elements are gated by the condition "if( more_data_in_payload( ) )". It is to be understood that embodiments may be similarly realized with any other syntax enabling controlled selection of input pictures to the NNPF inference.

[0623] The semantics of the additional syntax elements may be specified as follows:

[0624] nnpfa_input_selected_pics_flag equal to 1 specifies that the input pictures to the NNPF are selected from the list of candidate input pictures in a manner that some candidate input pictures are skipped. nnpfa_input_selected_pics_flag equal to 0 specifies that the input pictures to the NNPF are selected from the list of candidate input pictures without skipping. When not present, nnpfa_input_selected_pics_flag is inferred to be equal to 0.

[0625] nnpfa_num_input_pics_minusl plus 1 specifies the number of input pictures for the NNPF. When present, nnpfa_num_input_pics_minusl shall be equal to nnpfc_num_input_pics_minus 1 for an NNPF with nnpfc_id equal to nnpfa_target_id. When not present, nnpfa_num_input_pics_minusl is inferred to be equal to nnpfc_num_input_pics_minusl for an NNPF with nnpfc_id equal to nnpfa_target_id.

[0626] nnpfa_input_pic_skip_count[ i ] specifies the i-th picture count that is skipped in the list of candidate input pictures when selecting input pictures for the NNPF. When nnpfa_input_pic_skip_count[ i ] is not present, it is inferred to be equal to 0 for all values of i in the range of 0 to nnpfa_num_input_pics_minusl, inclusive.

[0627] NNPF extended activation (NNPFEA) SEI message

[0628] In an embodiment, an NNPF extended activation (NNPFEA) SEI message is specified. The NNPFEA SEI message includes syntax elements that correspond to those in the NNPFA SEI message and the syntax elements for controlled selection of input pictures to the NNPF inference. It is to be understood that the NNPFEA SEI message may interchangeably be called by another name and / or abbreviation.

[0629] In an embodiment, when an encoder chooses to control the selection of input pictures to the NNPF inference, the encoder encodes an NNPFEA SEI message.

[0630] In an embodiment, when a decoder parses an NNPFEA SEI message and activates an NNPF inference with selected of input pictures, as controlled by the NNPFEA SEI message.

[0631] In an embodiment, the NNPFEA SEI message has the following syntax. It is to be understood that embodiments may be similarly realized with any other syntax enabling controlled selection of input pictures to the NNPF inference.

[0632] nnpfea_target_id, nnpfea_cancel_flag, nnpfea_persistence_flag nnpfea_target_base_flag, nnpfea_no_pref_clvs_flag, nnpfea_no_foll_clvs_flag. nnpfea_num_output_entries, nnpfea_output_flag[ i ], nnpfea_input_selected_pics_flag. nnpfea_num_input_pics_minusl, and nnpfea_input_pic_skip_count[ i ] have the same semantics as nnpfa_target_id, nnpfa_cancel_flag, nnpfa_persistence_flag, nnpfa_target_base_flag, nnpfa_no_pref_clvs_flag, nnpfa_no_foll_clvs_flag, nnpfa_num_output_entries, nnpfa_output_flag[ i ], nnpfa_input_selected_pics_flag, nnpfa_num_input_pics_minusl, and nnpfa_input_pic_skip_count[ i ], respectively.

[0633] Input picture selection for NNPF inference

[0634] In an embodiment, the input pictures for an NNPF are selected in two phases: First, the candidate input pictures are selected the same way as input pictures for an NNPF are selected presently when applying NNPFC and NNPFA SEI messages. In many cases, a list of candidate input pictures includes the current picture for which the NNPF is activated followed by previous cropped decoded pictures in reverse output order. Second, the input pictures are selected from the candidate input pictures based on the syntax elements for controlled selection of input pictures, described above for the NNPFA and NNPFEA SEI messages.

[0635] In an embodiment, the selection of the input pictures for an NNPF in two phases may be specified as described in the following paragraphs. It is to be understood that embodiments may be similarly realized with other similar specifications of the algorithm. The algorithm is presented using the proposed new NNPFA syntax elements but would be likewise realized with the respective NNPFEA syntax elements.

[0636] The variable numCandlnputPics, which indicates the number of candidate input pictures to the NNPF, is derived as follows: numCandlnputPics = 0 for( i = 0; i <= nnpfa_num_input_pics_minusl; i++ )numCandlnputPics += 1 + nnpfa_input_pic_skip_count[ i ]

[0637] For each value of j in the range of 0 to numinferences - 1, inclusive, where numinferences is the count of inferences per activating an NNPF for a particular picture, the following applies. First, a list of candidate input pictures candlnputPic is derived similarly to how the list of input pictures is derived for the use of NNPFC and NNPFA SEI messages presently, as follows:- The arrays candlnputPic [ i ] and inputCandPresentFlag[ i ] for i in the range of 0 to numCandlnputPics - 1, inclusive, representing all the candidate input pictures and the presence of candidate input pictures, respectively, are specified as follows:- When j is greater than 0, for each value of k in the range of 0 to j - 1, inclusive, cand!nputPic[ k ] is set to be currPic and inputCandPresentFlag[ k ] is set equal to 0.- The j-th candidate input picture, cand!nputPic[ j ], is set to be currPic and inputCandPresentFlag[ j ] is set equal to 1.- When numCandlnputPics is greater than 1, the following applies for each value of i in the range of j + 1 to numCandlnputPics - 1, inclusive, in increasing order of i:- If both of the following conditions are true, candlnputPic [ i ] is set to be prevPic and inputPresentFlag[ i ] is set equal to 1:- Either of the following conditions is true:- pictureRateUpsamplingFlag is equal to 1 and currPic is associated with a frame packing arrangement SEI message with frame_packing_arrangement_type equal to 5 and a particular value of fp_current_frame_is_frameO_flag, and there is a cropped decoded output picture prevPic that is the last picture in output order among all cropped decoded output pictures that have nuh_layer_id equal to currLayerld, precede cand!nputPic[ i - 1 ] in output order, and are associated with a frame packing arrangement SEI message with frame_packing_arrangement_type equal to 5 and the same value of fp_current_frame_is_frameO_flag.- pictureRateUpsamplingFlag is equal to 0 or currPic is not associated with a frame packing arrangement SEI message with frame_packing_arrangement_type equal to 5, and there is a cropped decoded output picture prevPic that is the last picture in output order among all cropped decoded output pictures that have nuh_layer_id equal to currLayerld and precede cand!nputPic[ i - 1 ] in output order.- nnpfa_no_prev_clvs_flag is equal to 0 or the coded picture corresponding to prevPic and the current picture are present in the same CL VS.- Otherwise, the following applies:- candInputPic[ i ] is set to be the same picture as candInputPic[ i - 1 ] and inputCandPresentFlag[ i ] is set equal to 0.- It is a requirement of bitstream conformance that, when pictureRateUpsamplingFlag is equal to 1, nnpfc_interpolated_pics[ i - 1 ] shall be equal to 0.

[0638] Second, the input pictures are selected from the candidate input pictures based on the syntax elements for controlled selection of input pictures as follows:- The arrays inputPic[ i ] and inputPresentFlag[ i ] for i in the range of 0 to numlnputPics - 1, inclusive, representing all the input pictures and the presence of input pictures, respectively, are specified as follows: for( i = 0, candldx = 0; i <= nnpfa_num_input_pics_minusl; i++, candldx++ ) { candldx += nnpfa_input_pic_skip_count[ i ] inputPic[ i ] = cand!nputPic[ candldx ] (xx) inputPresentFlag[ i ] = inputCandPresentFlag[ candldx ]}- For each value of i in the range 0 to numlnputPics - 1, inclusive, it is a requirement of bitstream conformance that when inputPresentFlag[ i ] is equal to 0 and nnpfc_input_pic_output_flag[ i ] is equal to 1, the value of nnpfa_output_flag[ idx ] shall be equal to 0 for the value of idx such that Inp Idx [ idx ] is equal to i.

[0639] An NNPF controlled activation SEI message may be defined as an NNPFA SEI message including syntax elements for controlled selection of input pictures to the NNPF inference or an NNPFEA SEI message.

[0640] In an embodiment, an encoder selects a picture unit where an NNPF controlled activation SEI message is included in a manner that the following bitstream conformance constraint is obeyed: it is required for bitstream conformance that inputPic

[0000] shall correspond, in terms of output time, to the current picture, a temporally interpolated picture, or a temporally extrapolated picture and shall not correspond, in terms of output time, to any cropped decoded picture that is not the current picture.

[0641] In an embodiment, an encoder selects a picture unit where an NNPF controlled activation SEI message is included in a manner that the following bitstream conformance constraint is obeyed: it is required for bitstream conformance that inputPic

[0000] shall not correspond to or precede, in terms of output time, the last cropped decoded picture that precedes the current picture in output order.

[0642] In an embodiment, when an NNPFEA SEI message is not included in a PON SEI message, an encoder selects a picture unit where the NNPFEA SEI message is included in a manner that the following bitstream conformance constraint is obeyed: it is required for bitstream conformance that inputPic

[0000] shall be the cropped decoded picture corresponding to the current picture. Interchangeably, the requirement may be phrased as follows: when an NNPFEA SEI message is not included in a PON SEI message, nnpfea_input_pic_skip_count

[0000] shall be equal to 0.

[0643] Let a current NNPF activation be defined as the NNPF activation caused by a PON-nested NNPF controlled activation SEI message. In an embodiment, an encoder selects a picture unit where the PON-nested NNPF controlled activation SEI message causing the current NNPF activation is included in a manner that the NNPF activations of any previous processing stage precede, in output order, the current NNPF activation when an output of the previous processing stage is needed as input to the current processing stage. Such an encoder operation enables the decoding system to apply a processing chain in a depth first manner as described earlier.

[0644] In an embodiment, an encoder indicates, e.g., in an SPO SEI message, if a processing chain is intended to be applied in a depth first manner as described earlier. Without loss of generality, such an indication may be referred to as po_depth_first_flag. When po_depth_first_flag is equal to 1, a decoding system should apply the processing chain in a depth first manner as described earlier. In one embodiment, when po_depth_first_flag is equal to 0, a decoding system should apply the processing chain in a breadth first manner as described earlier. In another embodiment, when po_depth_first_flag is equal to 0, it is not specified whether a decoding system should apply the processing chain in a depth first manner or a breadth first manner or some other manner.

[0645] In an example, a processing chain includes the following ordered processing stages:1) Quality enhancement NNPF, which takes three consecutive cropped decoded pictures, in output order, as input and filters the mid-most picture. The quality enhancement NNPF is activated for every third picture in output order. The mid-most input picture has a lower quality than the surrounding input pictures and hence the quality of the mid-most input picture is enhanced to be closer to the quality of the surrounding input pictures.2) Picture rate upsampling NNPF, which takes two consecutive cropped decoded or processed pictures, in output order, as input and generates one intermediate picture as output.When three consecutive cropped decoded pictures in output order have output times 0, 2, and 4, and the quality enhancement is to be applied to picture 2, an encoder locates the NNPFA SEI message that activates the quality enhancement NNPF in picture unit 4. Moreover, the encoder locates twoPON-nested NNPF activation SEI messages for the picture rate upsampling NNPF in picture unit 4. The first PON-nested activation SEI message may be an NNPFA SEI message including syntax elements for controlled selection of input pictures to the NNPF inference or an NNPFEA SEI message and have nnpfa_input_skip_count

[0000] or nnpfea_input_skip_count

[0000] , respectively, equal to 1, indicating that picture 4 is skipped when selecting the O-th input picture from the candidate picture list and picture 2 is used instead. The second PON-nested activation SEI message may, for example, be an NNPFA SEI message without syntax elements for controlled selection of input pictures to the NNPF inference. An encoder may indicate, e.g., in an SPO SEI message, that the processing chain is intended to be applied in a depth first manner as described earlier. A decoding system may apply the processing chain in a depth first manner as described earlier. Applying the processing chain in a depth first manner performs the quality enhancement NNPF for pictures 0, 2, 4 to obtain a filtered picture 2, the picture rate upsampling NNPF with picture 0 and filtered picture 2 as inputs, and the picture rate upsampling NNPF with filtered picture 2 and picture 4 as inputs.

[0646] In an embodiment, an encoder indicates, e.g., in an SPO SEI message, one or more syntax elements indicative of a maximum latency required for the processing chain. The maximum latency may, for example, be defined as a value that is greater than or equal to the maximum number of pictures in any candlnputPicEist that correspond, in output time, to any cropped decoded picture. Alternatively, the maximum latency may, for example, be defined as a value that is greater than or equal to the maximum difference of output times of pictures that are present in any single candlnputPicEist and correspond, in output time, to any cropped decoded picture. It may be considered that the maximum latency indicates the latency required for the performing the processing chain without considering the processing time needed to perform the processing stages.

[0647] Examples of usage

[0648] Example 1 is described as follows: An interpolated picture is enhanced since the picture rate upsampling NNPF (also called the interpolator NNPF) has low complexity and produces relatively poor picture quality. Thus, a visual enhancement filter is to be applied only on the interpolated pictures, but not the cropped decoded pictures. The picture rate upsampling NNPF inputs two pictures are outputs one interpolated picture. The visual enhancement NNPF inputs one pictures and outputs the resulting filtered picture.

[0649] An encoding system according to an embodiment handles example 1 as follows:• An SPO SEI message is encoded, including: o NNPFC and / or NNPFA SEI message types for the picture rate upsampling NNPFo NNPFA SEI message type and optionally NNPFC SEI message type for the visual enhancement NNPF• NNPFC and / or NNPFA SEI messages for the picture rate upsampling NNPF are encoded.• NNPFC SEI message(s) for the visual enhancement NNPF are encoded.• PON SEI messages, each including an NNPFA SEI message for the visual enhancement NNPF are encoded. For each pair of consecutive pictures in output order that are used as input to the picture rate upsampling NNPF, the PON SEI message is present in the latter picture in output order. The PON-nested NNPFA SEI message includes: o nnpfa_input_selected_pics_flag equal to 1 o nnpfa_num_input_pics_minusl equal to 0 o nnpfa_input_pic_skip_count

[0000] equal to 1; thus, the NNPFA SEI message does not apply to the picture corresponding to the current picture where the NNPFA SEI message is present but to the previous picture in output order, which is the interpolated picture.

[0650] An encoding system according to another embodiment handles example 1 as follows:• An SPO SEI message is encoded, including: o NNPFC and / or NNPFA SEI message types for the picture rate upsampling NNPF o NNPFEA SEI message type and optionally NNPFC SEI message type for the visual enhancement NNPF• NNPFC and / or NNPFA SEI messages for the picture rate upsampling NNPF are encoded.• NNPFC SEI message(s) for the visual enhancement NNPF are encoded.• PON SEI messages, each including an NNPFEA SEI message for the visual enhancement NNPF are encoded. For each pair of consecutive pictures in output order that are used as input to the picture rate upsampling NNPF, the PON SEI message is present in the latter picture in output order. The PON-nested NNPFEA SEI message includes: o nnpfea_input_selected_pics_flag equal to 1 o nnpfea_num_input_pics_minus 1 equal to 0 o nnpfea_input_pic_skip_count

[0000] equal to 1; thus, the NNPFEA SEI message does not apply to the picture corresponding to the current picture where the NNPFEA SEI message is present but to the previous picture in output order, which is the interpolated picture.

[0651] Example 2 is similar to example 1 but the visual enhancement NNPF is applied selectively as follows: Consecutive pictures in output order belong to different temporal sublayers and have different qualities in a hierarchical fashion. An interpolator NNPF generates a picture 1, 3, 5, ...between each pair of cropped decoded pictures 0, 2, 4, 6, etc. Interpolated picture 1 will likely be of high quality and needs not be filtered by a quality enhancement NNPF. Interpolated picture 3 will likely be of lower quality than interpolated picture 1, and thus is to be visually enhanced with the quality enhancement NNPF.

[0652] An encoding system according to an embodiment operates like example 1, but the PON- nested NNPFA SEI message (for controlled input picture selection) for the visual enhancement NNPF is not present in coded picture unit 2 and is present in coded picture unit 4.

[0653] An encoding system according to another embodiment operates like example 2, but the PON-nested NNPFEA SEI message for the visual enhancement NNPF is not present in coded picture unit 2 and is present in coded picture unit 4.

[0654] Example 3 is described as follows: An encoder finetunes a quality-enhancement NNPF based on interpolated pictures of a random access segment and sends an NNPF update through an NNPFC SEI message. En encoder indicates the use of the base quality-enhancement NNPF for the cropped decoded pictures and the updated NNPF for the interpolated pictures. The processing order is visual quality NNPF for the cropped decoded pictures, followed by the picture rate upsampling, followed by the visual quality NNPF for the interpolated pictures.

[0655] An encoding system according to an embodiment handles example 3 as follows:• An SPO SEI message is encoded, including: o NNPF 1: NNPFC and / or NNPFA SEI message types for the visual enhancement NNPF o NNPF 2: NNPFA SEI message type and optionally NNPFC SEI message type for the picture rate upsampling NNPF o NNPF 3: NNPFA SEI message type and optionally NNPFC SEI message type for the visual enhancement NNPF.• An NNPFC SEI message specifying the base visual enhancement NNPF is encoded.• An NNPFC SEI message specifying the picture rate upsampling NNPF is encoded.• NNPFC SEI message(s) defining an update to the base visual enhancement NNPF is encoded.• NNPFA SEI message(s) for the visual enhancement NNPF are encoded. These NNPFA SEI messages are not included in a PON SEI message and thus apply to the cropped decoded pictures and have nnpfa_target_base_flag equal to 1.• PON SEI messages, each including an NNPFA SEI message for the picture rate upsampling are encoded. nnpfa_input_selected_pics_flag, nnpfa_num_input_pics_minusl, ornnpfa_input_pic_skip_count

[0000] are not present in the PON-nested NNPFA SEI message. Thus, the NNPFA SEI message applies the picture rate upsampling with quality-enhanced cropped decoded pictures as input.• PON SEI messages, each including an NNPFA SEI message for the visual enhancement NNPF are encoded. For each pair of consecutive pictures in output order that are used as input to the picture rate upsampling NNPF, the PON SEI message is present in the latter picture in output order. The PON-nested NNPFA SEI message has: o nnpfa_target_base_flag equal to 0. o nnpfa_input_selected_pics_flag equal to 1 o nnpfa_num_input_pics_minusl equal to 0 o nnpfa_input_pic_skip_count

[0000] equal to 1; thus, the NNPFA SEI message does not apply to the picture corresponding to the current picture where the NNPFA SEI message is present but to the previous picture in output order, which is the interpolated picture.

[0656] It is to be understood that the embodiment above may be likewise realized with NNPFEA SEI messages used for NNPF 3 respectively to what is described above for NNPFA SEI messages for NNPF 3.

[0657] Extending NNPF activation with SEI processing order related syntax elements

[0658] In an embodiment, an NNPFEA SEI message or alike may include one or more po_id values to indicate which processing chains it applies to. In an embodiment, an encoder includes po_id values of the processing chains in NNPFEA SEI message(s). In an embodiment, a decoder decodes po_id values of the processing chains from NNPFEA SEI message(s).

[0659] In an embodiment, an NNPFEA SEI message or alike may include one or more processing order values to indicate which position(s) in the processing chain(s) it applies to. In an embodiment, an encoder includes processing order values in NNPFEA SEI message(s). In an embodiment, a decoder decodes processing order values from NNPFEA SEI message(s).

[0660] In an embodiment, candidate input pictures may originate from any processing stage up to and inlcuding the current processing stage in a processing chain (where the NNPF is activated with an NNPFEA SEI message, and NNPFA SEI message or alike).

[0661] In an embodiment, the NNPFA SEI message is extended with, or an NNPFEA SEI message or alike includes indication(s) indicative of the processing stage that generated an input picture to the NNPF inference.

[0662] In an embodiment, candidate input pictures indexed in reverse output order, and candidate input pictures having the same output time have the same index. A candidate input picture selected to be an input picture to the NNPF inference is identified through its index and its processing stage. A processing stage may be identified through its index in a processing chain.

[0663] In an embodiment, the NNPFA SEI message is further extended with nnpfa_proc_stage_depth[ i ] syntax element as follows. It is to be understood that embodiments may be realized with other syntax similarly.

[0664] nnpfa_proc_stage_depth[ i ] equal to 0 specifies that the candidate input picture that resulted from the same processing order value as associated with this NNPFA SEI message is selected as the i-th input picture. Let the current NNPFA SEI message has a processing order value that has index currProcIdx, i.e., is the currProcIdx-th process in the processing chain. nnpfa_proc_stage_depth[ i ] greater than 0 specifies an differential index diffProcIdx, which specifies that the candidate input picture that resulted from the processing order value with index currProcIdx - diffProcIdx is selected as the i-th input picture.

[0665] GAIN-RELATED EMBODIMENTS

[0666] Expected gain

[0667] In one embodiment, the signalled information comprises one or more sets of expected gains associated to respective one or more post-processing filters in the group of post-processing filters, or comprises one set of expected gains associated to all the postfilters in the group of postfilters, for the case where a purpose of the one or more post-processing filters or all the postfilters, respectively, is visual enhancement, or enhancement of machine analysis task(s), or any other purpose where the goal is to obtain a gain in terms of one or more metrics. Such an expected gain may have been determined during or after a development stage of the postfilter, based at least on the performance of the postfilter on a validation dataset. In the set of expected gains associated to a certain postfilter or to the postfilters in a certain group of postfilters, there may be one or more expected gains associated to that postfilter or to the postfilters in that group of postfilters, where different expected gains may be expressed in terms of different metrics.

[0668] In one example, the expected gain is determined by another entity (usually a human being via a computer code, but may be an Al system or any other automated system) with respect to the encoder and decoder, such as during a development phase of the codec or after the codec has been developed. In another example, the expected gain may be determined by the encoder or a transmitter, where the encoder or the transmitter may evaluate the performance or gain of the postfilter(s) on a dataset that is available at the encoder side or transmitter side, respectively. In yet another example, the expected gain may be determined by the decoder or a receiver, where the decoder or the receiver may evaluate the performance or gain of the postfilter(s) on a dataset which is available at the decoder side or transmitter side, respectively. In this latter example, the signalled information may not comprise indications about the expected gain.

[0669] When the signalled information comprises one or more expected gains associated to all the postfilters in the group of postfilters, each of the one or more expected gains may be referred to also as a combined expected gain or an expected group gain, and it indicates the expected gain of using all the associated post-processing filters.

[0670] In one embodiment, an expected gain represents a gain that is expected to be obtained (although not necessarily precisely) for one or more data units (such as one or more pictures, or one or more CTUs) when using the associated post-processing filters on those one or more data units.

[0671] In one embodiment, an expected gain represents a gain that is expected to be obtained (although not necessarily precisely) for a portion of the video sequence when using the associatedpost-processing filters on at least a subset of that portion. The portion may be the whole video sequence, or may be expressed as a predetermined length of video sequence, or may be expressed as a number of frames, or may be expressed as a number of Groups Of Pictures (GOPs), and the like.

[0672] The signalled information may comprise an indication of whether the expected gain is a gain that is expected to be obtained for the one or more data units on which the postfilters are applied or a gain that is expected to be obtained for a portion of the video sequence (as for the two embodiments above), and what portion.

[0673] In one example, a postfilter has a purpose of objective visual enhancement in terms of PSNR; the expected gain is expressed in terms of expected PSNR gain, i.e., the expected difference between the PSNR of the data to be filtered by the postfilter and the PSNR of the data filtered by the postfilter.

[0674] In another example, a postfilter has a purpose of objective visual enhancement in terms of MS-SSIM; the expected gain is expressed in terms of expected MS-SSIM gain, i.e., the expected difference between the MS-SSIM of the data to be filtered by the postfilter and the MS-SSIM of the data filtered by the postfilter.

[0675] In another example, a postfilter has a purpose of subjective visual enhancement in terms of MOS (mean opinion score); the expected gain is expressed in terms of expected MOS gain, i.e., the expected difference between the MOS of the data to be filtered by the postfilter and the MOS of the data filtered by the postfilter.

[0676] In another example, a postfilter has a purpose of enhancement for an object detection task in terms of mAP (mean average precision); the expected gain is expressed in terms of expected mAP gain, i.e., the expected difference between the mAP obtained based at least on the data to be filtered by the postfilter and the mAP obtained based at least on the data filtered by the postfilter.

[0677] An example syntax table is as follows:

[0678] Where pfg_expected_gain_present_flag[ i ] indicates whether an expected gain is present for the i-th postfilter, pfg_expected_gain[ i ] indicates the expected gain for the i-th postfilter, pfg_expected_gain_type_idc[ i ] indicates the type of the expected gain information, i.e., whether the expected gain is a gain that is expected to be obtained for the one or more data units on which the postfilters are applied or a gain that is expected to be obtained for a portion of the video sequence (as for the two embodiments above), and what portion.

[0679] The metric in terms of which the expected gain is indicated may be predefined in a standard specification based on one or more syntax elements or derived variables for the postfilter. These one or more syntax elements or derived variables may, for example, comprise the purpose of the postfilter. For example, one or more of the metric associations based on the purpose as in the following table may be predefined, where it needs to be understood that the presented nnpfc_purpose values are merely examples, and any other values could likewise be specified for these purposes:

[0680] Values of pfg_expected_gain_type_idc may be as in the following table:

[0681] Other types may be defined, for example based on resolution, based on content nature (e.g., screen content, natural content, man-made structures, indoors, outdoors, etc.).

[0682] Another example syntax table, where the metrics are not predefined in a standard, is as follows:

[0683] Where pfg_expected_gain_metric[ i ] indicates the metric in terms of which the expected gain indicated by pfg_expected_gain[ i ] is expressed. For example, the possible values and interpretations of pfg_expected_gain_metric[ i ] may be as follows:

[0684] The information above may be alternatively signalled in a NNPFC SEI message, for example as follows:

[0685] The following is an example syntax table for the case where the signalled information comprises one set of expected gains associated to all the postfilters in the group of postfilters. In this example, the set comprises one combined expected gain (or, using a different terminology, one expected group gain):

[0686] Where pfg_expected_group_gain_present_flag indicates whether a combined expectedgain is present for the group of postfilters comprised or referred to in this PFG SEI message, pfg_expected_group_gain_metric indicates the metric in terms of which the combined expected gain is expressed, pfg_expected_group_gain indicates the combined expected gain for the group of postfilters comprised or referred to in this PFG SEI message. pfg_expected_group_gain_type_idc is defined similarly as for pfg_expected_gain_type_idc.

[0687] In one example, a PFG SEI message indicates that two postfilters whose purpose is objective visual enhancement are to be used in cascade. The following information would be included in a PFG SEI message in order to signal a combined expected gain that is obtainable by the cascade of postfilters for the data units (e.g., pictures) on which it is applied (1-4 as follows):

[0688] 1. pfg_expected_group_gain_present_flag equal to 1.

[0689] 2. pfg_expected_group_gain_type_idc equal to 0.

[0690] 3. pfg_expected_group_gain_metric equal to 0 (indicating PSNR metric).

[0691] 4. pfg_expected_group_gain equal to (for example) 0.5, where 0.5 represents an expected increase of 0.5 dB in PSNR when using the postfilter on one or more pictures.

[0692] In an embodiment, an expected post-filter gain SEI message is defined for indicating a filter group ID or a filter ID and syntax elements for the expected gain. An example syntax table is as follows:

[0693] Where epfg_expected_gain indicates the expected gain for the post-filter or post-filter group with ID equal to epfg_id (however, epfg_id may not be present in this SEI message, if this SEI message is included in a nesting postfilter group SEI message), and epfg_expected_gain_type_idc indicates the type of the expected gain information, i.e., whether the expected gain is a gain that is expected to be obtained for the one or more data units on which the postfilters are applied or a gain that is expected to be obtained for a portion of the video sequence (as for the two embodiments above), and what portion.

[0694] The metric in terms of which the expected gain is indicated may be predefined in astandard specification based on the purpose of the postfilter.

[0695] In an embodiment, an expected post-filter gain SEI message is intended to be used within a nesting SEI message that defines a filter group and hence its syntax need not include a filter group ID.

[0696] In an embodiment, a post-filter activation SEI message or a post-filter group activation SEI message is appended with expected gain information indicating the expected gain for the frames that the SEI message activates the filter or filter group.

[0697] An example syntax for extending a post-filter group activation SEI message is as follows, where the semantics of syntax elements is similar to what has been defined above.

[0698] An example syntax for extending a post-filter activation SEI message is as follows, where the semantics of syntax elements is similar to what has been defined above, and the syntax function more_data_in_payload( ) returns TRUE if the SEI message includes more data and FALSE if the SEI message does not include more data.

[0699] In an embodiment, a post-filter characteristics SEI message may be extended with expected gain syntax conditioned on a new mode indicator value defined for the NNPFC SEI message, which is used to define a filter group. In one example, the following syntax of the NNPFC SEI message may be used where the semantics of the additional syntax elements may be defined as above.

[0700] In one embodiment, an extension mechanism is included in an NNPFC SEI message, where the number of bits for an extension is indicated and where the extension may be skipped and ignored by a decoder. For example, the following syntax may be used:

[0701] Where nnpfc_metadata_extension_num_bits equal to 0 specifies that nnpfc_reserved_metadata_extension is not present. nnpfc_metadata_extension_num_bits greater than 0 specifies the length, in bits, of nnpfc_reserved_metadata_extension. Decoders may ignore the presence and value of nnpfc_reserved_metadata_extension. It may be required that nnpfc_metadata_extension_num_bits is equal to 0 and nnpfc_reserved_metadata_extension is not present, until syntax and semantics have been specified for bits within nnpfc_reserved_metadata_extension.

[0702] In an embodiment, NNPFC metadata extension carries syntax elements for expected gain, which may apply to a single post-filter or a group of filters depending on the mode indicator value (nnpfc_mode_idc) as discussed in other embodiments. In one example, the following syntax of the NNPFC SEI message may be used where the semantics of the additional syntax elements may be defined as above:

[0703] Where nnpfc_metadata_extension_num_bits equal to 0 specifies that nnpfc_gain_info_present_flag and nnpfc_reserved_metadata_extension are not present. nnpfc_metadata_extension_num_bits greater than 0 specifies the joint length, in bits, of nnpfc_gain_info_present_flag, nnpfc_expected_gain_type_idc (when present), nnpfc_exptected_gain (when present), and nnpfc_reserved_metadata_extension. Decoders may ignore the presence and value of nnpfc_reserved_metadata_extension. nnpfc_expected_gain_type_idc (when present) and nnpfc_exptected_gain (when present) are specified like in other embodiments. Let nnpfcGainExtensionLength be the joint length, in bits, of nnpfc_gain_info_present_flag, nnpfc_expected_gain_type_idc (when present), and nnpfc_exptected_gain (when present). When present, the length, in bits, of nnpfc_reserved_metadata_extension is nnpfc_metadata_extension_num_bits - nnpfcGainExtensionLength. It may be required that nnpfc_metadata_extension_num_bits - nnpfcGainExtensionLength is equal to 0 and nnpfc_reserved_metadata_extension is not present, until syntax and semantics have been specified for bits within nnpfc_reserved_metadata_extension.

[0704] Actual gain

[0705] In one embodiment, the signalled information comprises one or more sets of actual gains associated to respective one or more post-processing filters in the group of post-processing filters, or comprises one set of actual gains associated to all the postfilters in the group of postfilters, for the case where a purpose of the one or more post-processing filters or all the postfilters, respectively, isvisual enhancement, or enhancement of machine analysis task(s), or any other purpose where the goal is to obtain a gain in terms of one or more metrics. In the set of actual gains associated to a certain postfilter or to the postfilters in a certain group of postfilters, there may be one or more actual gains associated to that postfilter or to the postfilters in that group of postfilters, where different actual gains may be expressed in terms of different metrics.

[0706] When the signalled information comprises one or more actual gains associated to all the postfilters in the group of postfilters, each of the one or more actual gains may be referred to also as a combined actual gain or an actual group gain, and it indicates the actual gain of using all the associated post-processing filters.

[0707] An actual gain represents a gain which is actually obtained by a receiver when using the associated post-processing filter on at least one picture or other data unit (e.g., CTU) of the video sequence for which the postfilter is activated. However, as the actual gain may be computed at encoding side based on a slightly different process than at receiver side (e.g., using different size of the input to the postfilters), it is to be understood that there may be still some differences between the signalled actual gain and the gain which is obtained at receiver side.

[0708] At least some of the examples provided for the expected gain are applicable to the actual gain, such as the syntax tables and related semantics.

[0709] The following is an example syntax table for the case where the signalled information comprises one set of actual gains associated to all the postfilters in the group of postfilters, where the set comprises one combined actual gain (or, using a different terminology, one actual group gain):

[0710] Where pfg_actual_group_gain_present_flag indicates whether a combined actual gain is present for the group of postfilters comprised or referred to in this PFG SEI message, pfg_actual_group_gain_metric indicates the metric in terms of which the combined actual gain is expressed, pfg_actual_group_gain indicates the combined actual gain for the group of postfilters comprised or referred to in this PFG SEI message.

[0711] The actual gain which is obtainable for the data units (e.g., CTU, or pictures, etc.) on which the postfilters are applied may be signalled in an activation SEI message. The following is an example syntax table.

[0712] Where nnmpfa_actual_gain_present_flag indicates whether an actual gain is present, nnmpfa_actual_gain indicates the actual gain, nnmpfa_actual_gain_metric indicates the metric in terms of which the actual gain indicated by nnmpfa_actual_gain is expressed, nnmpfa_target_id identifies a postfilter, nnmpfa_cancel_flag indicates whether this SEI message cancels the persistence of the multiple postfilters identified by nnmpfa_target_id, nnmpfa_persistence_flag indicates the persistence of the multiple postfilters identified bynnmpfa_target_id .

[0713] Indicating two or more postfilters with same purpose, applied separately

[0714] In one embodiment, the signalled information may comprise one or more filter identifiers (filter IDs) that identify respective one or more post-processing filters with the same purpose, where only one of the identified one or more post-processing filters is used or activated for any input picture. The signalled information may indicate or may be considered to imply for a decoder that either all of the identified one or more post-processing filters are intended to be applied as activated or none of them are intended to be applied. In an additional embodiment, the signalled information may comprise an indication that a portion of the properties of the identified one or more postprocessing filters are shared (i.e., in common) for all of them.

[0715] In one embodiment, a postfilter group SEI message may comprise all the postfilters that are to be used for the video sequence associated to that PFG SEI message, even if two or more of those postfilters are not used for the same picture or same data unit. For example, for a video sequence, four different postfilters may be used for different pictures (i.e., a certain picture may be filtered only by one postfilter).

[0716] In this embodiment, the PFG SEI message may comprise signalling information that indicates an expected gain that is obtainable when using the postfilters for different data items.

[0717] The following is an example syntax table for this embodiment:

[0718] A new value may be defined for pfg_usage_idc:

[0719] In one embodiment, a postfilter group SEI message may comprise all the p...

Claims

CLAIMSWhat is claimed is:

1. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, causes the apparatus at least to: receive an encoding of at least one picture; receive signaling comprising information related to a group of at least one postprocessing filter; and use the information related to the group of the at least one post-processing filter to infer whether and how to use at least one post-processing filter in the group of the at least one postprocessing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one post-processing filter in the group when a previous post-processing filter in a cascade outputs a different number of pictures than what is used as input for a current post-processing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.

2. The apparatus of claim 1, wherein the at least one post-processing filter comprises a neural network post-filter.

3. The apparatus of any of claims 1 to 2, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: apply a picture rate upsampling post-processing filter to obtain a second frequency from an input having a first frequency; and apply the picture rate upsampling post-processing filter to obtain a third frequency froman input having the second frequency; wherein the signaled information indicates that the picture rate upsampling postprocessing filter is to be applied to obtain the second frequency from the input having the first frequency, and that the picture rate upsampling post-processing filter is to be applied to obtain the third frequency from the input having the second frequency.

4. The apparatus of any of claims 1 to 3, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: decode one or more indications identifying input pictures used as an input for a postprocessing filter in the group of the at least one post-processing filter.

5. The apparatus of any of claims 1 to 4, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: decode one or more indications identifying input pictures used as input for each postprocessing filter in the group of the at least one post-processing filter.

6. The apparatus of any of claims 1 to 5, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: infer one or more input pictures for an initial post-processing filter in the group of the at least one post-processing filter; and decode one or more indications identifying input pictures for each subsequent postprocessing filter in the group; wherein each subsequent post-processing filter in the group follows the initial postprocessing filter.

7. The apparatus of any of claims 1 to 6, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: decode one or more indications identifying input pictures for a post-processing filter in the group of the at least one post-processing filter from a neural-network post-filter activation supplemental enhancement information message that is comprised in a processing order nesting supplemental enhancement information message.

8. The apparatus of any of claims 1 to 7, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: decode one or more indications identifying input pictures for a post-processing filter in the group of the at least one post-processing filter from a neural-network post-filter extended activation supplemental enhancement information message.

9. The apparatus of any of claims 1 to 8, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: derive a list of candidate input pictures for a second or later post-processing filter in the group of the at least one post-processing filter from filtered pictures that are output by previous post-processing filters in the the group of the at least one post-processing filter, when any; interpolated pictures that are output by a post-processing filter process of previous postprocessing filters in the group of the at least one post-processing filter, when any; and candidate input pictures for a first post-processing filter in the the group of the at least one post-processing filter.

10. The apparatus of claim 9, wherein an order of pictures in the list of candidate input pictures is pre-defined.

11. The apparatus of any of claims 9 to 10, wherein an order of pictures in the list of candidate input pictures is an inverse output order.

12. The apparatus of any of claims 9 to 11, wherein the candidate input pictures within the list of candidate input pictures are non-overlapping in output time, the list comprising up to one picture per an output time.

13. The apparatus of any of claims 9 to 12, wherein a picture resulting from a subsequent postprocessing filter in the group of the at least one post-processing filter precedes a picture resulting from a previous post-processing filter in the group of the at least one post-processing filter in the list of candidate input pictures, when the picture resulting from the subsequent postprocessing filter in the the group of the at least one post-processing filter has a same output order as the picture resulting from the previous post-processing filter in the group of the at least one post-processing filter.

14. The apparatus of any of claims 1 to 13, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: decode, given a list of candidate input pictures for a post-processing filter in the group of the at least one post-processing filter, one or more indications that indicate which of the candidate input pictures in the list are selected as input pictures for the post-processing filter in the group of the at least one post-processing filter.

15. The apparatus of claim 14, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: decode, given the list of candidate input pictures for the post-processing filter in the group of the at least one post-processing filter, one or both of: an indication for the post-processing filter when all the pictures in the list of candidate input pictures are input pictures to the post-processing filter, or one or more skip counts corresponding to how many pictures in the list of candidate input pictures are skipped when selecting pictures from the list of candidate input pictures to be used as input pictures.

16. The apparatus of any of claims 14 to 15, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: decode, given the list of candidate input pictures for the post-processing filter in the group of the at least one post-processing filter, a bit mask, where a bit position in the bit mask corresponds to a picture in the list of candidate input pictures, and a value of a bit indicates whether a picture in a respective position within the list of candidate input pictures is selected as an input picture to the post-processing filter in the group of the at least one post-processing filter.

17. The apparatus of any of claims 1 to 16, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: decode a description of the group of the at least one post-processing filter that comprises more than one occurrence of the same post-processing filter.

18. An apparatus comprising: at least one processor; andat least one memory storing instructions that, when executed by the at least one processor, causes the apparatus at least to: include, in a bitstream, an encoding of at least one picture; determine information related to a group of at least one post-processing filter; and indicate, in the bitstream, the information related to the group of the at least one postprocessing filter; wherein the information related to the group of the at least one post-processing filter is configured to be used to infer whether and how to use at least one post-processing filter in the group of the at least one post-processing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one post-processing filter in the group when a previous post-processing filter in a cascade outputs a different number of pictures than what is used as input for a current post-processing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.

19. The apparatus of claim 18, wherein the at least one post-processing filter comprises a neural network post-filter.

20. The apparatus of any of claims 18 to 19, wherein the signaled information indicates that a picture rate upsampling post-processing filter is to be applied to obtain a second frequency from an input having a first frequency, and that the picture rate upsampling post-processing filter is to be applied to obtain a third frequency from an input having the second frequency.

21. The apparatus of any of claims 18 to 20, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: encode one or more indications identifying input pictures used as input for a postprocessing filter in the group of the at least one post-processing filter.

22. The apparatus of any of claims 18 to 21, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to:encode one or more indications identifying input pictures used as input for each postprocessing filter in the group of the at least one post-processing filter.

23. The apparatus of any of claims 18 to 22, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: infer one or more input pictures for an initial post-processing filter in the group of the at least one post-processing filter; and encode one or more indications identifying input pictures for each subsequent postprocessing filter in the group; wherein each subsequent post-processing filter in the group follows the initial postprocessing filter.

24. The apparatus of any of claims 18 to 23, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: encode one or more indications identifying input pictures for a post-processing filter in the group of the at least one post-processing filter from a neural-network post-filter activation supplemental enhancement information message that is comprised in a processing order nesting supplemental enhancement information message.

25. The apparatus of any of claims 18 to 24, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: encode one or more indications identifying input pictures for a post-processing filter in the group of the at least one post-processing filter from a neural-network post-filter extended activation supplemental enhancement information message.

26. The apparatus of any of claims 18 to 25, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: derive a list of candidate input pictures for a second or later post-processing filter in the group of the at least one post-processing filter from filtered pictures that are output by previous post-processing filters in the group of the at least one post-processing filter, when any, interpolated pictures that are output by a post-processing filter process of previous postprocessing filters in the group of the at least one post-processing filter, when any, andcandidate input pictures for a first post-processing filter in the the group of the at least one post-processing filter.

27. The apparatus of claim 26, wherein an order of pictures in the list of candidate input pictures is pre-defined.

28. The apparatus of any of claims 26 to 27, wherein an order of pictures in the list of candidate input pictures is an inverse output order.

29. The apparatus of any of claims 26 to 28, wherein the candidate input pictures within the list of candidate input pictures are non-overlapping in output time, the list comprising up to one picture per an output time.

30. The apparatus of any of claims 26 to 29, wherein a picture resulting from a subsequent postprocessing filter in the group of the at least one post-processing filter precedes a picture resulting from a previous post-processing filter in the group of the at least one post-processing filter in the list of candidate input pictures, when the picture resulting from the subsequent postprocessing filter in the group of the at least one post-processing filter has a same output order as the picture resulting from the previous post-processing filter in the group of the at least one post-processing filter.

31. The apparatus of any of claims 18 to 30, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: encode, given a list of candidate input pictures for a post-processing filter in the group of the at least one post-processing filter, one or more indications that indicate which of the candidate input pictures in the list are selected as input pictures for the post-processing filter in the group of the at least one post-processing filter.

32. The apparatus of claim 31, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: encode, given the list of candidate input pictures for the post-processing filter in the group of the at least one post-processing filter, one or both of: an indication for the post-processing filter when all the pictures in the list of candidate input pictures are input pictures to the post-processing filter, or one or more skip counts corresponding to how many pictures in the list of candidate input pictures are skipped when selecting pictures from the list of candidate input pictures to beused as input pictures.

33. The apparatus of any of claims 31 to 32, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: encode, given the list of candidate input pictures for the post-processing filter in the group of the at least one post-processing filter, a bit mask, where a bit position in the bit mask corresponds to a picture in the list of candidate input pictures, and a value of a bit indicates whether a picture in a respective position within the list of candidate input pictures is selected as an input picture to the post-processing filter in the group of the at least one post-processing filter.

34. The apparatus of any of claims 18 to 33, wherein the instructions, when executed by the at least one processor, cause the apparatus at least to: encode a description of the group of the at least one post-processing filter that comprises more than one occurrence of the same post-processing filter.

35. A method comprising: receiving an encoding of at least one picture; receiving signaling comprising information related to a group of at least one postprocessing filter; and using the information related to the group of the at least one post-processing filter to infer whether and how to use at least one post-processing filter in the group of the at least one post-processing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one post-processing filter in the group when a previous post-processing filter in a cascade outputs a different number of pictures than what is used as input for a current post-processing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.

36. A method comprising:including, in a bitstream, an encoding of at least one picture; determining information related to a group of at least one post-processing filter; and indicating, in the bitstream, the information related to the group of the at least one postprocessing filter; wherein the information related to the group of the at least one post-processing filter is configured to be used to infer whether and how to use at least one post-processing filter in the group of the at least one post-processing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one post-processing filter in the group when a previous post-processing filter in a cascade outputs a different number of pictures than what is used as input for a current post-processing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.

37. An apparatus comprising: means for receiving an encoding of at least one picture; means for receiving signaling comprising information related to a group of at least one post-processing filter; and means for using the information related to the group of the at least one post-processing filter to infer whether and how to use at least one post-processing filter in the group of the at least one post-processing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one post-processing filter in the group when a previous post-processing filter in a cascade outputs a different number of pictures than what is used as input for a current post-processing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.

38. An apparatus comprising: means for including, in a bitstream, an encoding of at least one picture; means for determining information related to a group of at least one post-processing filter; and means for indicating, in the bitstream, the information related to the group of the at least one post-processing filter; wherein the information related to the group of the at least one post-processing filter is configured to be used to infer whether and how to use at least one post-processing filter in the group of the at least one post-processing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one post-processing filter in the group when a previous post-processing filter in a cascade outputs a different number of pictures than what is used as input for a current post-processing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.

39. A non-transitory program storage device readable by a machine, tangibly embodying a program of instructions executable by the machine for performing operations, the operations comprising: receiving an encoding of at least one picture; receiving signaling comprising information related to a group of at least one postprocessing filter; and using the information related to the group of the at least one post-processing filter to infer whether and how to use at least one post-processing filter in the group of the at least one post-processing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one post-processing filter in the group when a previous post-processing filter in a cascade outputs a different number of picturesthan what is used as input for a current post-processing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.

40. A non-transitory program storage device readable by a machine, tangibly embodying a program of instructions executable by the machine for performing operations, the operations comprising: including, in a bitstream, an encoding of at least one picture; determining information related to a group of at least one post-processing filter; and indicating, in the bitstream, the information related to the group of the at least one postprocessing filter; wherein the information related to the group of the at least one post-processing filter is configured to be used to infer whether and how to use at least one post-processing filter in the group of the at least one post-processing filter for the at least one picture; wherein the information indicates at least one of: whether and how to use the at least one post-processing filter in the group when a previous post-processing filter in a cascade outputs a different number of pictures than what is used as input for a current post-processing filter in the cascade, or whether and how to apply the at least one post-processing filter in a hierarchical manner.