Quality metrics in coded video
By implementing quality metrics per video sequence and picture, the solution addresses the lack of effective quality assessment in video coding, enabling improved post-processing and filter selection for enhanced video quality.
Patent Information
- Application Number
- PCT/IB2025/053664
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-09
- Filing Date
- 2025-04-07
- Publication Date
- 2025-10-16
AI Technical Summary
Existing video coding technologies lack effective quality metrics for post-processing filters, making it difficult to determine and signal the quality gains of coded video effectively.
Implementing quality metrics per video sequence and picture, signaling these metrics as absolute or relative gains for post-processing filters, and performing decoder-side operations based on these metrics to enhance video quality.
Enables accurate determination and signaling of video quality gains, allowing for improved post-processing and selection of optimal filters, resulting in higher quality video output.
Smart Images

Figure IMGF000050_0001 
Figure IMGF000054_0001 
Figure IMGF000061_0001
Abstract
Description
QUALITY METRICS IN CODED VIDEOTECHNICAL FIELD
[0001] The examples and non-limiting embodiments relate generally to media coding, and more particularly to, implementing quality metrics in coded video.BACKGROUND
[0002] It is known to provide standardized formats for encoding, signaling, or decoding of media data.SUMMARY
[0003] Example 1: An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: determining a quality metrics per video sequence and / or picture; and signaling the quality metrics as an absolute or relative gains for post-processing filters or post-processing filter groups, relative to the cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group.
[0004] Example 2: The apparatus of example 1, wherein the quality metrics comprises user-specified quality metrics comprising one or more following: a text description of the metric; a first flag to indicate whether higher or lower values of the metrics correspond to increased quality; a second flag to indicate whether the quality metrics is a full reference metric; or signaling of number of bytes signaled for the metrics values.
[0005] Example 3: The apparatus of any of the examples 1 or 2, wherein the quality metrics is single color component, individually for 3 color components, or combined for the 3 color components.
[0006] Example 4: The apparatus of any of the previous examples, wherein the quality metrics is comprised in an information message, and wherein the information message comprises one or more of following: a quality metrics definition flag for indicating whether the information message comprises definitions of the quality metrics; a sequence level values flag for indicating whether a sequence level metric values are signaled and allowing the apparatus to choose whether to signal sequence level metrics values, picture level metric values, or both; a number of quality metrics syntax element for enabling theapparatus to send multiple metric values within the information message; a quality gain flag for enabling the apparatus to indicate whether the signaled metric values indicate a gain associated with the postprocessing filter, or whether the values indicate absolute quality of the picture or the video sequence; reference flag for indicating whether a reference for gain is an input picture to the apparatus or an input picture to a post-processing stage; quality metric type syntax element for indicating a metric type of each quality metric; an increasing flag for indicating whether a higher value of a metric corresponds to a higher quality value or a lower quality value; a full reference flag for indicating wherein the quality metric is a full reference metric; a length syntax element for specifying a number of bytes used to signal the quality metrics; description syntax element for signaling a text description of the quality metrics; or description present flag for indicating that the text description of the quality metrics is signaled using the description syntax element.
[0007] Example 5: The apparatus of any of the previous examples, wherein the quality metrics comprises values indicating one or more of following values: quality of the picture; a mean quality of pictures corresponding to the video sequence; a quality gain of the picture, which is a difference between the quality of the picture relative to a quality of a gain reference picture; or a mean quality gain of the pictures corresponding to the video sequence.
[0008] Example 6: The apparatus of example 5, wherein the apparatus is further caused to define one or more of the following variables for the information message: a chroma format indicator; a count of pictures; lists of picture widths and heights; a list of pictures; or a list of gain reference pictures.
[0009] Example 7 : The apparatus of any of the previous examples, wherein values the quality metrics are computed by using, as a reference, a filtered picture, and wherein the filtered picture is encoded by the apparatus.
[0010] Example 8: The apparatus of any of the examples 1 to 6, wherein values the quality metrics are computed by using, as a reference picture, a picture that has not been filtered, and wherein the picture is encoded by the apparatus.
[0011] Example 9: The apparatus of any of the examples 1 to 6, wherein values the quality metrics are computed by using, as a reference picture, a picture prior to any filtering, and wherein the picture is filtered to obtain a filtered picture, and where the filtered picture is encoded by the apparatus.
[0012] Example 10: The apparatus of any of the examples 7 to 9, wherein the apparatus is further caused to perform: signaling, in or along a bitstream information indicating one or more of thefollowing: whether the reference picture used for computing the quality metrics is a filtered picture; a type of filtering used for filtering the reference picture; or information that describes a filtering that is applied to the reference picture.
[0013] Example 11: The apparatus of any of the previous examples, wherein a value of the quality metrics indicates an average, absolute, or relative quality, for a set of pictures that comprises the picture to which an indication of the value is associated and following pictures, in output order or in decoding order in the video sequence.
[0014] Example 12: The apparatus of example 11, wherein the apparatus is further caused to perform: signaling, in or along a bitstream, information indicating whether the value of the quality metric refers to an average quality for the pictures in the video sequence or refers to an average quality for a set of pictures that comprises the picture to which the information is associated and the following pictures in output order or in decoding order in the video sequence.
[0015] Example 13: The apparatus of any of the examples 11 or 12, wherein information about quality is comprised in an information message, and wherein the information message further comprises a flag for indicating whether the value for the quality metrics is to be interpreted as an average value for the current picture and subsequent pictures.
[0016] Example 14: The apparatus of example 13, wherein, when the information indicating the value of the quality metrics is associated with a picture that is not a first picture in the video sequence, the value is interpreted as an average quality for the picture to which the information is associated and the subsequent pictures in either output order or decoding order.
[0017] Example 15: The apparatus of example 11, wherein, when one or more values of the quality metrics represents the absolute or relative quality for two or more pictures, the apparatus is further caused to perform: signaling, in or along a bitstream, information indicating how the one or more values were determined.
[0018] Example 16: The apparatus of any of the previous examples, wherein the video sequence comprises a coded layer video sequence.
[0019] Example 17: The apparatus of any of the previous examples, wherein the information message comprises a supplemental enhancement information message.
[0020] Example 18: An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving a quality metrics as an absolute or as relative gains for post-processing filters or post-processing filter groups, relative to the cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group, wherein the quality metrics is determined per video sequence and / or picture; and performing decoder side operations based on the quality metrics.
[0021] Example 19: The apparatus of example 18, wherein the quality metrics comprises user- specified quality metrics comprising one or more following: a text description of the metric; a first flag to indicate whether higher or lower values of the metrics correspond to increased quality; a second flag to indicate whether the quality metrics is a full reference metric; or signaling of number of bytes signaled for the metrics values.
[0022] Example 20: The apparatus of any of the examples 18 or 19, wherein the quality metrics is single color component, individually for 3 color components, or combined for the 3 color components.
[0023] Example 21: The apparatus of any of the previous examples 18 to 20, wherein the quality metrics is comprised in an information message, and wherein the information message comprises one or more of following: a quality metrics definition flag for indicating whether the information message comprises definitions of the quality metrics; a sequence level values flag for indicating whether a sequence level metric values are signaled and allows an encoder to choose whether to signal sequence level metrics values, picture level metric values, or both; a number of quality metrics syntax element for that enables the encoder to send multiple metric values within the information message; a quality gain flag that enables the encoder to indicate whether the signaled metric values indicate a gain associated with the post-processing filter, or whether the values indicate absolute quality of the picture or the video sequence; reference flag for indicating whether a reference for gain is an input picture to the encoder or an input picture to a post-processing stage; quality metric type syntax element for indicating a metric type of each quality metric; an increasing flag for indicating whether a higher value of a metric corresponds to a higher quality value or a lower quality value; a full reference flag for indicating wherein the quality metric is a full reference metric; a length syntax element for specifying a number of bytes used to signal the quality metrics; description syntax element for signaling a text description of the quality metrics; or description present flag for indicating that the text description of the quality metrics is signaled using the description syntax element.
[0024] Example 22: The apparatus of any of the previous examples 18 to 21, wherein the quality metrics comprises values indicating one or more of following: quality of the picture; a mean quality ofpictures corresponding to the video sequence; a quality gain of the picture, which is a difference between the quality of the picture relative to a quality of a gain reference picture; or a mean quality gain of the pictures corresponding to the video sequence.
[0025] Example 23: The apparatus of example 22, wherein the information message comprises one or more of following variables: a chroma format indicator; a count of pictures, lists of picture widths and heights; a list of pictures; or a list of gain reference pictures.
[0026] Example 24: The apparatus of any of the previous examples 18 to 23, wherein values the quality metrics are computed by using, as a reference, a filtered picture, and wherein the filtered picture is encoded by the encoder.
[0027] Example 25: The apparatus of any of the examples 18 to 23, wherein values the quality metrics are computed by using, as a reference picture, a picture that has not been filtered, and wherein the picture is encoded by the encoder.
[0028] Example 26: The apparatus of any of the examples 18 to 23, wherein values the quality metrics are computed by using, as a reference picture, a picture prior to any filtering, and wherein the picture is filtered to obtain a filtered picture, and where the filtered picture is encoded by the encoder.
[0029] Example 27: The apparatus of any of the examples 24 to 26, wherein the apparatus is further caused to perform: receiving, from or along a bitstream information indicating one or more of the following: whether the reference picture used for computing the quality metrics is a filtered picture; a type of filtering used for filtering the reference picture; or information that describes a filtering that is applied to the reference picture.
[0030] Example 28: The apparatus of any of the previous examples 18 to 27, wherein a value of the quality metrics indicates an average, absolute, or relative quality, for a set of pictures that comprises the picture to which an indication of the value is associated and following pictures, in output order or in decoding order in the video sequence.
[0031] Example 29: The apparatus of example 28, wherein the apparatus is further caused to perform: receiving, from or along a bitstream, information indicating whether the value of the quality metric refers to an average quality for the pictures in the video sequence or refers to an average quality for a set of pictures that comprises the picture to which the information is associated and the following pictures in output order or in decoding order in the video sequence.
[0032] Example 30: The apparatus of any of the examples 28 or 29, wherein the information about quality is comprised in an information message, and wherein the information message further comprises a flag for indicating whether the value for the quality metrics is to be interpreted as an average value for the current picture and subsequent pictures.
[0033] Example 31: The apparatus of example 30, wherein, when the information indicating the value of the quality metrics is associated with a picture that is not a first picture in the video sequence, the value is interpreted as an average quality for the picture to which the information is associated and the subsequent pictures in either output order or decoding order.
[0034] Example 32: The apparatus of example 28, wherein, when one or more values of the quality metrics represents the absolute or relative quality for two or more pictures, the apparatus is further caused to perform: receiving, from or along a bitstream, information indicating how the one or more values were determined.
[0035] Example 33: The apparatus of any of the previous examples 18 to 32, wherein the video sequence comprises a coded layer video sequence.
[0036] Example 34: The apparatus of any of the previous examples 18 to 32, wherein the information message comprises a supplemental enhancement information message.
[0037] Example 35: The apparatus of any of the previous examples 18 to 34, wherein the decoder side operations comprise one or more of the following: performing a high complexity post-processing operation; using the quality metrics for determining or selecting a post-processing filter or a postprocessing filter group; selecting highest quality frames of the video sequence to operate on; or determining not to display pictures with quality levels below a target quality level.
[0038] Example 36: A method comprising: determining a quality metrics per video sequence and / or picture; and signaling the quality metrics as an absolute or relative gains for post-processing filters or post-processing filter groups, relative to the cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group.
[0039] Example 37: The method of example 36, wherein the quality metrics comprises user- specified quality metrics comprising one or more following: a text description of the metric; a first flag to indicate whether higher or lower values of the metrics correspond to increased quality; a second flagto indicate whether the quality metrics is a full reference metric; or signaling of number of bytes signaled for the metrics values.
[0040] Example 38: The method of any of the examples 36 or 37, wherein the quality metrics is single color component, individually for 3 color components, or combined for the 3 color components.
[0041] Example 39: The method of any of the previous examples 36 to 38, wherein the quality metrics is comprised in an information message, and wherein the information message comprises one or more of following: a quality metrics definition flag for indicating whether the information message comprises definitions of the quality metrics; a sequence level values flag for indicating whether a sequence level metric values are signaled and allowing an encoder method to choose whether to signal sequence level metrics values, picture level metric values, or both; a number of quality metrics syntax element for enabling the encoder to send multiple metric values within the information message; a quality gain flag for enabling the encoder to indicate whether the signaled metric values indicate a gain associated with the post-processing filter, or whether the values indicate absolute quality of the picture or the video sequence; reference flag for indicating whether a reference for gain is an input picture to the encoder or an input picture to a post-processing stage; quality metric type syntax element for indicating a metric type of each quality metric; an increasing flag for indicating whether a higher value of a metric corresponds to a higher quality value or a lower quality value; a full reference flag for indicating wherein the quality metric is a full reference metric; a length syntax element for specifying a number of bytes used to signal the quality metrics; description syntax element for signaling a text description of the quality metrics; or description present flag for indicating that the text description of the quality metrics is signaled using the description syntax element.
[0042] Example 40: The method of any of the previous examples 36 to 39, wherein the quality metrics comprises values indicating one or more of following values: quality of the picture; a mean quality of pictures corresponding to the video sequence; a quality gain of the picture, which is a difference between the quality of the picture relative to a quality of a gain reference picture; or a mean quality gain of the pictures corresponding to the video sequence.
[0043] Example 41: The method of example 40 further comprising defining one or more of the following variables for the information message: a chroma format indicator; a count of pictures; lists of picture widths and heights; a list of pictures; or a list of gain reference pictures.
[0044] Example 42: The method of any of the previous examples 36 to 41, wherein values the quality metrics are computed by using, as a reference, a filtered picture, and wherein the filtered pictureis encoded by the encoder.
[0045] Example 43: The method of any of the examples 36 to 41, wherein values the quality metrics are computed by using, as a reference picture, a picture that has not been filtered, and wherein the picture is encoded by the encoder.
[0046] Example 44: method of any of the examples 36 to 41, wherein values the quality metrics are computed by using, as a reference picture, a picture prior to any filtering, and wherein the picture is filtered to obtain a filtered picture, and where the filtered picture is encoded by the encoder.
[0047] Example 45: The method of any of the examples 42 to 44 further comprising: signaling, in or along a bitstream information indicating one or more of the following: whether the reference picture used for computing the quality metrics is a filtered picture; a type of filtering used for filtering the reference picture; or information that describes a filtering that is applied to the reference picture.
[0048] Example 46: The method of any of the previous examples 36 to 45, wherein a value of the quality metrics indicates an average, absolute, or relative quality, for a set of pictures that comprises the picture to which an indication of the value is associated and following pictures, in output order or in decoding order in the video sequence.
[0049] Example 47: The method of example 46 further comprising: signaling, in or along a bitstream, information indicating whether the value of the quality metric refers to an average quality for the pictures in the video sequence or refers to an average quality for a set of pictures that comprises the picture to which the information is associated and the following pictures in output order or in decoding order in the video sequence.
[0050] Example 48: The method of any of the examples 46 or 47, wherein information about quality is comprised in an information message, and wherein the information message further comprises a flag for indicating whether the value for the quality metrics is to be interpreted as an average value for the current picture and subsequent pictures.
[0051] Example 49: The method of example 48, wherein, when the information indicating the value of the quality metrics is associated with a picture that is not a first picture in the video sequence, the value is interpreted as an average quality for the picture to which the information is associated and the subsequent pictures in either output order or decoding order.
[0052] Example 50: The method of example 48, wherein, when one or more values of the quality metrics represents the absolute or relative quality for two or more pictures, the method further comprises: signaling, in or along a bitstream, information indicating how the one or more values were determined.
[0053] Example 51: The method of any of the previous examples 36 to 50, wherein the video sequence comprises a coded layer video sequence.
[0054] Example 52: The method of any of the previous examples 36 to 51, wherein the information message comprises a supplemental enhancement information message.
[0055] Example 53: A method comprising: receiving a quality metrics as an absolute or as relative gains for post-processing filters or post-processing filter groups, relative to the cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group, wherein the quality metrics is determined per video sequence and / or picture; and performing decoder side operations based on the quality metrics.
[0056] Example 54: The method of example 53, wherein the quality metrics comprises user- specified quality metrics comprising one or more following: a text description of the metric; a first flag to indicate whether higher or lower values of the metrics correspond to increased quality; a second flag to indicate whether the quality metrics is a full reference metric; or signaling of number of bytes signaled for the metrics values.
[0057] Example 55: The method of any of the examples 53 or 54, wherein the quality metrics is single color component, individually for 3 color components, or combined for the 3 color components.
[0058] Example 56: The method of any of the previous examples 53 to 55, wherein the quality metrics is comprised in an information message, and wherein the information message comprises one or more of following: a quality metrics definition flag for indicating whether the information message comprises definitions of the quality metrics; a sequence level values flag for indicating whether a sequence level metric values are signaled and allows an encoder to choose whether to signal sequence level metrics values, picture level metric values, or both; a number of quality metrics syntax element that enables the encoder to send multiple metric values within the information message; a quality gain flag that enables the encoder to indicate whether the signaled metric values indicate a gain associated with the post-processing filter, or whether the values indicate absolute quality of the picture or the video sequence; reference flag for indicating whether a reference for gain is an input picture to the encoder oran input picture to a post-processing stage; quality metric type syntax element for indicating a metric type of each quality metric; an increasing flag for indicating whether a higher value of a metric corresponds to a higher quality value or a lower quality value; a full reference flag for indicating wherein the quality metric is a full reference metric; a length syntax element for specifying a number of bytes used to signal the quality metrics; description syntax element for signaling a text description of the quality metrics; or description present flag for indicating that the text description of the quality metrics is signaled using the description syntax element.
[0059] Example 57: The method of any of the previous examples 53 to 56, wherein the quality metrics comprises values indicating one or more of following: quality of the picture; a mean quality of pictures corresponding to the video sequence; a quality gain of the picture, which is a difference between the quality of the picture relative to a quality of a gain reference picture; or a mean quality gain of the pictures corresponding to the video sequence.
[0060] Example 58: The method of example 57, wherein the information message comprises one or more of following variables: a chroma format indicator; a count of pictures, lists of picture widths and heights; a list of pictures; or a list of gain reference pictures.
[0061] Example 59: The method of any of the previous examples 53 to 58, wherein values the quality metrics are computed by using, as a reference, a filtered picture, and wherein the filtered picture is encoded by the encoder.
[0062] Example 60: The method of any of the examples 53 to 58, wherein values the quality metrics are computed by using, as a reference picture, a picture that has not been filtered, and wherein the picture is encoded by the encoder.
[0063] Example 61: method of any of the examples 53 to 58, wherein values the quality metrics are computed by using, as a reference picture, a picture prior to any filtering, and wherein the picture is filtered to obtain a filtered picture, and where the filtered picture is encoded by the encoder.
[0064] Example 62: The method of any of the examples 59 to 61 further comprising: receiving, from or along a bitstream information indicating one or more of the following: whether the reference picture used for computing the quality metrics is a filtered picture; a type of filtering used for filtering the reference picture; or information that describes a filtering that is applied to the reference picture.
[0065] Example 63: The method of any of the previous examples 53 to 62, wherein a value of thequality metrics indicates an average, absolute, or relative quality, for a set of pictures that comprises the picture to which an indication of the value is associated and following pictures, in output order or in decoding order in the video sequence.
[0066] Example 64: The method of example 63 further comprising: receiving, from or along a bitstream, information indicating whether the value of the quality metric refers to an average quality for the pictures in the video sequence or refers to an average quality for a set of pictures that comprises the picture to which the information is associated and the following pictures in output order or in decoding order in the video sequence.
[0067] Example 65: The method of any of the examples 63 or 64, wherein the information about quality is comprised in an information message, and wherein the information message further comprises a flag for indicating whether the value for the quality metrics is to be interpreted as an average value for the current picture and subsequent pictures.
[0068] Example 66: The method of example 65, wherein, when the information indicating the value of the quality metrics is associated with a picture that is not a first picture in the video sequence, the value is interpreted as an average quality for the picture to which the information is associated and the subsequent pictures in either output order or decoding order.
[0069] Example 67: The method of example 63, wherein, when one or more values of the quality metrics represents the absolute or relative quality for two or more pictures, the further comprises: receiving, from or along a bitstream, information indicating how the one or more values were determined.
[0070] Example 68: The method of any of the previous examples 53 to 67, wherein the video sequence comprises a coded layer video sequence.
[0071] Example 69: The method of any of the previous examples 53 to 67, wherein the information message comprises a supplemental enhancement information message.
[0072] Example 70: The method of any of the previous examples 53 to 69, wherein the decoder side operations comprise one or more of the following: performing a high complexity post-processing operation; using the quality metrics for determining or selecting a post-processing filter or a postprocessing filter group; selecting highest quality frames of the video sequence to operate on; or determining not to display pictures with quality levels below a target quality level.
[0073] Example 71: An apparatus comprising: means for determining a quality metrics per video sequence and / or picture; and means for signaling the quality metrics as an absolute or relative gains for post-processing filters or post-processing filter groups, relative to the cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group.
[0074] Example 72: The apparatus of example 71, wherein the apparatus further comprises means for performing methods as described in any of the examples 37 to 52.
[0075] Example 73: An apparatus comprising: means for receiving a quality metrics as an absolute or as relative gains for post-processing filters or post-processing filter groups, relative to the cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group, wherein the quality metrics is determined per video sequence and / or picture; and means for performing decoder side operations based on the quality metrics.
[0076] Example 74: The apparatus of example 73, wherein the apparatus further comprises means for performing methods as described in any of the examples 54 to 70.
[0077] Example 75: A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform: determining a quality metrics per video sequence and / or picture; and signaling the quality metrics as an absolute or relative gains for postprocessing filters or post-processing filter groups, relative to the cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group.
[0078] Example 76: The apparatus of example 75, wherein the apparatus is further caused to perform methods as described in any of the examples 37 to 52.
[0079] Example 77: The computer readable medium of any of the examples 75 or 76, wherein the computer readable medium comprises a non-transitory computer readable medium.
[0080] Example 78: A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform: receiving a quality metrics as an absolute or as relative gains for post-processing filters or post-processing filter groups, relative to the cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group, wherein the quality metrics is determined per video sequence and / or picture; and performing decoder side operations based on the quality metrics.
[0081] Example 79: The apparatus of example 78, wherein the apparatus is further caused to perform methods as described in any of the examples 54 to 70.
[0082] Example 80: The computer readable medium of any of the examples 78 or 79, wherein the computer readable medium comprises a non-transitory computer readable medium.BRIEF DESCRIPTION OF THE DRAWINGS
[0083] The foregoing embodiments and other features are explained in the following description, taken in connection with the accompanying drawings, wherein:
[0084] FIG. 1 shows schematically an apparatus employing embodiments of the examples described herein.
[0085] FIG. 2 shows schematically a user equipment suitable for employing embodiments of the examples described herein.
[0086] FIG. 3 further shows schematically electronic devices employing embodiments of the examples described herein connected using wireless and wired network connections.
[0087] FIG. 4 is a block diagram illustrating a system in accordance with an example.
[0088] FIG. 5 is an example apparatus, which may be implemented in hardware, and is caused to, implement examples described herein.
[0089] FIG. 6 shows a representation of an example of non-volatile memory media used to store instructions that implement the examples described herein.
[0090] FIG. 7 is an example method performed with an encoder, based on the examples described herein.
[0091] FIG. 8 is an example method performed with a decoder, based on the examples describedherein.DETAILED DESCRIPTION OF EXAMPLE EMBODIMENTS
[0092] The following acronyms and abbreviations that may be found in the specification and / or the drawing figures are defined as follows (the abbreviations may be appended with each other or with other characters using e.g. a hyphen or dash (-), and may be case insensitive):4CC four character code5G fifth generation cellular network technology5GC 5G core network a.k.a. also known asAVC advanced video codingCU coding unitDSP digital signal processorDU distributed unit eNB (or eNodeB) evolved Node B (for example, an LTE base station)EN-DC E-UTRA -NR dual connectivity en-gNB or En-gNB node providing NR user plane and control plane protocol terminations towards the UE, and acting as secondary node in EN-DCE-UTRA evolved universal terrestrial radio access, for example, the LTE radio access technologyFl or Fl-C interface between CU and DU control interface gNB (or gNodeB) base station for 5G / NR, for example, a node providing NR user plane and control plane protocol terminations towards the UE, and connected via the NG interface to the 5GCIEC International Electrotechnical Commission loT internet of thingsISO International Organization for StandardizationISOBMFF ISO base media file formatJPEG joint photographic experts groupLTE long-term evolution mdat MediaDataBoxMIME Multipurpose Internet Mail ExtensionMME mobility management entity moov MovieBoxMP4 file format for MPEG-4 Part 14 filesMPEG moving picture experts groupMPEG-2 H.222 / H.262 as defined by the ITUMPEG-4 audio and video coding standard for ISO / IEC 14496MS-SSIM multi scale structural similarity ng or NG new generation ng-eNB or NG-eNB new generation eNBNR new radio (5G radio)N / W or NW networkPDCP packet data convergence protocolPHY physical layerPNG portable network graphicsPSNR peak signal-to-noise ratioRAN radio access networkRFC request for commentsREC radio link controlRRC radio resource controlRRH remote radio headRU radio unitRx receiverSDAP service data adaptation protocolSGW serving gatewaySI interface between eNodeBs and the EPCSMF session management functionSPS sequence parameter setSSIM structural similaritySVC scalable video coding trak TrackBoxTx transmitterUE user equipmentUICC Universal Integrated Circuit CardUPF user plane functionURL uniform resource locatorX2 interconnecting interface between two eNodeBs in LTE networkXn interface between two NG-RAN nodes
[0093] Some embodiments will now be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all, embodiments may be shown. Indeed, various embodiments of the invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Like reference numerals refer to like elements throughout. As used herein, the terms ‘data,’ ‘content,’ ‘information,’ and similar terms may be used interchangeably to refer to data capable of being transmitted, received and / or stored in accordance with embodiments of the present invention. Thus, use of any such terms should not be taken to limit the spirit and scope of embodiments.
[0094] Described herein is a method and apparatus for implementing quality metrics in coded video.
[0095] The following describes in detail a suitable apparatus and possible method for implementing quality metrics in coded video according to embodiments. In this regard reference is first made to FIG. 1 and FIG. 2, where FIG. 1 shows an example block diagram of an electronic device or apparatus 100. The apparatus 100 may be an Internet of Things (loT) apparatus configured to perform various functions, such as for example, gathering information by one or more sensors, receiving or transmitting information, analyzing information gathered or received by the apparatus, or the like. The apparatus may comprise a video coding system, which may incorporate a codec. FIG. 2 shows a layout of an apparatus according to an example embodiment. The elements of FIG. 1 and FIG. 2 are explained next.
[0096] The apparatus 100 may for example be a mobile terminal or user equipment of a wireless communication system, a sensor device, a tag, or other lower power device. However, it would be appreciated that embodiments of the examples described herein may be implemented within any electronic device or apparatus which may process data by neural networks.
[0097] The apparatus 100 may comprise a housing 101 for incorporating and protecting the device. The apparatus 100 further may comprise a display 102 in the form of a liquid crystal display. In other embodiments of the examples described herein the display may be any suitable display technology suitable to display an image or video. The apparatus 100 may further comprise a keypad 104. In other embodiments of the examples described herein any suitable data or user interface mechanism may be employed. For example the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display.
[0098] The apparatus may comprise a microphone 106 or any suitable audio input which may be a digital or analog signal input. The apparatus 100 may further comprise an audio output device which in embodiments of the examples described herein may be any one of: an earpiece 108, speaker, or an analog audio or digital audio output connection. The apparatus 100 may also comprise a battery (or in other embodiments of the examples described herein the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator). The apparatus 100 may further comprise a camera 109 capable of recording or capturing images and / or video. The apparatus 100 may further comprise an infrared port for short range line of sight communication to other devices. In other embodiments the apparatus 100 may further comprise any suitable short range communication solution such as for example a Bluetooth wireless connection or a USB / firewire wired connection.
[0099] The apparatus 100 may comprise a controller 110, processor or processor circuitry for controlling the apparatus 100. The controller 110 may be connected to memory 112 which in embodiments of the examples described herein may store both data in the form of image and audio data and / or may also store instructions for implementation on the controller 110. The controller 110 may further be connected to codec circuitry 114 suitable for carrying out coding and / or decoding of audio and / or video data or assisting in coding and / or decoding carried out by the controller.
[0100] The apparatus 100 may further comprise a card reader 118 and a smart card 116, for example a UICC and UICC reader for providing user information and being suitable for providing authentication information for authentication and authorization of the user at a network.
[0101] The apparatus 100 may comprise radio interface circuitry 120 connected to the controller and suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system or a wireless local area network. The apparatus 100 may further comprise an antenna 122 connected to the radio interface circuitry 120 for transmitting radio frequency signals generated at the radio interface circuitry 120 to other apparatus(es) and / or for receiving radio frequency signals from other apparatus(es).
[0102] The apparatus 100 may comprise a camera capable of recording or detecting individual frames which are then passed to the codec circuitry 114 or the controller for processing. The apparatus may receive the video image data for processing from another device prior to transmission and / or storage. The apparatus 100 may also receive either wirelessly or by a wired connection the image for coding / decoding. The structural elements of apparatus 100 described above represent examples of means for performing a corresponding function.
[0103] With respect to FIG. 3, an example of a system within which embodiments of the examples described herein can be utilized is shown. The system 300 comprises multiple communication devices which can communicate through one or more networks. The system 300 may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network (such as a GSM, UMTS, CDMA, LTE, 4G, 5G network, etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a Bluetooth personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and the Internet.
[0104] The system 300 may include both wired and wireless communication devices and / or apparatus 100 suitable for implementing embodiments of the examples described herein.
[0105] For example, the system shown in FIG. 3 shows a mobile telephone network 301 and a representation of the internet 302. Connectivity to the internet 302 may include, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.
[0106] The example communication devices shown in the system 300 may include, but are not limited to, an electronic device or apparatus 100, a combination of a personal digital assistant (PDA) and a mobile telephone 304, a PDA 306, an integrated messaging device (IMD) 308, a desktop computer 310, a notebook computer 312, or a head-mounted apparatus. The head-mounted apparatus may be a head-mounted display (HMD), or glasses having a device such as a camera configured to encode and / or decode images and / or video. The apparatus 100 may be stationary or mobile when carried by an individual who is moving. The apparatus 100 may also be located in a mode of transport including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle or any similar suitable mode of transport.
[0107] The embodiments may also be implemented in a set-top box; e.g., a digital TV receiver, which may / may not have a display or wireless capabilities, in tablets or (laptop) personal computers (PC), which have hardware and / or software to process neural network data, in various operating systems, and in chipsets, processors, DSPs and / or embedded systems offering hardware / software based coding.
[0108] Some or further apparatus may send and receive calls and messages and communicate with service providers through a wireless connection 314 to a base station 316. The base station 316 may be connected to a network server 318 that allows communication between the mobile telephone network301 and the internet 302. The system may include additional communication devices and communication devices of various types.
[0109] The communication devices may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocolinternet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11, 3GPP Narrowband loT and any similar wireless communication technology. A communications device involved in implementing various embodiments of the examples described herein may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.
[0110] In telecommunications and data networks, a channel may refer either to a physical channel or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire, whereas a logical channel may refer to a logical connection over a multiplexed medium, capable of conveying several logical channels. A channel may be used for conveying an information signal, for example a bitstream, from one or several senders (or transmitters) to one or several receivers.
[0111] The embodiments may also be implemented in so-called loT devices. The Internet of Things (loT) may be defined, for example, as an interconnection of uniquely identifiable embedded computing devices within the existing Internet infrastructure. The convergence of various technologies has and may enable many fields of embedded systems, such as wireless sensor networks, control systems, home / building automation, etc. to be included in the Internet of Things (loT). In order to utilize the Internet loT devices are provided with an IP address as a unique identifier. loT devices may be provided with a radio transmitter, such as a WLAN or Bluetooth transmitter or a RFID tag. Alternatively, loT devices may have access to an IP -based network via a wired network, such as an Ethernet-based network or a power-line connection (PLC).
[0112] FIG. 4 is a block diagram illustrating a system or apparatus 400 in accordance with several examples. In an example, the encoder 430 is used to encode an image or video from the scene 415, and the encoder 430 is implemented in a transmitting apparatus 480. The encoder 430 produces a bitstream 410 comprising signaling that is received by the receiving apparatus 482, which implements a decoder 440. The encoder 430 sends the bitstream 410 that comprises the herein described signaling. The decoder 440 forms the image / picture or video / video sequence for the scene 415-1, and the receiving apparatus 482 would present this to the user (e.g., via a smartphone, television, or projector amongmany other options) or to an artificial intelligence system, for example, a face detector, an object tracker, and the like.
[0113] In some examples, the transmitting apparatus 480 and the receiving apparatus 482 are at least partially within a common apparatus, and for example are located within a common housing 450. In other examples the transmitting apparatus 480 and the receiving apparatus 482 are at least partially not within a common apparatus and have at least partially different housings. Therefore in some examples, the encoder 430 and the decoder 440 are at least partially within a common apparatus, and for example are located within a common housing 450. For example, the common apparatus comprising the encoder 430 and decoder 440 implements a codec. In other examples the encoder 430 and the decoder 440 are at least partially not within a common apparatus and have at least partially different housings, but when together still implement a codec.
[0114] In some examples, 3D media from the capture (e.g., volumetric capture) at a viewpoint 412 of the scene 415, which includes a person 413) is converted via projection to a series of 2D representations with occupancy, geometry, attributes and / or displacements. Additional atlas information is also included in the bitstream to enable inverse reconstruction. For decoding, the received bitstream 410 is separated into its components with atlas information; occupancy, geometry, displacement, and attribute 2D representations. A 3D reconstruction is performed to reconstruct the scene 415-1 created looking at the viewpoint 412-1 with a “reconstructed” person 413-1. The “-1” are used to indicate that these are reconstructions of the original. As indicated at 420, the decoder 440 performs an action or actions based on the received signaling.
[0115] The transmitting apparatus 480 performs signaling of a quality metrics based on the examples described herein. The receiving apparatus performs receiving of the quality metrics and / or decoding of the picture and / or video sequence, based on the examples described herein.
[0116] Having thus introduced a suitable but non-limiting technical context for the practice of the example embodiments of the present disclosure, example embodiments will now be described in detail.
[0117] Video / image quality metrics such as PSNR, SSIM, and VMAF are commonly used in development and evaluation of video coding standards and video encoder implementations. These are full reference quality metrics that compare the input of the encoder with the output of the decoder.
[0118] Some video / image quality metrics are non-reference, meaning that they do not compare the decoder output with the encoder input. There are also limited reference image / video quality metrics.
[0119] Video coding
[0120] A video codec includes an encoder and a decoder. The encoder transforms the input video into a compressed representation suited for storage / transmission. The decoder can uncompress the compressed video representation back into a viewable form. A video encoder and / or a video decoder may also be separate from each other (e.g., need not form a codec). The encoder may discard some information in the original video sequence in order to represent the video in a more compact form (that is, at a lower bitrate).
[0121] Hybrid video encoders, for example many encoder implementations of ITU-T H.263 and H.264, may encode the video information in two phases. Firstly, pixel values in a certain picture area (or “block”) are predicted, for example by motion compensation means (e.g., finding and indicating an area in one of the previously coded video frames that corresponds closely to the block being coded) or by spatial means (e.g., using the pixel values around the block to be coded in a specified manner). Secondly, the prediction error (e.g., the difference between the predicted block of pixels and the original block of pixels) is coded. This may be done by transforming the difference in pixel values using a specified transform (e.g. discrete cosine transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (e.g., picture quality) and size of the resulting coded video representation (e.g., file size or transmission bitrate).
[0122] In temporal prediction, the sources of prediction are previously decoded pictures (a.k.a. reference pictures). In intra block copy (IBC; a.k.a. intra-block-copy prediction), prediction is applied similarly to temporal prediction, but the reference picture is the current picture and only previously decoded samples can be referred in the prediction process. Inter-layer or inter-view prediction may be applied similarly to temporal prediction, but the reference picture is a decoded picture from another scalable layer or from another view, respectively. In some cases, inter prediction may refer to temporal prediction only, while in other cases inter prediction may refer collectively to temporal prediction and any of intra block copy, inter-layer prediction, and inter-view prediction, provided that they are performed with the same or similar process as temporal prediction. Inter prediction or temporal prediction may sometimes be referred to as motion compensation or motion-compensated prediction.
[0123] Inter prediction, which may also be referred to as temporal prediction, motion compensation, or motion-compensated prediction, reduces temporal redundancy. In inter prediction, the sources of prediction are previously decoded pictures. Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in spatialor transform domain; either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.
[0124] One outcome of the coding procedure is a set of coding parameters, such as motion vectors and quantized transform coefficients. Many parameters can be entropy-coded more efficiently when they are predicted first from spatially or temporally neighboring parameters. For example, a motion vector may be predicted from spatially adjacent motion vectors, and only the difference relative to the motion vector predictor may be coded. Prediction of coding parameters and intra prediction may be collectively referred to as in-picture prediction.
[0125] The H.264 / AVC standard was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunications Standardization Sector of International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of International Organisation for Standardization (ISO) / International Electrotechnical Commission (IEC). The H.264 / AVC standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.264 and ISO / IEC International Standard 14496-10, also known as MPEG- 4 Part 10 Advanced Video Coding (AVC). There have been multiple versions of the H.264 / AVC standard, integrating new extensions or features to the specification. These extensions include Scalable Video Coding (SVC) and Multiview Video Coding (MVC).
[0126] The High Efficiency Video Coding standard (which may be abbreviated HE VC or H.265 / HEVC) was developed by the Joint Collaborative Team - Video Coding (JCT-VC) of VCEG and MPEG. The standard is published by both parent standardization organizations, and it is referred to as ITU-T Recommendation H.265 and ISO / IEC International Standard 23008-2, also known as MPEG-H Part 2 High Efficiency Video Coding (HEVC). Extensions to H.265 / HEVC include scalable, multiview, three-dimensional, and fidelity range extensions, which may be referred to as SHVC, MV- HEVC, 3D-HEVC, and REXT, respectively. The references in this description to H.265 / HEVC, SHVC, MV -HEVC, 3D-HEVC and REXT that have been made for the purpose of understanding definitions, structures or concepts of these standard specifications are to be understood to be references to the latest versions of these standards that were available before the date of this application, unless otherwise indicated.
[0127] SHVC, MV-HEVC, and 3D-HEVC use a common basis specification, specified in Annex F of the version 2 of the HEVC standard. This common basis comprises, for example, high-level syntax and semantics, for example specifying some of the characteristics of the layers of the bitstream, such as inter-layer dependencies, as well as decoding processes, such as reference picture list construction,including inter-layer reference pictures and picture order count derivation for multi-layer bitstream. Annex F may also be used in potential subsequent multi-layer extensions of HEVC.
[0128] It is to be understood that even though a video encoder, a video decoder, encoding methods, decoding methods, bitstream structures, and / or example embodiments may be described in the following with reference to specific extensions, such as SHVC and / or MV-HEVC, they are generally applicable to any multi-layer extensions of HEVC, and even more generally to any multilayer video coding scheme.
[0129] The Versatile Video Coding standard (which may be abbreviated VVC, H.266, or H.266 / VVC) was developed by the Joint Video Experts Team (JVET), which is a collaboration between the ISO / IEC MPEG and ITU-T VCEG. Extensions to VVC are presently under development.
[0130] A specification of the AV 1 bitstream format and decoding process were developed by the Alliance of Open Media (AOM). The AVI specification was published in 2018. AOM is reportedly working on the AV2 specification.
[0131] Some key definitions, bitstream and coding structures, and concepts of H.264 / AVC and HEVC are described in this section as an example of a video encoder, decoder, encoding method, decoding method, and a bitstream structure, wherein the example embodiments may be implemented. Some of the key definitions, bitstream and coding structures, and concepts of H.264 / AVC are the same as in HEVC; accordingly, they are described below jointly. The aspects of the present disclosure are not limited to H.264 / AVC or HEVC, but rather the description is given for one possible basis on top of which the example embodiments may be partly or fully realized. Many aspects described below in the context of H.264 / AVC or HEVC may apply to VVC, and the aspects of the present disclosure may hence be applied to VVC.
[0132] Similarly to many earlier video coding standards, the bitstream syntax and semantics, as well as the decoding process for error-free bitstreams, are specified in H.264 / AVC and HEVC. The encoding process is not specified, but encoders must generate conforming bitstreams. Bitstream and decoder conformance can be verified with the Hypothetical Reference Decoder (HRD). The standards include coding tools that help in coping with transmission errors and losses, but the use of the tools in encoding is optional, and no decoding process has been specified for erroneous bitstreams.
[0133] The elementary unit for the input to an H.264 / AVC or HEVC encoder, and the output of an H.264 / AVC or HEVC decoder, respectively, is a picture. A picture given as an input to an encodermay also be referred to as a source picture, and a picture decoded by a decoded may be referred to as a decoded picture.
[0134] The source and decoded pictures are each comprised of one or more sample arrays, such as one of the following sets of sample arrays: luma (Y) only (monochrome); luma and two chroma (YCbCr or YCgCo); green, blue, and red (GBR, also known as RGB); or arrays representing other unspecified monochrome or tri-stimulus color samplings (for example, YZX, also known as XYZ).
[0135] In the following, these arrays may be referred to as luma (or L or Y) and chroma, where the two chroma arrays may be referred to as Cb and Cr, regardless of the actual color representation method in use. The actual color representation method in use can be indicated, for example, in a coded bitstream (e.g. using the Video Usability Information (VUI) syntax of H.264 / AVC and / or HEVC). A component may be defined as an array or single sample from one of the three sample arrays (luma and two chroma), or an array or a single sample of the array that compose a picture in monochrome format.
[0136] In H.264 / AVC and HEVC, a picture may either be a frame or a field. A frame comprises a matrix of luma samples and possibly the corresponding chroma samples. A field is a set of alternate sample rows of a frame and may be used as encoder input, for example when the source signal is interlaced. Chroma sample arrays may be absent (and hence monochrome sampling may be in use), or chroma sample arrays may be subsampled when compared to luma sample arrays.
[0137] Chroma formats may be summarized as follows. In monochrome sampling, there is only one sample array, which may be nominally considered the luma array. In 4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array. In 4:2:2 sampling, each of the two chroma arrays has the same height and half the width of the luma array. In 4:4:4 sampling, when no separate color planes are in use, each of the two chroma arrays has the same height and width as the luma array.
[0138] In H.264 / AVC and HEVC, it is possible to code sample arrays as separate color planes into the bitstream and respectively decode separately coded color planes from the bitstream. When separate color planes are in use, each one of them is separately processed (by the encoder and / or the decoder) as a picture with monochrome sampling.
[0139] A partitioning may be defined as a division of a set into subsets such that each element of the set is in exactly one of the subsets.
[0140] When describing the operation of HEVC encoding and / or decoding, the following terms may be used. A coding block may be defined as an NxN block of samples for some value of N such that the division of a coding tree block into coding blocks is a partitioning. A coding tree block (CTB) may be defined as an NxN block of samples for some value of N such that the division of a component into coding tree blocks is a partitioning. A coding tree unit (CTU) may be defined as a coding tree block of luma samples, two corresponding coding tree blocks of chroma samples of a picture that has three sample arrays, or a coding tree block of samples of a monochrome picture or a picture that is coded using three separate color planes and syntax structures used to code the samples. A coding unit (CU) may be defined as a coding block of luma samples, two corresponding coding blocks of chroma samples of a picture that has three sample arrays, or a coding block of samples of a monochrome picture or a picture that is coded using three separate color planes and syntax structures used to code the samples. A CU with the maximum allowed size may be named as largest coding unit (LCU) or coding tree unit (CTU), and the video picture is divided into non-overlapping LCUs.
[0141] A CU includes one or more prediction units (PU) defining the prediction process for the samples within the CU and one or more transform units (TU) defining the prediction error coding process for the samples in the said CU. Typically, a CU includes a square block of samples with a size selectable from a predefined set of possible CU sizes. Each PU and TU can be further split into smaller PUs and TUs in order to increase granularity of the prediction and prediction error coding processes, respectively. Each PU has prediction information associated with it defining what kind of a prediction is to be applied for the pixels within that PU (e.g. motion vector information for inter predicted PUs and intra prediction directionality information for intra predicted PUs).
[0142] Each TU can be associated with information describing the prediction error decoding process for the samples within the said TU (including, e.g., DCT coefficient information). It may be signaled at the CU level whether prediction error coding is applied or not for each CU. In the case there is no prediction error residual associated with the CU, it can be considered there are no TUs for the said CU. The division of the image into CUs, and division of CUs into PUs and Tus, may be signaled in the bitstream, allowing the decoder to reproduce the intended structure of these units.
[0143] In HEVC, a picture can be partitioned in tiles, which are rectangular and include an integer number of LCUs. In HEVC, the partitioning to tiles forms a regular grid, where heights and widths of tiles differ from each other by one LCU at the maximum. In HEVC, a slice is defined to be an integer number of coding tree units contained in one independent slice segment and all subsequent dependent slice segments (if any) that precede the next independent slice segment (if any) within the same access unit. In HEVC, a slice segment is defined to be an integer number of coding tree units orderedconsecutively in the tile scan and contained in a single NAL unit. The division of each picture into slice segments is a partitioning. In HEVC, an independent slice segment is defined to be a slice segment for which the values of the syntax elements of the slice segment header are not inferred from the values for a preceding slice segment, and a dependent slice segment is defined to be a slice segment for which the values of some syntax elements of the slice segment header are inferred from the values for the preceding independent slice segment in decoding order. In HEVC, a slice header is defined to be the slice segment header of the independent slice segment that is a current slice segment or is the independent slice segment that precedes a current dependent slice segment, and a slice segment header is defined to be a part of a coded slice segment including the data elements pertaining to the first or all coding tree units represented in the slice segment. The CUs are scanned in the raster scan order of LCUs within tiles or within a picture, when tiles are not in use. Within an LCU, the CUs have a specific scan order.
[0144] An intra-coded slice (also called I slice) is a slice that only includes intra-coded blocks. The syntax of an I slice may exclude syntax elements that are related to inter prediction. An inter-coded slice is such where blocks can be intra- or inter-coded. Inter-coded slices may further be categorized into P and B slices, where P slices are such that blocks may be intra-coded or inter-coded but only using uni-prediction, and blocks in B slices may be intra-coded or inter-coded with uni- or bi-prediction.
[0145] The decoder reconstructs the output video by applying prediction means similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in the spatial pixel domain). After applying prediction and prediction error decoding means, the decoder sums up the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) can also apply additional filtering means to improve the quality of the output video before passing it for display and / or storing it as prediction reference for the forthcoming frames in the video sequence.
[0146] The filtering may for example include one more of the following: deblocking, sample adaptive offset (SAO), and / or adaptive loop filtering (ALF). H.264 / AVC includes a deblocking, whereas HEVC includes both deblocking and SAO.
[0147] In video codecs, the motion information may be indicated with motion vectors associated with each motion compensated image block, such as a prediction unit. Each of these motion vectors represents the displacement of the image block in the picture to be coded (in the encoder side) ordecoded (in the decoder side) and the prediction source block in one of the previously coded or decoded pictures. In order to represent motion vectors efficiently, those are typically coded differentially with respect to block specific predicted motion vectors.
[0148] In video codecs, the predicted motion vectors may be created in a predefined way, for example calculating the median of the encoded or decoded motion vectors of the adjacent blocks. Another way to create motion vector predictions is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures and signaling the chosen candidate as the motion vector predictor.
[0149] In addition to predicting the motion vector values, it can be predicted which reference picture(s) are used for motion-compensated prediction, and this prediction information may be represented, for example, by a reference index of previously coded / decoded picture. The reference index is typically predicted from adjacent blocks and / or co-located blocks in temporal reference picture. Moreover, high efficiency video codecs may employ an additional motion information coding / decoding mechanism, often called merging / merge mode, where all the motion field information, which includes motion vector and corresponding reference picture index for each available reference picture list, is predicted and used without any modification / correction. Similarly, predicting the motion field information is carried out using the motion field information of adjacent blocks and / or co-located blocks in temporal reference pictures, and the used motion field information is signaled among a list of motion field candidate list filled with motion field information of available adjacent / co-located blocks.
[0150] In video codecs, the prediction residual after motion compensation may be first transformed with a transform kernel (like DCT) and then coded. The reason for this is that, often, there still exists some correlation among the residual, and transform can in many cases help reduce this correlation and provide more efficient coding.
[0151] Video encoders may utilize Lagrangian cost functions to find optimal coding modes (e.g. the desired coding mode for a block and associated motion vectors). This kind of cost function uses a weighting factor X to tie together the (exact or estimated) image distortion due to lossy coding methods and the (exact or estimated) amount of information that is required to represent the pixel values in an image area:C = D + R,
[0152] where C is the Lagrangian cost to be minimized, D is the image distortion (e.g. Mean Squared Error) with the mode and motion vectors considered, and R the number of bits needed to represent the required data to reconstruct the image block in the decoder (including the amount of data to represent the candidate motion vectors).
[0153] Video coding standards and specifications may allow encoders to divide a coded picture to coded slices or alike. In-picture prediction is typically disabled across slice boundaries. Thus, slices can be regarded as a way to split a coded picture to independently decodable pieces. In H.264 / AVC and HEVC, in-picture prediction may be disabled across slice boundaries. Thus, slices can be regarded as a way to split a coded picture into independently decodable pieces, and slices are therefore often regarded as elementary units for transmission. In many cases, encoders may indicate in the bitstream which types of in-picture prediction are turned off across slice boundaries, and the decoder operation takes this information into account, for example when concluding which prediction sources are available. For example, samples from a neighboring CU may be regarded as unavailable for intra prediction, when the neighboring CU resides in a different slice.
[0154] An elementary unit for the output of an H.264 / AVC or HEVC encoder and the input of an H.264 / AVC or HEVC decoder, respectively, is a network abstraction layer (NAL) unit. For transport over packet-oriented networks or storage into structured files, NAL units may be encapsulated into packets or similar structures. A bytestream format has been specified in H.264 / AVC and HEVC for transmission or storage environments that do not provide framing structures. The bytestream format separates NAL units from each other by attaching a start code in front of each NAL unit. To avoid false detection of NAL unit boundaries, encoders run a byte-oriented start code emulation prevention algorithm, which adds an emulation prevention byte to the NAL unit payload when a start code would have occurred otherwise. In order to enable straightforward gateway operation between packet- and stream-oriented systems, start code emulation prevention may always be performed, regardless of whether the bytestream format is in use or not. A NAL unit (NALU) may be defined as a syntax structure including an indication of the type of data to follow and bytes including that data in the form of an RBSP interspersed as necessary with emulation prevention bytes. A raw byte sequence payload (RBSP) may be defined as a syntax structure including an integer number of bytes that is encapsulated in a NAL unit. An RBSP is either empty or has the form of a string of data bits including syntax elements followed by an RBSP stop bit and followed by zero or more subsequent bits equal to 0.
[0155] NAL units include a header and payload. In H.264 / AVC and HEVC, the NAL unit header indicates the type of the NAL unit.
[0156] In HEVC, a two-byte NAL unit header is used for all specified NAL unit types. The NAL unit header includes one reserved bit, a six-bit NAL unit type indication, a three-bit nuh_temporal_id_plusl indication for temporal level (may be required to be greater than or equal to 1) and a six-bit nuh_layer_id syntax element. The temporal_id_plusl syntax element may be regarded as a temporal identifier for the NAL unit, and a zero-based Temporalld variable may be derived as follows: Temporalld = temporal_id_plus 1 - 1. The abbreviation TID may be used interchangeably with the Temporalld variable. Temporalld equal to 0 corresponds to the lowest temporal level. The value of temporal_id_plusl is required to be non-zero in order to avoid start code emulation involving the two NAL unit header bytes. The bitstream created by excluding all VCL NAL units having a Temporalld greater than or equal to a selected value and including all other VCL NAL units remains conforming. Consequently, a picture having Temporalld equal to tid_value does not use any picture having a Temporalld greater than tid_value as inter prediction reference. A sub-layer or a temporal sub-layer may be defined to be a temporal scalable layer (or a temporal layer, TL) of a temporal scalable bitstream, including VCL NAL units with a particular value of the Temporalld variable and the associated non- VCL NAL units. The nuh_layer_id can be understood as a scalability layer identifier.
[0157] NAL units can be categorized into Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units may be coded slice NAL units. In HEVC, VCL NAL units include syntax elements representing one or more CU.
[0158] In HEVC, abbreviations for picture types may be defined as follows: trailing (TRAIL) picture, temporal sub-layer access (TSA), step-wise temporal sub-layer access (STSA), random access decodable leading (RADL) picture, random access skipped leading (RASL) picture, broken link access (BLA) picture, instantaneous decoding refresh (IDR) picture, clean random access (CRA) picture.
[0159] A random access point (RAP) picture, which may also be referred to as an intra random access point (IRAP) picture in an independent layer, includes only intra-coded slices. An IRAP picture belonging to a predicted layer may include P, B, and I slices, cannot use inter prediction from other pictures in the same predicted layer, and may use inter-layer prediction from its direct reference layers. In the present version of HEVC, an IRAP picture may be a BLA picture, a CRA picture or an IDR picture. The first picture in a bitstream including a base layer is an IRAP picture at the base layer. Provided the necessary parameter sets are available when they need to be activated, an IRAP picture at an independent layer and all subsequent non-RASL pictures at the independent layer in decoding order can be correctly decoded without performing the decoding process of any pictures that precede the IRAP picture in decoding order. The IRAP picture belonging to a predicted layer and all subsequent non-RASL pictures in decoding order within the same predicted layer can be correctly decoded withoutperforming the decoding process of any pictures of the same predicted layer that precede the IRAP picture in decoding order, when the necessary parameter sets are available when they need to be activated and when the decoding of each direct reference layer of the predicted layer has been initialized. There may be pictures in a bitstream that include only intra-coded slices that are not IRAP pictures.
[0160] A non-VCL NAL unit may be, for example, one of the following types: a sequence parameter set, a picture parameter set, a supplemental enhancement information (SEI) NAL unit, an access unit delimiter, an end of sequence NAL unit, an end of bitstream NAL unit, or a filler data NAL unit. Parameter sets may be needed for the reconstruction of decoded pictures, whereas many of the other non-VCL NAL units are not necessary for the reconstruction of decoded sample values.
[0161] Parameters that remain unchanged through a coded video sequence may be included in a sequence parameter set. In addition to the parameters that may be needed by the decoding process, the sequence parameter set may optionally include video usability information (VUI), which includes parameters that may be important for buffering, picture output timing, rendering, and resource reservation. In HEVC, a sequence parameter set RBSP includes parameters that can be referred to by one or more picture parameter set RBSPs or one or more SEI NAL units including a buffering period SEI message. A picture parameter set includes such parameters that are likely to be unchanged in several coded pictures. A picture parameter set RBSP may include parameters that can be referred to by the coded slice NAL units of one or more coded pictures.
[0162] In HEVC, a video parameter set (VPS) may be defined as a syntax structure including syntax elements that apply to zero or more entire coded video sequences as determined by the content of a syntax element, found in the SPS, referred to by a syntax element found in the PPS, referred to by a syntax element found in each slice segment header.
[0163] A video parameter set RBSP may include parameters that can be referred to by one or more sequence parameter set RBSPs.
[0164] The relationship and hierarchy between video parameter set (VPS), sequence parameter set (SPS), and picture parameter set (PPS) may be described as follows. VPS resides one level above SPS in the parameter set hierarchy and in the context of scalability and / or 3D video. VPS may include parameters that are common for all slices across all (scalability or view) layers in the entire coded video sequence. SPS includes the parameters that are common for all slices in a particular (scalability or view) layer in the entire coded video sequence, and may be shared by multiple (scalability or view) layers. PPS includes the parameters that are common for all slices in a particular layer representation(the representation of one scalability or view layer in one access unit) and are likely to be shared by all slices in multiple layer representations.
[0165] VPS may provide information about the dependency relationships of the layers in a bitstream, as well as many other information that are applicable to all slices across all (scalability or view) layers in the entire coded video sequence. VPS may be considered to comprise two parts, the base VPS and a VPS extension, where the VPS extension may be optionally present.
[0166] Out-of-band transmission, signaling or storage can additionally or alternatively be used for purposes other than tolerance against transmission errors, such as ease of access or session negotiation. For example, a sample entry of a track in a file conforming to the ISO Base Media File Format may comprise parameter sets, while the coded data in the bitstream is stored elsewhere in the file or in another file. The phrase along the bitstream (e.g. indicating along the bitstream) may be used in claims and described embodiments to refer to out-of-band transmission, signaling, or storage in a manner that the out-of-band data is associated with the bitstream. The phrase decoding along the bitstream, or alike, may refer to decoding the referred out-of-band data (which may be obtained from out-of-band transmission, signaling, or storage) that is associated with the bitstream.
[0167] A SEI NAL unit may include one or more SEI messages, which are not required for the decoding of output pictures but may assist in related processes, such as picture output timing, rendering, error detection, error concealment, and resource reservation. Several SEI messages are specified in H.264 / AVC and HEVC, and the user data SEI messages enable organizations and companies to specify SEI messages for their own use. H.264 / AVC and HEVC include the syntax and semantics for the specified SEI messages, but no process for handling the messages in the recipient is defined. Consequently, encoders are required to follow the H.264 / AVC standard or the HEVC standard when they create SEI messages, and decoders conforming to the H.264 / AVC standard or the HEVC standard, respectively, are not required to process SEI messages for output order conformance. One of the reasons to include the syntax and semantics of SEI messages in H.264 / AVC and HEVC is to allow different system specifications to interpret the supplemental information identically and hence interoperate. It is intended that system specifications can require the use of particular SEI messages both in the encoding end and in the decoding end, and additionally the process for handling particular SEI messages in the recipient can be specified.
[0168] In HEVC, there are two types of SEI NAL units, namely the suffix SEI NAL unit and the prefix SEI NAL unit, having a different nal_unit_type value from each other. The SEI message(s) contained in a suffix SEI NAL unit are associated with the VCL NAL unit preceding, in decoding order,the suffix SEI NAL unit. The SEI message(s) contained in a prefix SEI NAL unit are associated with the VCL NAL unit following, in decoding order, the prefix SEI NAL unit.
[0169] A coded picture is a coded representation of a picture. In HEVC, a coded picture may be defined as a coded representation of a picture including all coding tree units of the picture. In HEVC, an access unit (AU) may be defined as a set of NAL units that are associated with each other according to a specified classification rule, are consecutive in decoding order, and include at most one picture with any specific value of nuh_layer_id. In addition to including the VCL NAL units of the coded picture, an access unit may also include non-VCL NAL units. Said specified classification rule may for example associate pictures with the same output time or picture output count value into the same access unit.
[0170] A bitstream may be defined as a sequence of bits, in the form of a NAL unit stream or a byte stream, that forms the representation of coded pictures and associated data forming one or more coded video sequences. A first bitstream may be followed by a second bitstream in the same logical channel, such as in the same file or in the same connection of a communication protocol. An elementary stream (in the context of video coding) may be defined as a sequence of one or more bitstreams. The end of the first bitstream may be indicated by a specific NAL unit, which may be referred to as the end of bitstream (EOB) NAL unit and which is the last NAL unit of the bitstream. In HEVC and its current draft extensions, the EOB NAL unit is required to have nuh_layer_id equal to 0.
[0171] A coded video sequence may be defined as such a sequence of coded pictures in decoding order that is independently decodable and is followed by another coded video sequence or the end of the bitstream or an end of sequence NAL unit.
[0172] In HEVC, a coded video sequence may additionally or alternatively (to the specification above) be specified to end when a specific NAL unit, which may be referred to as an end of sequence (EOS) NAL unit, appears in the bitstream and has nuh_layer_id equal to 0.
[0173] A group of pictures (GOP) and its characteristics may be defined as follows. A GOP can be decoded regardless of whether any previous pictures were decoded. An open GOP is such a group of pictures in which pictures preceding the initial intra picture in output order might not be correctly decodable when the decoding starts from the initial intra picture of the open GOP. In other words, pictures of an open GOP may refer (in inter prediction) to pictures belonging to a previous GOP. An HEVC decoder can recognize an intra picture starting an open GOP, because a specific NAL unit type, CRA NAL unit type, may be used for its coded slices.
[0174] A closed GOP is such a group of pictures in which all pictures can be correctly decoded when the decoding starts from the initial intra picture of the closed GOP. In other words, no picture in a closed GOP refers to any pictures in previous GOPs. In H.264 / AVC and HEVC, a closed GOP may start from an IDR picture. In HEVC, a closed GOP may also start from a BLA_W_RADL or a BLA_N_LP picture. An open GOP coding structure is potentially more efficient in the compression compared to a closed GOP coding structure, due to a larger flexibility in selection of reference pictures.
[0175] A structure of pictures (SOP) may be defined as one or more coded pictures consecutive in decoding order, in which the first coded picture in decoding order is a reference picture at the lowest temporal sub-layer and no coded picture except, potentially, the first coded picture in decoding order is a RAP picture. All pictures in the previous SOP precede in decoding order all pictures in the current SOP, and all pictures in the next SOP succeed in decoding order all pictures in the current SOP. A SOP may represent a hierarchical and repetitive inter prediction structure. The term group of pictures (GOP) may sometimes be used interchangeably with the term SOP, and having the same semantics as the semantics of SOP.
[0176] A decoded picture buffer (DPB) may be used in the encoder and / or in the decoder. There are two reasons to buffer decoded pictures: for references in inter prediction, and for reordering decoded pictures into output order. As H.264 / AVC and HEVC provide a great deal of flexibility for both reference picture marking and output reordering, separate buffers for reference picture buffering and output picture buffering may waste memory resources. Hence, the DPB may include a unified decoded picture buffering process for reference pictures and output reordering. A decoded picture may be removed from the DPB when it is no longer used as a reference and is not needed for output.
[0177] In many coding modes of H.264 / AVC and HEVC, the reference picture for inter prediction is indicated with an index to a reference picture list. The index may be coded with variable length coding, which usually causes a smaller index to have a shorter value for the corresponding syntax element. In H.264 / AVC and HEVC, two reference picture lists (reference picture list 0 and reference picture list 1) are generated for each bi-predictive (B) slice, and one reference picture list (reference picture list 0) is formed for each inter-coded (P) slice.
[0178] A reference picture list, such as the reference picture list 0 and the reference picture list 1, may be constructed in two steps. First, an initial reference picture list is generated. The initial reference picture list may be generated, for example, on the basis of frame_num, picture order count (POC), temporal_id, or information on the prediction hierarchy such as a GOP structure, or any combination thereof. Second, the initial reference picture list may be reordered by reference picture listreordering (RPLR) syntax, also known as reference picture list modification syntax structure, which may be contained in slice headers. The initial reference picture lists may be modified through the reference picture list modification syntax structure, where pictures in the initial reference picture lists may be identified through an entry index to the list.
[0179] Many coding standards, including H.264 / AVC and HEVC, may have decoding process to derive a reference picture index to a reference picture list, which may be used to indicate which one of the multiple reference pictures is used for inter prediction for a particular block. A reference picture index may be coded by an encoder into the bitstream in some inter coding modes, or it may be derived (by an encoder and a decoder), for example using neighboring blocks, in some other inter coding modes.
[0180] Several candidate motion vectors may be derived for a single prediction unit. For example, motion vector prediction HEVC includes two motion vector prediction schemes, namely the advanced motion vector prediction (AMVP) and the merge mode. In the AMVP or the merge mode, a list of motion vector candidates is derived for a PU. There are two kinds of candidates: spatial candidates and temporal candidates, where temporal candidates may also be referred to as temporal motion vector prediction (TMVP) candidates.
[0181] A candidate list derivation may be performed, for example, as follows; it should be understood that other possibilities may exist for candidate list derivation. When the occupancy of the candidate list is not at a maximum, the spatial candidates are included in the candidate list first when they are available and not already in the candidate list. After that, when occupancy of the candidate list is not yet at the maximum, a temporal candidate is included in the candidate list. When the number of candidates still does not reach the maximum allowed number, the combined bi-predictive candidates (for B slices) and a zero motion vector are added in. After the candidate list has been constructed, the encoder decides the final motion information from candidates, for example based on a rate-distortion optimization (RDO) decision, and encodes the index of the selected candidate into the bitstream. Likewise, the decoder decodes the index of the selected candidate from the bitstream, constructs the candidate list, and uses the decoded index to select a motion vector predictor from the candidate list.
[0182] A motion vector anchor position may be defined as a position (e.g. horizontal and vertical coordinates) within a picture area relative to which the motion vector is applied. A horizontal offset and a vertical offset for the anchor position may be given in the slice header, slice parameter set, tile header, tile parameter set, or the like.
[0183] An example encoding method taking advantage of a motion vector anchor position may comprise: encoding an input picture into a coded constituent picture; reconstructing, as a part of said encoding, a decoded constituent picture corresponding to the coded constituent picture; encoding a spatial region into a coded tile, the encoding comprising: determining a horizontal offset and a vertical offset indicative of a region-wise anchor position of the spatial region within the decoded constituent picture; encoding the horizontal offset and the vertical offset; determining that a prediction unit at position of a first horizontal coordinate and a first vertical coordinate of the coded tile is predicted relative to the region-wise anchor position, wherein the first horizontal coordinate and the first vertical coordinate are horizontal and vertical coordinates, respectively, within the spatial region; indicating that the prediction unit is predicted relative to a prediction-unit anchor position that is relative to the regionwise anchor position; deriving a prediction-unit anchor position equal to sum of the first horizontal coordinate and the horizontal offset, and the first vertical coordinate and the vertical offset, respectively; determining a motion vector for the prediction unit; and applying the motion vector relative to the prediction-unit anchor position to obtain a prediction block.
[0184] An example decoding method wherein a motion vector anchor position is used may comprise: decoding a coded tile into a decoded tile, the decoding comprising: decoding a horizontal offset and a vertical offset; decoding an indication that a prediction unit at position of a first horizontal coordinate and a first vertical coordinate of the coded tile is predicted relative to a prediction-unit anchor position that is relative to the horizontal and vertical offset; deriving a prediction-unit anchor position equal to sum of the first horizontal coordinate and the horizontal offset, and the first vertical coordinate and the vertical offset, respectively; determining a motion vector for the prediction unit; and applying the motion vector relative to the prediction-unit anchor position to obtain a prediction block.
[0185] Features as described herein may generally relate to scalable video coding. Scalable video coding may refer to a coding structure where one bitstream can include multiple representations of the content, for example, at different bitrates, resolutions or frame rates. In these cases the receiver can extract the desired representation depending on its characteristics (e.g. resolution that matches best the display device). Alternatively, a server or a network element can extract the portions of the bitstream to be transmitted to the receiver depending on, for example, the network characteristics or processing capabilities of the receiver. A meaningful decoded representation can be produced by decoding only certain parts of a scalable bit stream. A scalable bitstream may include a “base layer” providing the lowest quality video available, and one or more enhancement layers that enhance the video quality when received and decoded together with the lower layers. In order to improve coding efficiency for the enhancement layers, the coded representation of that layer typically depends on the lower layers. For example, the motion and mode information of the enhancement layer can be predicted from lowerlayers. Similarly, the pixel data of the lower layers can be used to create prediction for the enhancement layer.
[0186] In some scalable video coding schemes, a video signal can be encoded into a base layer and one or more enhancement layers. An enhancement layer may enhance, for example, the temporal resolution (e.g., the frame rate), the spatial resolution, or simply the quality of the video content represented by another layer or part thereof. Each layer, together with all its dependent layers, is one representation of the video signal, for example, at a certain spatial resolution, temporal resolution and quality level. In the present disclosure, a scalable layer together with all of its dependent layers is referred to as a “scalable layer representation”. The portion of a scalable bitstream corresponding to a scalable layer representation can be extracted and decoded to produce a representation of the original signal at certain fidelity.
[0187] Scalability modes or scalability dimensions may include but are not limited to the following:
[0188] - Quality scalability: base layer pictures are coded at a lower quality than enhancement layer pictures, which may be achieved for example using a greater quantization parameter value (e.g., a greater quantization step size for transform coefficient quantization) in the base layer than in the enhancement layer. Quality scalability may be further categorized into fine-grain or fine-granularity scalability (FGS), medium-grain or medium-granularity scalability (MGS), and / or coarse-grain or coarse-granularity scalability (CGS), as described below.
[0189] - Spatial scalability: base layer pictures are coded at a lower resolution (e.g., have fewer samples) than enhancement layer pictures. Spatial scalability and quality scalability, particularly its coarse-grain scalability type, may sometimes be considered the same type of scalability.
[0190] - View scalability, which may also be referred to as multiview coding. The base layer represents a first view, whereas an enhancement layer represents a second view. A view may be defined as a sequence of pictures representing one camera or viewpoint. It may be considered that in stereoscopic or two-view video, one video sequence or view is presented for the left eye while a parallel view is presented for the right eye.
[0191] - Depth scalability, which may also be referred to as depth-enhanced coding. A layer or some layers of a bitstream may represent texture view(s), while other layer or layers may represent depth view(s).
[0192] It should be understood that many of the scalability types may be combined and applied together.
[0193] The term “layer” may be used in the context of any type of scalability, including view scalability and depth enhancements. An enhancement layer may refer to any type of an enhancement, such as SNR, spatial, multiview, and / or depth enhancement. A base layer may refer to any type of a base video sequence, such as a base view, a base layer for SNR / spatial scalability, or a texture base view for depth-enhanced video coding.
[0194] A sender, a gateway, a client, or another entity may select the transmitted layers and / or sub-layers of a scalable video bitstream. The terms layer extraction, extraction of layers, or layer downswitching may refer to transmitting fewer layers than what is available in the bitstream received by the sender, the gateway, the client, or another entity. Layer up-switching may refer to transmitting additional layer(s) compared to those transmitted prior to the layer up-switching by the sender, the gateway, the client, or another entity, e.g., restarting the transmission of one or more layers whose transmission was ceased earlier in layer down-switching. Similarly to layer down-switching and / or up- switching, the sender, the gateway, the client, or another entity may perform down- and / or up-switching of temporal sub-layers. The sender, the gateway, the client, or another entity may also perform both layer and sub-layer down-switching and / or up-switching. Layer and sub-layer down-switching and / or up-switching may be carried out in the same access unit or alike (e.g., virtually simultaneously) or may be carried out in different access units or alike (e.g., virtually at distinct times).
[0195] A scalable video encoder for quality scalability (also known as Signal-to-Noise or SNR) and / or spatial scalability may be implemented as follows. For a base layer, a conventional non-scalable video encoder and decoder may be used. The reconstructed / decoded pictures of the base layer are included in the reference picture buffer and / or reference picture lists for an enhancement layer. In case of spatial scalability, the reconstructed / decoded base-layer picture may be upsampled prior to its insertion into the reference picture lists for an enhancement-layer picture. The base layer decoded pictures may be inserted into a reference picture list(s) for coding / decoding of an enhancement layer picture, similarly to the decoded reference pictures of the enhancement layer. Consequently, the encoder may choose a base-layer reference picture as an inter prediction reference and indicate its use with a reference picture index in the coded bitstream. The decoder decodes from the bitstream, for example from a reference picture index, that a base-layer picture is used as an inter prediction reference for the enhancement layer. When a decoded base-layer picture is used as the prediction reference for an enhancement layer, it is referred to as an inter-layer reference picture.
[0196] While the previous paragraph described a scalable video codec with two scalability layers with an enhancement layer and a base layer, it needs to be understood that the description can be generalized to any two layers in a scalability hierarchy with more than two layers. In this case, a second enhancement layer may depend on a first enhancement layer in encoding and / or decoding processes, and the first enhancement layer may therefore be regarded as the base layer for the encoding and / or decoding of the second enhancement layer. Furthermore, it needs to be understood that there may be inter-layer reference pictures from more than one layer in a reference picture buffer or reference picture lists of an enhancement layer, and each of these inter-layer reference pictures may be considered to reside in a base layer or a reference layer for the enhancement layer being encoded and / or decoded. Furthermore, it needs to be understood that other types of inter-layer processing than reference-layer picture upsampling may take place instead or additionally. For example, the bit-depth of the samples of the reference-layer picture may be converted to the bit-depth of the enhancement layer and / or the sample values may undergo a mapping from the color space of the reference layer to the color space of the enhancement layer.
[0197] A scalable video coding and / or decoding scheme may use multi-loop coding and / or decoding, which may be characterized as follows. In the encoding / decoding, a base layer picture may be reconstructed / decoded to be used as a motion-compensation reference picture for subsequent pictures, in coding / decoding order, within the same layer or as a reference for inter-layer (or inter-view or inter-component) prediction. The reconstructed / decoded base layer picture may be stored in the DPB. An enhancement layer picture may likewise be reconstructed / decoded to be used as a motioncompensation reference picture for subsequent pictures, in coding / decoding order, within the same layer or as reference for inter-layer (or inter-view or inter-component) prediction for higher enhancement layers, when any. In addition to reconstructed / decoded sample values, syntax element values of the base / reference layer or variables derived from the syntax element values of the base / reference layer may be used in the inter-layer / inter-component / inter-view prediction.
[0198] Inter-layer prediction may be defined as prediction in a manner that is dependent on data elements (e.g. sample values or motion vectors) of reference pictures from a different layer than the layer of the current picture (being encoded or decoded). Many types of inter-layer prediction exist and may be applied in a scalable video encoder / decoder. The available types of inter-layer prediction may, for example, depend on the coding profile according to which the bitstream or a particular layer within the bitstream is being encoded or, when decoding, the coding profile that the bitstream or a particular layer within the bitstream is indicated to conform to. Alternatively or additionally, the available types of inter-layer prediction may depend on the types of scalability or the type of a scalable codec or video coding standard amendment (e.g. SHVC, MV-HEVC, or 3D-HEVC) being used.
[0199] A direct reference layer may be defined as a layer that may be used for inter-layer prediction of another layer for which the layer is the direct reference layer. A direct predicted layer may be defined as a layer for which another layer is a direct reference layer. An indirect reference layer may be defined as a layer that is not a direct reference layer of a second layer, but is a direct reference layer of a third layer that is a direct reference layer or indirect reference layer of a direct reference layer of the second layer for which the layer is the indirect reference layer. An indirect predicted layer may be defined as a layer for which another layer is an indirect reference layer. An independent layer may be defined as a layer that does not have direct reference layers. In other words, an independent layer is not predicted using inter-layer prediction. A non-base layer may be defined as any other layer than the base layer, and the base layer may be defined as the lowest layer in the bitstream. An independent nonbase layer may be defined as a layer that is both an independent layer and a non-base layer.
[0200] Similarly to MVC, in MV-HEVC, inter-view reference pictures can be included in the reference picture list(s) of the current picture being coded or decoded. SHVC uses multi-loop decoding operation (unlike the SVC extension of H.264 / AVC). SHVC may be considered to use a reference index based approach, e.g., an inter-layer reference picture can be included in one or more reference picture lists of the current picture being coded or decoded (as described above).
[0201] For the enhancement layer coding, the concepts and coding tools of HEVC base layer may be used in SHVC, MV-HEVC, and / or alike. However, the additional inter-layer prediction tools, which employ already coded data (including reconstructed picture samples and motion parameters a.k.a motion information) in reference layer for efficiently coding an enhancement layer, may be integrated to SHVC, MV-HEVC, and / or alike codec.
[0202] Video coding specifications may include a set of constraints for associating data units (e.g. NAL units in H.264 / AVC or HEVC) into access units. These constraints may be used to conclude access unit boundaries from a sequence of NAL units. For example, the following is specified in the HEVC standard:
[0203] - An access unit includes one coded picture with nuh_layer_id equal to 0, zero or moreVCL NAL units with nuh_layer_id greater than 0 and zero or more non-VCL NAL units.
[0204] - Let firstBIPicNalUnit be the first VCL NAL unit of a coded picture with nuh_layer_id equal to 0. The first of any of the following NAL units preceding firstBIPicNalUnit and succeeding the last VCL NAL unit preceding firstBIPicNalUnit, when any, specifies the start of a new access unit:access unit delimiter NAL unit with nuh_layer_id equal to 0 (when present); VPS NAL unit with nuh_layer_id equal to 0 (when present); SPS NAL unit with nuh_layer_id equal to 0 (when present); PPS NAL unit with nuh_layer_id equal to 0 (when present); Prefix SEI NAL unit with nuh_layer_id equal to 0 (when present); NAL units with nal_unit_type in the range of RSV_NVCL41..RSV_NVCL44 with nuh_layer_id equal to 0 (when present); NAL units with nal_unit_type in the range of UNSPEC48..UNSPEC55 with nuh_layer_id equal to 0 (when present).
[0205] The first NAL unit preceding firstBIPicNalUnit and succeeding the last VCL NAL unit preceding firstBIPicNalUnit, when any, can only be one of the above-listed NAL units. When there is none of the above NAL units preceding firstBIPicNalUnit and succeeding the last VCL NAL preceding firstBIPicNalUnit, when any, firstBIPicNalUnit starts a new access unit.
[0206] Access unit boundary detection may be based on, but may not be limited to, one or more of the following:
[0207] - Detecting that a VCL NAL unit of a base-layer picture is the first VCL NAL unit of an access unit, for example on the basis that: the VCL NAL unit includes a block address or alike that is the first block of the picture in decoding order; and / or the picture order count, picture number, or similar decoding or output order or timing indicator differs from that of the previous VCL NAL unit(s).
[0208] - Having detected the first VCL NAL unit of an access unit, concluding based on predefined rules, for example, based on nal_unit_type, which non-VCL NAL units that precede the first VCL NAL unit of an access unit and succeed the last VCL NAL unit of the previous access unit in decoding order belong to the access unit.
[0209] The versatile video coding (VVC) includes new coding tools compared to HEVC or H.264 / AVC. These coding tools are related to, for example, intra prediction; inter-picture prediction; transform, quantization and coefficients coding; entropy coding; in-loop filter; screen content coding; 360-degree video coding; high-level syntax and parallel processing. Some of these tools are briefly described in the following:
[0210] - Intra prediction: 67 intra mode with wide angles mode extension; block size and mode dependent 4 tap interpolation filter; position dependent intra prediction combination (PDPC); cross component linear model intra prediction (CCLM); multi-reference line intra prediction; intra subpartitions; weighted intra prediction with matrix multiplication.
[0211] - In ter -picture prediction: block motion copy with spatial, temporal, history-based, and pairwise average merging candidates; affine motion inter prediction; sub-block based temporal motion vector prediction; adaptive motion vector resolution; 8x8 block-based motion compression for temporal motion prediction; high precision (1 / 16 pel) motion vector storage and motion compensation with 8-tap interpolation filter for luma component and 4-tap interpolation filter for chroma component; triangular partitions; combined intra and inter prediction; merge with motion vector difference (MVD) (MMVD); symmetrical MVD coding; bi-directional optical flow; decoder side motion vector refinement; biprediction with CU-level weight.
[0212] - Transform, quantization and coefficients coding: multiple primary transform selection with DCT2, DST7 and DCT8; secondary transform for low frequency zone; sub-block transform for inter predicted residual; dependent quantization with max QP increased from 51 to 63; transform coefficient coding with sign data hiding; transform skip residual coding.
[0213] - Entropy coding: arithmetic coding engine with adaptive double windows probability update.
[0214] - In loop filter: in -loop reshaping; deblocking filter with strong longer filter; sample adaptive offset; adaptive loop filter.
[0215] - Screen content coding: current picture referencing with reference region restriction.
[0216] - 360-degree video coding: horizontal wrap-around motion compensation.
[0217] - High-level syntax and parallel processing: reference picture management with direct reference picture list signaling; tile groups with rectangular shape tile groups.
[0218] In VVC, each picture may be partitioned into coding tree units (CTUs) similar to HEVC. A CTU may be split into smaller CUs using quaternary tree structure. Each CU may be partitioned using quad-tree and nested multi-type tree, including ternary and binary split. There are specific rules to infer partitioning in picture boundaries. The redundant split patterns are disallowed in nested multitype partitioning.
[0219] In some video coding schemes, such as HEVC and VVC, a picture is divided into one or more tile rows and one or more tile columns. The partitioning of a picture to tiles forms a tile grid that may be characterized by a list of tile column widths and a list of tile row heights. A tile may be requiredto include an integer number of elementary coding blocks, such as CTUs in HEVC and VVC. Consequently, tile column widths and tile row heights may be expressed in the units of elementary coding blocks, such as CTUs in HEVC and VVC.
[0220] A tile may be defined as a sequence of elementary coding blocks, such as CTUs in HEVC and VVC, that covers one "cell" in the tile grid (e.g., a rectangular region of a picture). Elementary coding blocks, such as CTUs, may be ordered in the bitstream in raster scan order within a tile.
[0221] Some video coding schemes may allow further subdivision of a tile into one or more bricks, each including a number of CTU rows within the tile. A tile that is not partitioned into multiple bricks may also be referred to as a brick. However, a brick that is a true subset of a tile is not referred to as a tile.
[0222] In some video coding schemes, such as H.264 / AVC, HEVC and VVC, a coded picture may be partitioned into one or more slices. A slice may be decodable independently of other slices of a picture and hence a slice may be considered as a preferred unit for transmission. In some video coding schemes, such as H.264 / AVC, HEVC, and VVC, a video coding layer (VCL) NAL unit includes exactly one slice.
[0223] A slice may comprise an integer number of elementary coding blocks, such as CTUs in HEVC or VVC.
[0224] In some video coding schemes, such as VVC, a slice includes an integer number of tiles of a picture or an integer number of CTU rows of a tile.
[0225] In some video coding schemes, two modes of slices may be supported, namely the rasterscan slice mode and the rectangular slice mode. In the raster-scan slice mode, a slice includes a sequence of tiles in a tile raster scan of a picture. In the rectangular slice mode, a slice includes an integer number of tiles of a picture or an integer number of CTU rows of a tile that collectively form a rectangular region of the picture.
[0226] A non-VCL NAL unit may be, for example, one of the following types: a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a supplemental enhancement information (SEI) NAL unit, a picture header (PH) NAL unit, an end of sequence NAL unit, an end of bitstream NAL unit, or a filler data NAL unit. Some non-VCL NAL units, such as parameter sets and picture headers, may be needed for the reconstruction of decodedpictures, whereas many of the other non-VCL NAL units might not be necessary for the reconstruction of decoded sample values.
[0227] Some coding formats specify parameter sets that may carry parameter values needed for the decoding or reconstruction of decoded pictures. Some examples of different types of parameter sets are now briefly described. A video parameter set (VPS) may include parameters that are common across multiple layers in a coded video sequence or describe relations between layers. Parameters that remain unchanged through a coded video sequence (in a single-layer bitstream) or in a coded layer video sequence may be included in a sequence parameter set (SPS). In addition to the parameters that may be needed by the decoding process, the sequence parameter set may optionally include video usability information (VUI), which includes parameters that may be important for buffering, picture output timing, rendering, and resource reservation. A picture parameter set (PPS) includes such parameters that are likely to be unchanged in several coded pictures. A picture parameter set may include parameters that can be referred to by the coded image segments of one or more coded pictures. A header parameter set (HPS) has been proposed to include such parameters that may change on a picture basis. In VVC, an Adaptation Parameter Set (APS) may comprise parameters for decoding processes of different types, such as adaptive loop filtering or luma mapping with chroma scaling.
[0228] A parameter set may be activated when it is referenced e.g., through its identifier. For example, a header of an image segment, such as a slice header, may include an identifier of the PPS that is activated for decoding the coded picture including the image segment. A PPS may include an identifier of the SPS that is activated, when the PPS is activated. An activation of a parameter set of a particular type may cause the deactivation of the previously active parameter set of the same type.
[0229] Instead of or in addition to parameter sets at different hierarchy levels (e.g. sequence and picture), video coding formats may include header syntax structures, such as a sequence header or a picture header. A sequence header may precede any other data of the coded video sequence in the bitstream order. A picture header may precede any coded video data for the picture in the bitstream order.
[0230] Video coding specifications may enable the use of supplemental enhancement information (SEI) messages or alike. Some video coding specifications include SEI NAL units, and some video coding specifications include both prefix SEI NAL units and suffix SEI NAL units. A prefix SEI NAL unit can start a picture unit or alike; and a suffix SEI NAL unit can end a picture unit or alike. Hereafter, an SEI NAL unit may equivalently refer to a prefix SEI NAL unit or a suffix SEI NAL unit. An SEI NAL unit includes one or more SEI messages, which are not required for thedecoding of output pictures but may assist in related processes, such as picture output timing, postprocessing of decoded pictures, rendering, error detection, error concealment, and resource reservation.
[0231] Several SEI messages are specified in H.264 / AVC, H.265 / HEVC, H.266 / VVC, and H.274 / VSEI standards, and the user data SEI messages enable organizations and companies to specify SEI messages for specific use. The standards may include the syntax and semantics for the specified SEI messages, but a process for handling the messages in the recipient might not be defined. Consequently, encoders may be required to follow the standard specifying a SEI message when they create SEI message(s), and decoders might not be required to process SEI messages for output order conformance. One of the reasons to include the syntax and semantics of SEI messages in standards is to allow different system specifications to interpret the supplemental information identically and hence interoperate. It is intended that system specifications can require the use of particular SEI messages both in the encoding end and in the decoding end, and additionally the process for handling particular SEI messages in the recipient can be specified.
[0232] Features as described herein may generally relate to a bitstream. A bitstream may be defined as a sequence of bits, which may in some coding formats or standards be in the form of a NAL unit stream or a byte stream, that forms the representation of coded pictures and associated data forming one or more coded video sequences. A first bitstream may be followed by a second bitstream in the same logical channel, such as in the same file or in the same connection of a communication protocol. An elementary stream (in the context of video coding) may be defined as a sequence of one or more bitstreams. In some coding formats or standards, the end of the first bitstream may be indicated by a specific NAL unit, which may be referred to as the end of bitstream (EOB) NAL unit and is the last NAL unit of the bitstream.
[0233] A coded video sequence (CVS) may be defined as such a sequence of coded pictures in decoding order that is independently decodable and is followed by another coded video sequence or the end of the bitstream.
[0234] The subpicture feature of VVC allows for partitioning of the VVC bitstream in a flexible manner as multiple rectangles representing subpictures, where each subpicture comprises one or more slices. In other words, a subpicture may be defined as a rectangular region of one or more slices within a picture, wherein the one or more slices are complete. Consequently, a subpicture includes one or more slices that collectively cover a rectangular region of a picture. The slices of a subpicture may be required to be rectangular slices.
[0235] In VVC, the feature of subpictures enables efficient extraction of subpicture(s) from one or more bitstream and merging the extracted subpictures to form another bitstream without excessive penalty in compression efficiency and without modifications of VCL NAL units (e.g., slices).
[0236] The use of subpictures in a coded video sequence (CVS), however, requires appropriate configuration of the encoder and other parameters such as SPS / PPS and so on. In VVC, a layout of partitioning of a picture to subpictures may be indicated in and / or decoded from an SPS. A subpicture layout may be defined as a partitioning of a picture to subpictures. In VVC, the SPS syntax indicates the partitioning of a picture to subpictures by providing for each subpicture syntax elements indicative of: the x and y coordinates of the top-left corner of the subpicture, the width of the subpicture, and the height of the subpicture, in CTU units. One or more of the following properties may be indicated (e.g. by an encoder) or decoded (e.g. by a decoder) or inferred (e.g. by an encoder and / or a decoder) for the subpictures collectively or per each subpicture individually: i) whether or not a subpicture is treated like a picture in the decoding process (or equivalently, whether or not subpicture boundaries are treated like picture boundaries in the decoding process); in some cases, this property excludes in-loop filtering operations, which may be separately indicated / decoded / inferred; ii) whether or not in-loop filtering operations are performed across the subpicture boundaries. When a subpicture is treated like a picture in the decoding process, any references to sample locations outside the subpicture boundaries are saturated to be within the subpicture boundaries. This may be regarded being equivalent to padding samples outside subpicture boundaries with the boundary sample values for decoding the subpicture. Consequently, motion vectors may be allowed to cause references outside subpicture boundaries in a subpicture that is extractable.
[0237] An independent subpicture (a.k.a. an extractable subpicture) may be defined as a subpicture i) with subpicture boundaries that are treated as picture boundaries and ii) without loop filtering across the subpicture boundaries. A dependent subpicture may be defined as a subpicture that is not an independent subpicture.
[0238] In video coding, an isolated region may be defined as a picture region that is allowed to depend only on the corresponding isolated region in reference pictures and does not depend on any other picture regions in the current picture or in the reference pictures. The corresponding isolated region in reference pictures may be, for example, the picture region that collocates with the isolated region in a current picture. A coded isolated region may be decoded without the presence of any picture regions of the same coded picture.
[0239] A VVC subpicture with boundaries treated like picture boundaries may be regarded as an isolated region.
[0240] A motion-constrained tile set (MCTS) is a set of tiles such that the inter prediction process is constrained in encoding such that no sample value outside the MCTS, and no sample value at a fractional sample position that is derived using one or more sample values outside the motion- constrained tile set, is used for inter prediction of any sample within the motion-constrained tile set. Additionally, the encoding of an MCTS is constrained in a manner that no parameter prediction takes inputs from blocks outside the MCTS. For example, the encoding of an MCTS is constrained in a manner that motion vector candidates are not derived from blocks outside the MCTS. In HEVC, this may be enforced by turning off temporal motion vector prediction of HEVC, or by disallowing the encoder to use the temporal motion vector prediction (TMVP) candidate or any motion vector prediction candidate following the TMVP candidate in a motion vector candidate list for prediction units located directly left of the right tile boundary of the MCTS, except the last one at the bottom right of the MCTS.
[0241] In general, an MCTS may be defined to be a tile set that is independent of any sample values and coded data, such as motion vectors, that are outside the MCTS. An MCTS sequence may be defined as a sequence of respective MCTSs in one or more coded video sequences or alike. In some cases, an MCTS may be required to form a rectangular area. It should be understood that, depending on the context, an MCTS may refer to the tile set within a picture or to the respective tile set in a sequence of pictures. The respective tile set may be, but in general need not be, collocated in the sequence of pictures. A motion-constrained tile set may be regarded as an independently coded tile set, since it may be decoded without the other tile sets. An MCTS is an example of an isolated region.
[0242] Features as described herein may generally relate to VVC. The VVC has a functionality of subpictures which may be regarded to improve the motion constrained tiles. VVC support for realtime conversational and low latency use cases will be important to fully exploit the functionality and end user benefit with modern networks (e.g. ultra-reliable low-latency communication (ULLRC) 5G networks, over-the-top (OTT) delivery, etc.). VVC encoding and decoding is computationally complex. With increasing computational complexity, the end-user devices consuming the content are heterogeneous, for example devices supporting single decoding instances to devices supporting multiple decoding instances and more sophisticated devices having multiple decoders. Consequently, the system carrying the payload should be able to support a variety of scenarios for scalable deployments. There has been rapid growth in the resolution (e.g. 8K) of the video consumed via consumer electronics (CE) devices (e.g. TVs, mobile devices) which can benefit with the ability to execute multiple paralleldecoders. One example use case can be parallel decoding for low latency unicast or multicast delivery of 8K VVC encoded content.
[0243] The AV 1 codec supports input video signals in the 4:0:0 (monochrome), 4:2:0, 4:2:2, and 4:4:4 formats. The allowed pixel representations are 8, 10, and 12 bit. The AVI codec operates on pixel blocks. Each pixel block is processed in a predictive-transform coding scheme, where the prediction comes from either intraframe reference pixels, interframe motion compensation, or some combinations of the two. The residuals undergo a 2-D unitary transform to further remove the spatial correlations, and the transform coefficients are quantized. Both the prediction syntax elements and the quantized transform coefficient indexes are then entropy coded using arithmetic coding. There are three optional in-loop postprocessing filter stages to enhance the quality of the reconstructed frame for reference by subsequent coded frames. A normative film grain synthesis unit is also available to improve the perceptual quality of the displayed frames.
[0244] The AV 1 bitstream is packetized into open bitstream units (OBUs). An ordered sequence of OBUs is fed into the AVI decoding process, where each OBU comprises a variable length string of bytes. An OBU includes a header and a payload. The header identifies the OBU type and specifies the payload size. The OBU types may include the following.
[0245] 1) Sequence header includes information that applies to the entire sequence, for example sequence profile (see Section VIII) and whether to enable certain coding tools.
[0246] 2) Temporal delimiter indicates the frame presentation time stamp. All displayable frames following a temporal delimiter OBU will use this time stamp, until the next temporal delimiter OBU arrives. A temporal delimiter and its subsequent OBUs of the same time stamp are referred to as a temporal unit. In the context of scalable coding, the compression data associated with all representations of a frame at various spatial and fidelity resolutions will be in the same temporal unit.
[0247] 3) Frame header sets up the coding information for a given frame, including signaling inter or intraframe type, indicating the reference frames, and signaling probability model update method.
[0248] 4) Tile group includes the tile data associated with a frame. Each tile can be independently decoded. The collective reconstructions form the reconstructed frame after potential loop filtering.
[0249] 5) Frame includes the frame header and tile data. The frame OBU is largely equivalent to a frame header OBU and a tile group OBU, but allows less overhead cost.
[0250] 6) Metadata carries information, such as high dynamic range, scalability, and timecode.
[0251] 7) Tile list includes tile data similar to a tile group OBU. However, each tile here has an additional header that indicates its reference frame index and position in the current frame. This allows the decoder to process a subset of tiles and display the corresponding part of the frame, without the need to fully decode all the tiles in the frame.
[0252] In many video communication or transmission systems, transport mechanisms, and multimedia container file formats, there are mechanisms to transmit or store a scalability layer separately from another scalability layer of the same bitstream, e.g. to transmit or store the base layer separately from the enhancement layer(s). It may be considered that layers are stored in or transmitted through separate logical channels. For example in ISOBMFF, the base layer can be stored as a track and each enhancement layer can be stored in another track, which may be linked to the base-layer track using so-called track references.
[0253] SEI processing order (SPO) and processing order nesting (PON) SEI messages
[0254] Standardization is ongoing for specifying the SEI processing order (SPO) SEI message and the processing order nesting (PON) SEI message in VVC. At the time of writing this disclosure, the latest specification text for the SPO and PON SEI messages is available in document JVET- AG2027-V1.
[0255] The SEI processing order (SPO) SEI message carries information indicating the preferred processing order, as determined by the encoder (e.g., the content producer), for a group of types of SEI messages that may be present in a CVS.
[0256] The processing order nesting (PON) SEI message includes one or more SEI messages that should be applied only as parts of the processing chain identified by an associated SEI processing order SEI message and should not be applied in a manner that would contradict with the processing chain identified by the associated SEI processing order SEI message.
[0257] The syntax of SPO SEI message may be specified as follows:
[0258] The semantics of the SPO SEI message may be specified as follows:
[0259] The semantics of the SPO SEI message uses the concept of types of SEI messages. SEI messages that have different payloadType values are considered different types of SEI messages. Additionally, different SEI messages that have the same payloadType value but are differentiated by values of syntax elements in the SEI payload are considered different types of SEI messages. Such differentiation by values of syntax elements in the SEI payload is to be performed by comparing values sent using po_sei_prefix_data_bit[ i ][ j ] syntax elements (when present) or values sent as SEI messages within a processing order nesting SEI message (when present). For example, neural-network post-filter characteristics (NNPFC) SEI messages can be differentiated by having different nnpfc_id values.
[0260] When the i-th SEI message seiA in any SPO SEI message has po_sei_wrapping_flag[ i ] and po_sei_prefix_flag[ i ] both equal to 0, there shall be no other SEI message seiB included in thesame SPO SEI message or in a different SPO SEI message in the current CVS for which all of the following are true:
[0261] - The value of po_sei_payload_type[ i ] of seiB is the same as that for seiA.
[0262] - The value of po_sei_wrapping_flag[ i ] of seiB is equal to 0.
[0263] - The value of po_sei_prefix_flag[ i ] of seiB is equal to 1.
[0264] When an SPO SEI message with a particular value of po_id is present in any access unit of a CVS, an SPO SEI message with that particular value of po_id shall be present in the first access unit of the CVS in decoding order. The number of SEI messages and the payloadType codes of the SEI messages indicated within each SPO SEI message with the same value of po_id persist in decoding order from the current access unit until the end of the CVS in output order.
[0265] The SPO SEI message can carry one or more SEI prefix indications of a particular payloadType. When present, each SEI prefix indication is a bit string that follows the SEI pay load syntax of that value of payloadType and contains a number of complete syntax elements starting from the first syntax element in the SEI payload. These SEI prefix indications should provide sufficient information to determine the specific processing order for types of SEI messages having the same value of payloadType but a different preferred processing order.
[0266] po_id includes an identifying number to identify the SPO SEI message.
[0267] A processing chain includes a list of types of SEI messages identified by an SPO SEI message in the preferred processing order indicated in the SPO SEI message.
[0268] Each type of SEI message in the processing chain indicated by an SPO SEI message is identified by the syntax elements po_sei_payload_type[ i ], po_sei_wrapping_flag[ i ], po_sei_processing_order[ i ] and, when present, po_num_bits_in_prefix_indication_minusl[ i ] and po_prefix_data_bit[ i ][ j ].
[0269] An SEI message type is not required to belong to any processing chain and may belong to any number of processing chains identified by SPO SEI messages with different po_id values.
[0270] Each SEI message of an SEI message type identified within the SPO SEI message has the same persistence scope as if the SEI message was carried outside of the SPO SEI message and not identified within an SPO SEI message.
[0271] Processing chains can be alternatives to each other, e.g., such that at most processing chain is chosen to be applied, or they can be complementary, e.g., such that more than one processing chain is chosen and applied separately, with each processing chain generating one output.
[0272] po_num_sei_messages_minus2 plus 2 indicates the number of types of SEI messages for which the preferred order of processing is indicated in the SPO SEI message.
[0273] po_sei_wrapping_flag[ i ] equal to 1 specifies that one or more processing order nesting SEI messages with both of the following constraints should be present:
[0274] - pon_target_po_id[ j ] with any value of j is equal to po_id.
[0275] - There is a k-th loop entry in the processing order nesting SEI message such that the payloadType of the k-th nested SEI message is equal to po_sei_payload_type[ i ] and pon_processing_order[ k ] is equal to po_sei_processing_order[ i ].
[0276] When po_sei_wrapping_flag[ i ] is equal to 0, an SEI message with payloadType equal to po_sei_payload_type[ i ] (and, when po_sei_prefix_flag[ i ] equal to 1, prefix data that matches the values of po_sei_prefix_data_bit[ i ][ j ]) should be present outside of the processing order nesting SEI message.
[0277] NOTE 2 - po_sei_wrapping_flag[ i ] equal to 1 enables SEI messages to be carried within the processing order nesting SEI message to prevent such SEI messages from being incorrectly interpreted by decoders that do not process the SPO SEI message. Thus, po_sei_wrapping_flag[ i ] equal to 1 is intended to be used when po_sei_wrapping_flag[ i ] equal to 0 can lead to unintended results being produced by such decoders.
[0278] po_sei_importance_flag[ i ] indicates the degree of importance determined by the encoder for the type of SEI message with index i.
[0279] When the decoding system cannot interpret or does not support the functionality indicated by any indicated SEI message that has po_sei_importance_flag[ i ] equal to 1, it should ignore the entire SPO SEI message.
[0280] po_sei_payload_type[ i ] specifies the payloadType value of the i-th type of SEI message.
[0281] po_sei_prefix_flag[ i ] equal to 1 specifies that po_num_bits_in_prefix_indication_minus 1 [ i ] and some po_sei_prefix_data_bit[ i ] [ j ] syntax elements are present. po_sei_prefix_flag[ i ] equal to 0 specifies that these syntax elements are not present.
[0282] SeiProcessingOrderSeiList is set to include the payloadType values 3, 4, 5, 19, 137, 142, 144, 147, 148, 149, 165, 177, 210, and 211. The value of po_sei_payload_type[ i ] for each i in the range of 0 to po_num_sei_messages_minus2 + 1, inclusive, shall be equal to a value in SeiProcessingOrderSeiList.
[0283] po_sei_processing_order[ i ] indicates the preferred order of processing of the i-th type of SEI message for which preferred processing order information is provided in the SPO SEI message. For any two different integer values of m and n, po_sei_processing_order[ m ] less than po_sei_processing_order[ n ] indicates that the type of SEI message associated with index m should be processed before the type of SEI message associated with index n, and po_sei_processing_order[ m ] equal to po_sei_processing_order[ n ] indicates that there is no preferred order of processing between the types of SEI messages associated with indexes m and n (e.g., they can indicate different properties that are both applicable at that stage, or alternative processes that can be applied, or one can indicate a property and the other can indicate a process).
[0284] For i greater than 0, po_sei_processing_order[ i ] shall be greater than or equal to po_sei_processing_order[ i - 1 ].
[0285] po_num_bits_in_prefix_indication_minusl [ i ] and po_sei_prefix_data_bit[ i ] [ j ], when present, have the same semantics as the num_bits_in_prefix_indication_minusl[ i ] and sei_prefix_data_bit[ i ] [ j ] syntax elements of the SEI prefix indication SEI message, with prefix_sei_payload_type replaced by po_sei_payload_type[ i ].
[0286] When more than one SPO SEI message with a particular value of po_id is present in a CVS, the values of po_num_sei_messages_minus2 and, for each value of i, the values ofpo_sei_wrapping_flag[ i ], po_sei_prefix_flag[ i ], po_sei_importance_flag[ i ], po_sei_payload_type[ i ], po_sei_processing_order[ i ] shall be the same as in the other SPO SEI messages in the CVS with the same value of po_id.
[0287] po_byte_alignment_bit_equal_to_one shall be equal to 1.
[0288] The syntax of the PON SEI message may be specified as follows:
[0289] The semantics of the PON SEI message may be specified as follows.
[0290] The SEI messages included in a PON SEI message are referred to as PON-nested SEI messages.
[0291] An encoder may include multiple PON SEI messages in the same access unit. For example, a first PON SEI message in an access unit may include a PON-nested SEI message that applies to multiple processing chains and one or more other PON SEI messages in the same access unit that apply to a single processing chain only.
[0292] It is a requirement of bitstream conformance that the semantics and effect of an SEI message that is not a PON-nested SEI message shall not depend on any PON-nested SEI message. Consequences of this constraint include the following specific constraints, in which an associated SEImessage is considered to be an SEI message that affects the semantics or effect of a particular SEI message:
[0293] - When a neural-network post-filter characteristics SEI message is present with a particular value of nnpfc_id that is a PON-nested SEI message, any associated neural-network postfilter activation SEI messages with nnpfa_target_id equal to that particular value of nnpfc_id shall also be PON-nested SEI messages.
[0294] - When a neural-network post-filter activation (NNPFA) SEI message is present with nnpfa_persistence_flag equal to 1 and a particular value of nnpfa_target_id that is not a PON-nested SEI message, the next picture in output order in the same coded layer video sequence (CLVS) that has an NNPFA SEI message with the same value of nnpfa_target_id (if any) shall not have an associated NNPFA SEI message that is a PON-nested SEI message.
[0295] - When a film grain characteristics SEI message is present with fg_characteristics_persistence_flag equal to 1 that is not a PON-nested SEI message, there shall not be an associated film grain characteristics SEI message in the same CLVS that is a PON-nested SEI message.
[0296] - When a frame packing arrangement SEI message is present with fp_arrangement_persistence_flag equal to 1 that is not a PON-nested SEI message, there shall not be an associated frame packing arrangement SEI message in the same CLVS with fp_arrangement_cancel_flag equal to 1 or the same value of fp_arrangement_id that is a PON-nested SEI message.
[0297] - When a content colour volume SEI message is present with ccv_persistence_flag equal to 1 that is not a PON-nested SEI message, there shall not be an associated frame packing arrangement SEI message in the same CLVS that is a PON-nested SEI message.
[0298] - When an equirectangular projection SEI message is present with erp_persistence_flag equal to 1 that is not a PON-nested SEI message, there shall not be an associated equirectangular projection SEI message in the same CLVS that is a PON-nested SEI message.
[0299] - When a generalized cubemap project SEI message is present with gcmp_persistence_flag equal to 1 that is not a PON-nested SEI message, there shall not be an associated generalized cubemap project SEI message in the same CLVS that is a PON-nested SEI message.
[0300] - When a sphere rotation SEI message is present with sphere_rotation_persistence_flag equal to 1 that is not a PON-nested SEI message, there shall not be an associated sphere rotation SEI message in the same CLVS that is a PON-nested SEI message.
[0301] - When a region-wise packing SEI message is present with rwp_persistence_flag equal to 1 that is not a PON-nested SEI message, there shall not be an associated region-wise packing SEI message in the same CLVS that is a PON-nested SEI message.
[0302] - When an omnidirectional viewport SEI message is present with omni_viewport_persistence_flag equal to 1 that is not a PON-nested SEI message, there shall not be an associated omnidirectional viewport SEI message in the same CLVS that is a PON-nested SEI message.
[0303] - When a sample aspect ratio SEI message is present with sari_persistence_flag equal to 1 that is not a PON-nested SEI message, there shall not be an associated sample aspect ratio SEI message in the same CLVS that is a PON-nested SEI message.
[0304] - When an annotated regions SEI message is present that is not a PON-nested SEI message, there shall not be an associated annotated regions SEI message in the same CLVS that is a PON-nested SEI message.
[0305] - When an alpha channel information SEI message is present that is not a PON-nestedSEI message, there shall not be an associated alpha channel information SEI message in the same CLVS that is a PON-nested SEI message.
[0306] - When a display orientation SEI message is present that is not a PON-nested SEI message, there shall not be an associated display orientation SEI message in the same CLVS that is a PON-nested SEI message.
[0307] - When a colour transform indication SEI message is present with colour_transform_persistence_flag equal to 1 that is not a PON-nested SEI message, there shall not be an associated colour transform indication SEI message in the same CLVS with colour_transform_cancel_flag equal to 1 or the same value of colour_transform_id that is a PON-nested SEI message.
[0308] pon_num_po_ids_minusl plus 1 specifies the number of the SEI processing order SEI messages SEI associated with this PON SEI message.
[0309] pon_target_po_id[ i ] indicates the po_id of the i-th associated SEI processing order SEI message.
[0310] pon_num_seis_minusl plus 1 specifies the number of the PON-nested SEI messages that are included in this PON SEI message.
[0311] pon_processing_order[ i ] specifies the position of the i-th processing-order-nested SEI message within the processing order defined by the associated SEI processing order SEI message. When i is greater than 0, pon_processing_order[ i ] shall be greater than or equal to pon_processing_order[ i - 1 ].
[0312] For each associated SEI processing order SEI message there shall be at least one value of i in the range of 0 to pon_num_seis_minusl, inclusive, in the PON SEI message for which the associated SEI processing order SEI message has some entry k for which the following are true:
[0313] - po_sei_processing_order[ k ] is equal to pon_processing_order[ i ]; and
[0314] - po_sei_payload_type[ k ] is equal to the payloadType value of the i-th PON-nestedSEI message.
[0315] - When po_sei_prefix_flag[ k ] is equal to 1, po_sei_prefix_data_bit[ k ][ j ] for j in the range of 0 to po_num_bits_in_prefix_indication_minusl [ k ], inclusive, include the same content as the po_num_bits_in_prefix_indication_minus 1 [ k ] plus 1 initial bits of the SEI message payload of the i- th PON-nested SEI message.
[0316] The i-th PON-nested SEI message should be applied as the k-th loop entry of the associated SEI processing order SEI message.
[0317] ISO / IEC 23001-11 Energy-Efficient Media Consumption (Green Metadata) provides a mechanism to signal quality metrics for individual decoder output pictures in AVC, HEVC or VVC, or for subpictures in VVC, and provides definitions of the metrics, including formulas for calculation of metric values.
[0318] ISO / IEC 23001-10 Carriage of timed metadata metrics of media in ISO base media file format provides definitions for a specific list of quality metrics, including formulas for calculation of metric values, for inclusion in a file format.
[0319] The Post-processing filter gain SEI message, is in the Technologies under Consideration for future versions of VSEI, JVET-AF2032, available from [https: / / jvet- experts.org / doc_end_user / current_document.php?id= 13592 (last accessed on April 5, 2024)]. The postprocessing filter gain SEI provides signaling of the quality gain achieved by applying an indicated postprocessing filter (PPF) or an indicated PPF group.
[0320] Examples of gain signaling of a neural network post-filter is described in REZAZADEGAN TAVAKOLI et al., Pub No US 2022 / 0256227, which is hereby incorporated herein by reference in its entirety.
[0321] Applications can make use of different types of gain metrics, including absolute quality of decoded pictures, absolute quality of decoded pictures following post-processing filter application, and gain achieved by post-processing filter application. Both sequence (CLVS) quality metrics and per picture metrics are of interest.
[0322] A high complexity post-processing operation, e.g., an artificial intelligence (Al) task such as facial recognition, may select the highest quality frames of a video to operate on. A decoder may choose not to display pictures with quality levels below a target quality level.
[0323] Some commonly used quality metrics, such as VMAF, are not defined in a standard, such as an ISO / IEC or ITU-T, which may not allow them to be normatively referenced by a standard specification such as a JVET specification.
[0324] When a particular metric definition is unknown for a decoding system, it cannot be concluded if the quality increases or decreases when the metric value increases. Such knowledge would be beneficial, for example, to choose from multiple processing chains or to pick a good-quality decoded picture for machine analysis.
[0325] The ongoing standardization of the SEI processing order (SPO) SEI message enables indicating of several alternative post-processing chains. In order to provide decoding systems information that can be used to select between the processing chains, a mechanism would be neededto indicate picture quality or gain achieved by the processing chains described by the SPO SEI messages. Some processing stages described by the SPO SEI message may be indicated to be non- essential (po_sei_importance_flag[ i ] equal to 0). In order to provide decoding systems information that can be used to determine if a non-essential processing stage is worth performing, a mechanism would be needed to indicate a gain achieved by the processing stage described by an SPO SEI message.
[0326] In an example embodiment, a quality metric SEI message is proposed, which provides additional functionality than the post-processing filter gain SEI message in JVET-AF2032. The existing post-processing filter gain SEI message signals quality gain achieved using post-processing filters, while the proposed quality metric SEI adds signaling of absolute quality metrics per CLVS and / or individual picture and supports additional quality metrics, including user-specified metrics.
[0327] The proposed SEI message also supports use of user-specified metrics. For user-specified metrics, signaling includes a text description of the metric, a flag to indicate whether higher or lower values of the metric correspond to increased quality, and a flag to indicate when the quality metric is a full reference metric. Providing capability for user-specified quality metrics enables signaling of proprietary or newly developed metrics, while providing useful information about them. Optional signaling of a text description of user-specified quality metrics enables identification of metrics that are not specified in the VSEI standard, because there is not a referenceable standard, such as for VMAF, or for metrics that were defined after the VSEI version was published.
[0328] Some image quality metrics measure error rather than quality, for which lower values indicate increased quality. Non full reference quality metrics may represent no reference quality metrics, or may measure other uses, such as measuring the degree of realism of generative Al images or video, rather than measuring the impact of video compression.
[0329] The proposed functionality, as an example, is summarized below.
[0330] - Signal quality metrics per CLVS and / or cropped decoded picture;
[0331] - Signal quality metrics as absolute or as relative gains for PPFs / PPFGs, relative to the cropped decoded picture or relative to the input of the current PPF of the PPFG;
[0332] - User-specified quality metrics comprises at least:
[0333] - Text description of the metric;
[0334] - Flag to indicate whether higher or lower values of the metric correspond to increased quality;
[0335] - Flag to indicate when the quality metric is a full reference metric; and / or
[0336] - Signaling of number of bytes signaled for the metric values;
[0337] - Quality metrics can be signaled for a single color component, or for 3 components individually.
[0338] A flag, qm_metric_definitions_present_flag, is provided to indicate whether a particular SEI message includes definitions of the metrics. This allows the metric definitions to be signaled for the first picture of a sequence without requiring that the definitions be repeated in subsequent pictures in the sequence.
[0339] A flag, qm_clvs_values_present_flag, indicates whether sequence level metric values are signaled. A separate flag, qm_pic_values_present_flag, indicates whether per picture metric values are signaled. This allows an encoder to choose whether to signal sequence level metrics values, picture level metric values, or both.
[0340] The qm_num_metrics_minusl syntax element enables an encoder to send multiple metric values within the same SEI message. It is not required that all metric values be signaled for all pictures in the SEI message even when per picture metric signaling is used.
[0341] A flag, qm_gain_flag[ i ], enables an encoder to indicate whether the signaled i-th metric values indicate a gain associated with a post-processing filter, or whether the values indicate absolute quality of the picture or sequence.
[0342] When qm_gain_flag[ i ] is equal to 1, a flag qm_gain_reference_flag[ i ] indicates whether the reference for the gain is the input picture to the encoder or the input picture to the post-processing stage.
[0343] A table of allowable values for qm_metric_type[ i ] is provided to indicate the metric type of each metric. A value of 0 is used to indicate a user-specified metric. Reserved and unspecified values are also provided.
[0344] The PSNR-YUV metric type can be used to indicate a combination of the PSNR values for all 3 color components, Y, U, and V.
[0345] For a user-specified or unspecified metric type, additional information is signaled to provide information about the metric. A flag, qm_metric_increasing_flag[ i ], is used to indicate whether a higher value of the metric corresponds to a higher quality value or a lower quality value.
[0346] A flag, qm_full_reference_flag[ i ] , indicates whether the quality metric is a full reference metric, where a comparison is made between a test picture and a reference picture. Metrics that are not full reference metrics may refer to subjective quality, or may measure other characteristics, such as measuring the degree of realism of generative Al images or video, rather than measuring the impact of video compression.
[0347] When the SEI message is associated with a post-processing filter through a processing order nesting SEI message, the test picture is the output of the post processing filter, and if not, the test picture is the cropped decoded picture.
[0348] For user-specified or unspecified metric types, the qm_value_len_minusl_in_bytes[ i ] syntax element specifies the number of bytes used to signal metric values. For specified metric types, the number of bytes used to signal the metric values is defined for each type in a table.
[0349] For user-specified or unspecified metric types, a flag, qm_metric_description_present_flag[ i ], indicates that a text description of the metric is signaled, using the qm_metric_description[ i ] syntax element. This text description can be used to indicate a metric like video multi-method assessment fusion (VMAF), which is not normatively referenced, or can indicate other new metrics not specified in the standard.
[0350] The syntax table and semantics for the proposed SEI message are shown below.
[0351] Semantics
[0352] The quality metrics SEI message includes quality metric values indicating any of the following:
[0353] - The quality of a single picture;
[0354] - The mean quality of all the pictures corresponding to a CLVS;
[0355] - The quality gain of a single picture, which is difference between the quality of a single picture relative to the quality of a gain reference picture; or
[0356] - The mean quality gain of all the pictures corresponding to a CLVS.
[0357] Use of this SEI message requires the definition of the following variables:
[0358] - A chroma format indicator, denoted herein by ChromaFormatldc, as described in clause7.3 in H.274: VSEI, available from [https: / / www.itu.int / rec / T-REC-H.274-202309-I / en (last accessed on April 5, 2024)].
[0359] - A count of pictures NumPics;
[0360] - Lists of picture widths and heights in units of luma samples, denoted herein byPicWidth[ i ] and PicHeight[ i ], respectively, where i is in the range of 0 to NumPics - 1, inclusive;
[0361] - A list of pictures TestPicList[ i ] where i is in the range of 0 to NumPics - 1, inclusive; and
[0362] - When any qm_gain_flag[ i ] is equal to 1, a list of gain reference picturesGainRefPicList[ i ] for i is in the range of 0 to NumPics - 1, inclusive.
[0363] The variables SubWidthC and SubHeightC are derived from ChromaFormatldc as specified by Table 2. in H.274: VSEI, available from [https: / / www.itu.int / rec / T-REC-H.274-202309- I / en (last accessed on April 5, 2024)].
[0364] TestPicList[ i ][ cldx ] and ReferencePicList[ i ][ cldx ] denote the cldx-th sample array of the i-th picture among TestPicList and ReferencePicList, respectively.
[0365] TestPicList[ i ] [ cldx ] [ x ] [ y ] and ReferencePicList[ i ] [ cldx ] [ x ] [ y ] denote the sample at location ( x, y ) within the cldx-th sample array of the i-th picture among TestPicList andReferencePicList, respectively, where x is in the range of 0 to ( ( cldx = = 0 ) ? PicWidth[ i ] : PicWidth[ i ] / SubWidthC ) - 1, inclusive and y is in the range of 0 to ( ( cldx = = 0 ) ? PicHeight[ i ] : PicHeight[ i ] / SubHeightC ) - 1, inclusive.
[0366] Let currPicIdx have a value such that the output time of TestPicList[ currPicIdx ] is equal to the output time of the current picture.
[0367] qm_metric_definitions_present_flag equal to 1 specifies that information defining the quality metrics is present. qm_metric_definitions_present_flag equal to 0 specifies that information defining the quality metrics is not present.
[0368] When this SEI message is the first quality metrics SEI message in a CLVS, in decoding order, qm_metric_definitions_present_flag shall be equal to 1.
[0369] Otherwise, (this SEI message is not the first quality metrics SEI message in a CLVS, in decoding order,) it is a requirement of bitstream conformance that at least one of the two following conditions shall be met:
[0370] - qm_metric_definitions_present_flag is equal to 0; or
[0371] - the values of the qm_metric_type[i], qm_three_component_flag[ i ], qm_gain_flag[ i], qm_gain_reference_flag[ i ], qm_metric_increasing_flag[ i ], qm_full_reference_flag[ i ], qm_value_len_minusl_in_bytes[ i ], qm_metric_description_present_flag[ i ], and qm_metric_description[ i ] syntax elements, when present, shall be equal to the respective syntax elements in the first quality metric SEI message in the CLVS.
[0372] qm_clvs_values_present_flag equal to 1 specifies that qm_clvs_metric_value[ i ][ c ] syntax elements are present. qm_clvs_values_present_flag equal to 0 specifies qm_clvs_metric_value[ i ] [ c ] syntax elements are not present.
[0373] qm_pic_values_present_flag equal to 1 specifies that qm_pic_metric_value[ i ] [ c ] syntax elements are present. qm_pic_values_present_flag equal to 0 specifies that qm_pic_metric_value[ i ][ c ] syntax elements are not present.
[0374] qm_num_metrics_minusl plus 1 specifies the number of quality metric entries signaled.
[0375] qm_gain_enabled_flag equal to 1 specifies that the qm_gain_flag[ i ] syntax element is present. qm_gain_enabled_flag equal to 0 specifies that the qm_gain_flag[ i ] syntax element is not present.
[0376] qm_gain_flag[ i ] and qm_gain_reference_flag[ i ], when present, indicate the interpretation of the values of the qm_clvs_metric_value[ i ][ c ] and qm_pic_me tric_ value [ i ][ c ] syntax elements in specified below in section ‘Process for derivation of picMetricValue[ i ][ c ]’. When qm_gain_flag[ i ] is not present, it is inferred to be equal to 0.
[0377] qm_metric_type[ i ] specifies the type of the quality metric of the i-th entry, as specified in Table 1. The value of qm_metric_type[ i ] shall be in the range of 0..8, inclusive, or in the range of 128..255, inclusive, in bitstreams conforming to this version of the Specification. Values in the range of of 9..127, inclusive, for qm_metric_type[ i ] are reserved for future use by ITU-T I ISO / IEC and shall not be present in bitstreams conforming to this version of this Specification.
[0378] When the value of qm_metric_type [ i ] is in the range of 9 .. 127, decoders conforming to this version of this Specification shall ignore all the syntax elements for the i-th entry in this syntax structure.
[0379] When the value of qm_metric_type[ i ] is in the range of 128 .. 255, the quality metric type is unspecified or specified by other means not specified in this Specification.
[0380] Following Tablel, describes an example interpretation of qm_metric_type[ i ]Table 1
[0381] qm_three_component_flag[ i ] equal to 1 indicates that 3 component values are present for the i-th metric. qm_three_component_flag[ i ] equal to 0 indicates that a single value is present for the i-th metric. It is a requirement of bitstream conformance that when ChromaFormatldc is equal to 0, qm_three_component_flag[ i ] shall be equal to 0.
[0382] qm_metric_increasing_flag[ i ] equal to 1 indicates that a higher value of the i-th metric value represents an improvement in quality. qm_metric_increasing_flag[ i ] equal to 0 indicates that a lower value of the i-th metric value represents an improvement in quality. When not present, the value of qm_metric_increasing_flag[ i ] is inferred to be equal to IncreasingFlag[qm_metric_type[ i ]] in Table 1.
[0383] qm_full_reference_flag[ i ] equal to 1 indicates that the quality metric is a full reference quality metric, calculated by comparison of the pictures in TestPicList with the respective quality reference pictures. qm_full_reference_flag[ i ] equal to 0 indicates that the quality metric may or may not include comparison of the pictures in TestPicList with the respective quality reference picture. When not present, the value of qm_full_reference_flag[ i ] is inferred to be equal to FullReferenceFlag[qm_metric_type[ i ]] in Table 1.
[0384] qm_value_len_minusl_in_bytes[ i ] plus 1 specifies the length in bytes of the qm_pic_metric_value[ i ][ c ] syntax element. When not present, the qm_value_len_minusl_in_bytes[ i ] is inferred to be equal to NumBytes[ [qm_metric_type[ i ] ] - 1.
[0385] qm_metric_description_present_flag[ i ] equal to 1 specifies that qm_metric_description[ i ] is present. qm_metric_description_present_flag[ i ] equal to 0 specifies that qm_metric_description[ i ] is not present.
[0386] qm_bit_equal_to_zero shall be equal to 0.
[0387] qm_metric_description[ i ] specifies a text description of the i-th quality metric. The length of the syntax element shall be less than or equal to 4097 bytes, not including the null termination byte.
[0388] qm_clvs_metric_value[ i ][ c ] specifies the mean value of the i-th quality metric for the c-th component of the CLVS. The length of the syntax element is 8 * (qm_value_len_minusl_in_bytes[ i ] + 1) bits.
[0389] qm_pic_metric_value[ i ][ c ] specifies the value of the i-th quality metric for the c-th component of the current picture. The length of the syntax element is 8 * (qm_value_len_minusl_in_bytes[ i ] + 1) bits.
[0390] The semantics of the quality metric value are determined by the values of qm_metric_type[ i ], qm_gain_flag[ i ], and qm_gain_reference_flag[ i ].
[0391] When qm_pic_values_present_flag is equal to 1, qm_pic_metric_value[ i ][ c ] indicates the picture metric value, picMetric Value [ i ][ c ], of type qm_metric_type[ i ] as described in Table 1, and the following applies:
[0392] - When qm_gain_flag[ i ] is equal to 0, qm_pic_metric_value[ i ][ c ] has value derived by the process specified below in section ‘Process for derivation of picMetricValue[ i ][ c ]’ with testPic, picWidth, and picHeight assigned to be TestPicList[ currPicIdx ], PicWidth[ currPicIdx ], and PicHeight[ currPicIdx ], respectively.
[0393] - Otherwise (qm_gain_flag[ i ] is equal to 1), the following applies:
[0394] - picMetricValueTest[ i ][ c ] is set equal to picMetric Value [ i ][ c ] derived by the process specified below in section ‘Process for derivation of picMetricValue[ i ][ c ]’ with testPic, picWidth, and picHeight assigned to be TestPicList[ currPicIdx ], PicWidth[ currPicIdx ], and PicHeight[ currPicIdx ], respectively.
[0395] - picMetric ValueGainRef[ i ][ c ] is set equal to picM etric Value [ i ][ c ] derived by the process specified below in section ‘Process for derivation of picMetricValue[ i ][ c ]’ with testPic,picWidth, and picHeight assigned to be GainRefPicList[ currPieldx ], PicWidth[ currPieldx ], and PicHeight[ currPieldx ], respectively.
[0396] - qm_pic_metric_value[ i ][ c ] has the value picMetricValueTest[ i ][ c ] - picMetricV alueGainRef [ i ] [ c ] .
[0397] When qm_clvs_values_present_flag is equal to 1, qm_clvs_metric_value[ i ] [ c ] indicates the mean value of the picture metric values, listPicMetric Value [ j ][ i ][ c ], calculated over all pictures in TestPicList, of type qm_metric_type[ i ] as described in Table 1, where each listPicMetric Value[j ][ i ][ c ] is derived as follows for each value of j in the range of 0 to NumPics - 1, inclusive:
[0398] - When qm_gain_flag[ i ] is equal to 0, listPicMetricValue[ j ][ i ][ c ] is equal to the picMetrictValue[ i ][ c ] derived by the process specified below in section ‘Process for derivation of picMetricV alue[ i ] [ c ] ’ is performed with testPic, picWidth, and picHeight assigned to be TestPicList[ j ], PicWidth[ j ], and PicHeight[ j ], respectively.
[0399] - Otherwise (qm_gain_flag[ i ] is equal to 1), the following applies:
[0400] - listPicMetricValueTest[ j ][ i ][ c ] is the set equal to picMetricValue[ i ][ c ] derived by the process specified below in section ‘Process for derivation of picMetric Value [ i ][ c ]’ with testPic, picWidth, and picHeight assigned to be TestPicList[ currPieldx ], PicWidth[ currPieldx ], and PicHeight[ currPieldx ], respectively.
[0401] - picMetric ValueGainRef[ j ][ i ][ c ] is the set qual to picMetric Value [ i ][ c ] derived by the process specified below in section ‘Process for derivation of picMetricValue[ i ][ c ]’ with testPic, picWidth, and picHeight assigned to be GainRefPicList[ currPieldx ], PicWidth[ currPieldx ], and PicHeight[ currPieldx ], respectively.
[0402] - qm_pic_metric_value[ j ][ i ][ c ] has the value picMetricValueTest[ j ][ i ][ c ] - picMetricV alueGainRef [j ][ i ][ c ].
[0403] Process for derivation of picMetricValuej i ][ c ]
[0404] Inputs to this process are a tested picture testPic, a picture width picWidth in luma samples, and a picture height picHeight in luma samples.
[0405] Let a quality reference picture referencePic be the original picture that was given as input to the encoding system and has the output time equal to the output time of testPic.
[0406] testPic[ cldx ] and referencePic[ cldx ] denote the cldx-th sample array of the testPic and referencePic, respectively.
[0407] testPic[ cldx ][ x ][ y ] and referencePic[ cldx ][ x ][ y ] denote the sample at location ( x, y ) within the cldx-th sample array of testPic and referencePic, respectively.
[0408] The pictures quality metric, picMetricValue[ i ][ c ] is derived as follows:
[0409] When qm_metric_type[ i ] is equal to 0,
[0410] When qm_metric_increasing_flag[ i ] is equal 1, a higher value of picMetricValue[ i ][ c ] indicates testPic is of better quality than for a picture with a lower value of picMetricValue[ i ] [ c ].
[0411] When qm_full_reference_flag[ i ] is equal to 1, of picMetric Value[ i ][ c ] indicates a quality metric value calculated from a comparison of testPic with referencePic.
[0412] qm_metric_description[ i ] provides a text description of the quality metric indicated by picMetricValue[ i ][ c ].
[0413] Additional interpretation of picMetric Value [ i ][ c ] is determined by external means not specified in this Specification.
[0414] When qm_metric_type[ i ] is equal to 1,
[0415] picMetric Value [ i ]
[0000] is set equal to the PSNR value calculated using clauses 9.4.2 and D.2 of ISO / IEC 23001-11 [2] for the luma components of testPic and referencePic, with bit depth OrigBitDepth, width picWidth, and height picHeight.
[0416] When qm_three_component_flag[ i ] is equal to 1, picMetricValue[ i ]
[0001] and picMetricValue[ i ]
[0002] are set equal to the PSNR values calculated using clauses 9.4.2 and D.2 of ISO / IEC 23001-11 [2] for the Cb and Crcomponents, respectively, of testPic and referencePic, with bit depth OrigBitDepth width picWidth / SubWidthC, and height pic Heigh t / SubHeightC.
[0417] When qm_metric_type[ i ] is equal to 2, picMetricValue[ i ]
[0000] is set equal to the value of the variable psnrYUV calculated as follows:The variable psnrY is set equal to the PSNR value calculated using clauses 9.4.2 and D.2 of ISO / IEC 23001-11 [2] for the luma components of testPic and referencePic, with bit depth OrigBitDepth width picWidth, and height picHeight.The variables psnrU and psnrV are set equal to the PSNR values calculated using clauses 9.4.2 and D.2 of ISO / IEC 23001-11 [2] for the Cb and Cr components, respectively, of testPic and referencePic, with bit depth OrigBitDepth, width picWidth / SubWidthC, and height picHeight / SubHeightC.- psnrYUV = (10*psnrY + psnrU + psnrV ) -? 12
[0418] When qm_metric_type[ i ] is equal to 3, picMetricValue[ i ]
[0000] is set equal to the value of SSIM calculated from clauses 9.4.2 and D.5 of ISO / IEC 23001-11 [2] for the luma components of testPic and referencePic, with bit depth OrigBitDepth width picWidth, and height picHeight.When qm_three_component_flag[ i ] is equal to 1,- picMetric Value [ i ]
[0001] andpicMetricValue[i ]
[0002] are set equal to the SSIM values calculated from clauses 9.4.2 and D.5 of ISO / IEC 23001-11 [2] for the Cb and Cr components, respectively, of testPic and referencePic, with bit depth OrigBitDepth, width picWidth / SubWidthC, and height picHeight / SubHeightC.
[0419] When qm_metric_type[ i ] is equal to 4,
[0420] picMetric Value [ i ]
[0000] is set equal to the value of MS-SSIM calculated from clause 4.3.3 of ISO / IEC 23001-10 [1] for the luma components of testPic and referencePic, with bit depth OrigBitDepth.
[0421] When qm_three_component_flag[ i ] is equal to 1,
[0422] picMetric Value [ i ]
[0001] and picMetric Value [ i ]
[0002] are set equal to the MS-SSIM values calculated from clause 4.3.3 of ISO / IEC 23001-10 [1] for the Cb and Cr components, respectively, of testPic and referencePic, with bit depth OrigBitDepth.
[0423] When qm_metric_type[ i ] is equal to 5,
[0424] picMetric Value [ i ]
[0000] is set equal to the value of MOS specified in clause 4.3.6 ofISO / IEC 23001-10 [1].
[0425] When qm_metric_type[ i ] is equal to 6,
[0426] picMetric Value [ i ]
[0000] is set equal to the value of wPSNR calculated from clauses 9.4.2 and D.3 of ISO / IEC 23001-11 [2] for the luma components of testPic and referencePic, with bit depth OrigBitDepth width picWidth, and height picHeight.
[0427] When qm_three_component_flag[ i ] is equal to 1,
[0428] picMetric Value [ i ]
[0001] and picMetricValue[ i ]
[0002] are set equal to the wPSNR values calculated from clauses 9.4.2 and D.3 of ISO / IEC 23001-11 [2] for the Cb and Cr components, respectively, of testPic and referencePic, with bit depth OrigBitDepth, width picWidth / SubWidthC, and height picHeight / SubHeightC.
[0429] When qm_metric_type[ i ] is equal to 7,
[0430] picMetric Value [ i ]
[0000] is set equal to the value of WS-PSNR calculated from clauses 9.4.2 and D.4 of ISO / IEC 23001-11 [2] for the luma components of testPic and referencePic, with bit depth OrigBitDepth, width picWidth, and height picHeight.
[0431] When qm_three_component_flag[ i ] is equal to 1,
[0432] picMetric Value [ i ]
[0001] and picMetric Value[ i ]
[0002] are set equal to the WS-PSNR values calculated from clauses 9.4.2 and D.4 of ISO / IEC 23001-11 [2]for the Cb and Cr components, respectively, of testPic and referencePic, with bit depth OrigBitDepth, width picWidth / SubWidthC, and height picHeight / SubHeightC.
[0433] When qm_metric_type[ i ] is equal to 8,
[0434] picMetric Value [ i ]
[0000] is set equal to the value of lumaMse, interpreted as a floatingpoint value, derived as follows:
[0435] lumaSse = 0
[0436] for( y = 0; y < picHeight; y++)
[0437] for( x = 0; x < picWidth; X++ )
[0438] lumaSse += (testPic
[0000] [ x ][ y ] - referencePic
[0000] [ x ][ y ])2
[0439] lumaMse = ( lumaSse -? (CroppedHeight * CroppedWidth) ) / 100
[0440] When qm_three_component_flag[ i ] is equal to 1, picMetric Value[ i ]
[0001] and picMetricValue[ i ]
[0002] are set equal to the values of CbMse and CrMse, respectively, interpreted as floating-point values, derived as follows:
[0441] CbSse = 0
[0442] CrSse = 0
[0443] for( y = 0; y < picHeight / SubWidthC; y++)
[0444] for( x = 0; x < picWidth / SubWidthC; x++ ) {
[0445] CbSse += (testPic[l][y][x] - referencePic[l][y][x])2
[0446] CbSse += (testPic[l][y][x] - referencePic[l][y][x])2
[0447] }
[0448] CbMse = ( CbSse -? ( picHeight * CroppedWidth / (SubWidthC * SubWidthC) ) / 100
[0449] CrMse = ( CrSse -? ( picHeight * CroppedWidth / (SubWidthC * SubWidthC) ) / 100
[0450] Use of the quality metric SEI message
[0451] For purposes of interpretation of the quality metric SEI message, the derivation of the variables ChromaFormatldc, NumPics, TestPicList, PicWidth, PicHeight, and GainRefPicList is specified below.
[0452] Let TestPicList initially includes the cropped decoded pictures of the current CLVS in output order.
[0453] When the quality metric SEI message is included as a type of an SEI message in an SEI processing order SEI message, any quality metric SEI message that is associated with the SEI processing order SEI message shall be included in a processing order nesting SEI message.
[0454] When the quality metric SEI message is included as the i-th type of an SEI message in an SEI processing order SEI message, the following applies:
[0455] - It is a requirement of bitstream conformance that an SEI message seiB that implies post-processing to be performed shall be present as the j-th type of an SEI message in the same SEI processing order SEI message and po_sei_processing_order[ j ] shall be equal to po_sei_processing_order[ i ].
[0456] - The quality metric SEI message indicates the picture quality resulting from the postprocessing implied by seiB.
[0457] - TestPicList is updated as follows for each post-processing stage with po_sei_processing_order[ j ] less than or equal to po_sei_processing_order[ i ] in a non -decreasing order ofj.
[0458] - When the post-processing stage results into a picture picA with an output time that is equal to the output time of a picture picB in TestPicList, picB in TestPicList is replaced by picA.
[0459] - When the post-processing stage results into a picture picA with an output time that is not equal to the output time of any picture in TestPicList, picA is inserted in TestPicList in a manner that pictures in TestPicList remain in output order.
[0460] NumPics is set equal to the count of pictures in TestPicList.
[0461] PicWidth[ i ] and PicHeight[ i ] are set equal to the width and height of TestPicList[ i ], respectively, in luma samples.
[0462] When a quality metric SEI message is included as the k-th SEI message in a processing order nesting SEI message and qm_gain_flag[ i ] is equal to 1 for any value of i, the following applies:
[0463] - When qm_gain_reference_flag[ i ] is equal to 0, the i-th metric value in the quality metric SEI message represents a gain of the post-processing stage with po_sei_processing_order[ j ]equal to pon_processing_order[ k ] relative to the picture or pictures used as input to that postprocessing stage and GainRefPicList is set equal to TestPicList derived for the processing stages up to but not including the processing stage with po_sei_processing_order[ j ] .
[0464] - Otherwise (qm_gain_reference_flag[ i ] is equal to 1), the i-th metric value in the quality metric SEI message represents a cumulative gain of all the post-processing stages with po_sei_processing_order[ j ] less than or equal to pon_processing_order[ k ] relative to the cropped decoded picture or pictures and GainRefPicList consists of the cropped decoded pictures of the current CL VS in output order.
[0465] - It is a requirement of bitstream conformance that the count of pictures inGainRefPicList shall be equal to NumPics and the width and height of GainRefPicList[ i ] in luma samples shall be equal to PicWidth[ i ] and PicHeight[ i ], respectively.
[0466] The value of ChromaFormatldc is derived as follows:
[0467] - When the quality metric SEI message is not included in a processing order nesting SEI message, ChromaFormatldc is set equal to sps_chroma_format_idc.
[0468] - Otherwise (the quality metric SEI message is included in a processing order nestingSEI message), ChromaFormatldc is set to a value that matches the chroma format of the pictures in TestPicList and it is a requirement of bitstream conformance that the chroma format of all the pictures in TestPicList and GainRefPicList, when present, shall be identical.
[0469] When a quality metric SEI message qmSeiA is present in a picture unit that is not the first picture unit of a CLVS in decoding order and at least one value of qm_clvs_metric_value[ i ][ c ] is present, the following applies:
[0470] - When qmSeiA is not included in a processing order nesting SEI message, each value of qm_clvs_metric_value[ i ] [ c ] shall be equal to the value of qm_clvs_metric_value[ i ] [ c ] in the quality metric SEI message that is present in the first picture unit of the CLVS and is not is included in a processing order nesting SEI message.
[0471] - Otherwise (qmSeiA is included in a processing order nesting SEI message with a certain set of pon_target_po_id[ i ] values), the following applies:
[0472] - It is a requirement of bitstream conformance that there shall be quality metric SEI message qmSeiB that is present in the first picture unit of the CL VS and is included in a processing order nesting SEI message with the same set of pon_target_po_id[ i ] values.
[0473] - Each value of qm_clvs_metric_value[ i ] [ c ] in qmSeiA shall be equal to the value of qm_clvs_metric_value[ i ][ c ] in qmSeiB.
[0474] Additional or alternative embodiments on whether the reference is filtered
[0475] In an embodiment, a value of a quality metric may be computed by using, as a reference, a filtered picture, and where the filtered picture is encoded by an encoder. In one example, at encoder side, a picture is filtered by a Motion-Compensated Temporal Filter (MCTF), obtaining a MCTF- filtered picture; the MCTF-filtered picture is used as a reference for computing a value of a quality metric and it is encoded by an encoder.
[0476] In an alternative embodiment, a value of a quality metric may be computed by using, as a reference, a picture that has not been filtered, and where the picture is encoded by an encoder. In an example, at encoder side, a picture is used as reference for computing a value of a quality metric and it is encoded by an encoder.
[0477] In an alternative embodiment, a value of a quality metric may be computed by using, as a reference, a picture prior to any filtering, and where the picture is filtered to obtain a filtered picture, where the filtered picture is encoded by an encoder. In one example, at encoder side, a picture is used as a reference for computing a value of a quality metric, and the picture is filtered by a MCTF, obtaining a MCTF-filtered picture, where the MCTF-filtered picture is encoded by an encoder.
[0478] It is to be understood that, in the above embodiments and examples about whether a reference is filtered or not, the filtering is described as being performed outside of or prior to the encoding, for example as a pre-processing step that is not part of an encoder. However, in some encoders, a filtering (such as MCTF) may be considered to be part of an encoder, for example a preprocessing step that is part of an encoder.
[0479] In an embodiment, an encoder may signal, in or along a bitstream, or a decoder may receive, from or along a bitstream, information indicating one or more of the following:
[0480] - Whether a reference for computing a quality metric is a filtered picture.
[0481] - A type of filtering (for example, MCTF) that was applied to the reference.
[0482] - Any other information that describes a filtering that was applied to the reference.
[0483] Additional or alternative embodiments on temporal scope of a value of a quality metric
[0484] In an embodiment, a value of a quality metric may indicate an average (absolute or relative) quality for a set of pictures that comprises the picture to which an indication of the value is associated and all the following pictures (either in output / display order or in decoding order) in a CLVS.
[0485] In an embodiment, an encoder may signal, in or along a bitstream, or a decoder may receive, from or along a bitstream, information indicating whether a value of a quality metric refers to an average quality for all the pictures in a CLVS or refers to an average quality for a set of pictures that comprises the picture to which the information is associated and all the following pictures (either in output / display order or in decoding order) in a CLVS.
[0486] In an example, an SEI message carries information about quality. The SEI message is associated to a picture that is not the first picture of a CLVS. The SEI message comprises a flag that indicates whether a value for a certain quality metric is to be interpreted as an average value for the current picture and all subsequent pictures.
[0487] In an embodiment, when information indicating a value of a quality metric is associated with a picture that is not the first picture in a CVLS, the value is interpreted as an average quality for the picture to which the information is associated and all subsequent pictures in either output / display order or decoding order.
[0488] Additional or alternative embodiments on where the types of quality metric are carried
[0489] In one embodiment, a post-processing filter (PPF), such as a neural network postprocessing filter (NNPF) may be associated with one or more quality metrics, where the PPF or the NNPF improves those one or more quality metrics. An encoder may signal, in or along a bitstream, and a decoder may receive, from or along a bitstream, information indicating one or more values for the respective one or more quality metrics, where the information is associated with the PPF or the NNPF.
[0490] In one example, a NNPFC SEI message comprises indication of two quality metrics that the associated NNPF improves, such as a first quality metric being PSNR and a second quality metric being MS-SSIM. A quality metric SEI message is associated to the NNPFC SEI message or is associated to the NNPF specified in the NNPFC SEI message or anyway identifies the NNPF specified in the NNPFC SEI message, where the quality metric SEI message comprises two values, and where a first value represents a PSNR value and a second value represents a MS-SSIM value.
[0491] Additional or alternative embodiments on quality of more than one picture
[0492] In an embodiment, when one or more values of a quality metric represents a (absolute or relative) quality for two or more pictures, an encoder may signal, in or along a bitstream, or a decoder may receive, from or along a bitstream, information indicating how the one or more values were determined, such as by averaging the quality values of the two or more pictures, or by computing a median value of the quality values of the two or more pictures, or by computing the mode (e.g., the most frequent value), or by computing an average and a standard deviation of the quality values of the two or more pictures, and the like.
[0493] Accordingly, features of the embodiments described herein include: provide indication of increasing vs decreasing metric; provide indication of relative or absolute gain; usage with the SEI processing order SEI message and the processing order nesting SEI message to determine the tested picture and reference picture, indication of reference picture for relative gain, e.g., output picture or input to postprocessing stage; use a user defined metrics with text description; signal quality for 3 combined components for some metrics; signal metric definitions once per CLVS, while metric values are be signaled once per picture; provide indication that quality is for a full reference picture or not.
[0494] FIG. 5 is an example apparatus 500, which may be implemented in hardware, configured to implement the examples described herein. The apparatus 500 comprises at least one processor 502 (e.g., an FPGA and / or CPU), at least one memory 504 including computer program code 505, the computer program code 505 having instructions to carry out the methods described herein, wherein the at least one memory 504 and the computer program code 505 are configured to, with the at least one processor 502, cause the apparatus 500 to implement circuitry, a process, component, module, or function (implemented with control module 506) to implement the examples described herein, including implementing quality metrics in coded media, for example, videos or images. Optionally included encoder 508 of the control module 506 implements encoding based on the examples described herein, and optionally included decoder 510 implements decoding based on the examples described herein. The at least one memory 504 may be a non-transitory memory, a transitory memory, a volatile memory (e.g. RAM), or a non-volatile memory (e.g., ROM).
[0495] The apparatus 500 includes a display and / or I / O interface 512, which includes user interface (UI) circuitry and elements, that may be used to display features or a status of the methods described herein (e.g., as one of the methods is being performed or at a subsequent time), or to receive input from a user such as with using a keypad, camera, touchscreen, touch area, microphone, biometric recognition, one or more sensors, etc. The apparatus 500 includes one or more communication e.g. network (N / W) interfaces (I / F(s)) 514. The communication I / F(s) 514 may be wired and / or wireless and communicate over the Internet / other network(s) via any communication technique including via one or more links 516. The communication I / F(s) 514 may comprise one or more transmitters or one or more receivers.
[0496] The transceiver 518 comprises one or more transmitters 520 and one or more receivers 522. The transceiver 518 and / or communication I / F(s) 514 may comprise standard well-known components such as an amplifier, filter, frequency-converter, (de)modulator, and encoder / decoder circuitries and one or more antennas, such as antennas 524 used for communication over wireless link 526.
[0497] The control module 506 of the apparatus 500 comprises one of or both parts 506-1 and / or 506-2, which may be implemented in a number of ways. The control module 506 may be implemented in hardware as control module 506-1, such as being implemented as part of the at least one processor 502. The control module 506-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the control module 506 may be implemented as control module 506-2, which is implemented as computer program code (having corresponding instructions) 505 and is executed by the at least one processor 502. For instance, the at least one memory 504 store instructions that, when executed by the at least one processor 502, cause the apparatus 500 to perform one or more of the operations as described herein. Furthermore, the at least one processor 502, the at least one memory 504, and example algorithms (e.g., as flowcharts and / or signaling diagrams), encoded as instructions, programs, or code, are means for causing performance of the operations described herein.
[0498] The apparatus 500 to implement the functionality of control module 506 may correspond to any of the apparatuses depicted herein. Alternatively, apparatus 500 and its elements may not correspond to any of the other apparatuses depicted herein, as apparatus 500 may be part of a self- organizing / optimizing network (SON) node or other node, such as a node in a cloud.
[0499] The apparatus 500 may also be distributed throughout the network including within and between apparatus 500 and any network element (such as a base station and / or terminal device and / or user equipment).
[0500] Interface 528 enables data communication and signaling between the various items of apparatus 500, as shown in FIG. 5. For example, the interface 528 may be one or more buses such as address, data, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. Computer program code (e.g. instructions) 505, including control module 506 may comprise object-oriented software configured to pass data or messages between objects within computer program code 505. The apparatus 500 need not comprise each of the features mentioned, or may comprise other features as well. The various components of apparatus 500 may at least partially reside in a housing 530, or a subset of the various components of apparatus 500 may at least partially be located in different housings, which different housings may include housing 530.
[0501] FIG. 6 shows a schematic representation of non-volatile memory media 600a (e.g. computer / compact disc (CD) or digital versatile disc (DVD)) and 600b (e.g. universal serial bus (USB) memory stick) and 600c (e.g. cloud storage for downloading instructions and / or parameters 602 or receiving emailed instructions and / or parameters 602) storing instructions and / or parameters 602 which when executed by a processor allows the processor to perform one or more of the operations of the methods described herein. Instructions and / or parameters 602 may represent or correspond to a non- transitory computer readable medium.
[0502] FIG. 7 is an example method 700 performed with an encoder, based on the example embodiments described herein. At 702, the method 700 includes determining a quality metrics per video sequence and / or picture. At 704, the method 700 includes signaling the quality metrics as an absolute or relative gains for post-processing filters or post-processing filter groups, relative to the cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group.
[0503] The method 700 may be performed with an encoding apparatus, such as apparatus 100, apparatuses depicted in FIG. 3 and FIG. 4, transmitting apparatus 480 with an encoder 430, or an apparatus 400 with the encoder 430.
[0504] FIG. 8 is an example method 800 performed with a decoder, based on the example embodiments described herein. At 802, the method 800 includes receiving a quality metrics as an absolute or as relative gains for post-processing filters or post-processing filter groups, relative to the cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group, wherein the quality metrics is determined per video sequence and / or picture. At 804, the method 800 includes performing decoder side operations based on the quality metrics.
[0505] Some examples of the decoder side operations include, but are not limited to, performing a high complexity post-processing operation, using the quality metrics for determining or selecting a postprocessing filter or a post-processing filter group, selecting highest quality frames of the video sequence to operate on, or determining not to display pictures with quality levels below a target quality level. Some examples of the high complexity post-processing operation include, but are not limited to, face recognition, face detection, image segmentation, and the like.
[0506] The method 800 may be performed with an encoding apparatus, such as apparatus 100, apparatuses depicted in FIG. 3 and FIG. 4, receiving apparatus 482 with a decoder 440, or an apparatus 400 with the decoder 440.
[0507] As described above, FIGs. 7 and 8 include flowcharts of an apparatus (e.g. 100, 500, or any other apparatuses described herein), method, and computer program product according to certain example embodiments. It will be understood that each block of the flowcharts, and combinations of blocks in the flowcharts, may be implemented by various means, such as hardware, firmware, processor, circuitry, and / or other devices associated with execution of software including one or more computer program instructions. For example, one or more of the procedures described above may be embodied by computer program instructions. In this regard, the computer program instructions which embody the procedures described above may be stored by a memory (e.g. 58, 125, or 504) of an apparatus employing an embodiment of the present invention and executed by processing circuitry (e.g., 56, 120, or 502) of the apparatus. As will be appreciated, any such computer program instructions may be loaded onto a computer or other programmable apparatus (e.g., hardware) to produce a machine, such that the resulting computer or other programmable apparatus implements the functions specified in the flowchart blocks. These computer program instructions may also be stored in a computer-readable memory that may direct a computer or other programmable apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture, the execution of which implements the function specified in the flowchart blocks. The computer program instructions may also be loaded onto a computer or other programmable apparatus to cause a series of operations to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide operations for implementing the functions specified in the flowchart blocks.
[0508] A computer program product is therefore defined in those instances in which the computer program instructions, such as computer-readable program code portions, are stored by at least one non- transitory computer-readable storage medium with the computer program instructions, such as thecomputer-readable program code portions, being configured, upon execution, to perform the functions described above, such as in conjunction with the flowchart(s) of FIGs. 7 and 8. In other embodiments, the computer program instructions, such as the computer-readable program code portions, need not be stored or otherwise embodied by a non-transitory computer-readable storage medium, but may, instead, be embodied by a transitory medium with the computer program instructions, such as the computer- readable program code portions, still being configured, upon execution, to perform the functions described above.
[0509] Accordingly, blocks of the flowcharts support combinations of means for performing the specified functions and combinations of operations for performing the specified functions for performing the specified functions. It will also be understood that one or more blocks of the flowcharts, and combinations of blocks in the flowcharts, may be implemented by special purpose hardware-based computer systems which perform the specified functions, or combinations of special purpose hardware and computer instructions.
[0510] In some embodiments, certain ones of the operations above may be modified or further amplified. Furthermore, in some embodiments, additional optional operations may be included. Modifications, additions, or amplifications to the operations above may be performed in any order and in any combination.
[0511] Some embodiments have been described in relation to one or more neural networks performing visual temporal extrapolation. It is to be understood that embodiments can be realized with any generative modelling neural networks.
[0512] In the above, some example embodiments have been described with the help of syntax of the bitstream. It needs to be understood, however, that the corresponding structure and / or computer program may reside at the encoder for generating the bitstream and / or at the decoder for decoding the bitstream.
[0513] In the above, where example embodiments have been described with reference to an encoder, it needs to be understood that the resulting bitstream and the decoder have corresponding elements in them. Likewise, where example embodiments have been described with reference to a decoder, it needs to be understood that the encoder has structure and / or computer program for generating the bitstream to be decoded by the decoder.
[0514] Many modifications and other embodiments of the inventions set forth herein will come tomind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the inventions are not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims. Moreover, although the foregoing descriptions and the associated drawings describe example embodiments in the context of certain example combinations of elements and / or functions, it should be appreciated that different combinations of elements and / or functions may be provided by alternative embodiments without departing from the scope of the appended claims. In this regard, for example, different combinations of elements and / or functions than those explicitly described above are also contemplated as may be set forth in some of the appended claims. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
[0515] It should be understood that the foregoing description is only illustrative. Various alternatives and modifications may be devised by those skilled in the art. For example, features recited in the various dependent claims could be combined with each other in any suitable combination(s). In addition, features from different embodiments described above could be selectively combined into a new embodiment. Accordingly, the description is intended to embrace all such alternatives, modifications and variances which fall within the scope of the appended claims.
[0516] References to a ‘computer’, ‘processor’, etc. should be understood to encompass not only computers having different architectures such as single / multi-processor architectures and sequential (Von Neumann) / parallel architectures but also specialized circuits such as field-programmable gate arrays (FPGA), application specific circuits (ASIC), signal processing devices and other processing circuitry. References to computer program, instructions, code etc. should be understood to encompass software for a programmable processor or firmware such as, for example, the programmable content of a hardware device such as instructions for a processor, or configuration settings for a fixed-function device, gate array or programmable logic device, and the like.
[0517] As used herein, the term ‘circuitry’ may refer to any of the following: (a) hardware circuit implementations, such as implementations in analog and / or digital circuitry, and (b) combinations of circuits and software (and / or firmware), such as (as applicable): (i) a combination of processor(s) or (ii) portions of processor(s) / software including digital signal processor(s), software, and memory(ies) that work together to cause an apparatus to perform various functions, and (c) circuits, such as a microprocessor(s) or a portion of a microprocessor(s), that require software or firmware for operation,even when the software or firmware is not physically present. This description of ‘circuitry’ applies to uses of this term in this application. As a further example, as used herein, the term ‘circuitry’ would also cover an implementation of merely a processor (or multiple processors) or a portion of a processor and its (or their) accompanying software and / or firmware. The term ‘circuitry’ would also cover, for example and when applicable to the particular element, a baseband integrated circuit or applications processor integrated circuit for a mobile phone or a similar integrated circuit in a server, a cellular network device, or another network device.
[0518] Circuitry or Circuit: As used in this application, the term ‘circuitry’ or ‘circuit’ may refer to one or more or all of the following:(a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); and(b) combinations of hardware circuits and software, such as (as applicable):(i) a combination of analog and / or digital hardware circuit(s) with software / firmware; and(ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and(c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.
[0519] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example, and when applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
Claims
CLAIMSWhat is claimed is:
1. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: determining a quality metrics per video sequence and / or picture; and signaling the quality metrics as an absolute or relative gains for post-processing filters or post-processing filter groups, relative to a cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group.
2. The apparatus of claim 1, wherein the quality metrics comprises user-specified quality metrics comprising one or more following: a text description of the metric; a first flag to indicate whether higher or lower values of the metrics correspond to increased quality; a second flag to indicate whether the quality metrics is a full reference metric; or signaling of number of bytes signaled for the metrics values.
3. The apparatus of any of the claims 1 or 2, wherein the quality metrics is single color component, individually for 3 color components, or combined for the 3 color components.
4. The apparatus of any of the previous claims, wherein the quality metrics is comprised in an information message, and wherein the information message comprises one or more of following: a quality metrics definition flag for indicating whether the information message comprises definitions of the quality metrics; a sequence level values flag for indicating whether a sequence level metric values are signaled and allowing the apparatus to choose whether to signal sequence level metrics values, picture level metric values, or both; a number of quality metrics syntax element for enabling the apparatus to send multiple metric values within the information message; a quality gain flag for enabling the apparatus to indicate whether the signaled metric values indicate a gain associated with the post-processing filter, or whether the values indicateabsolute quality of the picture or the video sequence; reference flag for indicating whether a reference for gain is an input picture to the apparatus or an input picture to a post-processing stage; quality metric type syntax element for indicating a metric type of each quality metric; an increasing flag for indicating whether a higher value of a metric corresponds to a higher quality value or a lower quality value; a full reference flag for indicating wherein the quality metric is a full reference metric; a length syntax element for specifying a number of bytes used to signal the quality metrics; description syntax element for signaling a text description of the quality metrics; or description present flag for indicating that the text description of the quality metrics is signaled using the description syntax element.
5. The apparatus of any of the previous claims, wherein the quality metrics comprises values indicating one or more of following values: quality of the picture; a mean quality of pictures corresponding to the video sequence; a quality gain of the picture, which is a difference between the quality of the picture relative to a quality of a gain reference picture; or a mean quality gain of the pictures corresponding to the video sequence.
6. The apparatus of claim 5, wherein the apparatus is further caused to define one or more of the following variables for the information message: a chroma format indicator; a count of pictures; lists of picture widths and heights; a list of pictures; or a list of gain reference pictures.
7. The apparatus of any of the previous claims, wherein values the quality metrics are computed by using, as a reference, a filtered picture, and wherein the filtered picture is encoded by the apparatus.
8. The apparatus of any of the claims 1 to 6, wherein values the quality metrics are computed by using, as a reference picture, a picture that has not been filtered, and wherein the picture is encoded by the apparatus.
49. The apparatus of any of the claims 1 to 6, wherein values the quality metrics are computed by using, as a reference picture, a picture prior to any filtering, and wherein the picture is filtered to obtain a filtered picture, and where the filtered picture is encoded by the apparatus.
10. The apparatus of any of the claims 7 to 9, wherein the apparatus is further caused to perform: signaling, in or along a bitstream information indicating one or more of the following: whether the reference picture used for computing the quality metrics is the filtered picture; a type of filtering used for filtering the reference picture; or information that describes a filtering that is applied to the reference picture.
11. The apparatus of any of the previous claims, wherein a value of the quality metrics indicates an average, absolute, or relative quality, for a set of pictures that comprises the picture to which an indication of the value is associated and following pictures, in output order or in decoding order in the video sequence.
12. The apparatus of claim 11, wherein the apparatus is further caused to perform: signaling, in or along a bitstream, information indicating whether the value of the quality metric refers to an average quality for the pictures in the video sequence or refers to an average quality for a set of pictures that comprises the picture to which the information is associated and the following pictures in output order or in decoding order in the video sequence.
13. The apparatus of any of the claims 11 or 12, wherein information about quality is comprised in an information message, and wherein the information message further comprises a flag for indicating whether the value for the quality metrics is to be interpreted as an average value for current picture and subsequent pictures.
14. The apparatus of claim 13, wherein, when the information indicating the value of the quality metrics is associated with a picture that is not a first picture in the video sequence, the value is interpreted as an average quality for the picture to which the information is associated and the subsequent pictures in either output order or decoding order.
15. The apparatus of claim 11, wherein, when one or more values of the quality metrics represents the absolute or relative quality for two or more pictures, the apparatus is furthercaused to perform: signaling, in or along a bitstream, information indicating how the one or more values were determined.
16. The apparatus of any of the previous claims, wherein the video sequence comprises a coded layer video sequence.
17. The apparatus of any of the previous claims, wherein the information message comprises a supplemental enhancement information message.
18. An apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving a quality metrics as an absolute or as relative gains for post-processing filters or post-processing filter groups, relative to a cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group, wherein the quality metrics is determined per video sequence and / or picture; and performing decoder side operations based on the quality metrics.
19. The apparatus of claim 18, wherein the quality metrics comprises user-specified quality metrics comprising one or more following: a text description of the metric; a first flag to indicate whether higher or lower values of the metrics correspond to increased quality; a second flag to indicate whether the quality metrics is a full reference metric; or signaling of number of bytes signaled for the metrics values.
20. The apparatus of any of the claims 18 or 19, wherein the quality metrics is single color component, individually for 3 color components, or combined for the 3 color components.
21. The apparatus of any of the previous claims 18 to 20, wherein the quality metrics is comprised in an information message, and wherein the information message comprises one or more of following: a quality metrics definition flag for indicating whether the information message comprises definitions of the quality metrics; a sequence level values flag for indicating whether a sequence level metric values aresignaled and allowing the apparatus to choose whether to signal sequence level metrics values, picture level metric values, or both; a number of quality metrics syntax element for enabling the apparatus to send multiple metric values within the information message; a quality gain flag for enabling the apparatus to indicate whether the signaled metric values indicate a gain associated with the post-processing filter, or whether the values indicate absolute quality of the picture or the video sequence; reference flag for indicating whether a reference for gain is an input picture to the apparatus or an input picture to a post-processing stage; quality metric type syntax element for indicating a metric type of each quality metric; an increasing flag for indicating whether a higher value of a metric corresponds to a higher quality value or a lower quality value; a full reference flag for indicating wherein the quality metric is a full reference metric; a length syntax element for specifying a number of bytes used to signal the quality metrics; description syntax element for signaling a text description of the quality metrics; or description present flag for indicating that the text description of the quality metrics is signaled using the description syntax element.
22. The apparatus of any of the previous claims 18 to 21, wherein the quality metrics comprises values indicating one or more of following: quality of the picture; a mean quality of pictures corresponding to the video sequence; a quality gain of the picture, which is a difference between the quality of the picture relative to a quality of a gain reference picture; or a mean quality gain of the pictures corresponding to the video sequence.
23. The apparatus of claim 22, wherein the information message comprises one or more of following variables: a chroma format indicator; a count of pictures, lists of picture widths and heights; a list of pictures; or a list of gain reference pictures.
24. The apparatus of any of the previous claims 18 to 23, wherein values the quality metrics$7are computed by using, as a reference, a filtered picture, and wherein the filtered picture is encoded by an encoder.
25. The apparatus of any of the claims 18 to 23, wherein values the quality metrics are computed by using, as a reference picture, a picture that has not been filtered, and wherein the picture is encoded by an encoder.
26. The apparatus of any of the claims 18 to 23, wherein values the quality metrics are computed by using, as a reference picture, a picture prior to any filtering, and wherein the picture is filtered to obtain a filtered picture, and where the filtered picture is encoded by an encoder.
27. The apparatus of any of the claims 24 to 26, wherein the apparatus is further caused to perform: receiving, from or along a bitstream information indicating one or more of the following: whether the reference picture used for computing the quality metrics is a filtered picture; a type of filtering used for filtering the reference picture; or information that describes a filtering that is applied to the reference picture.
28. The apparatus of any of the previous claims 18 to 27, wherein a value of the quality metrics indicates an average, absolute, or relative quality, for a set of pictures that comprises the picture to which an indication of the value is associated and following pictures, in output order or in decoding order in the video sequence.
29. The apparatus of claim 28, wherein the apparatus is further caused to perform: receiving, from or along a bitstream, information indicating whether the value of the quality metric refers to an average quality for the pictures in the video sequence or refers to an average quality for a set of pictures that comprises the picture to which the information is associated and the following pictures in output order or in decoding order in the video sequence.
30. The apparatus of any of the claims 28 or 29, wherein the information about quality is comprised in an information message, and wherein the information message further comprises a flag for indicating whether the value for the quality metrics is to be interpreted as an average value for a current picture and subsequent pictures.
831. The apparatus of claim 30, wherein, when the information indicating the value of the quality metrics is associated with a picture that is not a first picture in the video sequence, the value is interpreted as an average quality for the picture to which the information is associated and the subsequent pictures in either output order or decoding order.
32. The apparatus of claim 28, wherein, when one or more values of the quality metrics represents the absolute or relative quality for two or more pictures, the apparatus is further caused to perform: receiving, from or along a bitstream, information indicating how the one or more values were determined.
33. The apparatus of any of the previous claims 18 to 32, wherein the video sequence comprises a coded layer video sequence.
34. The apparatus of any of the previous claims 18 to 32, wherein the information message comprises a supplemental enhancement information message.
35. The apparatus of any of the previous claims 18 to 34, wherein the decoder side operations comprise one or more of the following; performing a high complexity post-processing operation; using the quality metrics for determining or selecting a post-processing filter or a postprocessing filter group; selecting highest quality frames of the video sequence to operate on; or determining not to display pictures with quality levels below a target quality level.
36. A method comprising: determining a quality metrics per video sequence and / or picture; and signaling the quality metrics as an absolute or relative gains for post-processing filters or post-processing filter groups, relative to the cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group.
37. The method of claim 36, wherein the quality metrics comprises user-specified quality metrics comprising one or more following: a text description of the metric; a first flag to indicate whether higher or lower values of the metrics correspond to increased quality; a second flag to indicate whether the quality metrics is a full reference metric; or19signaling of number of bytes signaled for the metrics values.
38. The method of any of the claims 36 or 37, wherein the quality metrics is single color component, individually for 3 color components, or combined for the 3 color components.
39. The method of any of the previous claims 36 to 38, wherein the quality metrics is comprised in an information message, and wherein the information message comprises one or more of following: a quality metrics definition flag for indicating whether the information message comprises definitions of the quality metrics; a sequence level values flag for indicating whether a sequence level metric values are signaled and allowing an encoder method to choose whether to signal sequence level metrics values, picture level metric values, or both; a number of quality metrics syntax element for enabling the encoder to send multiple metric values within the information message; a quality gain flag for enabling the encoder to indicate whether the signaled metric values indicate a gain associated with the post-processing filter, or whether the values indicate absolute quality of the picture or the video sequence; reference flag for indicating whether a reference for gain is an input picture to the encoder or an input picture to a post-processing stage; quality metric type syntax element for indicating a metric type of each quality metric; an increasing flag for indicating whether a higher value of a metric corresponds to a higher quality value or a lower quality value; a full reference flag for indicating wherein the quality metric is a full reference metric; a length syntax element for specifying a number of bytes used to signal the quality metrics; description syntax element for signaling a text description of the quality metrics; or description present flag for indicating that the text description of the quality metrics is signaled using the description syntax element.
40. The method of any of the previous claims 36 to 39, wherein the quality metrics comprises values indicating one or more of following values: quality of the picture; a mean quality of pictures corresponding to the video sequence; a quality gain of the picture, which is a difference between the quality of the picture relative to a quality of a gain reference picture; or10a mean quality gain of the pictures corresponding to the video sequence.
41. The method of claim 40 further comprising defining one or more of the following variables for the information message: a chroma format indicator; a count of pictures; lists of picture widths and heights; a list of pictures; or a list of gain reference pictures.
42. The method of any of the previous claims 36 to 41, wherein values the quality metrics are computed by using, as a reference, a filtered picture, and wherein the filtered picture is encoded by the encoder.
43. The method of any of the claims 36 to 41, wherein values the quality metrics are computed by using, as a reference picture, a picture that has not been filtered, and wherein the picture is encoded by the encoder.
44. method of any of the claims 36 to 41, wherein values the quality metrics are computed by using, as a reference picture, a picture prior to any filtering, and wherein the picture is filtered to obtain a filtered picture, and where the filtered picture is encoded by the encoder.
45. The method of any of the claims 42 to 44 further comprising: signaling, in or along a bitstream information indicating one or more of the following: whether the reference picture used for computing the quality metrics is a filtered picture; a type of filtering used for filtering the reference picture; or information that describes a filtering that is applied to the reference picture.
46. The method of any of the previous claims 36 to 45, wherein a value of the quality metrics indicates an average, absolute, or relative quality, for a set of pictures that comprises the picture to which an indication of the value is associated and following pictures, in output order or in decoding order in the video sequence.
47. The method of claim 46 further comprising: signaling, in or along a bitstream, information indicating whether the value of the quality metric refers to an average quality for11the pictures in the video sequence or refers to an average quality for a set of pictures that comprises the picture to which the information is associated and the following pictures in output order or in decoding order in the video sequence.
48. The method of any of the claims 46 or 47, wherein information about quality is comprised in an information message, and wherein the information message further comprises a flag for indicating whether the value for the quality metrics is to be interpreted as an average value for the current picture and subsequent pictures.
49. The method of claim 48, wherein, when the information indicating the value of the quality metrics is associated with a picture that is not a first picture in the video sequence, the value is interpreted as an average quality for the picture to which the information is associated and the subsequent pictures in either output order or decoding order.
50. The method of claim 48, wherein, when one or more values of the quality metrics represents the absolute or relative quality for two or more pictures, the method further comprises: signaling, in or along a bitstream, information indicating how the one or more values were determined.
51. The method of any of the previous claims 36 to 50, wherein the video sequence comprises a coded layer video sequence.
52. The method of any of the previous claims 36 to 51, wherein the information message comprises a supplemental enhancement information message.
53. A method comprising: receiving a quality metrics as an absolute or as relative gains for post-processing filters or post-processing filter groups, relative to the cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group, wherein the quality metrics is determined per video sequence and / or picture; and performing decoder side operations based on the quality metrics.
54. The method of claim 53, wherein the quality metrics comprises user-specified quality metrics comprising one or more following: a text description of the metric; a first flag to indicate whether higher or lower values of the metrics correspond to12increased quality; a second flag to indicate whether the quality metrics is a full reference metric; or signaling of number of bytes signaled for the metrics values.
55. The method of any of the claims 53 or 54, wherein the quality metrics is single color component, individually for 3 color components, or combined for the 3 color components.
56. The method of any of the previous claims 53 to 55, wherein the quality metrics is comprised in an information message, and wherein the information message comprises one or more of following: a quality metrics definition flag for indicating whether the information message comprises definitions of the quality metrics; a sequence level values flag for indicating whether a sequence level metric values are signaled and allows an encoder to choose whether to signal sequence level metrics values, picture level metric values, or both; a number of quality metrics syntax element that enables the encoder to send multiple metric values within the information message; a quality gain flag that enables the encoder to indicate whether the signaled metric values indicate a gain associated with the post-processing filter, or whether the values indicate absolute quality of the picture or the video sequence; reference flag for indicating whether a reference for gain is an input picture to the encoder or an input picture to a post-processing stage; quality metric type syntax element for indicating a metric type of each quality metric; an increasing flag for indicating whether a higher value of a metric corresponds to a higher quality value or a lower quality value; a full reference flag for indicating wherein the quality metric is a full reference metric; a length syntax element for specifying a number of bytes used to signal the quality metrics; description syntax element for signaling a text description of the quality metrics; or description present flag for indicating that the text description of the quality metrics is signaled using the description syntax element.
57. The method of any of the previous claims 53 to 56, wherein the quality metrics comprises values indicating one or more of following: quality of the picture; a mean quality of pictures corresponding to the video sequence;13a quality gain of the picture, which is a difference between the quality of the picture relative to a quality of a gain reference picture; or a mean quality gain of the pictures corresponding to the video sequence.
58. The method of claim 57, wherein the information message comprises one or more of following variables: a chroma format indicator; a count of pictures, lists of picture widths and heights; a list of pictures; or a list of gain reference pictures.
59. The method of any of the previous claims 53 to 58, wherein values the quality metrics are computed by using, as a reference, a filtered picture, and wherein the filtered picture is encoded by the encoder.
60. The method of any of the claims 53 to 58, wherein values the quality metrics are computed by using, as a reference picture, a picture that has not been filtered, and wherein the picture is encoded by the encoder.
61. method of any of the claims 53 to 58, wherein values the quality metrics are computed by using, as a reference picture, a picture prior to any filtering, and wherein the picture is filtered to obtain a filtered picture, and where the filtered picture is encoded by the encoder.
62. The method of any of the claims 59 to 61 further comprising: receiving, from or along a bitstream information indicating one or more of the following: whether the reference picture used for computing the quality metrics is a filtered picture; a type of filtering used for filtering the reference picture; or information that describes a filtering that is applied to the reference picture.
63. The method of any of the previous claims 53 to 62, wherein a value of the quality metrics indicates an average, absolute, or relative quality, for a set of pictures that comprises the picture to which an indication of the value is associated and following pictures, in output order or in decoding order in the video sequence.1464. The method of claim 63 further comprising: receiving, from or along a bitstream, information indicating whether the value of the quality metric refers to an average quality for the pictures in the video sequence or refers to an average quality for a set of pictures that comprises the picture to which the information is associated and the following pictures in output order or in decoding order in the video sequence.
65. The method of any of the claims 63 or 64, wherein the information about quality is comprised in an information message, and wherein the information message further comprises a flag for indicating whether the value for the quality metrics is to be interpreted as an average value for the current picture and subsequent pictures.
66. The method of claim 65, wherein, when the information indicating the value of the quality metrics is associated with a picture that is not a first picture in the video sequence, the value is interpreted as an average quality for the picture to which the information is associated and the subsequent pictures in either output order or decoding order.
67. The method of claim 63, wherein, when one or more values of the quality metrics represents the absolute or relative quality for two or more pictures, the further comprises: receiving, from or along a bitstream, information indicating how the one or more values were determined.
68. The method of any of the previous claims 53 to 67, wherein the video sequence comprises a coded layer video sequence.
69. The method of any of the previous claims 53 to 67, wherein the information message comprises a supplemental enhancement information message.
70. The method of any of the previous claims 53 to 69, wherein the decoder side operations comprise one or more of the following: performing a high complexity post-processing operation; using the quality metrics for determining or selecting a post-processing filter or a postprocessing filter group; selecting highest quality frames of the video sequence to operate on; or determining not to display pictures with quality levels below a target quality level.
71. An apparatus comprising:15means for determining a quality metrics per video sequence and / or picture; and means for signaling the quality metrics as an absolute or relative gains for postprocessing filters or post-processing filter groups, relative to the cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group.
72. The apparatus of claim 71, wherein the apparatus further comprises means for performing methods as claimed in any of the claims 37 to 52.
73. An apparatus comprising: means for receiving a quality metrics as an absolute or as relative gains for postprocessing filters or post-processing filter groups, relative to the cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group, wherein the quality metrics is determined per video sequence and / or picture; and means for performing decoder side operations based on the quality metrics.
74. The apparatus of claim 73, wherein the apparatus further comprises means for performing methods as claimed in any of the claims 54 to 70.
75. A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform: determining a quality metrics per video sequence and / or picture; and signaling the quality metrics as an absolute or relative gains for post-processing filters or post-processing filter groups, relative to the cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group.
76. The apparatus of claim 75, wherein the apparatus is further caused to perform methods as claimed in any of the claims 37 to 52.
77. The computer readable medium of any of the claims 75 or 76, wherein the computer readable medium comprises a non-transitory computer readable medium.
78. A computer readable medium comprising program instructions that, when executed by an apparatus, cause the apparatus to perform: receiving a quality metrics as an absolute or as relative gains for post-processing filters or post-processing filter groups, relative to the cropped decoded picture or relate to an input of a current post-processing filter of a post-processing filter group, wherein the quality metrics is16determined per video sequence and / or picture; and performing decoder side operations based on the quality metrics.
79. The apparatus of claim 78, wherein the apparatus is further caused to perform methods as claimed in any of the claims 54 to 70.
80. The computer readable medium of any of the claims 78 or 79, wherein the computer readable medium comprises a non-transitory computer readable medium.
Citation Information
Patent Citations
High-level syntax for signaling neural networks within a media bitstream
US20220256227A1
Picture orientation and quality metrics supplemental enhancement information message for video coding
US20220321918A1
Signaling information about multiple post processing filters
WO2024214011A1