Indication of use suitability of bitstream or post-processing filter and removal of SEI messages in sub-bitstream extraction

By removing the neural network post-processing filter SEI messages in the video encoding and codec standard, bitstream processing is optimized, and the encoding and codec complexity and bandwidth requirements caused by SEI messages are solved, thereby achieving more efficient video data conversion.

CN120513636APending Publication Date: 2025-08-19DOUYIN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480007436.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-10
Filing Date
2024-01-09
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

In the existing video encoding and decoding standards, the use of SEI messages leads to inconsistent bitstream consistency and output timing, and some SEI messages are not required for the decoding process, which increases the encoding and decoding complexity and bandwidth requirements.

Method used

By removing the neural network postprocessing filter (NNPF) SEI message, signaling indication applicability is used to optimize bitstream processing to realize the conversion of video data and bitstream.

Benefits of technology

It improves the efficiency of the video encoding and decoding process, reduces unnecessary signaling information, optimizes bandwidth usage, and ensures the consistency of the bitstream and the accuracy of the decoding output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120513636A_ABST
    Figure CN120513636A_ABST
Patent Text Reader

Abstract

A mechanism for processing video data is disclosed. The mechanism includes determining an indication indicative of information related to use of a neural network post-processing filter (NNPF). The indication may indicate that appropriate use of the NNPF or video bitstream is unknown. A conversion between the visual media data and the bitstream is performed based on the indication.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This patent application claims the benefit of U.S. Provisional Patent Application No. 63 / 438,194, filed January 10, 2023, the teachings and disclosures of which are incorporated herein by reference in their entirety. Technical Field

[0003] This patent document relates to the generation, storage and use of digital audio and video media information in file format. Background Art

[0004] Digital video accounts for the largest share of bandwidth used on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, the demand for bandwidth used by digital video is likely to continue to grow. Summary of the Invention

[0005] A first aspect relates to a method for processing video data, comprising: determining an indication indicating information related to use of a neural network post-processing filter (NNPF); and performing conversion between visual media data and a bitstream based on the indication.

[0006] A second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform any one of the preceding aspects.

[0007] A third aspect relates to a non-transitory computer-readable medium, comprising a computer program product for use with a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, so that when executed by a processor, the video codec device performs the method of any one of the preceding aspects.

[0008] A fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method includes: determining an indication indicating information related to the use of a neural network post-processing filter (NNPF); and generating a bitstream based on the determination.

[0009] A fifth aspect relates to a method for storing a bitstream of a video, comprising: determining an indication indicating information related to the use of a neural network post-processing filter (NNPF); generating a bitstream based on the determination; and storing the bitstream in a non-temporary computer-readable recording medium.

[0010] For purposes of clarity, any of the foregoing embodiments may be combined with one or more other foregoing embodiments to create new embodiments within the scope of the present disclosure.

[0011] These and other features will be more clearly understood from the following detailed description with reference to the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] For a more complete understanding of this disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.

[0013] Figure 1 An example of deriving a luma channel from a luminance component is shown.

[0014] Figure 2 is a block diagram illustrating an example video processing system.

[0015] Figure 3 is a block diagram of an example video processing device.

[0016] Figure 4 is a flow chart of an example method of video processing.

[0017] Figure 5 is a block diagram illustrating an example video encoding and decoding system.

[0018] Figure 6 is a block diagram illustrating an example encoder.

[0019] Figure 7 is a block diagram illustrating an example decoder.

[0020] Figure 8 is a schematic diagram of an example encoder.

[0021] Figure 9 is a flow chart of an example method of video processing. DETAILED DESCRIPTION

[0022] It should be understood at the outset that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or yet to be developed. The present disclosure should not be limited to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations shown and described herein, but may be modified within the scope of the appended claims and their full scope of equivalents.

[0023] The section headings used in this document are for ease of understanding and do not limit the applicability of the techniques and embodiments disclosed in each section to only that section. In addition, the use of H.266 terminology in some descriptions is solely for ease of understanding and is not intended to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs. In this document, editorial changes relative to the Versatile Video Codec (VVC) specification and / or the Versatile SEI Messages for Encoding and Decoding Video Bitstreams (VSEI) standard are shown above the text by bold italics (indicating deleted text) and bold (indicating added text).

[0024] 1. Preliminary Discussion

[0025] This document relates to image / video codec technology. In particular, the present disclosure relates to the signaling of an indication of the suitability of a bitstream or post-processing filter for use, and the removal of Neural Network Post-Processing Filter (NNPF) SEI messages in sub-bitstream extraction. These ideas can be applied alone or in various combinations to video bitstreams encoded or decoded by any codec, such as the Versatile Video Codec (VVC) standard and / or the Versatile Supplemental Enhancement Information (SEI) message (VSEI) standard for encoding and decoding video bitstreams.

[0026] 2. Abbreviation

[0027] Adaptation Parameter Set (APS), Access Unit (AU), Codec Video Sequence (CLVS), Codec Video Sequence Start (CLVSS), Cyclic Redundancy Check (CRC), Codec Video Sequence (CVS), Finite Impulse Response (FIR), Intra-frame Random Access Point (IRAP), Network Abstraction Layer (NAL), Picture Parameter Set (PPS), Picture Unit (PU), Random Access Skipped Leading (RASL) Picture, Supplemental Enhancement Information (SEI), Step-by-Step Temporal Sublayer Access (STSA), Video Codec Layer (VCL), Versatile Supplemental Enhancement Information described in ITU-T Rec. H.274 | ISO / IEC 23002-7 (VSEI), Video Usability Information (VUI), Versatile Video Codec described in ITU-T Rec. H.266 | ISO / IEC 23090-3 (VVC)

[0028] 3. Further Discussion

[0029] 3.1 Video Codec Standards

[0030] Video codec standards have evolved primarily through the development of standards within the International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) and the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). ITU-T developed the H.261 and H.263 standards, ISO / IEC developed the Moving Picture Experts Group (MPEG)-1 and MPEG-4 Visual standards, and the two organizations jointly developed the H.262 / MPEG-2 Video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and the H.265 / High Efficiency Video Coding (HEVC) [1] standard. Starting with H.262, video codec standards are based on a hybrid video codec structure that utilizes temporal prediction plus transform coding. To explore video codec technologies beyond High Efficiency Video Coding (HEVC), the Video Coding Experts Group (VCEG) and the Moving Picture Experts Group (MPEG) established the Joint Video Exploration Team (JVET). Furthermore, JVET adopted some of the methods and incorporated them into a reference software called the Joint Exploration Model (JEM) [2]. JVET was later renamed the Joint Video Experts Team (JVET) when the Versatile Video Codec (VVC) project was officially launched. VVC[3] is a codec standard that aims to reduce bitrate by 50% compared to HEVC.

[0031] The Versatile Video Codec (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) [3] and the associated Versatile Supplementary Enhancement Information (VSEI) standard for codec video bitstreams (ITU-T H.274 | ISO / IEC 23002-7) [4] are designed for the widest range of applications, including simple uses such as television broadcasting, video conferencing, or playback from storage media, as well as more advanced use cases such as adaptive bitrate streaming, video region extraction, composition and merging of content from multiple codec video bitstreams, multi-view video, scalable layered codecs, and viewport-adaptive 360° immersive media.

[0032] The Essential Video Codec (EVC) standard (ISO / IEC 23094-1) is another video codec standard developed by MPEG.

[0033] 3.2 SEI messages in general and in VVC and VSEI

[0034] SEI messages assist processes related to decoding, display, or other purposes. However, SEI messages are not required to construct luma or chroma samples through the decoding process. Standard-compliant decoders do not need to process this information to achieve output order consistency. Some SEI messages are required to check bitstream consistency and output timing decoder consistency. Other SEI messages are not required to check bitstream consistency.

[0035] Annex D of VVC specifies the syntax and semantics of the SEI message payload of some SEI messages, and specifies the use of SEI messages and VUI parameters whose syntax and semantics are specified in ITU-T H.274 | ISO / IEC 23002-7.

[0036] 3.3 Signaling of Neural Network Post-Processing Filters

[0037] WG 05 output documents N0158 [5] and JVET-AB2006 [6] include specifications for two SEI messages for signaling of neural network post-processing filters, as shown below.

[0038] 8.28 Neural Network Post-Processing Filter Characteristics SEI Message

[0039] 8.28.1 Neural Network Post-Processing Filter Characteristics SEI Message Syntax

[0040]

[0041]

[0042]

[0043] 8.28.2 Neural Network Post-Processing Filter Characteristics SEI Message Semantics

[0044] The Neural Network Post-Processing Filter Characteristic (NNPFC) SEI message specifies a neural network that can be used as a post-processing filter. The Neural Network Post-Processing Filter Activation SEI message is used to indicate the use of a specified post-processing filter for a particular picture.

[0045] Using this SEI message requires defining the following variables:

[0046] – The width and height of the cropped decoded output picture, in units of luma samples, denoted as CroppedWidth and CroppedHeight in this document.

[0047] – The luma sample array CroppedYPic[idx] and chroma sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] of the cropped decoded output picture, when present, where idx is in the range 0 to numInputPics-1 (inclusive), which are used as input to the post-processing filters.

[0048] – The bit depth BitDepthY of the luma sample array of the cropped decoded output picture.

[0049] – The bit depth BitDepthC of the chroma sample array (if any) of the cropped decoded output picture.

[0050] – a chroma format indicator, denoted herein as ChromaFormatIdc, as described in subclause 7.3.

[0051] – When nnpfc_auxiliary_inp_idc is equal to 1, the filter strength control value StrengthControlVal shall be a real number in the range of 0 to 1 (inclusive).

[0052] The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc as specified in Table 2. NOTE 1 – More than one NNPFC SEI message may be present for the same picture. When more than one NNPFC SEI message with different values of nnpfc_id is present or activated for the same picture, they may have the same or different values of nnpfc_purpose and nnpfc_mode_idc.

[0053] nnpfc_id contains an identification number that can be used to identify the post-processing filter. The value of nnpfc_id shall be in the range of 0 to 232-2 (inclusive). The values of nnpfc_id from 256 to 511 (inclusive) and from 231 to 232-2 (inclusive) are reserved for future use by ITU-T | ISO / IEC. A decoder conforming to this version of this document that encounters an NNPFC SEI message with a nnpfc_id in the range of 256 to 511 (inclusive) or in the range of 231 to 232-2 (inclusive) shall ignore the SEI message.

[0054] When the NNPFC SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, the following applies:

[0055] – This SEI message specifies the basic post-processing filters.

[0056] – This SEI message applies to the current decoded picture and all subsequent decoded pictures of the current layer (in output order) until the end of the current CLVS.

[0057] When an NNPFC SEI message is a repetition of a previous NNPFC SEI message in decoding order in the current CLVS, subsequent semantics apply as if the SEI message is the only NNPFC SEI message with the same content within the current CLVS.

[0058] When the NNPFC SEI message is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, the following applies:

[0059] – This SEI message defines updates relative to the preceding base post-processing filter in decoding order with the same nnpfc_id value.

[0060] – This SEI message applies to the current decoded picture and all subsequent decoded pictures of the current layer (in output order) until the end of the current CLVS or the next NNPFC SEI message (in output order) with this specific nnpfc_id value within the current CLVS.

[0061] nnpfc_mode_idc equal to 0 indicates that the SEI message contains an ISO / IEC 15938-17 bitstream that specifies a base post-processing filter or an update relative to a base post-processing filter with the same nnpfc_id value.

[0062] When the NNPFC SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, nnpfc_mode_idc equal to 1 specifies that the base post-processing filter associated with the nnpfc_id value is a neural network identified by the uniform resource identifier (URI) indicated by nnpfc_uri using the format identified by the tag URI nnpfc_tag_uri.

[0063] When the NNPFC SEI message is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, nnpfc_mode_idc equal to 1 specifies that updates relative to the base post-processing filter with the same nnpfc_id value are defined by the URI indicated by nnpfc_uri using the format identified by the tag URI nnpfc_tag_uri.

[0064] In bitstreams conforming to this version of this document, the value of nnpfc_mode_idc shall be in the range of 0 to 1, inclusive. Values of nnpfc_mode_idc from 2 to 255, inclusive, are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_mode_idc in the range of 2 to 255, inclusive. Values of nnpfc_mode_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use.

[0065] When the SEI message is the first NNPFC SEI message in decoding order with a specific nnpfc_id value within the current CLVS, the post-processing filter PostProcessingFilter() is allocated to be the same as the basic post-processing filter.

[0066] When the SEI message is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, the post-processing filter PostProcessingFilter() is obtained by applying the updates defined by the SEI message to the base post-processing filter.

[0067] Updates are not cumulative, but each update is applied on the base post-processing filter specified by the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS.

[0068] nnpfc_reserved_zero_bit_a shall be equal to 0 in bitstreams conforming to this version of this document. Decoders shall ignore NNPFC SEI messages with nnpfc_reserved_zero_bit_a not equal to 0.

[0069] nnpfc_tag_uri contains a tag URI with syntax and semantics as specified by Internet Engineering Task Force (IETF) Request for Comments (RFC) 4151, identifying the format and related information of a neural network used as a base post-processing filter or an update relative to the base post-processing filter with the same nnpfc_id value specified by nnpfc_uri.

[0070] NOTE 2 – nnpfc_tag_uri is able to uniquely identify the format of neural network data specified by nnrpf_uri without the need for a central registration authority.

[0071] nnpfc_tag_uri equal to "tag:iso.org,2023:15938-17" indicates that the neural network data identified by nnpfc_uri complies with ISO / IEC 15938-17.

[0072] nnpfc_uri contains a URI with syntax and semantics as specified in IETF Internet Standard 66 that identifies a neural network used as a base post-processing filter or an update relative to a base post-processing filter with the same nnpfc_id value.

[0073] nnpfc_formatting_and_purpose_flag equal to 1 specifies that syntax elements related to filter purpose, input format, output format and complexity are present. nnpfc_formatting_and_purpose_flag equal to 0 specifies that syntax elements related to filter purpose, input format, output format and complexity are not present.

[0074] When this SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, nnpfc_formatting_and_purpose_flag shall be equal to 1. When this SEI message is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, nnpfc_formatting_and_purpose_flag shall be equal to 0.

[0075] nnpfc_purpose indicates the purpose of the post-processing filter, as specified in Table 20.

[0076] In bitstreams conforming to this version of this document, the value of nnpfc_purpose shall be in the range of 0 to 5, inclusive. Values of nnpfc_purpose from 6 to 1023, inclusive, are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_purpose in the range of 6 to 1203, inclusive. Values of nnpfc_purpose greater than 1023 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use.

[0077] Table 20 - Definition of nnpfc_purpose

[0078]

[0079]

[0080] NOTE 3 – When the reserved value of nnpfc_purpose is used by ITU-T | ISO / IEC in the future, the syntax of this SEI message may be extended with syntax elements whose presence is conditional on nnpfc_purpose being equal to this value.

[0081] When SubWidthC is equal to 1 and SubHeightC is equal to 1, nnpfc_purpose should not be equal to 2 or 4.

[0082] nnpfc_out_sub_c_flag equal to 1 specifies that outSubWidthC is equal to 1 and outSubHeightC is equal to 1. nnpfc_out_sub_c_flag equal to 0 specifies that outSubWidthC is equal to 2 and outSubHeightC is equal to 1. When nnpfc_out_sub_c_flag is not present, outSubWidthC is inferred to be equal to SubWidthC, and outSubHeightC is inferred to be equal to SubHeightC. When ChromaFormatIdc is equal to 2 and nnpfc_out_sub_c_flag is present, the value of nnpfc_out_sub_c_flag shall be equal to 1.

[0083] nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples specify the width and height, respectively, of the luma sample array of the picture produced by applying the post-processing filter identified by nnpfc_id to the cropped decoded output picture. When nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples are not present, they are inferred to be equal to CroppedWidth and CroppedHeight, respectively. The value of nnpfc_pic_width_in_luma_samples shall be in the range of CroppedWidth to CroppedWidth*16-1 (inclusive). The value of nnpfc_pic_height_in_luma_samples shall be in the range of CroppedHeight to CroppedHeight*16-1 (inclusive).

[0084] nnpfc_num_input_pics_minus2 plus 2 specifies the number of decoded output pictures used as input to the post-processing filter.

[0085] nnpfc_interpolated_pics[i] specifies the number of interpolated pictures generated by the post-processing filter between the i-th picture and the (i+1)-th picture used as input to the post-processing filter.

[0086] The variable numInputPics, which specifies the number of pictures used as input to the post-processing filter, and the variable numOutputPics, which specifies the total number of pictures produced by the post-processing filter, are derived as follows:

[0087]

[0088]

[0089] nnpfc_component_last_flag is equal to 1 to indicate that the last dimension in the input tensor inputTensor of the post-processing filter and the output tensor outputTensor produced by the post-processing filter is used for the current channel. nnpfc_component_last_flag is equal to 0 to indicate that the third dimension in the input tensor inputTensor of the post-processing filter and the output tensor outputTensor produced by the post-processing filter is used for the current channel.

[0090] NOTE 4 – The first dimension in the input and output tensors is used for batch indexing, which is a practice in some neural network frameworks. Although the formulas in the semantics of this SEI message use a batch size corresponding to a batch index equal to 0, it is up to the post-processing implementation to determine the batch size used as input for neural network inference.

[0091] Note 5 – For example, when nnpfc_inp_order_idc is equal to 3 and nnpfc_auxiliary_inp_idc is equal to 1, there are 7 channels in the input tensor, including four luma matrices, two chroma matrices, and one auxiliary input matrix. In this case, the process DeriveInputTensors() will derive each of these 7 channels of the input tensor one by one, and when processing a specific channel of these channels, that channel is called the current channel during the process.

[0092] nnpfc_inp_format_idc indicates a method for converting the sample values of the cropped decoded output picture into the input values of the post-processing filter. When nnpfc_inp_format_idc is equal to 0, the input values of the post-processing filter are real numbers, and the functions InpY() and InpC() are defined as follows:

[0093] InpY( x ) = x ÷ ( ( 1 << BitDepthY ) - 1 ) (77)

[0094] InpC( x )= x ÷ ( ( 1 << BitDepthC ) - 1 ) (78)

[0095] When nnpfc_inp_format_idc is equal to 1, the input values of the post-processing filter are unsigned integers, and the functions InpY() and InpC() are defined as follows:

[0096]

[0097] The variable inpTensorBitDepth is derived from the syntax element nnpfc_inp_tensor_bitdepth_minus8 as specified below.

[0098] Values of nnpfc_inp_format_idc greater than 1 are reserved for future specification by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages containing reserved values of nnpfc_inp_format_idc.

[0099] nnpfc_inp_tensor_bitdepth_minus8 specifies the bit depth of the luma sample values in the input integer tensor plus 8. The value of inpTensorBitDepth is derived as follows:

[0100] inpTensorBitDepth = nnpfc_inp_tensor_bitdepth_minus8 + 8 (81)

[0101] A bitstream conformance requirement is that the value of nnpfc_inp_tensor_bitdepth_minus8 must be in the range 0 to 24 (inclusive).

[0102] nnpfc_inp_order_idc indicates a method of ordering the sample array of the cropped decoded output picture as one of the input pictures of the post-processing filter.

[0103] In bitstreams conforming to this version of this document, the value of nnpfc_inp_order_idc shall be in the range of 0 to 3 (inclusive). Values of nnpfc_inp_order_idc from 4 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 4 to 255 (inclusive). Values of nnpfc_inp_order_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use.

[0104] When ChromaFormatIdc is not equal to 1, nnpfc_inp_order_idc shall not be equal to 3.

[0105] Table 21 contains the informative description of the nnpfc_inp_order_idc values.

[0106] Table 21 - Description of nnpfc_inp_order_idc values

[0107]

[0108]

[0109] Figure 1 An example of deriving the luma channel from the luma component is shown, for example when nnpfc_inp_order_idc is equal to 3.

[0110] A tile is a rectangular array of samples of a component (eg, luma or chroma components) from a picture.

[0111] nnpfc_auxiliary_inp_idc greater than 0 indicates that auxiliary input data is present in the input tensor of the neural network post-processing filter. nnpfc_auxiliary_inp_idc equal to 0 indicates that auxiliary input data is not present in the input tensor. nnpfc_auxiliary_inp_idc equal to 1 specifies that the auxiliary input data is derived as specified in Equation 82.

[0112] In bitstreams conforming to this version of this document, the value of nnpfc_auxiliary_inp_idc shall be in the range of 0 to 1, inclusive. Values of nnpfc_inp_order_idc from 2 to 255, inclusive, are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 2 to 255, inclusive. Values of nnpfc_inp_order_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use.

[0113] The procedure DeriveInputTensors() for deriving an input tensor inputTensor specifying the given vertical sample coordinates cTop and horizontal sample coordinates cLeft of the top-left sample location of a patch of samples included in the input tensor is specified as follows:

[0114]

[0115]

[0116]

[0117]

[0118]

[0119] nnpfc_separate_colour_description_present_flag equal to 1 indicates that a different combination of color primaries, transfer characteristics, and matrix coefficients for the picture produced by the post-processing filters is specified in the SEI message syntax structure. nnfpc_separate_colour_description_present_flag equal to 0 indicates that the combination of color primaries, transfer characteristics, and matrix coefficients for the picture produced by the post-processing filters is the same as indicated in the VUI parameters of the CLVS.

[0120] nnpfc_colour_primaries has the same semantics as specified in subclause 7.3 for the vui_colour_primaries syntax element, except as follows:

[0121] –nnpfc_colour_primaries specifies the color primaries of the picture resulting from applying the neural network post-processing filters specified in the SEI message, instead of the color primaries used for CLVS.

[0122] – When nnpfc_colour_primaries is not present in the NNPFC SEI message, the value of nnpfc_colour_primaries is inferred to be equal to vui_colour_primaries.

[0123] nnpfc_transfer_characteristics has the same semantics as specified in subclause 7.3 for the vui_transfer_characteristics syntax element, except as follows:

[0124] –nnpfc_transfer_characteristics specifies the transfer characteristics of the picture resulting from applying the neural network post-processing filters specified in the SEI message, other than the transfer characteristics used for CLVS.

[0125] – When nnpfc_transfer_characteristics is not present in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is presumed to be equal to vui_transfer_characteristics.

[0126] nnpfc_matrix_coeffs has the same semantics as specified for the vui_matrix_coeffs syntax element in Subentry 7.3, except as follows:

[0127] – nnpfc_matrix_coeffs specifies the matrix coefficients of the picture produced by the neural network post - processing filter specified in the application SEI message, rather than the matrix coefficients for CLVS.

[0128] – When nnpfc_matrix_coeffs is not present in the NNPFC SEI message, the value of nnpfc_matrix_coeffs is presumed to be equal to vui_matrix_coeffs.

[0129] – The values allowed for nnpfc_matrix_coeffs are not restricted by the chroma format of the decoded video picture indicated by the value of ChromaFormatIdc of the semantics of the VUI parameters.

[0130] – When nnpfc_matrix_coeffs is equal to 0, nnpfc_out_order_idc shall not be equal to 1 or 3.

[0131] nnpfc_out_format_idc equal to 0 indicates that the sample values output by the post - processing filter are real numbers, where the value range from 0 to 1 (including the boundary values) is linearly mapped to the unsigned integer value range from 0 to (1 << bitDepth) – 1 (including the boundary values) for any desired bit depth bitDepth for subsequent post - processing or display.

[0132] nnpfc_out_format_flag equal to 1 indicates that the sample values output by the post - processing filter are unsigned integers within the range from 0 to (1 << (nnpfc_out_tensor_bitdepth_minus8 + 8)) - 1 (including the boundary values).

[0133] Values of nnpfc_out_format_idc greater than 1 are reserved for future specification by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages containing reserved values of nnpfc_out_format_idc.

[0134] nnpfc_out_tensor_bitdepth_minus8 specifies the bit depth of the sample values in the output integer tensor plus 8. The value of nnpfc_out_tensor_bitdepth_minus8 must be in the range of 0 to 24 (inclusive).

[0135] nnpfc_out_order_idc indicates the output order of samples produced by the post-processing filter.

[0136] In bitstreams conforming to this version of this document, the value of nnpfc_out_order_idc shall be in the range of 0 to 3, inclusive. Values of nnpfc_out_order_idc from 4 to 255, inclusive, are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_out_order_idc in the range of 4 to 255, inclusive. Values of nnpfc_out_order_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use.

[0137] When nnpfc_purpose is equal to 2 or 4, nnpfc_out_order_idc should not be equal to 3.

[0138] Table 22 contains an informative description of the nnpfc_out_order_idc values.

[0139] Table 22 - Description of nnpfc_out_order_idc values

[0140]

[0141] The procedure StoreOutputTensors() is used to derive the sample values in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor, where the output tensor outputTensor specifies the top-left sample position of a small block of samples included in the input tensor for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft. The procedure StoreOutputTensors() is specified as follows:

[0142]

[0143]

[0144]

[0145]

[0146] nnpfc_constant_patch_size_flag equal to 1 indicates that the post-processing filter accepts as input the exact patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. nnpfc_constant_patch_size_flag equal to 0 indicates that the post-processing filter accepts as input any patch size that is a positive integer multiple of the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1.

[0147] nnpfc_patch_width_minus1+1 indicates the horizontal sample count of the required patch size as input to the post-processing filter when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_width_minus1 shall be in the range of 0 to Min(32766, CroppedWidth-1), inclusive.

[0148] nnpfc_patch_height_minus1+1, indicates the vertical sample count of the patch size required for input to the post-processing filter when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_height_minus1 shall be in the range of 0 to Min(32766, CroppedHeight-1), inclusive.

[0149] Let the variables inpPatchWidth and inpPatchHeight be the width and height of the patch size respectively.

[0150] If nnpfc_constant_patch_size_flag is equal to 0, the following applies:

[0151] – The values of inpPatchWidth and inpPatchHeight are provided by external means not specified in this document, or are set by the post-process itself.

[0152] –inpPatchWidth must be a positive integer multiple of nnpfc_patch_width_minus1+1 and must be less than or equal to CroppedWidth. inpPatchHeight must be a positive integer multiple of nnpfc_patch_height_minus1+1 and must be less than or equal to CroppedHeight.

[0153] Otherwise (nnpfc_constant_patch_size_flag is equal to 1), the value of inpPatchWidth is set equal to nnpfc_patch_width_minus1+1, and the value of inpPatchHeight is set equal to nnpfc_patch_height_minus1+1.

[0154] nnpfc_overlap indicates the horizontal and vertical sample counts of overlap of adjacent input tensors to the post-processing filter. The value of nnpfc_overlap must be in the range of 0 to 16383 (inclusive).

[0155] The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows:

[0156] outPatchWidth=(nnpfc_pic_width_in_luma_samples*inpPatchWidth) / CroppedWidth(84)

[0157] outPatchHeight=(nnpfc_pic_height_in_luma_samples*inpPatchHeight) / CroppedHeight(85)

[0158] horCScaling = SubWidthC / outSubWidthC (86)

[0159] verCScaling = SubHeightC / outSubHeightC (87)

[0160] outPatchCWidth = outPatchWidth * horCScaling (88)

[0161] outPatchCHeight = outPatchHeight * verCScaling (89)

[0162] overlapSize = nnpfc_overlap (90)

[0163] The bitstream conformance requirement is that outPatchWidth*CroppedWidth shall be equal to nnpfc_pic_width_in_luma_samples*inpPatchWidth, and outPatchHeight*CroppedHeight shall be equal to nnpfc_pic_height_in_luma_samples*inpPatchHeight.

[0164] nnpfc_padding_type indicates the padding process when referring to sample positions outside the boundaries of the cropped decoded output picture, as described in Table 23. The value of nnpfc_padding_type shall be in the range of 0 to 15 (inclusive).

[0165] Table 23 - Informative description of nnpfc_padding_type values

[0166] nnpfc_padding_type describe 0 Zero padding 1 Copy Fill 2 Reflection Fill 3 Surround Fill 4 Fixed padding 5..15 reserve

[0167] nnpfc_luma_padding_val indicates the luma value to be used for padding when nnpfc_padding_type is equal to 4.

[0168] nnpfc_cb_padding_val indicates the Cb value to be used for padding when nnpfc_padding_type is equal to 4.

[0169] nnpfc_cr_padding_val indicates the Cr value to be used for padding when nnpfc_padding_type is equal to 4.

[0170] The function InpSampleVal(y,x,picHeight,picWidth,croppedPic) takes as input the vertical sample position y, the horizontal sample position x, the picture height picHeight, the picture width picWidth, and the sample array croppedPic. The function returns the value of sampleVal derived as follows:

[0171] NOTE 6 – For the input to the function InpSampleVal(), the vertical positions are listed before the horizontal positions to be compatible with the input tensor convention of some inference engines.

[0172]

[0173] The following example process may be used to filter the cropped decoded output picture on a tile-by-tile basis using a post-processing filter PostProcessingFilter() to generate a filtered picture containing Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated by nnpfc_out_order_idc.

[0174]

[0175]

[0176] nnpfc_complexity_info_present_flag equal to 1 specifies that one or more syntax elements indicating the complexity of the post-processing filter associated with nnpfc_id are present. nnpfc_complexity_info_present_flag equal to 0 specifies that no syntax elements indicating the complexity of the post-processing filter associated with nnpfc_id are present.

[0177] nnpfc_parameter_type_idc equal to 0 indicates that the neural network uses only integer parameters. nnpfc_parameter_type_flag equal to 1 indicates that the neural network can use floating-point parameters or integer parameters. nnpfc_parameter_type_idc equal to 2 indicates that the neural network uses only binary parameters. nnpfc_parameter_type_idc equal to 3 is reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_parameter_type_idc equal to 3.

[0178] nnpfc_log2_parameter_bit_length_minus3 equal to 0, 1, 2, and 3 respectively indicates that the neural network does not use parameters with bit lengths greater than 8, 16, 32, and 64. When nnpfc_parameter_type_idc is present and nnpfc_log2_parameter_bit_length_minus3 is not present, the neural network does not use parameters with bit lengths greater than 1.

[0179] nnpfc_num_parameters_idc indicates the maximum number of neural network parameters for the post-processing filters, in powers of 2048. nnpfc_num_parameters_idc equal to 0 indicates that the maximum number of neural network parameters is unknown. The value of nnpfc_num_parameters_idc shall be in the range of 0 to 52, inclusive. Values of nnpfc_num_parameters_idc greater than 52 are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_num_parameters_idc greater than 52.

[0180] If the value of nnpfc_num_parameters_idc is greater than 0, the variable maxNumParameters is derived as follows:

[0181] maxNumParameters = (2048 << nnpfc_num_parameters_idc) - 1 (93)

[0182] The bitstream conformance requirement is that the number of neural network parameters of the post-processing filters must be less than or equal to maxNumParameters.

[0183] nnpfc_num_kmac_operations_idc greater than 0 indicates that the maximum number of multiply-accumulate operations per sample of the post-processing filter is less than or equal to nnpfc_num_kmac_operations_idc * 1000. nnpfc_num_kmac_operations_idc equal to 0 indicates that the maximum number of multiply-accumulate operations of the network is unknown. The value of nnpfc_num_kmac_operations_idc shall be between 0 and 2. 32 The range is -1 (including the boundary value).

[0184] nnpfc_total_kilobyte_size is greater than 0 to indicate the total size in kilobytes required to store the uncompressed parameters of the neural network. The total size in bits is the number of bits equal to or greater than the sum of the bits used to store each parameter. nnpfc_total_kilobyte_size is the total size in bits divided by 8000, rounded up. nnpfc_total_kilobyte_size is equal to 0 to indicate that the total size required to store the parameters of the neural network is unknown. The value of nnpfc_total_kilobyte_size must be between 0 and 2. 32 The range is -1 (including the boundary value).

[0185] nnpfc_reserved_zero_bit_b shall be equal to 0 in bitstreams conforming to this version of this document. Decoders shall ignore NNPFC SEI messages with nnpfc_reserved_zero_bit_b not equal to 0.

[0186] nnpfc_payload_byte[i] contains the i-th byte of a bitstream conforming to ISO / IEC 15938-17. The byte sequence nnpfc_payload_byte[i] for all present values of i shall be a complete bitstream conforming to ISO / IEC 15938-17.

[0187] 8.29 Neural Network Post-Processing Filter Activation SEI Message

[0188] 8.29.1 Neural Network Post-Processing Filter Activation SEI Message Syntax

[0189]

[0190] 8.29.2 Neural Network Post-Processing Filter Activation SEI Message Semantics

[0191] The Neural Network Post-Processing Filter Activation (NNPFA) SEI message activates or deactivates the possible use of the target neural network post-processing filter identified by nnpfa_target_id for post-processing filtering on a set of pictures.

[0192] NOTE 1 – Multiple NNPFA SEI messages may be present for the same picture, for example when the post-processing filters are used for different purposes or filter different color components.

[0193] nnpfa_target_id indicates the target neural network post-processing filter, which is specified by one or more neural network post-processing filter characteristics SEI messages related to the current picture and with nnpfc_id equal to nnfpa_target_id.

[0194] The value of nnpfa_target_id must be between 0 and 2 32 -2 (including the boundary value). The value of nnpfa_target_id is from 256 to 511 (including the boundary value) and from 2 31 to 2 32 -2 (inclusive) is reserved for future use by ITU-T|ISO / IEC. A decoder conforming to this version of this document shall encounter a nnpfa_target_id in the range 256 to 511 (inclusive) or in the range 256 to 511 (inclusive). 31 to 2 32 When an NNPFA SEI message is received in the range of -2 (including the boundary value), the SEI message shall be ignored.

[0195] An NNPFA SEI message with a specific value of nnpfa_target_id shall not be present in the current PU unless one or both of the following conditions are true:

[0196] – Within the current CLVS, an NNPFC SEI message with nnpfc_id equal to a specific value of nnpfa_target_id exists in the PU preceding the current PU in decoding order.

[0197] – There is an NNPFC SEI message with nnpfc_id equal to a specific value of nnpfa_target_id in the current PU.

[0198] When a PU includes both an NNPFC SEI message with a specific value of nnpfc_id and an NNPFA SEI message with nnpfa_target_id equal to the specific value of nnpfc_id, the NNPFC SEI message shall precede the NNPFA SEI message in decoding order.

[0199] nnpfa_cancel_flag equal to 1 indicates that the persistence of the target neural network post-processing filter established by any previous NNPFA SEI message with the same nnpfa_target_id as the current SEI message is cancelled, i.e., the target neural network post-processing filter is not used again unless it is activated by another NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 0. nnpfa_cancel_flag equal to 0 indicates that nnpfa_persistence_flag follows.

[0200] nnpfa_persistence_flag specifies the persistence of the target neural network post-processing filters of the current layer.

[0201] nnpfa_persistence_flag equal to 0 specifies that the target neural network post-processing filter can only be used for post-processing filtering of the current image.

[0202] nnpfa_persistence_flag equal to 1 specifies that the target neural network post-processing filters can be used for post-processing filtering of the current picture and all subsequent pictures of the current layer (in output order) until one or more of the following conditions are true:

[0203] – A new CLVS starts for the current layer.

[0204] – End of bitstream.

[0205] – Output the picture in the current layer that is associated with an NNPFA SEI message that has the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1, and that follows the current picture in output order.

[0206] NOTE 2 - The target neural network post-processing filter is not applied to the subsequent pictures in the current layer associated with an NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1.

[0207] 3.4 Signaling of the target usage of the neural network post-processing filter (NNPF)

[0208] JVET-AC0076 proposes adding a target usage for post-processing filters in the NNPFC SEI message to indicate whether the filtered video is suitable for any use, intended for user viewing, or expected to be provided as input for machine analysis.

[0209] The syntax and semantic changes proposed in JVET-AC0076 are as follows:

[0210]

[0211] nnpfc_target_usage indicates the intended usage of the filtered output sample array produced by the post-processing filter, as specified in Table XX. The filtered output sample array may undergo further processing, such as color space conversion, before its intended use.

[0212] In bitstreams conforming to this version of this document, the value of nnpfc_target_usage shall be in the range of 0 to 2 (inclusive). Values of nnpfc_target_usage from 3 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_target_usage in the range of 3 to 255 (inclusive). Values of nnpfc_target_usage greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use.

[0213] Table XX-nnpfc_target_usage definition

[0214] value Intended use of the filtered output sample array produced by the post-processing filter 0 Any use 1 User viewing 2 Machine Analysis

[0215] NOTE YY1 – When nnpfc_target_usage is equal to 1, the post-processing filters are intended to improve fidelity but may negatively impact machine analysis accuracy. When nnpfc_target_usage is equal to 2, the post-processing filters are intended to improve machine analysis accuracy but may negatively impact subjective quality.

[0216] NOTE YY2 – When the decoding device displays the video for user viewing and does not perform machine analysis, it is recommended to omit any post-processing filters with nnpfc_target_usage equal to 2. When the decoding device performs machine analysis and does not display the video, it is recommended to omit any post-processing filters with nnpfc_target_usage equal to 1.

[0217] 3.5 Signaling of video bitstream availability

[0218] JVET-AC0077 proposes an SEI message that indicates that the video produced by decoding the current CLVS is not optimized for user viewing and may therefore appear choppy. The proposed SEI message is said to be useful for avoiding accidental display of a bitstream intended for a machine.

[0219] The syntax and semantics of the proposed SEI message are as follows:

[0220] suboptimal_user_viewing_indication(payloadSize){ Descriptor }

[0221] The suboptimal user viewing indication SEI message indicates that the video produced by decoding the current CLVS is not optimized for user viewing and therefore may appear choppy.

[0222] Note 1 – Videos are optimized for purposes other than user viewing, such as machine analysis tasks.

[0223] When a suboptimal user viewing indication SEI message is present for any picture of the CLVS, the suboptimal user viewing indication SEI message shall be present in the first picture unit in decoding order within the CLVS with a specific temporal sub-layer identifier equal to tId, and shall not be present in any other picture unit of the CLVS. When a suboptimal user viewing indication SEI message is present in a picture unit with tId greater than 0, only the sequence of cropped decoded output pictures decoded from picture units with temporal sub-layer identifiers less than tId shall be suitable for user viewing.

[0224] 4. Technical problems solved by the disclosed technical solution

[0225] The example design for signaling the target usage of NNPF and the availability of video bitstreams and the existing designs for sub-bitstream extraction have the following problems:

[0226] First, a value for nnpfc_target_usage should be added to indicate that the appropriate usage is unknown to avoid burdening the encoder.

[0227] Second, a value of nnpfc_target_usage indicating that it is applicable to any usage is of little use, because there are many possible usages, including unknown usage scenarios, and it is almost impossible to guarantee that NNPF can be applied to any usage.

[0228] Third, for the value of nnpfc_target_usage indicating that it is expected to be provided as input to machine analysis, since there are different types of machine analysis, it is necessary to further signal the type or types of machine analysis for which the filtered content is suitable for use as input.

[0229] Fourth, for signaling the availability of a video bitstream, it is not sufficient to simply signal that the video produced by decoding the current CLVS is not optimized for user viewing and may therefore appear choppy. It would be more complete to signal the appropriate or target use of the entire bitstream, certain operating points of the bitstream, the CVS, certain operating points within the CVS, the CLVS, or certain temporal subsets of the CLVS.

[0230] Fifth, the semantics in JVET-AC0077 include the following: When a suboptimal user viewing indication SEI message is present for any picture of a CLVS, the suboptimal user viewing indication SEI message shall be present in the first picture unit in decoding order within the CLVS with a specific temporal sub-layer identifier equal to tId, and shall not be present in any other picture unit of the CLVS. However, it is unclear what the values of the temporal sub-layer identifiers of the SEI message and the picture should be.

[0231] Sixth, the purpose of NNPF can be for picture rate upsampling. However, when the picture rate of the bitstream is reduced by dropping some of the highest temporal sub-layer pictures, the NNPFC SEI message may be retained in the bitstream. However, does it make sense to retain such SEI messages for picture rate upsampling while reducing the picture rate by dropping some of the highest temporal sub-layer pictures? Naturally, such SEI messages should be deleted when extracting the temporal subset from the temporally scalable bitstream.

[0232] 5. List of solutions and implementation examples

[0233] In order to solve the above problems, the following methods are disclosed. These aspects should be considered as examples to explain general concepts and should not be interpreted in a narrow sense. In addition, these examples can be applied alone or in any combination.

[0234] 1) To address issue 1, an indication may be signaled, for example, in an SEI message, and the indication indicates that the appropriate use of the NNPF or the video bitstream is unknown.

[0235] 2) To solve issue 2, the value of nnpfc_target_usage indicating that it is applicable to any usage is not specified.

[0236] 3) To address issue 3, when the value of nnpfc_target_usage indicates that it is desired to be provided as input to machine analysis, further signaling one or more types of machine analysis for which the filtered content is suitable for use as input.

[0237] 4) To address issue 4, an indication is signaled in an SEI message that indicates the appropriate use or target use of a video bitstream, certain operating points of a bitstream, a CVS, certain operating points within a CVS, a CLVS, or certain temporal subsets of a CLVS.

[0238] For example, the operating point (OP) defined by VVC is as follows: a temporal subset of the output layer set (OLS) identified by the highest value of the OLS index and TemporalId. OLS refers to the output layer set, which is defined in VVC as: a set of layers where one or more layers are specified as output layers.

[0239] 5) To solve question 5, the following semantics:

[0240] When a suboptimal user viewing indication SEI message is present for any picture of the CLVS, the suboptimal user viewing indication SEI message shall be present in the first picture unit in decoding order within the CLVS with the specific temporal sub-layer identifier equal to tId, and shall not be present in any other picture unit of the CLVS.

[0241] is changed as follows:

[0242] When a sub-optimal user viewing indication SEI message is present for any picture of the CLVS with a specific temporal sub-layer identifier equal to tId, the sub-optimal user viewing indication SEI message shall be present in the first picture unit in decoding order within the CLVS with a temporal sub-layer identifier equal to tId, and shall not be present in any picture unit of the CLVS with a temporal sub-layer identifier not equal to tId.

[0243] 6) To address issue 6, one or more of the following is provided:

[0244] a. Specifies that when a sub-bitstream is extracted from a temporal scalable bitstream and does not include one or more pictures of the highest temporal sub-layer, all NNPFC SEI messages with the purpose of indicating the use of picture rate upsampling and all corresponding NNPFA SEI messages shall not be included in the extracted sub-bitstream.

[0245] b. It is specified that whenever the sub-picture sub-bitstream extraction process is applied and at least one sub-picture is deleted, all NNPFC SEI messages and all NNPFA SEI messages are also deleted, since NNPF is only applicable to the entire picture.

[0246] c. Change the general sub-bitstream extraction process specified in VVC so that when a sub-bitstream is extracted from a temporal scalable bitstream and does not include one or more pictures of the highest temporal sub-layer, all NNPFC SEI messages with the purpose of indicating the use of picture rate upsampling and all corresponding NNPFA SEI messages are not included in the extracted sub-bitstream.

[0247] d. Change the sub-picture sub-bitstream extraction process specified in VVC so that when at least one sub-picture is deleted, all NNPFC SEI messages and all NNPFA SEI messages are also deleted.

[0248] e. To make it easier to remove the NNPFC and NNPFA SEI messages during sub-bitstream extraction, one or more of the following constraints are specified:

[0249] i. A SEI NAL unit containing an NNPFC SEI message or an NNPFA SEI message shall not contain any other SEI message that is not a scalable nesting SEI message, an NNPFC SEI message, or an NNPFA SEI message.

[0250] ii. An SEI NAL unit containing an NNPFC SEI message or an NNPFA SEI message associated with an NNPF whose purpose includes picture rate upsampling shall not contain any other SEI message that is not a scalable nesting SEI message, an NNPFC SEI message associated with an NNPF whose purpose includes picture rate upsampling, or an NNPFA SEI message associated with an NNPF whose purpose includes picture rate upsampling.

[0251] 6. References

[0252] [1] ITU-T and ISO / IEC, “High efficiency video coding”, ITU-T Recommendation H.265 | ISO / IEC 23008-2 (current version).

[0253] [2] J.Chen, E.Alshina, GJSullivan, J.-R.Ohm, J.Boyce, "Algorithmdescription of Joint Exploration Test Model 7(JEM7)", JVET-G1001, August 2017.

[0254] [3] ITU-T Recommendation H.266 | ISO / IEC 23090-3, “Versatile Video Coding”, 2022.

[0255] [4] ITU-T Recommendation H.274 | ISO / IEC 23002-7, “Versatile Supplemental Enhancement Information Messages for Coded Video Bitstreams”, 2022.

[0256] [5]ISO / IEC JTC 1 / SC 29 / WG 05 output document N0158, "Text ofISO / IEC 23002-7:202x(2nd Ed.)DAM 1Information technology—MPEG video technologies—Part 7: Versatile supplemental enhancement information messages for coded videobitstreams, AMENDMENT 1:Additional SEI messages", October 2022.

[0257] [6] S. McCarthy, T. Chujoh, M. Hannuksela, G. Sullivan, and Y.-K. Wang (eds.), “Additional SEI messages for VSEI (Draft 3)”, JVET output document JVET-AB 2006, publicly available online: https: / / www.jvet-experts.org / doc_end_user / current_document.php?id=12215.

[0258] Figure 2 is a block diagram illustrating an example video processing system 4000 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or may be in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, passive optical networks (PONs), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).

[0259] System 4000 may include a codec component 4004 that can implement the various codecs or encoding methods described in this document. Codec component 4004 can reduce the average bit rate of the video from input 4002 to the output of codec component 4004 to generate a codec representation of the video. Codec technology is therefore sometimes referred to as video compression or video transcoding technology. The output of codec component 4004 can be stored or transmitted via a communication connection such as represented by component 4006. The bitstream (or codec) representation of the video received at input 4002 or the communication transmission can be used by component 4008 to generate pixel values or displayable video that is sent to display interface 4010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it is understood that the codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the codec results will be performed by the decoder.

[0260] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High-Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE) interface, etc. The technology described in this document may be embodied in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.

[0261] Figure 3 is a block diagram of an example video processing device 4100. Device 4100 can be used to implement one or more methods described herein. Device 4100 can be embodied in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. Device 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. Processor(s) 4102 can be configured to implement one or more methods described in this document. Memory(s) 4104 can be used to store data and code for implementing the methods and techniques described herein. Video processing circuitry 4106 can be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, video processing circuitry 4106 can be at least partially included in processor 4102, such as a graphics coprocessor.

[0262] Figure 44 is a flow chart of an example method 4200 for video processing. At step 4202, the method 4200 determines an indication indicating that appropriate use of a neural network post-processing filter (NNPF) or a video bitstream is unknown. At step 4204, conversion between visual media data and a bitstream is performed based on the indication. The conversion may include encoding at an encoder, decoding at a decoder, or a combination thereof.

[0263] It should be noted that method 4200 can be implemented in an apparatus for processing video data that includes a processor and a non-transitory memory having instructions thereon, such as video encoder 4400, video decoder 4500, and / or encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform method 4200. Furthermore, method 4200 can be performed by a non-transitory computer-readable medium comprising a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on a non-transitory computer-readable medium, such that when executed by a processor, the video codec device performs method 4200.

[0264] Figure 5 4 is a block diagram illustrating an example video codec system 4300 that can utilize the techniques of this disclosure. Video codec system 4300 can include a source device 4310 and a destination device 4320. Source device 4310 generates encoded video data, where source device 4310 can be referred to as a video encoding device. Destination device 4320 can decode the encoded video data generated by source device 4310, where destination device 4320 can be referred to as a video decoding device.

[0265] Source device 4310 may include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. Video source 4312 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a codec representation of the picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The coded video data may be transmitted directly to target device 4320 via network 4330 via I / O interface 4316. The coded video data may also be stored on storage medium / server 4340 for access by target device 4320.

[0266] The target device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may obtain encoded video data from the source device 4310 or the storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the target device 4320, or may be external to the target device 4320, wherein the target device 4320 may be configured to interface with an external display device.

[0267] The video encoder 4314 and the video decoder 4324 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or further standards.

[0268] Figure 6 is a block diagram illustrating an example of a video encoder 4400, which may be Figure 5 Video encoder 4314 in system 4300 is shown. Video encoder 4400 can be configured to perform any or all of the techniques of this disclosure. Video encoder 4400 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video encoder 4400. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0269] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402, a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a cache 4413 and an entropy coding unit 4414. The prediction unit 4402 may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405 and an intra-frame prediction unit 4406.

[0270] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in accordance with an IBC mode, where at least one reference picture is a picture in which the current video block is located.

[0271] Furthermore, some components, such as the motion estimation unit 4404 and the motion compensation unit 4405 , may be highly integrated, but are represented separately in the example of the video encoder 4400 for purposes of explanation.

[0272] The segmentation unit 4401 may segment a picture into one or more video blocks. The video encoder 4400 and the video decoder 4500 may support various video block sizes.

[0273] The mode selection unit 4403 can, for example, select one of a plurality of codec modes (intra-frame codec or inter-frame codec) based on the error result, and provide the generated intra-frame codec block or inter-frame codec block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 4403 can select a joint intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 4403 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).

[0274] To perform inter-frame prediction on the current video block, the motion estimation unit 4404 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 4413. The motion compensation unit 4405 may determine a predicted video block for the current video block based on the motion information and decoded samples of a picture from the buffer 4413 (other than the picture associated with the current video block).

[0275] The motion estimation unit 4404 and the motion compensation unit 4405 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0276] In some examples, motion estimation unit 4404 may perform unidirectional prediction on the current video block, and motion estimation unit 4404 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 4404 may then generate a reference index indicating a reference picture in list 0 or list 1, where the reference picture contains the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 may output the reference index, the prediction direction indicator, and the motion vector as motion information for the current video block. Motion compensation unit 4405 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.

[0277] In other examples, motion estimation unit 4404 may perform bidirectional prediction on the current video block. Motion estimation unit 4404 may search the reference pictures in list 0 for a reference video block for the current video block and may also search the reference pictures in list 1 for another reference video block for the current video block. Motion estimation unit 4404 may then generate reference indexes indicating the reference pictures in list 0 and list 1 containing the reference video block and a motion vector indicating the spatial displacement between the reference video block and the current video block. Motion estimation unit 4404 may output the reference index and motion vector for the current video block as motion information for the current video block. Motion compensation unit 4405 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0278] In some examples, motion estimation unit 4404 can output a complete set of motion information for use in the decoding process of a decoder. In some examples, motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, motion estimation unit 4404 can reference motion information of another video block to signal the motion information of the current video block. For example, motion estimation unit 4404 can determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0279] In one example, the motion estimation unit 4404 may indicate to the video decoder 4500 a value in a syntax structure associated with the current video block that indicates that the current video block has the same motion information as another video block.

[0280] In another example, the motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0281] As discussed above, the video encoder 4400 can signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.

[0282] Intra-frame prediction unit 4406 can perform intra-frame prediction on the current video block. When intra-frame prediction unit 4406 performs intra-frame prediction on the current video block, intra-frame prediction unit 4406 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0283] The residual generation unit 4407 can generate residual data for the current video block by subtracting the predicted video block(s) of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0284] In other examples, such as in skip mode, there may be no residual data for the current video block and the residual generation unit 4407 may not perform a subtraction operation.

[0285] Transform processing unit 4408 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to a residual video block associated with the current video block.

[0286] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0287] Inverse quantization unit 4410 and inverse transform unit 4411 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 4412 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by prediction unit 4402 to generate a reconstructed video block associated with the current block for storage in buffer 4413.

[0288] After the reconstruction unit 4412 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.

[0289] The entropy coding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives data, the entropy coding unit 4414 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[0290] Figure 7 is a block diagram illustrating an example of a video decoder 4500, which may be Figure 5 Video decoder 4324 in system 4300 is shown. Video decoder 4500 can be configured to perform any or all of the techniques of this disclosure. In the example shown, video decoder 4500 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video decoder 4500. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0291] In the example shown, video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, video decoder 4500 may perform a decoding process that is generally the inverse of the encoding process described with respect to video encoder 4400.

[0292] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 4501 can decode the entropy-encoded video data, and based on the entropy-encoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 4502 can determine this information, for example, by performing AMVP and Merge modes.

[0293] The motion compensation unit 4502 may generate a motion compensated block and may perform interpolation based on an interpolation filter. An identifier of an interpolation filter to be used with sub-pixel precision may be included in a syntax element.

[0294] The motion compensation unit 4502 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters as used by the video encoder 4400 during encoding of the video block. The motion compensation unit 4502 may determine the interpolation filters used by the video encoder 4400 based on received syntax information, and the motion compensation unit 4502 may use the interpolation filters to produce a prediction block.

[0295] The motion compensation unit 4502 can use some syntax information to determine the size of the blocks used to encode (one or more) frames and / or (one or more) slices of the encoded video sequence, partitioning information describing how each macroblock of the pictures of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame codec block, and other information for decoding the encoded video sequence.

[0296] The intra-frame prediction unit 4503 can form a prediction block from spatially neighboring blocks using, for example, an intra-frame prediction mode received in the bitstream. The inverse quantization unit 4504 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 4501. The inverse transform unit 4505 applies an inverse transform.

[0297] The reconstruction unit 4506 can add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be used to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 4507, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.

[0298] Figure 8 is a schematic diagram of an example encoder 4600. The encoder 4600 is suitable for implementing VVC technology. The encoder 4600 includes three loop filters, namely a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike the DF 4602, which uses a predefined filter, the SAO 4604 and the ALF 4606 use the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter, respectively, and using the side information of the codec to signal the offset and filter coefficients. The ALF 4606 is located at the last processing stage of each picture and can be regarded as a tool that attempts to capture and repair artifacts caused by previous stages.

[0299] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610, which are configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using reference pictures obtained from a reference picture cache 4612. The residual block from the inter-frame prediction or intra-frame prediction is fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed into an entropy codec component 4618. The entropy codec component 4618 performs entropy coding and decoding on the prediction results and quantized transform coefficients and transmits them to a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 can output images to the DF 4602 , SAO 4604 , and ALF 4606 for filtering before these images are stored in the reference picture cache 4612 .

[0300] Figure 94700 is a flow chart of an example method 4700 for video processing. At step 4702, method 4700 determines an indication indicating information related to the use of an NNPF. At step 4704, conversion between visual media data and a bitstream is performed based on the indication. The conversion may include encoding at an encoder, decoding at a decoder, or a combination thereof.

[0301] It should be noted that method 4700 can be implemented in an apparatus for processing video data that includes a processor and non-transitory memory having instructions thereon, such as video encoder 4400, video decoder 4500, and / or encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform method 4700. Furthermore, method 4700 can be performed by a non-transitory computer-readable medium comprising a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on a non-transitory computer-readable medium, such that when executed by a processor, the video codec device performs method 4700.

[0302] A list of some example preferred solutions is provided below.

[0303] The following solutions illustrate examples of the techniques discussed herein.

[0304] 1. A method for processing media data, comprising: determining an indication that appropriate use of a neural network post-processing filter (NNPF) or a video bitstream is unknown; and performing conversion between visual media data and a bitstream based on the indication.

[0305] 2. The method of solution 1, wherein the indication is included in a Supplemental Enhancement Information (SEI) message.

[0306] 3. A method according to any of solutions 1-2, wherein the indication is a value of a neural network post-processing filter characteristic (NNPFC) target usage (nnpfc_target_usage).

[0307] 4. The method according to any of solutions 1-3, wherein the value of nnpfc_target_usage indicates one or more types of machine analysis that are expected to be provided as input to machine analysis, and wherein the filtered content is suitable for use as input is signaled in the bitstream.

[0308] 5. A method according to any of solutions 1-4, wherein the indication indicates suitable use or target use of a video bitstream, certain operation points of a bitstream, a codec video sequence (CVS), certain operation points within a CVS, a codec layer video sequence (CLVS), or certain temporal subsets of a CLVS.

[0309] 6. A method according to any of solutions 1-5, wherein the operating point is a time domain subset of the output layer set (OLS), identified by the highest value of the OLS index and the time domain identifier (TemporalId), and wherein the OLS is a set of layers in which one or more layers are specified as output layers.

[0310] 7. A method according to any one of solutions 1-6, wherein when a suboptimal user viewing indication SEI message exists for any picture of the CLVS with a specific temporal sub-layer identifier equal to tId, the suboptimal user viewing indication SEI message shall exist in the first picture unit in the decoding order within the CLVS with a temporal sub-layer identifier equal to tId, and shall not exist in any picture unit of the CLVS whose temporal sub-layer identifier is not equal to tId.

[0311] 8. A method according to any of solutions 1-7, wherein when a sub-bitstream is extracted from a time-domain scalable bitstream without including one or more pictures of the highest time-domain sub-layer, all NNPFC SEI messages with the purpose of indicating the use of picture rate upsampling and all corresponding neural network post-processing filter activation (NNPFA) SEI messages should not be included in the extracted sub-bitstream.

[0312] 9. The method according to any of solutions 1-8, wherein whenever the sub-picture sub-bitstream extraction process is applied and at least one sub-picture is deleted, since NNPF only applies to the entire picture, all NNPFC SEI messages and all NNPFA SEI messages are also deleted.

[0313] 10. A method according to any of solutions 1-9, wherein when a sub-bitstream is extracted from a time-domain scalable bitstream without including one or more pictures of the highest time-domain sub-layer, all NNPFC SEI messages having the purpose of indicating the use of picture rate upsampling and all corresponding NNPFA SEI messages are not included in the extracted sub-bitstream.

[0314] 11. The method according to any of solutions 1-10, wherein when at least one sub-picture is deleted, all NNPFC SEI messages and all NNPFA SEI messages are also deleted.

[0315] 12. A method according to any one of solutions 1-11, wherein the SEI network abstraction layer (NAL) unit containing the NNPFC SEI message or the NNPFASEI message should not contain any other SEI message that is not a scalable nesting SEI message, an NNPFC SEI message, or an NNPFA SEI message.

[0316] 13. A method according to any one of solutions 1-12, wherein a SEINAL unit containing an NNPFC SEI message or an NNPFA SEI message associated with an NNPF whose purpose includes picture rate upsampling should not contain any other SEI message that is not a scalable nesting SEI message, an NNPFC SEI message associated with an NNPF whose purpose includes picture rate upsampling, or an NNPFA SEI message associated with an NNPF whose purpose includes picture rate upsampling.

[0317] 14. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of solutions 1-13.

[0318] 15. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, so that when executed by a processor, the video codec device performs the method of any one of solutions 1-13.

[0319] 16. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method includes: determining an indication that appropriate use of a neural network post-processing filter (NNPF) or a video bitstream is unknown; and generating a bitstream based on the determination.

[0320] 17. A method for storing a bitstream of a video, comprising: determining an indication that appropriate use of a neural network post-processing filter (NNPF) or a video bitstream is unknown; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0321] 18. A method, apparatus or system as described in this document.

[0322] The following solutions illustrate further examples of the techniques discussed herein.

[0323] 1. A method for processing media data, comprising: determining an indication indicating information related to use of a neural network post-processing filter (NNPF); and performing conversion between visual media data and a bitstream based on the indication.

[0324] 2. The method of solution 1, wherein the indication indicates that the appropriate use of the NNPF or video bitstream is unknown.

[0325] 3. The method according to any of solutions 1-2, wherein the indication is included in a Supplemental Enhancement Information (SEI) message.

[0326] 4. A method according to any of solutions 1-3, wherein the value of no neural network post-processing filter characteristic (NNPFC) target usage (nnpfc_target_usage) is specified to indicate that it is applicable to any usage.

[0327] 5. The method of any of solutions 1-4, wherein when the value of nnpfc_target_usage indicates that it is desired to be provided as input to machine analysis, one or more types of machine analysis for which the filtered content is suitable for use as input is signaled in the bitstream.

[0328] 6. A method according to any of solutions 1-5, wherein the indication indicates appropriate use or target use of a video bitstream, certain operation points of a bitstream, a codec video sequence (CVS), certain operation points within a CVS, a codec layer video sequence (CLVS), or certain temporal subsets of a CLVS, and wherein the indication is included in an SEI message.

[0329] 7. A method according to any of solutions 1-6, wherein the operating point is a time domain subset of an output layer set (OLS), wherein the operating point is identified by the highest value of the OLS index and the time domain identifier (TemporalId), and wherein the OLS is a set of layers in which one or more layers are specified as output layers.

[0330] 8. A method according to any one of solutions 1-7, wherein when a suboptimal user viewing indication SEI message exists for any picture of the CLVS with a specific temporal sublayer identifier equal to the temporal identifier (tId), the suboptimal user viewing indication SEI message shall exist in the first picture unit in the decoding order within the CLVS with the temporal sublayer identifier equal to tId, and shall not exist in any picture unit of the CLVS whose temporal sublayer identifier is not equal to tId.

[0331] 9. A method according to any of solutions 1-8, wherein when a sub-bitstream is extracted from a time-domain scalable bitstream without including one or more pictures of one or more highest time-domain sub-layers, all NNPFC SEI messages with the purpose of indicating the use of picture rate upsampling and all corresponding neural network post-processing filter activation (NNPFA) SEI messages should not be included in the extracted sub-bitstream.

[0332] 10. The method according to any of solutions 1-9, wherein whenever the sub-picture sub-bitstream extraction process is applied and at least one sub-picture is deleted, since NNPF only applies to the entire picture, all NNPFC SEI messages and all NNPFA SEI messages are also deleted.

[0333] 11. A method according to any of solutions 1-10, wherein when a sub-bitstream is extracted from a time-domain scalable bitstream without including one or more pictures of one or more highest time-domain sub-layers, all NNPFC SEI messages having the purpose of indicating the use of picture rate upsampling and all corresponding NNPFA SEI messages are not included in the extracted sub-bitstream.

[0334] 12. The method according to any of solutions 1-11, wherein when at least one sub-picture is deleted, all NNPFC SEI messages and all NNPFA SEI messages are also deleted.

[0335] 13. A method according to any one of solutions 1-12, wherein the SEI network abstraction layer (NAL) unit containing the NNPFC SEI message or the NNPFASEI message should not contain any SEI messages other than the scalable nesting SEI message, the NNPFC SEI message, and the NNPFA SEI message.

[0336] 14. A method according to any one of solutions 1-13, wherein a SEINAL unit containing an NNPFC SEI message or an NNPFA SEI message associated with an NNPF whose purpose includes picture rate upsampling should not contain any SEI messages other than a scalable nesting SEI message, an NNPFC SEI message associated with an NNPF whose purpose includes picture rate upsampling, and an NNPFA SEI message associated with an NNPF whose purpose includes picture rate upsampling.

[0337] 15. The method of any of solutions 1-14, wherein the converting comprises encoding the visual media data into the bitstream.

[0338] 16. The method of any of solutions 1-14, wherein the converting comprises decoding the visual media data from the bitstream.

[0339] 17. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of solutions 1-16.

[0340] 18. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, so that when executed by a processor, the video codec device performs the method of any one of solutions 1-16.

[0341] 19. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method includes: determining an indication indicating information related to the use of a neural network post-processing filter (NNPF); and generating a bitstream based on the determination.

[0342] 20. A method for storing a bitstream of a video, comprising:

[0343] Determining an indication indicating information related to use of a neural network post-processing filter (NNPF); generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0344] In the described solution, an encoder can conform to the format rules by generating a codec representation according to the format rules. In the described solution, a decoder can parse syntax elements in the codec representation according to the format rules using known information about the presence and absence of syntax elements to produce decoded video.

[0345] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. For example, the bitstream representation of a current video block may correspond to bits spread across the same position in the bitstream or at different positions as defined by the syntax. For example, a macroblock may be encoded based on error residual values after transformation and encoding, and bits from headers and other fields in the bitstream may also be used. Furthermore, during conversion, the decoder may parse the bitstream knowing that some fields may or may not be present, based on the determination, as described in the above solution. Similarly, the encoder may determine whether to include or not include particular syntax fields, and generate the codec representation accordingly by including the syntax fields or excluding the syntax fields from the codec representation.

[0346] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of matter that effects a machine-readable propagated signal, or a combination of one or more. The term "data processing apparatus" includes all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus may also include code that creates an execution environment for an associated computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.

[0347] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including stand-alone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that preserves other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to a related program, or in multiple collaborative files (e.g., files storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers that are located at a site or are distributed across multiple sites and interconnected by a communication network.

[0348] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC).

[0349] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor that executes instructions and one or more memory devices that store instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices; magnetic disks, such as internal or removable hard disks; magneto-optical disks; and compact disc read-only memory (CD ROM) and digital versatile disc read-only memory (DVD-ROM) disks. The processor and memory may be supplemented by, or incorporated into, dedicated logic circuitry.

[0350] Although this patent document contains many details, these details should not be construed as limitations on any subject matter or the scope of what may be claimed, but rather as descriptions of features unique to particular embodiments of particular technologies. In this patent document, certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable subcombination. In addition, although features may function in certain combinations as described above and may even be initially claimed in this manner, in some cases, one or more features in a claimed combination may be omitted from the combination, and the claimed combination may be directed to a subcombination or a variation of a subcombination.

[0351] Similarly, while operations may be depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed sequentially in the particular order or sequence shown, or that all illustrated operations be performed to achieve desired results. Furthermore, the partitioning of various system components in the embodiments described in this patent document should not be understood as requiring such partitioning in all embodiments.

[0352] Only a few implementations and examples are described, and other implementations, improvements, and variations can be made based on what is described and illustrated in this patent document.

[0353] A first component is directly coupled to a second component when there are no intervening components other than a line, trace, or other medium between the first and second components. A first component is indirectly coupled to a second component when there are intervening components other than a line, trace, or other medium between the first and second components. The term "coupled" and its variations encompass both direct and indirect couplings. The use of the term "about" is intended to encompass a range of ±10% of the subsequent figure unless otherwise indicated.

[0354] Although a number of embodiments are provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples are to be considered illustrative rather than restrictive, and the present invention is not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be omitted or not implemented.

[0355] In addition, the techniques, systems, subsystems, and methods described and illustrated as discrete or separate in various embodiments may be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of this disclosure. Other items illustrated or discussed as coupled may be directly connected, or may be indirectly coupled or communicated through some interface, device, or intermediate component, whether electrically, mechanically, or otherwise. Other examples of variations, substitutions, and modifications are ascertainable to those skilled in the art and may be made without departing from the spirit and scope of this disclosure.

Claims

1. A method for processing media data, comprising: determining an indication indicative of information related to use of a neural network post-processing filter (NNPF); as well as Conversion between the visual media data and the bitstream is performed based on the indication.

2. The method according to claim 1, wherein The indication indicates that suitable usage of the NNPF or video bitstream is unknown.

3. The method according to any one of claims 1 to 2, wherein The indication is included in a Supplemental Enhancement Information (SEI) message.

4. The method according to any one of claims 1 to 3, wherein Specifying a value for no neural network post-processing filter characteristics (NNPFC) target usage (nnpfc_target_usage) indicates that any usage applies.

5. The method according to any one of claims 1 to 4, wherein When the value of nnpfc_target_usage indicates that it is desired to be provided as input to machine analysis, one or more types of machine analysis for which the filtered content is suitable for use as input is signaled in the bitstream.

6. The method according to any one of claims 1 to 5, wherein The indication indicates appropriate use or target use of a video bitstream, certain operation points of a bitstream, a codec video sequence (CVS), certain operation points within a CVS, a codec layer video sequence (CLVS), or certain temporal subsets of a CLVS, and wherein the indication is included in an SEI message.

7. The method according to any one of claims 1 to 6, wherein An operating point is a temporal subset of an output layer set (OLS), wherein the operating point is identified by the highest value of the OLS index and the temporal identifier (TemporalId), and wherein the OLS is a set of layers in which one or more layers are specified as output layers.

8. The method according to any one of claims 1 to 7, wherein When a suboptimal user viewing indication SEI message is present for any picture of the CLVS with a specific temporal sub-layer identifier equal to the temporal identifier (tId), the suboptimal user viewing indication SEI message shall be present in the first picture unit in decoding order within the CLVS with the temporal sub-layer identifier equal to tId, and shall not be present in any picture unit of the CLVS with a temporal sub-layer identifier not equal to tId.

9. The method according to any one of claims 1 to 8, wherein When a sub-bitstream is extracted from a time-domain scalable bitstream and does not include one or more pictures of one or more highest time-domain sub-layers, all NNPFC SEI messages with the purpose of indicating the use of picture rate upsampling and all corresponding neural network post-processing filter activation (NNPFA) SEI messages should not be included in the extracted sub-bitstream.

10. The method according to any one of claims 1 to 9, wherein Whenever the sub-picture sub-bitstream extraction process is applied and at least one sub-picture is deleted, all NNPFC SEI messages and all NNPFA SEI messages are also deleted since NNPF is only applicable to the entire picture.

11. The method according to any one of claims 1 to 10, wherein When a sub-bitstream is extracted from a time-domain scalable bitstream and does not include one or more pictures of one or more highest temporal sub-layers, all NNPFC SEI messages with the purpose of indicating the use of picture rate upsampling and all said corresponding NNPFA SEI messages are not included in the extracted sub-bitstream.

12. The method according to any one of claims 1 to 11, wherein When at least one sub-picture is deleted, all NNPFC SEI messages and all NNPFA SEI messages are also deleted.

13. The method according to any one of claims 1 to 12, wherein An SEI network abstraction layer (NAL) unit containing an NNPFC SEI message or an NNPFA SEI message shall not contain any SEI messages other than the scalable nesting SEI message, the NNPFC SEI message, and the NNPFA SEI message.

14. The method according to any one of claims 1 to 13, wherein: An SEI NAL unit containing an NNPFC SEI message or an NNPFA SEI message associated with an NNPF whose purpose includes picture rate upsampling shall not contain any SEI messages other than the scalable nesting SEI message, the NNPFC SEI message associated with an NNPF whose purpose includes picture rate upsampling, and the NNPFA SEI message associated with an NNPF whose purpose includes picture rate upsampling.

15. The method according to any one of claims 1 to 14, wherein The converting includes encoding the visual media data into the bitstream.

16. The method according to any one of claims 1 to 14, wherein The converting includes decoding the visual media data from the bitstream.

17. An apparatus for processing video data, comprising: processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1-16.

18. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, the video codec device performs the method of any one of claims 1-16.

19. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by a video processing apparatus, wherein The method comprises: determining an indication indicative of information related to use of a neural network post-processing filter (NNPF); and A bitstream is generated based on the determination.

20. A method for storing a bitstream of a video, comprising: determining an indication indicative of information related to use of a neural network post-processing filter (NNPF); generating a bitstream based on the determination; as well as The bitstream is stored in a non-transitory computer-readable recording medium.