Video availability information related indications and other items in neural network post-processing filter SEI messages
By introducing the neural network post-processing filter feature SEI message into the video bitstream, the problem of insufficient image integrity indication in NNPF output in video encoding and decoding technology is solved, and more efficient and reliable video data transmission is achieved.
Patent Information
- Application Number
- CN202480015704.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-01
- Filing Date
- 2024-03-01
- Publication Date
- 2025-10-17
AI Technical Summary
Existing video encoding and decoding technologies lack an effective signaling mechanism to indicate whether the image is within the complete range when processing images output by neural network post-processing filters, resulting in insufficient reliability and efficiency of video data transmission.
By introducing the neural network post-processing filter feature SEI message into the video bitstream, a signaling mechanism is provided to indicate whether the color space of the NNPF output image and the color space of the decoded image are within the full range, thereby achieving effective transmission of video availability information.
It improves the reliability and efficiency of video data transmission, ensures the integrity and consistency of video data in different color spaces, and is suitable for various video encoding and decoding standards and application scenarios.
Smart Images

Figure CN120814224A_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 487,814, filed March 1, 2023, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD
[0003] The present disclosure relates to generating, storing, and using digital audiovisual media information in a file format. BACKGROUND
[0004] Digital video accounts for the largest bandwidth use on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video usage can continue to grow. SUMMARY
[0005] A first aspect relates to a method of processing video data, comprising: determining to signal a syntax element to indicate whether a neural network post filter (NNPF) output picture is in a full range when a color space of the NNPF output picture and a color space of a decoded picture or a color space of a cropped decoded output picture are different; and performing a conversion between a video media data and a bitstream based on the NNPF.
[0006] A second aspect relates to an apparatus of processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform any one of the preceding aspects.
[0007] A third aspect relates to a non-transitory computer readable medium comprising a computer program product for use by a video coding device, the computer program product including computer executable instructions stored on the non-transitory computer readable medium such that when executed by a processor cause the video coding device to perform the method of any one of the preceding aspects.
[0008] A fourth aspect relates to a non-transitory computer readable recording medium storing a bitstream of a video generated by a method performed by a video processing apparatus, wherein the method comprises: determining to signal a syntax element to indicate whether a neural network post filter (NNPF) output picture is in a full range when a color space of the NNPF output picture and a color space of a decoded picture or a color space of a cropped decoded output picture are different; and generating the bitstream based on the determining.
[0009] A fifth aspect relates to a method for storing a bitstream of a video, including: when a color space of a neural network post filter (NNPF) output picture and a color space of a decoded picture or a color space of a cropped decoded output picture are different, determining to signal a syntax element to indicate whether the NNPF output picture is in full range; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] A sixth aspect relates to the methods, apparatuses, or systems described in this disclosure.
[0011] For the sake of clarity, any of the foregoing embodiments can be combined with one or more other foregoing embodiments to create new embodiments within the scope of the present disclosure.
[0012] These and other features will become more apparent from the following detailed description, taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF DRAWINGS
[0013] For a more complete understanding of the present disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, and wherein like reference numerals represent like parts.
[0014] Figure 1 An example of deriving a luma channel from a luma component is shown.
[0015] Figure 2 is a block diagram illustrating an example video processing system.
[0016] Figure 3 is a block diagram of an example video processing apparatus.
[0017] Figure 4 is a flowchart of an example method of video processing.
[0018] Figure 5 is a block diagram illustrating an example video coding system.
[0019] Figure 6 is a block diagram illustrating an example encoder.
[0020] Figure 7 is a block diagram illustrating an example decoder.
[0021] Figure 8 is a schematic diagram of an example encoder. DETAILED DESCRIPTION
[0022] It should be understood at the outset that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of embodiments, whether currently known or in development. The disclosure should in no way be limited to the illustrative implementations, drawings, and embodiments shown, including exemplary designs and implementations set forth herein, but can be modified in various ways within the scope of the appended claims and their equivalents.
[0023] The use of section headings in this disclosure is for ease of understanding only and is not intended to be used to limit the applicability of the teachings and embodiments disclosed in each section to only the section. Also, the use of H.266 terminology in some descriptions is merely for ease of understanding and is not intended to limit the scope of the disclosed embodiments. Thus, the embodiments described herein are also applicable to other video codec protocols and designs. In this disclosure, editing changes to text are shown by indicating deleted text in bold slant and added text in bold.
[0024] 1. Preliminary Discussion
[0025] This disclosure relates to image / video coding techniques. In particular, this disclosure relates to video usability information (VUI) related information communicated or changed by a neural network post filter (NNPF) message; and other items in a NNPF SEI message. These ideas can be applied individually or in various combinations for video bitstreams coded by any codec, such as the Versatile Video Coding (VVC) standard and / or the Versatile Supplemental Enhancement Information (SEI) message for coded video bitstreams (VSEI) standard.
[0026] 2. Abbreviations
[0027] The following acronyms can be used in this disclosure: adaptive parameter set (APS), access unit (AU), coded layer video sequence (CLVS), coded layer video sequence start (CLVSS), cyclic redundancy check (CRC), coded video sequence (CVS), finite impulse response (FIR), intra random access point (IRAP), network abstraction layer (NAL), neural network post filter (NNPF), neural network post filter activation (NNPFA), neural network post filter characteristics (NNPFC), picture parameter set (PPS), picture unit (PU), random access skipped leading (RASL) picture, supplemental enhancement information (SEI), stepping through temporal sub-layers access (STSA), uniform resource identifier (URI), video coding layer (VCL), versatile supplemental enhancement information (VSEI) described in Recommendation ITU-T H.274 | ISO / IEC 23002-7, video usability information (VUI), versatile video coding (VVC) described in Recommendation ITU-T H.266 | ISO / IEC 23090-3.
[0028] 3. Further discussion
[0029] 3.1 Video coding standards
[0030] Video coding standards have evolved primarily through the development of international standards by the International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) and the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). The ITU-T produced the H.261 and H.263 standards, the ISO / IEC produced the Motion Picture Expert Group (MPEG)-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video standard and the H.264 / MPEG-4 Advanced Video Coding (AVC) standard and the H.265 / High Efficiency Video Coding (HEVC) [1] standard. Since H.262, the video coding standards are based on the hybrid video coding structure, where temporal prediction is combined with transform coding. To explore video coding technologies beyond high efficiency video coding (HEVC), the Joint Video Exploration Team (JVET) was established by the Video Coding Experts Group (VCEG) and the Motion Picture Experts Group (MPEG). In addition, the JVET adopted some methods and incorporated them into a reference software named Joint Exploration Model (JEM) [2]. When the project of versatile video coding (VVC) was officially started, the JVET was later renamed as Joint Video Team (JVT). VVC [3] is a coding standard targeting at 50% bitrate reduction compared to HEVC.
[0031] The Versatile Video Coding (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) [3] and the associated Versatile Supplemental Enhancement Information (VSEI) standard for coded video bitstreams (ITU-T H.274 | ISO / IEC 23002-7) [4] are designed for a wide range of applications, including simple use cases such as television broadcast, video conferencing or playback from storage media, as well as more advanced use cases such as adaptive bitrate streaming, video region extraction, composition and merging of content from multiple coded video bitstreams, multi-view video, scalable layered coding and viewport-adaptive 360° immersive media.
[0032] The Essential Video Coding (EVC) standard (ISO / IEC 23094-1) is another video coding standard developed by MPEG.
[0033] 3.2 General and SEI messages in VVC and VSEI
[0034] SEI messages assist processes related to decoding, display or other purposes. However, SEI messages are not necessary for constructing luma or chroma samples by the decoding process. A conforming decoder does not need to process this information to achieve output order conformance. Some SEI messages are necessary for bitstream conformance and output timing decoder conformance. Other SEI messages are not necessary for bitstream conformance.
[0035] Annex D of VVC specifies the syntax and semantics of SEI message payload for some SEI messages and specifies the usage of SEI messages and VUI parameters whose syntax and semantics are specified in ITU-T H.274 | ISO / IEC 23002-7.
[0036] 3.3 Signaling of neural network post-processing filters
[0037] JVET-AC2032 [5] includes the specification of two SEI messages for the signaling of neural network post-processing filters, as follows.
[0038] 8.28 Neural network post-processing filter characteristics SEI message
[0039] 8.28.1 Neural network post-processing filter characteristics SEI message syntax
[0040]
[0041]
[0042]
[0043] 8.28.2 Neural network post-filter characteristics SEI message semantics
[0044] The neural network post-filter characteristics (NNPFC) SEI message specifies a neural network that can be used as a post-filter. The neural network post-filter activation (NNPFA) SEI message is used to indicate the use of the specified neural network post-filter (NNPF) for a particular picture.
[0045] The following variables need to be defined for the use of this SEI message:
[0046] - the input picture width and height in luma samples, denoted CroppedWidth and CroppedHeight, respectively, in this document.
[0047] - the luma sample array CroppedYPic[idx] and the chroma sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] (when present) of the input pictures, with index idx in the range 0 to numInputPics - 1, inclusive, which are used as input to the NNPF.
[0048] - the bit depth BitDepth for the luma sample array of the input pictures Y .
[0049] - the bit depth BitDepth for the chroma sample arrays (if any) of the input pictures C .
[0050] - the chroma format indicator, denoted ChromaFormatIdc in this document, as specified in subclause 7.3.
[0051] - when nnpfc_auxiliary_inp_idc is equal to 1, the filter strength control value StrengthControlVal shall be a real number in the range 0 to 1, inclusive.
[0052] The input picture with index 0 corresponds to the picture for which the NNPF defined by this NNPFC SEI message is activated by the NNPFA SEI message. The input pictures with index i in the range 1 to numInputPics - 1, inclusive, precede the input picture with index i - 1 in output order.
[0053] When ( nnpfc_purpose & 0x08 ) is not equal to 0 and the input picture with index 0 is associated with a frame packing arrangement SEI message with fp_arrangement_type equal to 5, all input pictures are associated with a frame packing arrangement SEI message with fp_arrangement_type equal to 5 and the same value of fp_current_frame_is_frame0_flag.
[0054] The variables SubWidthC and SubHeightC are derived from ChromaFormatldc as specified in Table 2.
[0055] NOTE 1 - More than one NNPFC SEI message can exist for the same picture. When more than one NNPFC SEI message with different values of nnpfc_id exist or are active for the same picture, they can have the same or different values of nnpfc_purpose and nnpfc_mode_idc.
[0056] nnpfc_purpose indicates the purpose of the NNPFC as specified in Table 20.
[0057] In bitstreams conforming to this version of this document, the value of nnpfc_purpose shall be in the range of 0 to 63, inclusive. The values of nnpfc_purpose from 64 to 65535, inclusive, are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_purpose in the range of 64 to 65535, inclusive.
[0058] Table 20 - Definition of nnpfc_purpose
[0059]
[0060]
[0061] NOTE 2 - When a reserved value of nnpfc_purpose is used in the future by ITU-T | ISO / IEC, the syntax of this SEI message can be extended with syntax elements that are present if nnpfc_purpose is equal to that value.
[0062] When ChromaFormatldc is equal to 3, ( nnpfc_purpose & 0x02 ) shall be equal to 0.
[0063] When ChromaFormatIdc or nnpfc_purpose&0x02 is not equal to 0, nnpfc_purpose&0x20 must be equal to 0.
[0064] nnpfc_id contains an identification number that can be used to identify the NNPF. The value of nnpfc_id must be between 0 and 2. 32 The value of nnpfc_id is in the range of 256 to 511 (inclusive) and 2 31 to 2 32 –2 (inclusive) are reserved for future use by ITU-T|ISO / IEC. A decoder conforming to this version of this document shall encounter an nnpfc_id in the range 256 to 511 (inclusive) or 2 31 to 2 32 When the NNPFC SEI message is in the range –2 (inclusive), the SEI message shall be ignored.
[0065] When the NNPFC SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, the following applies:
[0066] – This SEI message specifies the base NNPF.
[0067] – This SEI message applies to the current decoded picture of the current layer and all subsequent decoded pictures in output order until the end of the current CLVS.
[0068] nnpfc_mode_idc equal to 0 indicates that this SEI message contains the ISO / IEC 15938-17 bitstream of the specified base NNPF, or is an update relative to a base NNPF with the same nnpfc_id value.
[0069] When the NNPFC SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, nnpfc_mode_idc equal to 1 specifies that the base NNPF associated with the nnpfc_id value is a neural network identified by the URI indicated by nnpfc_uri, where the URI has a format identified by the tag URI nnpfc_tag_uri.
[0070] When the NNPFC SEI message is neither the first NNPFC SEI message in decoding order with a particular nnpfc id value within the current CLVS, nor a repetition of the first NNPFC SEI message in decoding order with that particular nnpfc id value, nnpfc mode idc equal to 1 specifies that the update relative to the underlying NNPF with the same nnpfc id value is defined by the URI indicated by nnpfc uri, where the URI has the format identified by the label URI nnpfc tag uri.
[0071] In bitstreams conforming to this version of this document, the value of nnpfc mode idc shall be in the range of 0 to 1, inclusive. Values of nnpfc mode idc from 2 to 255, inclusive, are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages for which nnpfc mode idc is in the range of 2 to 255, inclusive. Values of nnpfc mode idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use.
[0072] When the SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc id value within the current CLVS, NNPF PostProcessingFilter() is assigned to be the same as the underlying NNPF.
[0073] When the SEI message is neither the first NNPFC SEI message in decoding order with a particular nnpfc id value within the current CLVS, nor a repetition of the first NNPFC SEI message in decoding order with that particular nnpfc id value, NNPF PostProcessingFilter() is obtained by applying the update defined by the SEI message to the underlying NNPF.
[0074] The updates are not cumulative, but each update is applied to the underlying NNPF, which is the NNPF specified by the first NNPFC SEI message in decoding order with a particular nnpfc id value within the current CLVS.
[0075] nnpfc reserved zero bit a shall be equal to 0 in bitstreams conforming to this version of this document. Decoders shall ignore NNPFC SEI messages for which nnpfc reserved zero bit a is not equal to 0.
[0076] nnpfc_tag_uri contains a tag URI with the syntax and semantics as specified in IETF RFC 4151 that identifies the format and associated information of the neural network used as the underlying NNPF or an update relative to the underlying NNPF with the same nnpfc_id value specified by nnpfc_uri.
[0077] NOTE 3 - nnpfc_tag_uri is able to uniquely identify the format of the neural network data specified by nnrpf_uri without the need for a central registry authority.
[0078] nnpfc_tag_uri equal to "tag:iso.org,2023:15938-17" indicates that the neural network data identified by nnpfc_uri conforms to ISO / IEC 15938-17.
[0079] nnpfc_uri contains a URI with the syntax and semantics as specified in IETF Internet Standard 66 that identifies the neural network used as the underlying NNPF or an update relative to the underlying NNPF with the same nnpfc_id value.
[0080] nnpfc_property_present_flag equal to 1 specifies that the syntax elements related to filter purpose, input format, output format, and complexity are present. nnpfc_property_present_flag equal to 0 specifies that the syntax elements related to filter purpose, input format, output format, and complexity are not present.
[0081] When this SEI message is the first NNPFC SEI message in decoding order with the specific nnpfc_id value within the current CLVS, nnpfc_property_present_flag shall be equal to 1.
[0082] When nnpfc_property_present_flag is equal to 0, the values of all syntax elements that can be present and for which no respective inference value is specified are inferred to be equal to the corresponding syntax elements in the NNPFC SEI message of the underlying NNPF for which this SEI provides an update only when nnpfc_property_present_flag is equal to 1.
[0083] nnpfc_base_flag equal to 1 specifies that the SEI message specifies the underlying NNPF. nnpf_base_flag equal to 0 specifies that the SEI message specifies an update relative to the underlying NNPF. When not present, the value of nnpfc_base_flag is inferred to be equal to 0.
[0084] The following constraints apply to the value of nnpfc_base_flag:
[0085] - When the NNPFC SEI message is the first NNPFC SEI message with the particular nnpfc_id value within the current CLVS in decoding order, the value of nnpfc_base_flag shall be equal to 1.
[0086] - When the NNPFC SEI message nnpfcB is not the first NNPFC SEI message with the particular nnpfc_id value within the current CLVS in decoding order, and the value of nnpfc_base_flag is equal to 1, the NNPFC SEI message shall be a repetition of the first NNPFC SEI message nnpfcA with the same nnpfc_id in decoding order, i.e., the payload content of nnpfcB shall be the same as nnpfcA.
[0087] When the NNPFC SEI message is not the first NNPFC SEI message with the particular nnpfc_id value within the current CLVS in decoding order, and is not a repetition of the first NNPFC SEI message with that particular nnpfc_id, the following applies:
[0088] - The SEI message defines an update relative to the base NNPF that precedes it in decoding order with the same nnpfc_id value.
[0089] - The SEI message applies to the current decoded picture of the current layer and all subsequent decoded pictures in output order until the end of the current CLVS, or until but not including the decoded picture that follows the current decoded picture in output order within the current CLVS and that is associated with a subsequent NNPFC SEI message in decoding order with the particular nnpfc_id value within the current CLVS, whichever is earlier.
[0090] When the NNPFC SEI message nnpfcCurr is not the first NNPFC SEI message with the particular nnpfc_id value within the current CLVS in decoding order, is not a repetition of the first NNPFC SEI message with that particular nnpfc_id (i.e., the value of nnpfc_base_flag is equal to 0), and the value of nnpfc_property_present_flag is equal to 1, the following constraint applies:
[0091] – The value of nnpfc_purpose in the NNPFC SEI message shall be the same as the value of nnpfc_purpose in the first NNPFC SEI message in decoding order with the particular nnpfc id value within the current CLVS.
[0092] – The values of the syntax elements in the NNPFC SEI message in decoding order following nnpfc_base_flag and preceding nnpfc_complexity_info_present_flag shall be the same as the values of the corresponding syntax elements in the first NNPFC SEI message in decoding order with the particular nnpfc id value within the current CLVS.
[0093] – nnpfc_complexity_info_present_flag shall be equal to 0 or nnpfc_complexity_info_present_flag in the first NNPFC SEI message (denoted as nnpfcBase below) in decoding order with the particular nnpfc id value within the current CLVS shall be equal to 1 and all the following apply:
[0094] – nnpfc_parameter_parameter_type_idc in nnpfcCurr shall be equal to nnpfc_parameter_parameter_type_idc in nnpfcBase.
[0095] – nnpfc_log2_parameter_bit_length_minus3 in nnpfcCurr (when present) shall be less than or equal to nnpfc_log2_parameter_bit_length_minus3 in nnpfcBase.
[0096] – If nnpfc_num_parameters_idc in nnpfcBase is equal to 0, nnpfc_num_parameters_idc in nnpfcCurr shall be equal to 0.
[0097] – Otherwise (nnpfc_num_parameters_idc in nnpfcBase is greater than 0), nnpfc_num_parameters_idc in nnpfcCurr shall be greater than 0 and less than or equal to nnpfc_num_parameters_idc in nnpfcBase.
[0098] - If nnpfc_num_kmac_operations_idc in nnpfcBase is equal to 0, nnpfc_num_kmac_operations_idc in nnpfcCurr shall be equal to 0.
[0099] - Otherwise (nnpfc_num_kmac_operations_idc in nnpfcBase is greater than 0), nnpfc_num_kmac_operations_idc in nnpfcCurr shall be greater than 0 and less than or equal to nnpfc_num_kmac_operations_idc in nnpfcBase.
[0100] - If nnpfc_total_kilobyte_size in nnpfcBase is equal to 0, nnpfc_total_kilobyte_size in nnpfcCurr shall be equal to 0.
[0101] - Otherwise (nnpfc_total_kilobyte_size in nnpfcBase is greater than 0), nnpfc_total_kilobyte_size in nnpfcCurr shall be greater than 0 and less than or equal to nnpfc_total_kilobyte_size in nnpfcBase.
[0102] When nnpfc_purpose & 0x02 is not equal to 0, nnpfc_out_sub_c_flag specifies the values of the variables outSubWidthC and outSubHeightC. nnpfc_out_sub_c_flag equal to 1 specifies that outSubWidthC is equal to 1 and outSubHeightC is equal to 1. nnpfc_out_sub_c_flag equal to 0 specifies that outSubWidthC is equal to 2 and outSubHeightC is equal to 1. When ChromaFormatldc is equal to 2 and nnpfc_out_sub_c_flag is present, the value of nnpfc_out_sub_c_flag shall be equal to 1.
[0103] When nnpfc_purpose & 0x20 is not equal to 0, nnpfc_out_colour_format_idc specifies the colour format of the NNPF output, thereby specifying the values of the variables outSubWidthC and outSubHeightC. nnpfc_out_colour_format_idc equal to 1 specifies that the colour format of the NNPF output is 4:2:0 format and both outSubWidthC and outSubHeightC are equal to 2. nnpfc_out_colour_format_idc equal to 2 specifies that the colour format of the NNPF output is 4:2:2 format and outSubWidthC is equal to 2 and outSubHeightC is equal to 1. nnpfc_out_colour_format_idc equal to 3 specifies that the colour format of the NNPF output is 4:2:4 format and both outSubWidthC and outSubHeightC are equal to 1. The value of nnpfc_out_colour_format_idc shall not be equal to 0.
[0104] When both nnpfc_purpose & 0x02 and nnpfc_purpose & 0x20 are equal to 0, outSubWidthC and outSubHeightC are inferred to be equal to SubWidthC and SubHeightC, respectively.
[0105] nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples specify the width and height, respectively, of the array of luma samples of the picture produced by the NNPF identified by the nnpfc_id to the cropped decoded output picture. When nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples are not present, they are inferred to be equal to CroppedWidth and CroppedHeight, respectively. The value of nnpfc_pic_width_in_luma_samples shall be in the range of CroppedWidth to CroppedWidth * 16 - 1, inclusive. The value of nnpfc_pic_height_in_luma_samples shall be in the range of CroppedHeight to CroppedHeight * 16 - 1, inclusive.
[0106] nnpfc num input pics minusl plus 1 specifies the number of decoded output pictures used as input to the NNPF. The value of nnpfc num input pics minusl shall be in the range of 0 to 63, inclusive. When nnpfc purpose & 0x08 is not equal to 0, the value of nnpfc num input pics minusl shall be greater than 0.
[0107] nnpfc interpolated pics [ i ] specifies the number of interpolated pictures generated by the NNPF between the i-th picture and the (i+1)-th picture used as input to the NNPF. The value of nnpfc interpolated pics [ i ] shall be in the range of 0 to 63, inclusive. For at least one i in the range of 0 to nnpfc num input pics minusl - 1, inclusive, the value of nnpfc interpolated pics [ i ] shall be greater than 0.
[0108] nnpfc input pic output flag [ i ] equal to 1 specifies that for the i-th input picture, the NNPF generates a corresponding output picture. nnpfc input pic output flag [ i ] equal to 0 specifies that for the i-th input picture, the NNPF does not generate a corresponding output picture.
[0109] The variable numInputPics specifying the number of pictures used as input to the NNPF and the variable numOutputPics specifying the total number of pictures produced by the NNPF are derived as follows:
[0110]
[0111]
[0112] nnpfc component last flag equal to 1 specifies that the last dimension in the input tensor inputTensor to the NNPF and the output tensor outputTensor produced by the NNPF is used for the current channel. nnpfc component last flag equal to 0 specifies that the third dimension in the input tensor inputTensor to the NNPF and the output tensor outputTensor produced by the NNPF is used for the current channel.
[0113] NOTE 4 - The first dimension in the input tensors and the output tensors is used for batch indexing, which is a practice in some neural network frameworks. While the formula in the semantics of this SEI message uses the batch size corresponding to the batch index equal to 0, the batch size used as input for neural network inference is determined by the post-processing implementation.
[0114] NOTE 5 - For example, when nnpfc_inp_order_idc is equal to 3 and nnpfc_auxiliary_inp_idc is equal to 1, there are 7 channels in the input tensors, including four luma matrices, two chroma matrices, and one auxiliary input matrix. In this case, the process DeriveInputTensors() will derive each of these 7 channels of the input tensors one by one, and when a particular channel among these channels is being processed, this channel is referred to as the current channel during the process.
[0115] nnpfc_inp_format_idc specifies the method to convert the sample values of the cropped decoded output picture into the input values of the NNPF. When nnpfc_inp_format_idc is equal to 0, the input values of the NNPF are real numbers, and the functions InpY() and InpC() are specified as follows:
[0116] InpY( x ) = x ÷ ( ( 1 « BitDepth Y ) - 1 ) (77)
[0117] InpC( x ) = x ÷ ( ( 1 « BitDepth C ) - 1 ) (78)
[0118] When nnpfc_inp_format_idc is equal to 1, the input values of the NNPF are unsigned integers, and the functions InpY() and InpC() are specified as follows:
[0119]
[0120] The variable inpTensorBitDepthY is derived from the syntax element nnpfc_inp_tensor_luma_bitdepth_minus8 specified as follows. The variable inpTensorBitDepthC is derived from the syntax element nnpfc_inp_tensor_chroma_bitdepth_minus8 specified as follows.
[0121] Values of nnpfc_inp_format_idc greater than 1 are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages containing reserved values of nnpfc_inp_format_idc.
[0122] nnpfc_inp_tensor_luma_bitdepth_minus8 plus 8 specifies the bit depth of luma sample values in the input integer tensor. The value of inpTensorBitDepthY is derived as follows:
[0123] inpTensorBitDepth Y = nnpfc_inp_tensor_luma_bitdepth_minus8 + 8 (81)
[0124] It is a requirement of bitstream conformance that the value of nnpfc_inp_tensor_luma_bitdepth_minus8 be in the range of 0 to 24, inclusive.
[0125] nnpfc_inp_tensor_chroma_bitdepth_minus8 plus 8 specifies the bit depth of chroma sample values in the input integer tensor. The value of inpTensorBitDepthC is derived as follows:
[0126] inpTensorBitDepth C = nnpfc_inp_tensor_chroma_bitdepth_minus8 + 8 (82)
[0127] It is a requirement of bitstream conformance that the value of nnpfc_inp_tensor_chroma_bitdepth_minus8 be in the range of 0 to 24, inclusive.
[0128] nnpfc_inp_order_idc indicates the method of ordering the sample array of the cropped decoded output picture as one of the input pictures of the NNPF.
[0129] In bitstreams conforming to this version of this document, the value of nnpfc_inp_order_idc shall be in the range of 0 to 3, inclusive. Values of nnpfc_inp_order_idc from 4 to 255, inclusive, are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages for which nnpfc_inp_order_idc is in the range of 4 to 255, inclusive. Values of nnpfc_inp_order_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use.
[0130] When ChromaFormatldc is not equal to 1, nnpfc_inp_order_idc shall not be equal to 3.
[0131] Table 21 contains an informative description of the nnpfc_inp_order_idc values.
[0132] Table 21 - Description of nnpfc_inp_order_idc values
[0133]
[0134]
[0135] Figure 1 An example of deriving the luma channel from the luma component is shown.
[0136] A tile is a rectangular array of samples from a component (e.g., luma or chroma component) of a picture.
[0137] nnpfc_auxiliary_inp_idc greater than 0 indicates that auxiliary input data is present in the input tensors of the NNPFC. nnpfc_auxiliary_inp_idc equal to 0 indicates that auxiliary input data is not present in the input tensors. nnpfc_auxiliary_inp_idc equal to 1 specifies that the auxiliary input data is derived as specified by formula 84.
[0138] In bitstreams conforming to this version of this document, the value of nnpfc auxiliary inp idc shall be in the range of 0 to 1, inclusive. The values of nnpfc inp order idc from 2 to 255, inclusive, are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages for which nnpfc inp order idc is in the range of 2 to 255, inclusive. Values of nnpfc inp order idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use.
[0139] When nnpfc auxiliary inp idc is equal to 1, the variable strengthControlScaledVal is derived as follows:
[0140]
[0141] The process DeriveInputTensors() for deriving given vertical sample coordinate cTop and horizontal sample coordinate cLeft for the top-left sample position of a tile for samples included in a specified input tensor is specified as follows:
[0142] for(i = 0; i < numInputPics; i++) {
[0143]
[0144]
[0145]
[0146]
[0147] nnpfc separate colour description present flag equal to 1 specifies that different combinations of colour primaries, transfer characteristics, and matrix coefficients for pictures produced by the NNP are specified in the SEI message syntax structure. nnpfc separate colour description present flag equal to 0 specifies that the combination of colour primaries, transfer characteristics, and matrix coefficients for pictures produced by the NNP are the same as indicated in the VUI parameters of the CLVS.
[0148] nnpfc_colour_primaries has the same semantics as specified in subclause 7.3 for the vui_colour_primaries syntax element, except as follows:
[0149] – nnpfc_colour_primaries specifies the colour primaries of pictures produced by the NNPF specified in the applying SEI message, instead of the colour primaries for the CLVS.
[0150] – When nnpfc_colour_primaries is not present in the NNPFC SEI message, the value of nnpfc_colour_primaries is inferred to be equal to vui_colour_primaries.
[0151] nnpfc_transfer_characteristics has the same semantics as specified in subclause 7.3 for the vui_transfer_characteristics syntax element, except as follows:
[0152] – nnpfc_transfer_characteristics specifies the transfer characteristics of pictures produced by the NNPF specified in the applying SEI message, instead of the transfer characteristics for the CLVS.
[0153] – When nnpfc_transfer_characteristics is not present in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is inferred to be equal to vui_transfer_characteristics.
[0154] nnpfc_matrix_coeffs has the same semantics as specified in subclause 7.3 for the vui_matrix_coeffs syntax element, except as follows:
[0155] – nnpfc_matrix_coeffs specifies the matrix coefficients of pictures produced by the NNPF specified in the applying SEI message, instead of the matrix coefficients for the CLVS.
[0156] – When nnpfc_matrix_coeffs is not present in the NNPFC SEI message, the value of nnpfc_matrix_coeffs is inferred to be equal to vui_matrix_coeffs.
[0157] – The allowed values of nnpfc_matrix_coeffs are not constrained by the chroma format of the decoded video pictures indicated by the value of ChromaFormatldc of the semantics of the VUI parameter.
[0158] – When nnpfc_matrix_coeffs is equal to 0, nnpfc_out_order_idc shall not be equal to 1 or 3.
[0159] nnpfc_out_format_idc equal to 0 indicates that the sample values of the NNPF output are real numbers, with a value range of 0 to 1 (inclusive) linearly mapped to a range of unsigned integer values of 0 to (1 « bitDepth) - 1 (inclusive), where bitDepth is any desired bit depth for subsequent post-processing or display.
[0160] nnpfc_out_format_idc equal to 1 indicates that the luma sample values of the NNPF output are unsigned integers in the range of 0 to (1 « (nnpfc_out_tensor_luma_bitdepth_minus8 + 8)) - 1 (inclusive), and the chroma sample values of the NNPF output are unsigned integers in the range of 0 to (1 « (nnpfc_out_tensor_chroma_bitdepth_minus8 + 8)) - 1 (inclusive).
[0161] Values of nnpfc_out_format_idc greater than 1 are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages containing reserved values of nnpfc_out_format_idc.
[0162] nnpfc_out_tensor_luma_bitdepth_minus8 plus 8 specifies the bit depth of the luma sample values in the output integer tensor. The value of nnpfc_out_tensor_luma_bitdepth_minus8 shall be in the range of 0 to 24, inclusive.
[0163] nnpfc_out_tensor_chroma_bitdepth_minus8 plus 8 specifies the bit depth of the chroma sample values in the output integer tensor. The value of nnpfc_out_tensor_chroma_bitdepth_minus8 shall be in the range of 0 to 24, inclusive.
[0164] When nnpfc_purpose & 0x10 is not equal to 0, the value of nnpfc_out_format_idc shall be equal to 1 and at least one of the following conditions shall be true:
[0165] - nnpfc_out_tensor_luma_bitdepth_minus8 + 8 is greater than BitDepth Y .
[0166] - nnpfc_out_tensor_chroma_bitdepth_minus8 + 8 is greater than BitDepth C .
[0167] nnpfc_out_order_idc indicates the output order of the samples produced by the NNPF.
[0168] In bitstreams conforming to this version of this Document, the value of nnpfc_out_order_idc shall be in the range of 0 to 3, inclusive. Values of nnpfc_out_order_idc from 4 to 255, inclusive, are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this Document. Decoders conforming to this version of this Document shall ignore NNPFC SEI messages with nnpfc_out_order_idc in the range of 4 to 255, inclusive. Values of nnpfc_out_order_idc greater than 255 shall not be present in bitstreams conforming to this version of this Document and are not reserved for future use.
[0169] When nnpfc_purpose & 0x02 is not equal to 0, nnpfc_out_order_idc shall not be equal to 3.
[0170] Table 22 contains the informative description of the nnpfc_out_order_idc values.
[0171] Table 22 - Description of nnpfc_out_order_idc values
[0172]
[0173] A process StoreOutputTensors() for deriving the sample values in the filtered output sample arrays FilteredYPic, FilteredCbPic and FilteredCrPic from the output tensor outputTensor for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft specifying the top-left sample position of the block of samples included in the input tensor is specified as follows:
[0174]
[0175]
[0176]
[0177] nnpfc overlap indicates the horizontal and vertical sample counts of the overlap of the neighboring input tensors of the NNPF. The value of nnpfc overlap shall be in the range of 0 to 16383, inclusive.
[0178] nnpfc_constant_patch_size_flag equal to 1 indicates that the NNPF accepts as input the exact patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. nnpfc_constant_patch_size_flag equal to 0 indicates that the NNPF accepts as input any patch size with a width of inpPatchWidth and a height of inpPatchHeight such that the width of the extended patch (i.e., the patch plus the overlap region) which is equal to inpPatchWidth + 2*nnpfc_overlap is a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2*nnpfc_overlap and the height of the extended patch which is equal to inpPatchHeight + 2*nnpfc_overlap is a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2*nnpfc_overlap.
[0179] When nnpfc_constant_patch_size_flag is equal to 1, nnpfc_patch_width_minus1 plus 1 specifies the horizontal sample count of the patch size required for the input of the NNPF. The value of nnpfc_patch_width_minus1 shall be in the range of 0 to Min(32766, CroppedWidth - 1), inclusive.
[0180] When nnpfc_constant_patch_size_flag is equal to 1, nnpfc_patch_height_minus1 plus 1 specifies the vertical sample count of the patch size required for the input of the NNPF. The value of nnpfc_patch_height_minus1 shall be in the range of 0 to Min(32766, CroppedHeight - 1), inclusive.
[0181] When nnpfc_constant_patch_size_flag is equal to 0, nnpfc_extended_patch_width_cd_delta_minus1 plus 1 plus 2 * nnpfc_overlap specifies the greatest common divisor of all allowed values of the width of the extended patch required for the input of the NNPF. The value of nnpfc_extended_patch_width_cd_delta_minus1 shall be in the range of 0 to Min(32766, CroppedWidth - 1), inclusive.
[0182] When nnpfc_constant_patch_size_flag is equal to 0, nnpfc_extended_patch_height_cd_delta_minus1 plus 1 plus 2 * nnpfc_overlap specifies the greatest common divisor of all allowed values of the height of the extended patch required for the input of the NNPF. The value of nnpfc_extended_patch_height_cd_delta_minus1 shall be in the range of 0 to Min(32766, CroppedHeight - 1), inclusive.
[0183] Let the variables inpPatchWidth and inpPatchHeight be the patch size width and patch size height, respectively.
[0184] If nnpfc_constant_patch_size_flag is equal to 0, the following applies:
[0185] – The values of inpPatchWidth and inpPatchHeight are provided by external means not specified in this document or set by the post-filter itself.
[0186] – The value of inpPatchWidth + 2 * nnpfc overlap must be a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc overlap and inpPatchWidth must be less than or equal to CroppedWidth. The value of inpPatchHeight + 2 * nnpfc overlap must be a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc overlap and inpPatchHeight must be less than or equal to CroppedHeight.
[0187] Otherwise (nnpfc_constant_patch_size_flag is equal to 1), the value of inpPatchWidth is set equal to nnpfc_patch_width_minus1 + 1 and the value of inpPatchHeight is set equal to nnpfc_patch_height_minus1 + 1.
[0188] The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth and outPatchCHeight are derived as follows:
[0189] outPatchWidth = ( nnpfc_pic_width_in_luma_samples * inpPatchWidth ) / CroppedWidth (86)
[0191] outPatchHeight = ( nnpfc_pic_height_in_luma_samples * inpPatchHeight ) / CroppedHeight (87)
[0193] horCScaling = SubWidthC / outSubWidthC (88)
[0194] verCScaling = SubHeightC / outSubHeightC (89)
[0195] outPatchCWidth = outPatchWidth * horCScaling (90)
[0196] outPatchCHeight = outPatchHeight * verCScaling (91)
[0197] The requirement of bitstream conformance is that outPatchWidth * CroppedWidth shall be equal to nnpfc_pic_width_in_luma_samples * inpPatchWidth, and outPatchHeight * CroppedHeight shall be equal to nnpfc_pic_height_in_luma_samples * inpPatchHeight.
[0198] nnpfc_padding_type indicates the padding process when referring to sample positions that are outside the boundaries of the cropped decoded output picture, as described in Table 23. The value of nnpfc_padding_type shall be in the range of 0 to 15, inclusive.
[0199] Table 23 - Informative description of the values of nnpfc_padding_type
[0200] nnpfc_padding_type Description 0 Zero padding 1 Copy padding 2 Reflective padding 3 Wrap padding 4 Fixed padding 5..15 Preserve
[0201] nnpfc_luma_padding_val indicates the luma value to be used for padding when nnpfc_padding_type is equal to 4.
[0202] nnpfc_cb_padding_val indicates the Cb value to be used for padding when nnpfc_padding_type is equal to 4.
[0203] nnpfc_cr_padding_val indicates the Cr value to be used for padding when nnpfc_padding_type is equal to 4.
[0204] The function InpSampleVal(y, x, picHeight, picWidth, croppedPic) returns the value of sampleVal derived as follows:
[0205] NOTE 6 - For the input to the function InpSampleVal(), the vertical position is listed before the horizontal position to be compatible with the input tensor convention of some inference engines.
[0206]
[0207]
[0208] The following example process can be used with the NNPF PostProcessingFilter() to generate filtered and / or interpolated picture(s) in a small block manner, containing Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated by nnpfc_out_order_idc:
[0209]
[0210]
[0211] The order of pictures in the output tensor stored is the output order, and the output order generated by applying the NNPF in the output order is interpreted as the output order (and does not conflict with the output order of the input pictures).
[0212] nnpfc_complexity_info_present_flag equal to 1 specifies that one or more syntax elements indicating the complexity of the NNPF associated with nnpfc_id are present. nnpfc_complexity_info_present_flag equal to 0 specifies that no syntax elements indicating the complexity of the NNPF associated with nnpfc_id are present.
[0213] nnpfc_parameter_type_idc equal to 0 indicates that the neural network uses only integer parameters. nnpfc_parameter_type_flag equal to 1 indicates that the neural network can use floating-point or integer parameters. nnpfc_parameter_type_idc equal to 2 indicates that the neural network uses only binary parameters. nnpfc_parameter_type_idc equal to 3 is reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_parameter_type_idc equal to 3.
[0214] nnpfc_log2_parameter_bit_length_minus3 equal to 0, 1, 2, and 3 indicate that the neural network does not use parameters with bit length greater than 8, 16, 32, and 64, respectively. When nnpfc_parameter_type_idc is present and nnpfc_log2_parameter_bit_length_minus3 is not present, the neural network does not use parameters with bit length greater than 1.
[0215] nnpfc_num_parameters_idc indicates the maximum number of neural network parameters of the NNPF, in units of powers of 2048. nnpfc_num_parameters_idc equal to 0 indicates that the maximum number of neural network parameters is unknown. The value of nnpfc_num_parameters_idc shall be in the range of 0 to 52, inclusive. Values of nnpfc_num_parameters_idc greater than 52 are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_num_parameters_idc greater than 52.
[0216] If the value of nnpfc_num_parameters_idc is greater than 0, the variable maxNumParameters is derived as follows:
[0217] maxNumParameters = ( 2048 « nnpfc_num_parameters_idc ) - 1 (94)
[0218] A requirement for bitstream conformance is that the number of neural network parameters of the NNPF shall be less than or equal to maxNumParameters.
[0219] nnpfc_num_kmac_operations_idc greater than 0 indicates that the maximum number of multiply-accumulate operations per sample of the NNPF is less than or equal to nnpfc_num_kmac_operations_idc * 1000. nnpfc_num_kmac_operations_idc equal to 0 indicates that the maximum number of multiply-accumulate operations of the network is not known. The value of nnpfc_num_kmac_operations_idc shall be in the range of 0 to 2 32 – 2, inclusive.
[0220] nnpfc_total_kilobyte_size greater than 0 indicates the total size in kilobytes required to store the uncompressed parameters of the neural network. The total size in bits is a number equal to or greater than the sum of bits used to store each parameter. nnpfc_total_kilobyte_size is the total size in bits divided by 8000, rounded down. nnpfc_total_kilobyte_size equal to 0 indicates that the total size required to store the parameters of the neural network is not known. The value of nnpfc_total_kilobyte_size shall be in the range of 0 to 2 32 – 2, inclusive.
[0221] nnpfc_reserved_zero_bit_b shall be equal to 0 in bitstreams conforming to this version of this document. Decoders shall ignore NNPFC SEI messages for which nnpfc_reserved_zero_bit_b is not equal to 0.
[0222] nnpfc_payload_byte[ i ] contains the i-th byte of a bitstream conforming to ISO / IEC 15938-17. The sequence of bytes nnpfc_payload_byte[ i ] for all i values present shall be a complete bitstream conforming to ISO / IEC 15938-17.
[0223] 8.29 Neural network post-processing filter activation SEI message
[0224] 8.29.1 Neural network post-processing filter activation SEI message syntax
[0225]
[0226] 8.29.2 Neural network post-processing filter activation SEI message semantics
[0227] A neural network post-filter activation (NNPFA) SEI message activates or deactivates the possible use of a target neural network post-filter (NNPF) identified by nnpfa_target_id for post-filtering of a set of pictures. For a particular picture where a NNPF is activated, the target NNPF is the NNPF specified by the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id that precedes, in decoding order, the first VCL NAL unit of the current picture and that is not a repetition of the NNPFC SEI message containing the underlying NNPF.
[0228] NOTE 1 - For example, when NNPFs are used for different purposes or for filtering different color components, there can be multiple NNPFA SEI messages for the same picture.
[0229] nnpfa_target_id indicates the target NNPF, which is specified by one or more NNPFC SEI messages related to the current picture and with nnpfc_id equal to nnpfa_target_id.
[0230] The value of nnpfa_target_id shall be in the range of 0 to 2 32 - 2, inclusive. The values of nnpfa_target_id from 256 to 511, inclusive, and 2 31 to 2 32 - 2, inclusive, are reserved for future use by ITU-T | ISO / IEC. Decoders conforming to this version of this document shall ignore an NNPFA SEI message when nnpfa_target_id is in the range of 256 to 511, inclusive, or 2 31 to 2 32 - 2, inclusive.
[0231] An NNPFA SEI message with a particular value of nnpfa_target_id shall not be present in the current PU unless one or both of the following conditions are true:
[0232] - There is an NNPFC SEI message with nnpfc_id equal to the particular value of nnpfa_target_id present in a PU that precedes, in decoding order, the current PU within the current CLVS.
[0233] - There is an NNPFC SEI message with nnpfc_id equal to the particular value of nnpfa_target_id present in the current PU.
[0234] When a PU contains a NNPFC SEI message with a particular value of nnpfc_id and a NNPFA SEI message with nnpfa_target_id equal to the particular value of nnpfc_id, the NNPFC SEI message shall precede the NNPFA SEI message in decoding order.
[0235] nnpfa_cancel_flag equal to 1 indicates that the persistence of the target NNPF established by any previous NNPFA SEI message with the same nnpfa_target_id as the current SEI message is cancelled, i.e. the target NNPF is no longer used, unless it is activated by another NNPFA SEI message with the same nnpfa_target_id and nnpfa_cancel_flag equal to 0 as the current SEI message. nnpfa_cancel_flag equal to 0 indicates that nnpfa_persistence_flag follows.
[0236] nnpfa_persistence_flag specifies the persistence of the target NNPF for the current layer.
[0237] nnpfa_persistence_flag equal to 0 specifies that the target NNPF is only used for post-processing filtering of the current picture.
[0238] nnpfa_persistence_flag equal to 1 specifies that the target NNPF is available for post-processing filtering of the current picture and all subsequent pictures of the current layer in output order until one or more of the following conditions are true:
[0239] - a new CLVS of the current layer starts.
[0240] - the bitstream ends.
[0241] - the picture in the current layer that follows the current picture in output order and that is associated with a NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1 is output.
[0242] NOTE 2 - the target NNPF is not applied to the subsequent picture in the current layer that is associated with a NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1.
[0243] nnpfaTargetPictures is the set of pictures related to the last NNPFC SEI message with nnpfa_id equal to nnpfa_target_id that precedes the current NNPFA SEI message in decoding order. nnpfaTargetPictures is the set of pictures that the target NNPFA activates through the current NNPFA SEI message. The requirement for bitstream conformance is that any picture included in nnpfaTargetPictures must also be included in nnpfcTargetPictures.
[0244] 4. Technical problems solved by the disclosed embodiments
[0245] Example designs of the neural network post-processing filter characteristics (NNPFC) SEI message and the neural network post-processing filter activation (NNPFA) SEI message have the following problems:
[0246] First, the NNP(s)F can change the VUI-related information of the output picture. There is a lack of various VUI-related information in the NNPF syntax table.
[0247] Second, the NNPF process can require a StrengthControlVal that is set by the decoder or system. The current design can send an invalid StrengthControlVal to the NNPF process.
[0248] Third, when the NNPF purpose indicates chroma (and possibly other types of format changes or upsampling), nnpfc_out_order_idc should be constrained to avoid or reduce meaningless cases.
[0249] Fourth, in some cases, padding values are signaled while not using padding values.
[0250] 5. List of solutions and embodiments
[0251] To solve the above problems, the following methods are disclosed as summarized below. These aspects should be considered as examples to explain the general concept and should not be interpreted in a narrow way. In addition, these examples can be applied individually or combined in any way.
[0252] 1) To solve problem 1, one or more of the following aspects are specified:
[0253] a. When the NNPF output picture is in 4:2:0 color format, one or more syntax elements can be signaled to indicate the location of the chroma samples.
[0254] i. In one example, when the NNPF purpose indicates colorization (and possibly other types of format changes or upsampling) and the NNPF output picture(s) are in 4:2:0 color format, syntax elements are signaled to specify the location of the chroma samples of the NNPF output picture.
[0255] b. When the color space of the NNPF output picture is different from the color space of the decoded picture or the cropped decoded output picture, a syntax element can be signaled to indicate whether the NNPF output picture is in full range.
[0256] c. A different set of aspect ratio related parameters can be signaled in the NNPF SEI message.
[0257] i. In one example, a first flag is signaled to indicate whether the NNPF process changes the aspect ratio.
[0258] ii. In one example, a different set of aspect ratio parameters, such as aspect_ratio_idc, sar_width, sar_height, can be signaled.
[0259] 1. In one example, whether these parameters are signaled depends on the value of the first flag.
[0260] 2. In one example, when these parameters are not present, the aspect ratio properties are inferred to be the same as the decoded picture or the cropped decoded output picture in the CLVS.
[0261] d. One or more syntax elements can be signaled to indicate the source scan type of the output of the NNPF process.
[0262] i. When not present, the source scan type is inferred to be the same as the decoded picture or the cropped decoded output picture in the CLVS.
[0263] e. One or more syntax elements can be signaled to indicate the preferred display method of the output of the NNPF process.
[0264] i. When not present, the preferred display method is inferred to be the same as the decoded picture or the cropped decoded output picture in the CLVS.
[0265] 2) To address problem 2, one or more of the following aspects are specified:
[0266] a. When NNPF is used, StrengthControlVal is set to a value equal to (SliceQp Y + QpBdOffset) ÷ (63 + QpBdOffset), where SliceQp YSliceQp of the first slice of currCodedPic Y .
[0267] b. Alternatively, when NNPF is used, StrengthControlVal is set to a value equal to (SliceQp Y + A) ÷ B, where SliceQp Y is the SliceQp of the first slice of currCodedPic Y , A is the maximum possible value of QpBdOffset, and B is the maximum possible value of SliceQp Y .
[0268] i. In one example, A is equal to 48 and B is equal to 111.
[0269] c. Alternatively, StrengthControlVal can be signaled in the NNPF SEI message.
[0270] i. In one example, StrengthControlVal can be signaled in the NNPFC SEI message.
[0271] ii. In one example, StrengthControlVal can be signaled in the NNPFA SEI message to activate NNPF.
[0272] 1. In one example, StrengthControlVal can be signaled within a specific range and converted to the input value range of NNPF.
[0273] 3) To address problem 3, one or more of the following aspects are specified:
[0274] a. When NNPF purpose indicates colorization (and possibly other types of format change or upsampling), chroma matrices must be present in the output tensor.
[0275] i. In one example, it is specified that nnpfc_out_order_idc shall not be equal to 0 when nnpfc_purpose & 0x20 is not equal to 0.
[0276] 4) To address problem 4, one or more of the following aspects are specified:
[0277] a. When the input tensor does not contain luma matrices, signaling of luma padding values can be skipped.
[0278] b. When the input tensor does not contain chroma matrices, signaling of chroma padding values can be skipped.
[0279] 6. Embodiments
[0280] Below are some example embodiments of the aspects outlined in Section 5. Most of the relevant parts that have been added or modified are shown in bold font, and some of the parts that have been deleted are shown in italic bold font. There can be some other changes as a matter of editing, so they are not highlighted.
[0281] 6.1 Embodiment 1
[0282] This embodiment covers aspects of items 1, 1a. and 1.b. and all their sub-items, as outlined in Section 5 above. Text changes are based on JVET-AC2032-v2.
[0283] 8.28.1 Neural network post-filtering characteristics SEI message syntax
[0284]
[0285]
[0286] 8.28.2 Neural network post-filtering characteristics SEI message semantics ...
[0288] When both nnpfc_purpose & 0x02 and nnpfc_purpose & 0x20 are equal to 0, outSubWidthC and outSubHeightC are inferred to be equal to SubWidthC and SubHeightC, respectively.
[0289] ...
[0291] nnpfc_separate_colour_description_present_flag equal to 1 indicates that different combinations of colour primaries, transfer characteristics, matrix coefficients for pictures produced by the NNPF are specified in the SEI message syntax structure. nnpfc_separate_colour_description_present_flag equal to 0 indicates that the combination of colour primaries, transfer characteristics, matrix coefficients for pictures produced by the NNPF are the same as indicated in the VUI parameters of the CLVS.
[0292] nnpfc_colour_primaries has the same semantics as specified in subclause 7.3 for the vui_colour_primaries syntax element, except as follows:
[0293] - nnpfc_colour_primaries specifies the colour primaries of pictures produced by the NNPF specified in the applying SEI message, instead of the colour primaries for the CLVS.
[0294] - When nnpfc_colour_primaries is not present in the NNPFC SEI message, the value of nnpfc_colour_primaries is inferred to be equal to vui_colour_primaries.
[0295] nnpfc_transfer_characteristics has the same semantics as specified in subclause 7.3 for the vui_transfer_characteristics syntax element, except as follows:
[0296] - nnpfc_transfer_characteristics specifies the transfer characteristics of pictures produced by the NNPF specified in the applying SEI message, instead of the transfer characteristics for the CLVS.
[0297] - When nnpfc_transfer_characteristics is not present in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is inferred to be equal to vui_transfer_characteristics.
[0298] nnpfc_matrix_coeffs has the same semantics as specified in subclause 7.3 for the vui_matrix_coeffs syntax element, except as follows:
[0299] - nnpfc_matrix_coeffs specifies the matrix coefficients of pictures produced by the NNPF specified in the applying SEI message, instead of the matrix coefficients for the CLVS.
[0300] - When nnpfc_matrix_coeffs is not present in the NNPFC SEI message, the value of nnpfc_matrix_coeffs is inferred to be equal to vui_matrix_coeffs.
[0301] - The value for nnpfc_matrix_coeffs is allowed to be unconstrained by the chroma format of the decoded video picture indicated by the value of ChromaFormatldc of the semantics of the VUI parameter.
[0302] - When nnpfc_matrix_coeffs is equal to 0, nnpfc_out_order_idc shall not be equal to 1 or 3.
[0303] ...
[0305] 6.2 Embodiment 2
[0306] This embodiment covers aspects of item 2 and 2.b. and all of its sub-items, as outlined in section 5 above. The text changes are based on JVET-AC2005-vl.
[0307] D.12.11 Usage of neural network post-filtering filter characteristics SEI message ...
[0309] For the purpose of interpreting the NNPFC SEI message, the following variables are specified:
[0310] - If pictureRateUpsamplingFlag is equal to 1 and there exists a second NNPF defined by at least one NNPFC SEI message, activated by the NNPFA SEI message for currCodedPic and nnpfc_purpose is equal to 4, the following applies:
[0311] - CroppedWidth is set equal to nnpfc_pic_width_in_luma_samples defined for the second NNPF.
[0312] - CroppedHeight is set equal to nnpfc_pic_height_in_luma_samples defined for the second NNPF.
[0313] - Otherwise, the following applies:
[0314] - CroppedWidth is set equal to the value of pps_pic_width_in_luma_samples - SubWidthC * (pps_conf_win_left_offset + pps_conf_win_right_offset) of currCodedPic.
[0315] - CroppedHeight is set equal to the value of pps_pic_height_in_luma_samples of currCodedPic - SubHeightC * (pps_conf_win_top_offset + pps_conf_win_bottom_offset).
[0316] - The luma sample array CroppedYPic[i] and the chroma sample arrays CroppedCbPic[i] and CroppedCrPic[i] (when present) are derived for each value of i in the range of 0 to numInputPics - 1, inclusive, as follows:
[0317] - such that sourcePic is the cropped decoded output picture having a PicOrderCntVal equal to the inputPicPoc[i] in the CLVS containing currCodedPic.
[0318] - If pictureRateUpsamplingFlag is equal to 0, the following applies:
[0319] - The luma sample array CroppedYPic[i] and the chroma sample arrays CroppedCbPic[i] and CroppedCrPic[i] (when present) are set to two-dimensional arrays of decoded sample values of the Y, Cb and Cr components of sourcePic.
[0320] - Otherwise (pictureRateUpsamplingFlag is equal to 1), the following applies:
[0321] - The variable sourceWidth is set equal to the value of pps_pic_width_in_luma_samples of sourcePic - SubWidthC * (pps_conf_win_left_offset + pps_conf_win_right_offset).
[0322] - The variable sourceHeight is set equal to the value of pps_pic_height_in_luma_samples of sourcePic - SubHeightC * (pps_conf_win_top_offset + pps_conf_win_bottom_offset).
[0323] - If sourceWidth is equal to CroppedWidth and sourceHeight is equal to CroppedHeight, inputPic is set to be the same as sourcePic.
[0324] - Otherwise (sourceWidth is not equal to CroppedWidth or sourceHeight is not equal to CroppedHeight), the following applies:
[0325] - There shall be a NNPFC SEI message defined by at least one NNPFA SEI message that is activated for sourcePic and nnpfc_purpose is equal to 4, nnpfc_pic_width_in_luma_samples is equal to CroppedWidth, and nnpfc_pic_height_in_luma_samples is equal to CroppedHeight, hereafter referred to as the super-resolution NNPFC.
[0326] - inputPic is set to be the output of the neural network inference of the super-resolution NNPFC with sourcePic as input.
[0327] - The luma sample array CroppedYPic[i] and the chroma sample arrays CroppedCbPic[i] and CroppedCrPic[i] (when present) are set to be two-dimensional arrays of decoded sample values of the Y, Cb and Cr components of inputPic.
[0328] - BitDepth Y and BitDepth C are set to be equal to BitDepth.
[0329] - ChromaFormatldc is set to be equal to sps_chroma_format_idc.
[0330] - StrengthControlVal is set to be equal to the value of (SliceQp Y + 48) ÷ 63 111 of the first slice of currCodedPic. ...
[0332] 6.3 Embodiment 3
[0333] This embodiment covers aspects of items 4, 4a. and 4.b. and all their sub-items, as outlined in section 5 above. The text changes are based on JVET-AC2032-v2.
[0334] 8.28.1 Neural network post-processing filter characteristics SEI message syntax
[0335]
[0336] 7. References
[0337] [1] ITU-T and ISO / IEC, “High efficiency video coding”, Rec. ITU-T H.265 | ISO / IEC 23008-2 (in force edition).
[0338] [2] J. Chen, E. Alshina, G. J. Sullivan, J.-R. Ohm, J. Boyce, “Algorithm description of Joint Exploration Test Model 7 (JEM7),” JVET-G1001, Aug. 2017.
[0339] [3] Rec. ITU-T H.266 | ISO / IEC 23090-3, “Versatile Video Coding”, 2022.
[0340] [4] Rec. ITU-T Rec. H.274 | ISO / IEC 23002-7, “Versatile Supplemental Enhancement Information Messages for Coded Video Bitstreams”, 2022.
[0341] [5] S. McCarthy, S. Deshpande, M. Hannuksela, Hendry, G. Sullivan, and Y.-K. Wang (editors), “Improvements under consideration for neural network postfilter SEI messages,” JVET output document JVET-AC2032, publicly available online herein: https: / / jvet-experts.org / doc_end_user / current_document.php?id=12585.
[0342] Figure 2is a block diagram illustrating an example video processing system 4000 in which various embodiments disclosed herein can be implemented. Various implementations can include some or all of the components of the system 4000. The system 4000 can include an input 4002 for receiving video content. The video content can be received in a raw or uncompressed format, e.g., 8 or 10 bit multi-component pixel values, or can be received in a compressed or encoded format. The input 4002 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0343] The system 4000 can include an encoding component 4004 that can implement various coding or encoding methods described in this disclosure. The encoding component 4004 can reduce the average bitrate of video from the input 4002 to the output of the encoding component 4004 to produce an encoded representation of the video. The encoding techniques are thus sometimes referred to as video compression or video transcoding techniques. The output of the encoding component 4004 can be stored, or transmitted via a communication connection as represented by component 4006. The bitstream (or encoded) representation of the video received at the input 4002 for storage or communication transmission can be used by component 4008 to generate pixel values or displayable video that is sent to a display interface 4010. The process of generating user- visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although certain video processing operations are referred to as “encoding” operations or tools, it should be understood that the encoding tools or operations are used at an encoder, and corresponding decoding tools or operations that reverse the encoding results will be performed by a decoder.
[0344] Examples of peripheral bus interfaces or display interfaces can include Universal Serial Bus (USB) or High-Definition Multimedia Interface (HDMI) or Displayport, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE) interfaces, etc. Embodiments described in this disclosure can be embodied in various electronic devices such as mobile telephones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.
[0345] Figure 3is a block diagram of an example video processing device 4100. The device 4100 can be used to implement one or more of the methods described herein. The device 4100 can be embodied in a smartphone, tablet computer, computer, Internet of Things (IoT) receiver, etc. The device 4100 can include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processor(s) 4102 can be configured to implement one or more methods described in the present disclosure. The memory(ies) 4104 can be used for storing data and code used for implementing the methods and embodiments described herein. The video processing circuitry 4106 can be used to implement some embodiments described in the present disclosure in hardware circuitry. In some embodiments, the video processing circuitry 4106 can be included at least in part in the processor 4102, for example, a graphics co-processor.
[0346] Figure 4 is a flowchart of an example method 4200 of video processing. The method 4200 determines to signal a syntax element to indicate whether a neural network post filter (NNPF) output picture is in a full range when a color space of the NNPF output picture and a color space of a decoded picture or a color space of a cropped decoded output picture are different, at step 4202. The method 4200 performs a conversion between the visual media data and the bitstream based on the NNPF, at step 4204. The conversion can include encoding at an encoder, decoding at a decoder, or a combination thereof.
[0347] It should be noted that the method 4200 can be implemented in a device that processes video data, such as the video encoder 4400, the video decoder 4500, and / or the encoder 4600, including a processor and a non-transitory memory having instructions executed by the processor. In this case, the instructions, when executed by the processor, cause the processor to perform the method 4200. In addition, the method 4200 can be performed by a non-transitory computer-readable medium comprising a computer program product for use by a video coding device. The computer program product includes computer executable instructions stored on the non-transitory computer-readable medium such that when executed by a processor cause the video coding device to perform the method 4200.
[0348] Figure 5 is a block diagram illustrating an example video coding system 4300 that can utilize embodiments of the present disclosure. The video coding system 4300 can include a source device 4310 and a destination device 4320. The source device 4310 generates encoded video data, where the source device 4310 can be referred to as a video encoding device. The destination device 4320 can decode the encoded video data generated by the source device 4310, where the destination device 4320 can be referred to as a video decoding device.
[0349] Source device 4310 may include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. Video source 4312 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include encoded pictures and associated data. An encoded picture is a coded representation of the picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be transmitted directly to target device 4320 via network 4330 via I / O interface 4316. The encoded video data may also be stored on storage medium / server 4340 for access by target device 4320.
[0350] The target device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may obtain encoded video data from the source device 4310 or the storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the target device 4320, or may be external to the target device 4320, wherein the target device 4320 may be configured to interface with an external display device.
[0351] The video encoder 4314 and the video decoder 4324 may operate according to a video compression standard, such as the HEVC standard, the VVC standard, and other existing and / or further standards.
[0352] Figure 6 is a block diagram illustrating an example of a video encoder 4400, which may be Figure 5 Video encoder 4314 in system 4300 is shown. Video encoder 4400 can be configured to perform any or all embodiments of the present disclosure. Video encoder 4400 includes multiple functional components. The embodiments described in this disclosure can be shared between the various components of video encoder 4400. In some examples, a processor can be configured to perform any or all embodiments described in this disclosure.
[0353] The functional components of video encoder 4400 can include a partitioning unit 4401, a prediction unit 4402, which can include a mode select unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra-prediction unit 4406, a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy encoding unit 4414.
[0354] In other examples, video encoder 4400 can include more, less, or different functional components. In one example, prediction unit 4402 can include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, in which at least one reference picture is the picture in which the current video block is located.
[0355] Furthermore, some components, such as motion estimation unit 4404 and motion compensation unit 4405, can be highly integrated, but are represented separately in the example of video encoder 4400 for explanatory purposes.
[0356] Partitioning unit 4401 can partition a picture into one or more video blocks. Video encoder 4400 and video decoder 4500 can support various video block sizes.
[0357] Mode select unit 4403 can select one of a plurality of coding modes (intra-coding or inter-coding), e.g., based on error results, and provide a resulting intra-coded or inter-coded block to residual generation unit 4407 for generation of residual block data and to reconstruction unit 4412 for reconstruction of the coded block for use as reference picture. In some examples, mode select unit 4403 can select a combined inter-intra prediction (CIIP) mode, in which the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, mode select unit 4403 can also select a resolution for a motion vector for the block (e.g., sub-pixel precision or integer pixel precision).
[0358] To perform inter prediction for a current video block, motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from buffer 4413 to the current video block. Motion compensation unit 4405 can determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 4413 other than the picture in which the current video block is associated.
[0359] Motion estimation unit 4404 and motion compensation unit 4405 can perform different operations, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0360] In some examples, the motion estimation unit 4404 can perform uni-prediction for the current video block, and the motion estimation unit 4404 can search for a reference video block for the current video block in a reference picture in List 0 or List 1. The motion estimation unit 4404 can then generate a reference index indicating the reference picture in List 0 or List 1 containing the reference video block, and a motion vector indicating a spatial displacement between the current video block and the reference video block. The motion estimation unit 4404 can output the reference index, the prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0361] In other examples, the motion estimation unit 4404 can perform bi-prediction for the current video block, the motion estimation unit 4404 can search for a reference video block for the current video block in a reference picture in List 0, and can also search for another reference video block for the current video block in a reference picture in List 1. The motion estimation unit 4404 can then generate a reference index indicating the reference pictures in List 0 and List 1 containing the reference video blocks, and a motion vector indicating a spatial displacement between the reference video blocks and the current video block. The motion estimation unit 4404 can output the reference index and the motion vector for the current video block as the motion information for the current video block. The motion compensation unit 4405 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information for the current video block.
[0362] In some examples, the motion estimation unit 4404 can output a complete set of motion information for a decoder to use in decoding processing. In some examples, the motion estimation unit 4404 can not output a complete set of motion information for the current video. Instead, the motion estimation unit 4404 can signal the motion information for the current video block by reference to the motion information of another video block. For example, the motion estimation unit 4404 can determine that the motion information for the current video block is sufficiently similar to the motion information of a neighboring video block.
[0363] In one example, the motion estimation unit 4404 can indicate to the video decoder 4500 in a syntax structure associated with the current video block a value indicating that the current video block has the same motion information as another video block.
[0364] In another example, motion estimation unit 4404 can identify, in a syntax structure associated with the current video block, another video block and a motion vector difference (MVD). The motion vector difference indicates a difference between a motion vector of the current video block and a motion vector of the indicated video block. Video decoder 4500 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0365] As discussed above, video encoder 4400 can signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that can be implemented by video encoder 4400 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0366] Intra prediction unit 4406 can perform intra prediction on the current video block. When intra prediction unit 4406 performs intra prediction on the current video block, intra prediction unit 4406 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0367] Residual generation unit 4407 can generate residual data for the current video block by subtracting the prediction video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of samples in the current video block.
[0368] In other examples, such as in skip mode, there can be no residual data for the current video block, and residual generation unit 4407 can not perform a subtraction operation.
[0369] Transform processing unit 4408 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0370] After transform processing unit 4408 generates transform coefficient video blocks associated with the current video block, quantization unit 4409 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0371] Inverse quantization unit 4410 and inverse transform unit 4411 can apply inverse quantization and inverse transforms, respectively, to the transform coefficient video blocks to reconstruct the residual video blocks from the transform coefficient video blocks. Reconstruct unit 4412 can add the reconstructed residual video blocks to corresponding samples from the one or more prediction video blocks generated by prediction unit 4402 to produce a reconstructed video block associated with the current block for storage in buffer 4413.
[0372] After the video block is reconstructed at the reconstruction unit 4412, an in-loop filtering operation can be performed to reduce video blockiness artifacts in the video block.
[0373] The entropy encoding unit 4414 can receive data from other functional components of the video encoder 4400. When the entropy encoding unit 4414 receives data, the entropy encoding unit 4414 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.
[0374] Figure 7 FIG. 45 is a block diagram illustrating an example of a video decoder 4500 that can be Figure 5 the system 4300 shown. The video decoder 4500 can be configured to perform any or all of the embodiments of the disclosure. In the example shown, the video decoder 4500 includes a number of functional components. The embodiments described in this disclosure can be shared among the various components of the video decoder 4500. In some examples, a processor can be configured to perform any or all of the embodiments described in this disclosure.
[0375] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra prediction unit 4503, an inverse quantization unit 4504, an inverse transformation unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 can perform a decoding process generally reciprocal to the encoding process described with respect to the video encoder 4400.
[0376] The entropy decoding unit 4501 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 4501 can decode the entropy encoded video data, and the motion compensation unit 4502 can determine motion information from the entropy decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 4502 can determine this information, for example, by performing AMVP and Merge modes.
[0377] The motion compensation unit 4502 can generate a motion compensated block, and can perform interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision can be included in a syntax element.
[0378] Motion compensation unit 4502 can calculate the interpolation for sub-integer pixels of the reference block using the interpolation filter used by video encoder 4400 during encoding of the video block. Motion compensation unit 4502 can determine the interpolation filter used by video encoder 4400 from the syntax information received and can use that interpolation filter to generate the prediction block.
[0379] Motion compensation unit 4502 can use some of the syntax information to determine the size of the blocks used to encode the frame(s) and / or slice(s) of the encoded video sequence, partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and lists of reference frames) for each inter-coded block, and other information to decode the encoded video sequence.
[0380] Intra prediction unit 4503 can use, for example, intra prediction modes received in the bitstream to form the prediction block from spatial neighboring blocks. Inverse quantization unit 4504 inverse quantizes, i.e., de-quantizes, quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501. Inverse transform unit 4505 applies an inverse transform.
[0381] Reconstruction unit 4506 can add the residual block to the corresponding prediction block generated by motion compensation unit 4502 or intra prediction unit 4503 to form a decoded block. If desired, a deblocking filter can also be used to filter the decoded block in order to remove blockiness artifacts. The decoded video blocks are then stored in buffer 4507, which supplies reference blocks for subsequent motion compensation / intra prediction and which also produces decoded video for presentation on a display device.
[0382] Figure 8 is a schematic diagram of an example encoder 4600. Encoder 4600 is suitable for implementing techniques of VVC. Encoder 4600 includes three in-loop filters, namely a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike DF 4602, which uses a pre-defined filter, SAO 4604 and ALF 4606 exploit original samples of the current picture by adding an offset and applying a finite impulse response (FIR) filter, respectively, and exploit coded side information to signal the offset and filter coefficients, to reduce the mean squared error between the original samples and the reconstructed samples. ALF 4606 is located at the last processing stage of each picture and can be viewed as a tool that attempts to capture and fix artifacts caused by previous stages.
[0383] The encoder 4600 also includes an intra prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive the input video. The intra prediction component 4608 is configured to perform intra prediction while the ME / MC component 4610 is configured to perform inter prediction with reference pictures obtained from a reference picture cache 4612. Residual blocks from either inter or intra prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients that are fed into an entropy encoding component 4618. The entropy encoding component 4618 entropy encodes the prediction results and quantized transform coefficients and transmits them to a video decoder (not shown). Quantized components output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 is capable of outputting the pictures to the DF 4602, SAO 4604, and ALF 4606 for filtering before these pictures are stored in the reference picture cache 4612.
[0384] Next, a list of some example preferred solutions is provided.
[0385] The following solutions show examples of embodiments discussed herein.
[0386] 1. A method of processing media data, comprising: determining to signal a syntax element to indicate whether a neural network post filter (NNPF) output picture is in full range when a color space of the NNPF output picture and a color space of a decoded picture or a color space of a cropped decoded output picture are different; and performing a conversion between visual media data and a bitstream based on the NNPF output picture.
[0387] 2. The method of solution 1, wherein one or more syntax elements are signaled to indicate a location of chroma samples when the NNPF output picture is a 4:2:0 color format.
[0388] 3. The method of solution 1 or 2, wherein a syntax element is signaled to specify a location of chroma samples of the NNPF output picture when a NNPF purpose indicates colorization and the NNPF output picture is a 4:2:0 color format.
[0389] 4. The method of any of solutions 1-3, wherein a set of aspect ratio related parameters are signaled in a NNPF supplemental enhancement information (SEI) message.
[0390] 5. The method of any of solutions 1-4, wherein a flag is signaled to indicate whether NNPF processing changes an aspect ratio.
[0391] 6. The method of any of solutions 1-5, wherein aspect_ratio_idc, sar_width, sar_height, or a combination thereof are included in the NNPF SEI message according to the value of the flag.
[0392] 7. The method of any of solutions 1-6, wherein when aspect_ratio_idc, sar_width, and sar_height are not present in the bitstream, the aspect ratio properties for the NNPF output picture are inferred to be the same as the aspect ratio properties for the decoded picture in the coded layer video sequence (CLVS) or the aspect ratio properties for the cropped decoded output picture.
[0393] 8. The method of any of solutions 1-7, wherein one or more syntax elements indicating the source scan type of the output of the NNPF process are conditionally signaled, and when the syntax elements indicating the source scan type are not present, the source scan type for the NNPF output picture is inferred to be the same as the source scan type for the decoded picture in the CLVS or the source scan type for the cropped decoded output picture.
[0394] 9. The method of any of solutions 1-8, wherein one or more syntax elements indicating the preferred display mechanism of the output of the NNPF process are conditionally signaled, and when the syntax elements indicating the preferred display mechanism are not present, the preferred display mechanism for the NNPF output picture is inferred to be the same as the preferred display mechanism for the decoded picture in the CLVS or the preferred display mechanism for the cropped decoded output picture.
[0395] 10. The method of any of solutions 1-9, wherein when NNPF is used, StrengthControlVal is set to a value equal to (SliceQpY + QpBdOffset) ÷ (63 + QpBdOffset), where SliceQpY is the SliceQpY of the first slice of currCodedPic.
[0396] 11. The method of any of solutions 1-10, wherein when NNPF is used, StrengthControlVal is set to a value equal to (SliceQpY + A) ÷ B, where SliceQpY is the SliceQpY of the first slice of currCodedPic, A is the maximum possible value of QpBdOffset, and B is the maximum possible value of SliceQpY.
[0397] 12. A method according to any of solutions 1-11, wherein StrengthControlVal is transmitted by signal in an NNPF SEI message or a Neural Network Post-Processing Filter Activation (NNPFA) SEI message.
[0398] 13. The method according to any one of solutions 1-12, wherein StrengthControlVal is transmitted via a signal within a specific range and converted into an input value range of the NNPF.
[0399] 14. A method according to any of solutions 1-13, wherein when the NNPF purpose indicates colorization, format change, or upsampling, the chroma matrix must be present in the output tensor.
[0400] 15. The method according to any of solutions 1-14, wherein when nnpfc_purpose&0x20 is not equal to 0, nnpfc_out_order_idc shall not be equal to 0.
[0401] 16. A method according to any of solutions 1-15, wherein signaling of luma fill values is skipped when the input tensor does not contain a luma matrix.
[0402] 17. A method according to any of solutions 1-16, wherein signaling of chroma fill values is skipped when the input tensor does not contain a chroma matrix.
[0403] 18. A device for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of solutions 1-17.
[0404] 19. A non-transitory computer-readable medium, comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, such that when the computer-executable instructions are executed by a processor, the video codec device performs a method according to any one of solutions 1-17.
[0405] 20. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method includes: when the color space of a neural network post-processing filter (NNPF) output picture and the color space of a decoded picture or a cropped decoded output picture are different, determining whether a syntax element is transmitted through a signal to indicate whether the NNPF output picture is within a full range; and generating a bitstream based on the determination.
[0406] 21. A method for storing a bitstream of a video, comprising: determining to signal a syntax element to indicate whether a neural network post filter (NNPF) output picture is in full range when a color space of the NNPF output picture and a color space of a decoded picture or a color space of a cropped decoded output picture are different; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0407] 22. A method, apparatus or system described in the disclosure.
[0408] In the described solutions, an encoder can conform to a format rule by producing an encoded representation according to the format rule. In the described solutions, a decoder can parse syntax elements in an encoded representation according to the format rule with known information of presence and absence of the syntax elements to produce a decoded video.
[0409] In the disclosure, the term “video processing” can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during a conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. For example, a bitstream representation of a current video block can correspond to a bit position in a bitstream defined by syntax or bits that are spread in different positions. For example, a macroblock can be encoded according to transformed and coded error residual values, and also using bits in headers and other fields in the bitstream. Furthermore, during the conversion, a decoder can parse the bitstream based on the determination, knowing that some fields can be present or absent, as described in the above solutions. Similarly, an encoder can determine to include or not include particular syntax fields, and generate a coded representation accordingly by including or excluding the syntax fields from the coded representation.
[0410] The disclosed and other solutions, examples, embodiments, modules and functional operations described in this disclosure can be realized in digital electronic circuitry, or in computer software, firmware, or hardware, including the structural equivalents of such disclosure as set forth herein, or in combinations of one or more of them. The disclosed and other embodiments can be realized as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also include, in addition to a hardware component, code that creates an execution environment for the associated computer program in the form of a firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. The propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus.
[0411] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or code portions). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.
[0412] The processes and logic flows described in this disclosure can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, and that apparatus can also be implemented as special purpose logic circuitry, e.g., an FPGA or an ASIC.
[0413] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0414] Although the present disclosure includes a number of details, these should not be construed as limitations on any subject matter or scope of protection that can be required to practice the subject matter. In the present disclosure, certain features that are described in the context of separate embodiments can also be implemented in combination with each other. Conversely, various features that are described in the context of a single embodiment can also be implemented separately from each other. Furthermore, although features can be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.
[0415] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such order, nor that all illustrated operations be performed, to achieve desirable results. Moreover, the separation of various system components in the embodiments described in this disclosure should not be understood as requiring such separation in all embodiments.
[0416] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this disclosure.
[0417] A first component is directly coupled to a second component when there are no intervening components between the first component and the second component other than a wire, trace, or other medium. A first component is indirectly coupled to a second component when there are intervening components between the first component and the second component other than a wire, trace, or other medium. The term “coupled” and variations thereof include both direct and indirect coupling. The use of the term “about” means a range of ±10% of the subsequent number, unless otherwise indicated.
[0418] While several embodiments have been provided in the present disclosure, it should be understood that the disclosed systems and methods might be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples are to be considered as illustrative and not restrictive, and the intention is not to limit the disclosure to the details given herein. For example, the various elements or components can be combined or integrated in another system or certain features can be omitted, without departing from the scope of the disclosure.
[0419] Also, techniques, systems, subsystems, and methods described and illustrated in the various embodiments as discrete or separate can be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of the present disclosure. Other items shown or discussed as separate from the other items or hierarchical from other items can be, in actuality, merged into a single item, and the terms “processor” and / or “module” as used herein can encompass both a standalone and / or integrated hardware-based component. Other examples of changes, substitutions, and alterations are ascertainable by one skilled in the art and can be made without departing from the spirit and scope disclosed herein.
Claims
1. A method for processing media data, comprising: When a color space of a neural network post-processing filter (NNPF) output picture is different from a color space of a decoded picture or a color space of a cropped decoded output picture, determining to signal a syntax element to indicate whether the NNPF output picture is within a full range; as well as Conversion between visual media data and a bitstream is performed based on the NNPF output picture.
2. The method according to claim 1, wherein When the NNPF output picture is in 4:2:0 color format, one or more syntax elements are signaled to indicate the positions of chroma samples.
3. The method according to claim 1 or 2, wherein: When the NNPF destination indicates colorization and the NNPF output picture is in 4:2:0 color format, syntax elements are signaled to specify the locations of the chroma samples of the NNPF output picture.
4. The method according to any one of claims 1 to 3, wherein When the NNPF purpose indicates a format change or upsampling and the NNPF output picture is in 4:2:0 color format, syntax elements are signaled to specify the locations of the chroma samples of the NNPF output picture.
5. The method according to any one of claims 1 to 4, wherein A set of aspect ratio related parameters are signaled in the NNPF Supplemental Enhancement Information (SEI) message.
6. The method according to any one of claims 1 to 5, wherein A flag is signaled to indicate whether the NNPF processing changes the aspect ratio.
7. The method according to any one of claims 1 to 6, wherein aspect_ratio_idc, sar_width, sar_height, or a combination thereof is included in the NNPF SEI message according to the value of the flag.
8. The method according to any one of claims 1 to 7, wherein When aspect_ratio_idc, sar_width, and sar_height are not present in the bitstream, the aspect ratio attributes for the NNPF output picture are inferred to be the same as the aspect ratio attributes for the decoded picture in the codec layer video sequence (CLVS) or the aspect ratio attributes for the cropped decoded output picture.
9. The method according to any one of claims 1 to 8, wherein One or more syntax elements indicating a source scan type of an output of the NNPF process are conditionally transmitted via a signal, and when the syntax element indicating the source scan type is not present, the source scan type for the NNPF output picture is presumed to be the same as the source scan type for the decoded picture in the CLVS or the source scan type for the cropped decoded output picture.
10. The method according to any one of claims 1 to 9, wherein One or more syntax elements indicating a preferred display mechanism for the output of the NNPF process are conditionally transmitted via a signal, and when the syntax element indicating the preferred display mechanism is not present, the preferred display mechanism for the NNPF output picture is presumed to be the same as the preferred display mechanism for the decoded picture in the CLVS or the preferred display mechanism for the cropped decoded output picture.
11. The method according to any one of claims 1 to 10, wherein When using NNPF, StrengthControlVal is set to a value equal to (SliceQpY+QpBdOffset)÷(63+QpBdOffset), where SliceQpY is the SliceQpY of the first slice of currCodedPic.
12. The method according to any one of claims 1 to 10, wherein When using NNPF, StrengthControlVal is set to a value equal to (SliceQpY+A)÷B, where SliceQpY is the SliceQpY of the first slice of currCodedPic, A is the maximum possible value of QpBdOffset, and B is the maximum possible value of SliceQpY.
13. The method according to claim 12, wherein: A equals 48, B equals 111.
14. The method according to any one of claims 1 to 13, wherein: StrengthControlVal is signaled in a NNPF SEI message, a Neural Network Post-Processing Filter Characteristics (NNPFC) SEI message, or a Neural Network Post-Processing Filter Activation (NNPFA) SEI message.
15. The method according to claim 14, wherein StrengthControlVal is signaled in the NNPFA SEI message to activate NNPF processing.
16. The method according to any one of claims 1 to 15, wherein StrengthControlVal is signaled within a specific range and converted into an input value range for the NNPF.
17. The method according to any one of claims 1 to 16, wherein: When the NNPF destination indicates colorization, format conversion, or upsampling, the chroma matrix must be present in the output tensor.
18. The method according to any one of claims 1 to 17, wherein When nnpfc_purpose&0x20 is not equal to 0, nnpfc_out_order_idc is not equal to 0.
19. The method according to any one of claims 1 to 18, wherein When the input tensor does not contain a luma matrix, signaling of luma padding values is skipped.
20. The method according to any one of claims 1 to 19, wherein When the input tensor does not contain a chroma matrix, signaling of chroma padding values is skipped.
21. A device for processing video data, comprising: processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-20.
22. A non-transitory computer-readable medium comprising a computer program product for use with a video encoding and decoding device, wherein: The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, such that when the computer-executable instructions are executed by a processor, the video coding device is caused to perform the method according to any one of claims 1 to 20.
23. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by a video processing apparatus, wherein The method comprises: When a color space of a neural network post-processing filter (NNPF) output picture is different from a color space of a decoded picture or a color space of a cropped decoded output picture, determining to signal a syntax element to indicate whether the NNPF output picture is within a full range; and A bitstream is generated based on the determination.
24. A method for storing a bitstream of a video, comprising: When a color space of a neural network post-processing filter (NNPF) output picture is different from a color space of a decoded picture or a color space of a cropped decoded output picture, determining to signal a syntax element to indicate whether the NNPF output picture is within a full range; generating a bitstream based on the determination; as well as The bitstream is stored in a non-transitory computer-readable recording medium.