Neural network post-processing filter video usability information related indication and miscellaneous in sei messages
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DOUYIN CO LTD
- Filing Date
- 2024-03-20
- Publication Date
- 2026-08-07
AI Technical Summary
随着能够接收和显示视频的连接用户设备的数量增加,对数字视频使用的带宽需求可能继续增长
Smart Images

Figure CN122535947A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority and benefit to U.S. Provisional Application No. 63 / 619,255, filed January 9, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to the generation, storage, and use of digital audio and video media information in file formats. Background Technology
[0004] Digital video accounts for the largest share of bandwidth used on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is likely to continue to grow. Summary of the Invention
[0005] The first aspect relates to a method for processing video data, comprising: determining whether the color space of the output image of a neural network post-processing filter (NNPF) is in the full range when the color space of the output image differs from that of the decoded image or the color space of the cropped decoded output image; and performing a conversion between visual media data and a bitstream based on the NNPF.
[0006] The second aspect relates to an apparatus for processing video data, comprising: a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform any of the aforementioned aspects.
[0007] The third aspect relates to a non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec device to perform the methods of any of the preceding aspects when executed by a processor.
[0008] The fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining, when the color space of a neural network post-processing filter (NNPF) output image differs from the color space of a decoded image or a cropped decoded output image, whether the NNPF output image is in the full range by signal transmission syntax elements; and generating a bitstream based on the determination.
[0009] The fifth aspect relates to a method for storing a bitstream of video, comprising: determining, when the color space of a neural network post-processing filter (NNPF) output image differs from the color space of a decoded image or the color space of a cropped decoded output image, whether the NNPF output image is in the full range by signaling a syntax element to indicate that the NNPF output image is in the full range; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] The sixth aspect relates to the methods, apparatus, or systems described in this disclosure.
[0011] For clarity, any of the embodiments described above may be combined with one or more other embodiments described above to create new embodiments within the scope of this disclosure.
[0012] These and other features will become clearer through the following detailed description of the embodiments with reference to the accompanying drawings and claims. Attached Figure Description
[0013] To gain a more complete understanding of this disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals denote like parts.
[0014] Figure 1 An example of deriving the luminance channel from the luminance component is shown.
[0015] Figure 2 This is a block diagram illustrating an example video processing system.
[0016] Figure 3 This is a block diagram of an example video processing device.
[0017] Figure 4 This is a flowchart of an example method for video processing.
[0018] Figure 5 This is a block diagram illustrating an example video codec system.
[0019] Figure 6 This is a block diagram showing an example encoder.
[0020] Figure 7 This is a block diagram showing an example decoder.
[0021] Figure 8 This is a schematic diagram of an example encoder. Detailed Implementation
[0022] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of embodiments, whether currently known or yet to be developed. This disclosure should not be limited in any way to the illustrative implementations, drawings, and embodiments described below, including the exemplary designs and implementations shown and described herein, but rather to modifications within the scope of the appended claims and their full equivalents.
[0023] Chapter headings are used in this disclosure for ease of understanding and not to limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter only. Furthermore, H.266 terminology is used in some descriptions merely for ease of understanding and not to limit the scope of the disclosed embodiments. Therefore, the embodiments described herein are also applicable to other video codec protocols and designs. In this disclosure, editorial modifications relative to the Multi-Functional Video Codec (VVC) specification are indicated in the text by bold italics (indicating deleted text) and bold text (indicating added text).
[0024] 1. Preliminary Discussion
[0025] This disclosure relates to image / video codec techniques. Specifically, this disclosure relates to Video Availability Information (VUI) related information transmitted or modified via Neural Network Post-Processing Filter (NNPF) messages; and miscellaneous information in NNPF SEI messages. This idea can be applied alone or in various combinations to video bitstreams encoded and decoded by any codec, such as the Multi-Function Video Codec (VVC) standard and / or the Multi-Function Supplemental Enhancement Information (SEI) message (VSEI) standard for encoding and decoding video bitstreams.
[0026] 2. Abbreviation
[0027] The following abbreviations may be used in this disclosure: Adaptive Parameter Set (APS), Access Unit (AU), Codec Layer Video Sequence (CLVS), Codec Layer Video Sequence Start (CLVSS), Cyclic Redundancy Check (CRC), Codec Video Sequence (CVS), Finite Impulse Response (FIR), Intra-Frame Random Access Point (IRAP), Network Abstraction Layer (NAL), Neural Network Post-Processing Filter (NNPF), Neural Network Post-Processing Filter Activation (NNPFA), Neural Network Post-Processing Filter Feature (NNPFC), Picture Parameter Set (PPS), Picture Unit (PU), Random Access Skip Preamble (RASL) Picture, Supplemental Enhancement Information (SEI), Stepped Temporal Sublayer Access (STSA), Uniform Resource Identifier (URI), Video Codec Layer (VCL), Multifunctional Supplemental Enhancement Information (VSEI) as described in Recommendation ITU-T H.274 | ISO / IEC 23002-7, Video Availability Information (VUI), and Multifunctional Video Codec (VVC) as described in Recommendation ITU-T H.266 | ISO / IEC 23090-3.
[0028] 3. Further discussion
[0029] 3.1 Video Coding and Decoding Standards
[0030] Video coding standards have evolved primarily through the development of standards by the International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) and the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). ITU-T produced the H.261 and H.263 standards, ISO / IEC produced the Moving Picture Experts Group (MPEG)-1 and MPEG-4 Vision standards, and the two organizations jointly produced the H.262 / MPEG-2 video standard and the H.264 / MPEG-4 Advanced Video Coding (AVC) standard and the H.265 / High Efficiency Video Coding (HEVC) standard[1]. Starting with H.262, video coding standards are based on a hybrid video coding architecture, which utilizes temporal prediction plus transform coding. In order to explore video coding technologies other than High Efficiency Video Coding (HEVC), the Joint Video Exploration Team (JVET) was established by the Video Coding Experts Group (VCEG) and the Moving Picture Experts Group (MPEG). In addition, JVET adopted some methods and incorporated them into a reference software called the Joint Exploration Model (JEM)[2]. When the Multi-Functional Video Coding (VVC) project was officially launched, JVET was later renamed the Joint Video Experts Group (JVET). VVC[3] is a coding standard that aims to reduce the bit rate by 50% compared to HEVC.
[0031] The Multi-Functional Video Coding (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) [3] and the related Multi-Functional Supplemental Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) [4] are designed for use in the widest range of applications, including simple uses such as television broadcasting, video conferencing or playback from storage media, as well as more advanced use cases such as adaptive bitrate streaming, video region extraction, synthesis and merging of content from multiple codec video bitstreams, multi-view video, scalable layered coding and decoding and viewport adaptive 360° immersive media.
[0032] The Basic Video Codec (EVC) standard (ISO / IEC 23094-1) is another video codec standard that MPEG is developing.
[0033] 3.2 General SEI messages and SEI messages in VVC and VSEI
[0034] SEI messages assist in processes related to decoding, display, or other purposes. However, SEI messages are not essential for constructing luma or chroma samples during the decoding process. Standard-compliant decoders do not need to process this information to achieve output order consistency. Some SEI messages are necessary for checking bitstream consistency and output timing decoder consistency. Other SEI messages are not necessary for checking bitstream consistency.
[0035] Appendix D of VVC specifies the syntax and semantics of the SEI message payload for some SEI messages, and specifies the use of SEI messages and VUI parameters with the syntax and semantics specified in ITU-T H.274 | ISO / IEC 23002-7.
[0036] 3.3 Signaling of Neural Network Post-processing Filters
[0037] JVET-AC2032[5] includes specifications for two SEI messages for signaling used in neural network post-processing filters, as shown below.
[0038] 8.28 Neural Network Post-Processing Filter Characteristics SEI Message
[0039] 8.28.1 Characteristics of Neural Network Post-Processing Filters and SEI Message Syntax
[0040] nn_post_filter_characteristics( payloadSize ) { descriptor nnpfc_purpose u(16) nnpfc_id ue(v) nnpfc_mode_idc ue(v) if( nnpfc_mode_idc == 1 ) { while( !byte_aligned( ) ) nnpfc_reserved_zero_bit_a u(1) nnpfc_tag_uri st(v) nnpfc_uri st(v) } nnpfc_property_present_flag u(1) if( nnpfc_property_present_flag ) { nnpfc_base_flag u(1) / * Input and output format * / nnpfc_num_input_pics_minus1 ue(v) if( ( nnpfc_purpose & 0x02 ) != 0 ) nnpfc_out_sub_c_flag u(1) if( ( nnpfc_purpose & 0x20 ) != 0 ) nnpfc_out_colour_format_idc u(2) if( ( nnpfc_purpose & 0x04 ) != 0 ) { nnpfc_pic_width_in_luma_samples ue(v) nnpfc_pic_height_in_luma_samples ue(v) } for( i = 0; i < nnpfc_num_input_pics_minus1; i++ ) if( ( nnpfc_purpose & 0x08 ) != 0 ) { nnpfc_interpolated_pics[ i ] for( i = 0; i <= nnpfc_num_input_pics_minus1; i++ ) ue(v) u(1) } nnpfc_input_pic_output_flag[ i ] u(1) nnpfc_component_last_flag nnpfc_inp_format_idc ue(v) if( nnpfc_inp_format_idc == 1 ) { nnpfc_inp_tensor_luma_bitdepth_minus8 ue(v) nnpfc_inp_tensor_chroma_bitdepth_minus8 } ue(v) nnpfc_inp_order_idc ue(v) nnpfc_auxiliary_inp_idc ue(v) u(1) nnpfc_separate_colour_description_present_flag if( nnpfc_separate_colour_description_present_flag ) { nnpfc_colour_primaries u(8) nnpfc_transfer_characteristics u(8) nnpfc_matrix_coeffs u(8) } nnpfc_out_format_idc ue(v) if( nnpfc_out_format_idc == 1 ) { nnpfc_out_tensor_luma_bitdepth_minus8 ue(v) nnpfc_out_tensor_chroma_bitdepth_minus8 ue(v) } nnpfc_out_order_idc ue(v) nnpfc_overlap ue(v) nnpfc_constant_patch_size_flag u(1) if( nnpfc_constant_patch_size_flag ) { nnpfc_patch_width_minus1 ue(v) nnpfc_patch_height_minus1 ue(v) } else { nnpfc_extended_patch_width_cd_delta_minus1 ue(v) nnpfc_extended_patch_height_cd_delta_minus1 ue(v) } nnpfc_padding_type ue(v) if( nnpfc_padding_type == 4 ) { nnpfc_luma_padding_val ue(v) nnpfc_cb_padding_val ue(v) nnpfc_cr_padding_val ue(v) } nnpfc_complexity_info_present_flag u(1) if( nnpfc_complexity_info_present_flag ) { nnpfc_parameter_type_idc u(2) if( nnpfc_parameter_type_idc != 2 ) nnpfc_log2_parameter_bit_length_minus3 u(2) nnpfc_num_parameters_idc u(6) nnpfc_num_kmac_operations_idc ue(v) nnpfc_total_kilobyte_size ue(v) } } / * ISO / IEC 15938-17 bitstream * / if( nnpfc_mode_idc == 0 ) { while( !byte_aligned( ) ) nnpfc_reserved_zero_bit_b u(1) for( i = 0; more_data_in_payload( ); i++ ) nnpfc_payload_byte[i] b(8) } }
[0041] 8.28.2 Characteristics of Neural Network Post-Processing Filters and Semantics of SEI Messages
[0042] The Neural Network Post-Processing Filter Characteristics (NNPFC) SEI message specifies the neural networks that can be used as post-processing filters. The use of a specified Neural Network Post-Processing Filter (NNPF) for a specific image is indicated by the Neural Network Post-Processing Filter Activation (NNPFA) SEI message.
[0043] To use this SEI message, the following variables need to be defined:
[0044] – Input the width and height of the image, in units of brightness samples, which are denoted as CroppedWidth and CroppedHeight in this article.
[0045] – The input image contains a luminance sample array CroppedYPic[idx] and chrominance sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] (if present), where the index idx ranges from 0 to numInputPics − 1 (inclusive), which are used as inputs to NNPF.
[0046] – BitDepth for the luminance sample array of the input image Y .
[0047] – BitDepth for the chroma sample array (if any) of the input image C .
[0048] – Chroma format indicator, referred to herein as ChromaFormatIdc, as described in sub-entry 7.3.
[0049] – When nnpfc_auxiliary_inp_idc equals 1, the filter strength control value StrengthControlVal must be a real number in the range of 0 to 1 (inclusive).
[0050] The input image at index 0 corresponds to the image whose NNPF is activated by the NNPFA SEI message defined by that NNPFC SEI message. Input images with indices i in the range of 1 to numInputPics − 1 (inclusive) precede the input image at index i − 1 in the output order.
[0051] When an input image with index 0 and nnpfc_purpose & 0x08 not equal to 0 is associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5, all input images are associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5 and the same value as fp_current_frame_is_frame0_flag.
[0052] The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc, as specified in Table 2.
[0053] Note 1 – More than one NNPFC SEI message can exist for the same image. When more than one NNPFC SEI message with different values of nnpfc_id exists or is activated for the same image, they can have the same or different values of nnpfc_purpose and nnpfc_mode_idc.
[0054] nnpfc_purpose indicates the purpose of NNPF, as specified in Table 20.
[0055] The value of nnpfc_purpose in this version of the bitstream must be in the range of 0 to 63 (inclusive). Values of nnpfc_purpose from 64 to 65535 (inclusive) are reserved for future use by ITU-T | ISO / IEC and should not exist in this version of the bitstream. Decoders conforming to this version of the document must ignore NNPFC SEI messages with nnpfc_purpose in the range of 64 to 65535 (inclusive).
[0056] Table 20 – Definition of nnpfc_purpose
[0057] value explain nnpfc_purpose == 0 It can be used depending on the application. nnpfc_purpose > 0 &&( nnpfc_purpose & 0x01 ) = =0 No general improvement in visual quality. (nnpfc_purpose & 0x01) != 0 With general visual quality improvement nnpfc_purpose > 0 &&( nnpfc_purpose & 0x02 ) = = 0 No chroma upsampling (from 4:2:0 chroma format to 4:2:2 or 4:4:4 chroma format, or from 4:2:2 chroma format to 4:4:4 chroma format). (nnpfc_purpose & 0x02) != 0 With chromaticity upsampling nnpfc_purpose > 0 &&( nnpfc_purpose & 0x04 ) = = 0 No resolution upsampling (increase width or height) (nnpfc_purpose & 0x04) != 0 With resolution upsampling nnpfc_purpose > 0 &&( nnpfc_purpose & 0x08 ) = = 0 No image rate upsampling (nnpfc_purpose & 0x08) != 0 With image rate upsampling nnpfc_purpose > 0 &&( nnpfc_purpose & 0x10 ) = = 0 No bit depth upsampling (increase luma bit depth or chroma bit depth) (nnpfc_purpose & 0x10) != 0 Bit depth upsampling nnpfc_purpose > 0 &&( nnpfc_purpose & 0x20 ) = = 0 No coloring (from 4:0:0 chroma format to 4:2:0, 4:2:2, or 4:4:4 chroma format) (nnpfc_purpose & 0x20) != 0 With color
[0058] Note 2 – When the reserved value of nnpfc_purpose is used by ITU-T | ISO / IEC in the future, the syntax of this SEI message can be extended using the following syntax elements, provided that nnpfc_purpose is equal to that value.
[0059] When ChromaFormatIdc equals 3, nnpfc_purpose & 0x02 must equal 0.
[0060] When ChromaFormatIdc or nnpfc_purpose & 0x02 is not equal to 0, nnpfc_purpose & 0x20 must be equal to 0.
[0061] The `nnpfc_id` contains an identifier that can be used to identify NNPF. The value of `nnpfc_id` must be in the range of 0 to 2^32 − 2 (inclusive). Values of `nnpfc_id` from 256 to 511 (inclusive) and from 231 to 232 − 2 (inclusive) are reserved for future use by ITU-T | ISO / IEC. Decoders conforming to this version of this document must ignore NNPFC SEI messages when they encounter `nnpfc_id` in the range of 256 to 511 (inclusive) or in the range of 231 to 232 − 2 (inclusive).
[0062] The following applies when the NNPFC SEI message is the first NNPFC SEI message in the decoding order with a specific nnpfc_id value within the current CLVS:
[0063] – This SEI message specifies the underlying NNPF.
[0064] – This SEI message applies to the currently decoded image and all subsequent decoded images of the current layer (in output order) until the current CLVS ends.
[0065] An nnpfc_mode_idc value of 0 indicates that the SEI message contains an ISO / IEC 15938-17 bitstream that specifies the underlying NNPF or an update relative to the underlying NNPF with the same nnpfc_id value.
[0066] When an NNPFC SEI message is the first NNPFC SEI message in the decoding order with a specific nnpfc_id value within the current CLVS, nnpfc_mode_idc equals 1, indicating that the underlying NNPF associated with the nnpfc_id value is a neural network identified by a URI indicated by nnpfc_uri, where nnpfc_uri has a format identified by the tag URI nnpfc_tag_uri.
[0067] When an NNPFC SEI message is neither the first NNPFC SEI message in the decoding order with a specific nnpfc_id value within the current CLVS, nor a duplicate of the first NNPFC SEI message in the decoding order with a specific nnpfc_id value within the current CLVS, an nnpfc_mode_idc of 1 indicates that the update relative to the underlying NNPF with the same nnpfc_id value is defined by the URI indicated by nnpfc_uri, where nnpfc_uri has a format identified by the tag URI nnpfc_tag_uri.
[0068] The value of nnpfc_mode_idc must be in the range of 0 to 1 (inclusive) in the bitstream conforming to this version of the document. Values of nnpfc_mode_idc from 2 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and should not exist in the bitstream conforming to this version of the document. Decoders conforming to this version of the document must ignore NNPFC SEI messages with nnpfc_mode_idc in the range of 2 to 255 (inclusive). Values of nnpfc_mode_idc greater than 255 should not exist in the bitstream conforming to this version of the document and are not reserved for future use.
[0069] When the SEI message is the first NNPFC SEI message in the decoding order with a specific nnpfc_id value within the current CLVS, the NNPF PostProcessingFilter() is assigned the same as the underlying NNPF.
[0070] When the SEI message is neither the first NNPFC SEI message in the decoding order with a specific nnpfc_id value within the current CLVS, nor a duplicate of the first NNPFC SEI message in the decoding order with a specific nnpfc_id value within the current CLVS, the NNPFPostProcessingFilter() is obtained by applying the update defined by the SEI message to the underlying NNPF.
[0071] Updates are not cumulative; instead, each update is applied to the underlying NNPF, which is defined by the first NNPFC SEI message in the decoding order that has a specific nnpfc_id value within the current CLVS.
[0072] nnpfc_reserved_zero_bit_a must be equal to 0 in the bitstream conforming to this version of the document. The decoder must ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_a is not equal to 0.
[0073] The nnpfc_tag_uri contains a tag URI, which has the syntax and semantics as specified in IETF RFC 4151, identifying the format and related information of the neural network used as the underlying NNPF or an update relative to the underlying NNPF with the same nnpfc_id value as specified by the nnpfc_uri.
[0074] Note 3 – nnpfc_tag_uri enables the unique identification of neural network data in the format specified by nnrpf_uri without requiring a central registry.
[0075] The nnpfc_tag_uri equal to “tag:iso.org,2023:15938-17” indicates that the neural network data identified by the nnpfc_uri conforms to ISO / IEC 15938-17.
[0076] nnpfc_uri contains a URI, which has the syntax and semantics as specified in IETF Internet Standard 66, identifying the neural network used as the underlying NNPF or an update relative to the underlying NNPF with the same nnpfc_id value.
[0077] The value of nnpfc_property_present_flag equal to 1 indicates the existence of syntax elements related to the filter's purpose, input format, output format, and complexity. The value of nnpfc_property_present_flag equal to 0 indicates the absence of syntax elements related to the filter's purpose, input format, output format, and complexity.
[0078] When the SEI message is the first NNPFC SEI message in the decoding order with a specific nnpfc_id value within the current CLVS, nnpfc_property_present_flag must be equal to 1.
[0079] When nnpfc_property_present_flag equals 0, all syntax elements that can exist only when nnpfc_property_present_flag equals 1 and for which no presumed value is specified are presumed to be equal to the corresponding syntax element in the NNPFC SEI message containing the underlying NNPF for which the SEI provides updates.
[0080] An nnpfc_base_flag value of 1 indicates that the SEI message specifies the underlying NNPF. An nnpf_base_flag value of 0 indicates that the SEI message specifies an update relative to the underlying NNPF. When it does not exist, the value of nnpfc_base_flag is presumed to be 0.
[0081] The following constraints apply to the value of nnpfc_base_flag:
[0082] – When the NNPFC SEI message is the first NNPFC SEI message in the decoding order with a specific nnpfc_id value within the current CLVS, the value of nnpfc_base_flag must be equal to 1.
[0083] – When NNPFC SEI message nnpfcB is not the first NNPFC SEI message in the decoding order with a specific nnpfc_id value within the current CLVS, and the value of nnpfc_base_flag is equal to 1, the NNPFC SEI message must be a duplicate of the first NNPFC SEI message nnpfcA in the decoding order with the same nnpfc_id, that is, the payload content of nnpfcB must be the same as that of nnpfcA.
[0084] The following applies when the NNPFC SEI message is neither the first NNPFC SEI message in the decoding order with a specific nnpfc_id value within the current CLVS, nor a duplicate of the first NNPFC SEI message with that specific nnpfc_id:
[0085] – This SEI message defines an update to the underlying NNPF that comes first in the decoding order, relative to the nnpfc_id value.
[0086] – This SEI message applies to the current decoded picture and all subsequent decoded pictures of the current layer (in output order) until the end of the current CLVS, or until, but not including, the decoded picture within the current CLVS that is after the current decoded picture in output order and is associated with a subsequent NNPFC SEI message in decoding order that has that particular nnpfc_id value within the current CLVS, whichever is earlier.
[0087] The following constraints apply when the NNPFC SEI message nnpfcCurr is not the first NNPFC SEI message in the decoding order with a specific nnpfc_id value within the current CLVS, is not a duplicate of the first NNPFC SEI message with that specific nnpfc_id (i.e., the value of nnpfc_base_flag is equal to 0), and the value of nnpfc_property_present_flag is equal to 1:
[0088] – The value of nnpfc_purpose in the NNPFC SEI message must be the same as the value of nnpfc_purpose in the first NNPFC SEI message in the decoding order that has that particular nnpfc_id value within the current CLVS.
[0089] – The value of the syntax element in the NNPFC SEI message that is after nnpfc_base_flag and before nnpfc_complexity_info_present_flag in the decoding order must be the same as the value of the corresponding syntax element in the first NNPFC SEI message in the decoding order that has that particular nnpfc_id value within the current CLVS.
[0090] – nnpfc_complexity_info_present_flag must be equal to 0, or in the first NNPFC SEI message (hereinafter referred to as nnpfcBase) with that particular nnpfc_id value in the current CLVS, nnpfc_complexity_info_present_flag must be equal to 1, and all of the following apply:
[0091] – The nnpfc_parameter_parameter_type_idc in nnpfcCurr must be equal to the nnpfc_parameter_parameter_type_idc in nnpfcBase.
[0092] – nnpfc_log2_parameter_bit_length_minus3 in nnpfcCurr (if it exists) must be less than or equal to nnpfc_log2_parameter_bit_length_minus3 in nnpfcBase.
[0093] – If nnpfc_num_parameters_idc in nnpfcBase is equal to 0, then nnpfc_num_parameters_idc in nnpfcCurr must be equal to 0.
[0094] Otherwise (nnpfc_num_parameters_idc in nnpfcBase is greater than 0), nnpfc_num_parameters_idc in nnpfcCurr must be greater than 0 and less than or equal to nnpfc_num_parameters_idc in nnpfcBase.
[0095] – If nnpfc_num_kmac_operations_idc in nnpfcBase is equal to 0, then nnpfc_num_kmac_operations_idc in nnpfcCurr must be equal to 0.
[0096] – Otherwise (nnpfc_num_kmac_operations_idc in nnpfcBase is greater than 0), nnpfc_num_kmac_operations_idc in nnpfcCurr must be greater than 0 and less than or equal to nnpfc_num_kmac_operations_idc in nnpfcBase.
[0097] – If nnpfc_total_kilobyte_size in nnpfcBase is equal to 0, then nnpfc_total_kilobyte_size in nnpfcCurr must be equal to 0.
[0098] – Otherwise (nnpfc_total_kilobyte_size in nnpfcBase is greater than 0), nnpfc_total_kilobyte_size in nnpfcCurr must be greater than 0 and less than or equal to nnpfc_total_kilobyte_size in nnpfcBase.
[0099] When `nnpfc_purpose & 0x02` is not equal to 0, `nnpfc_out_sub_c_flag` specifies the values of variables `outSubWidthC` and `outSubHeightC`. `nnpfc_out_sub_c_flag` equal to 1 specifies that `outSubWidthC` and `outSubHeightC` are both equal to 1. `nnpfc_out_sub_c_flag` equal to 0 specifies that `outSubWidthC` is equal to 2 and `outSubHeightC` is equal to 1. When `ChromaFormatIdc` is equal to 2 and `nnpfc_out_sub_c_flag` exists, the value of `nnpfc_out_sub_c_flag` must be equal to 1.
[0100] When `nnpfc_purpose & 0x20` is not equal to 0, `nnpfc_out_colour_format_idc` specifies the color format of the NNPF output, thus defining the values of the variables `outSubWidthC` and `outSubHeightC`. `nnpfc_out_colour_format_idc` equal to 1 specifies that the NNPF output color format is 4:2:0, and both `outSubWidthC` and `outSubHeightC` are equal to 2. `nnpfc_out_colour_format_idc` equal to 2 specifies that the NNPF output color format is 4:2:2, and both `outSubWidthC` and `outSubHeightC` are equal to 1. `nnpfc_out_colour_format_idc` equal to 3 specifies that the NNPF output color format is 4:2:4, and both `outSubWidthC` and `outSubHeightC` are equal to 1. The value of `nnpfc_out_colour_format_idc` should not be equal to 0.
[0101] When both nnpfc_purpose & 0x02 and nnpfc_purpose & 0x20 are equal to 0, outSubWidthC and outSubHeightC are presumed to be equal to SubWidthC and SubHeightC, respectively.
[0102] `nnpfc_pic_width_in_luma_samples` and `nnpfc_pic_height_in_luma_samples` specify the width and height of the luminance sample array of the image generated by applying the NNPF identified by `nnpfc_id` to the cropped decoded output image, respectively. When `nnpfc_pic_width_in_luma_samples` and `nnpfc_pic_height_in_luma_samples` are not present, they are presumed to be equal to `CroppedWidth` and `CroppedHeight`, respectively. The value of `nnpfc_pic_width_in_luma_samples` must be in the range from `CroppedWidth` to `CroppedWidth * 16 – 1` (inclusive). The value of `nnpfc_pic_height_in_luma_samples` must be in the range from `CroppedHeight` to `CroppedHeight * 16 – 1` (inclusive).
[0103] The increment of 1 in `nnpfc_num_input_pics_minus1` specifies the number of decoded output images used as input to NNPF. The value of `nnpfc_num_input_pics_minus1` must be in the range of 0 to 63 (inclusive). When `nnpfc_purpose & 0x08` is not equal to 0, the value of `nnpfc_num_input_pics_minus1` must be greater than 0.
[0104] `nnpfc_interpolated_pics[i]` specifies the number of interpolated pictures generated by NNPF between the i-th picture and the (i+1)-th picture used as input to NNPF. The value of `nnpfc_interpolated_pics[i]` must be in the range of 0 to 63 (inclusive). The value of `nnpfc_interpolated_pics[i]` must be greater than 0 for at least one `i` in the range of 0 to `nnpfc_num_input_pics_minus1 − 1` (inclusive).
[0105] `nnpfc_input_pic_output_flag[i]` equal to 1 indicates that NNPF generates the corresponding output image for the i-th input image. `nnpfc_input_pic_output_flag[i]` equal to 0 indicates that NNPF does not generate the corresponding output image for the i-th input image.
[0106] The variables numInputPics, which specify the number of images used as input to NNPF, and numOutputPics, which specify the total number of images generated by NNPF, are derived as follows:
[0107] numInputPics = nnpfc_num_input_pics_minus1 + 1
[0108] if( ( nnpfc_purpose & 0x08 ) != 0 ) {
[0109] for( i = 0, numOutputPics = 0; i < numInputPics; i++ )
[0110] if( nnpfc_input_pic_output_flag[ i ] )
[0111] numOutputPics++
[0112] for( i = 0; i <= numInputPics − 2; i++ )(76)
[0113] numOutputPics += nnpfc_interpolated_pics[ i ]
[0114] } else
[0115] numOutputPics = 1
[0116] `nnpfc_component_last_flag` equal to 1 indicates that the last dimension in the input tensor `inputTensor` and the output tensor `outputTensor` generated by `NNPF` is used for the current channel. `nnpfc_component_last_flag` equal to 0 indicates that the third dimension in the input tensor `inputTensor` and the output tensor `outputTensor` generated by `NNPF` is used for the current channel.
[0117] Note 4 – The first dimension in both the input and output tensors is used for batch indexing, which is a practice in some neural network frameworks. Although the formula in the semantics of this SEI message uses a batch size corresponding to a batch index of 0, the batch size used as input for neural network inference is determined by the post-processing implementation.
[0118] Note 5 – For example, when nnpfc_inp_order_idc equals 3 and nnpfc_auxiliary_inp_idc equals 1, the input tensor has 7 channels, including four luminance matrices, two chrominance matrices, and one auxiliary input matrix. In this case, the procedure DeriveInputTensors() will derive each of these 7 channels of the input tensor one by one, and when processing a particular channel of these channels, that channel is referred to as the current channel during processing.
[0119] `nnpfc_inp_format_idc` specifies the method for converting the sample values of the cropped decoded output image into NNPF input values. When `nnpfc_inp_format_idc` equals 0, the NNPF input values are real numbers, and the functions `InpY()` and `InpC()` are defined as follows:
[0120] InpY( x ) = x ÷ ( ( 1 << BitDepth Y ) − 1 ) (77)
[0121] InpC( x )= x ÷ ( ( 1 << BitDepth C ) − 1 ) (78)
[0122] When nnpfc_inp_format_idc equals 1, the input values of NNPF are unsigned integers, and the functions InpY() and InpC() are defined as follows:
[0123] shiftY = BitDepthY − inpTensorBitDepthY
[0124] if( inpTensorBitDepthY >= BitDepthY)
[0125] InpY( x ) = x << ( inpTensorBitDepthY − BitDepthY ) (79)
[0126] else
[0127] InpY( x ) = Clip3(0, ( 1 << inpTensorBitDepthY ) − 1, ( x + ( 1 << (shiftY
[0128] -1))) >> shiftY)
[0129] shiftC = BitDepthC − inpTensorBitDepthC
[0130] if( inpTensorBitDepthC >= BitDepthC )
[0131] InpC( x ) = x << ( inpTensorBitDepthC − BitDepthC ) (80)
[0132] else
[0133] InpC( x ) = Clip3(0, ( 1 << inpTensorBitDepthC ) − 1, ( x + ( 1 << (shiftC
[0134] - 1 ) ) ) >> shiftC )
[0135] The variable inpTensorBitDepthY is inferred from the syntax element nnpfc_inp_tensor_luma_bitdepth_minus8, as specified below. The variable inpTensorBitDepthC is inferred from the syntax element nnpfc_inp_tensor_chroma_bitdepth_minus8, as specified below.
[0136] Values greater than 1 for nnpfc_inp_format_idc are reserved for future ITU-T | ISO / IEC specifications and should not be present in bitstreams conforming to this version of the document. Decoders conforming to this version of the document must ignore NNPFC SEI messages containing reserved values for nnpfc_inp_format_idc.
[0137] `nnpfc_inp_tensor_luma_bitdepth_minus8` plus 8 specifies the bit depth of the luminance sample values in the input integer tensor. The value of `inpTensorBitDepthY` is derived as follows:
[0138] inpTensorBitDepth Y = nnpfc_inp_tensor_luma_bitdepth_minus8 + 8 (81)
[0139] The value of nnpfc_inp_tensor_luma_bitdepth_minus8 must be in the range of 0 to 24 (inclusive), which is a requirement for bitstream consistency.
[0140] `nnpfc_inp_tensor_chroma_bitdepth_minus8` plus 8 specifies the bit depth of the chroma sample values in the input integer tensor. The value of `inpTensorBitDepthC` is derived as follows:
[0141] inpTensorBitDepth C = nnpfc_inp_tensor_chroma_bitdepth_minus8 + 8 (82)
[0142] The value of nnpfc_inp_tensor_chroma_bitdepth_minus8 must be in the range of 0 to 24 (inclusive), which is a requirement for bitstream consistency.
[0143] nnpfc_inp_order_idc indicates a method for ordering the sample arrays of the cropped decoded output image into one of the input images for NNPF.
[0144] The value of nnpfc_inp_order_idc must be in the range of 0 to 3 (inclusive) in the bitstream conforming to this version of the document. Values of nnpfc_inp_order_idc from 4 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and should not exist in the bitstream conforming to this version of the document. Decoders conforming to this version of the document must ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 4 to 255 (inclusive). Values of nnpfc_inp_order_idc greater than 255 should not exist in the bitstream conforming to this version of the document and are not reserved for future use.
[0145] When ChromaFormatIdc is not equal to 1, nnpfc_inp_order_idc must not be equal to 3.
[0146] Table 21 contains informative descriptions of the nnpfc_inp_order_idc values.
[0147] Table 21 – Description of nnpfc_inp_order_idc values
[0148] nnpfc_inp_order_idc describe 0 If nnpfc_auxiliary_inp_idc equals 0, then a luminance matrix exists in the input tensor for each input image, and the number of channels is 1. Otherwise, when nnpfc_auxiliary_inp_idc equals 1, a luminance matrix and an auxiliary input matrix exist, and the number of channels is 2. 1 If nnpfc_auxiliary_inp_idc equals 0, then two chroma matrices exist in the input tensor, and the number of channels is 2. Otherwise, when nnpfc_auxiliary_inp_idc equals 1, two chroma matrices and one auxiliary input matrix exist, and the number of channels is 3. 2 If nnpfc_auxiliary_inp_idc equals 0, then a luminance matrix and two chrominance matrices exist in the input tensor, and the number of channels is 3. Otherwise, when nnpfc_auxiliary_inp_idc equals 1, a luminance matrix, two chrominance matrices, and an auxiliary input matrix exist, and the number of channels is 4. 3 If nnpfc_auxiliary_inp_idc equals 0, then four luminance matrices and two chrominance matrices exist in the input tensor, and the number of channels is 6. Otherwise, when nnpfc_auxiliary_inp_idc equals 1, four luminance matrices, two chrominance matrices, and one auxiliary input matrix exist in the input tensor, and the number of channels is 7. The luminance channels are derived in an interleaved manner as shown in Figure 12. nnpfc_inp_order_idc can only be used when the input chrominance format is 4:2:0. 4..255 reserve
[0149] Figure 1 An example of deriving the luminance channel from the luminance component is shown.
[0150] Small blocks are rectangular arrays of samples from the components of an image (e.g., luminance or chrominance components).
[0151] A value greater than 0 for nnpfc_auxiliary_inp_idc indicates that the auxiliary input data exists in the input tensor of NNPF. A value equal to 0 for nnpfc_auxiliary_inp_idc indicates that the auxiliary input data does not exist in the input tensor. A value equal to 1 for nnpfc_auxiliary_inp_idc indicates that the auxiliary input data is derived according to Equation 84.
[0152] The value of nnpfc_auxiliary_inp_idc must be in the range of 0 to 1 (inclusive) in the bitstream conforming to this version of the document. Values of nnpfc_inp_order_idc from 2 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and should not exist in the bitstream conforming to this version of the document. Decoders conforming to this version of the document must ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 2 to 255 (inclusive). Values of nnpfc_inp_order_idc greater than 255 should not exist in the bitstream conforming to this version of the document and are not reserved for future use.
[0153] When nnpfc_auxiliary_inp_idc equals 1, the variable strengthControlScaledVal is derived as follows:
[0154] if(nnpfc_inp_format_idc = 1)
[0155] strengthControlScaledVal = Floor ( StrengthControlVal * ( ( 1 < <inpTensorBitDepthY ) − 1 ) ) (83)
[0156] else
[0157] strengthControlScaledVal = StrengthControlVal
[0158] The procedure DeriveInputTensors() is used to derive an input tensor inputTensor for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft of a block of samples included in the input tensor. The procedure DeriveInputTensors() is defined as follows:
[0159] for( i = 0; i < numInputPics; i++ ) {if( nnpfc_inp_order_idc = = 0 )for( yP = −nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)for( xP =−nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {inpVal = InpY(InpSampleVal( cTop + yP, cLeft + xP, CroppedHeight,CroppedWidth, CroppedYPic[ i ] ) )yPovlp = yP + nnpfc_overlapxPovlp = xP + nnpfc_overlapif( !nnpfc_component_last_flag )inputTensor[ 0 ][ i ][ 0 ][ yPovlp ][ xPovlp ] =inpValelseinputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 0 ] = inpValif( nnpfc_auxiliary_inp_idc = = 1 )if( !nnpfc_component_last_flag )inputTensor[ 0 ][ i][ 1 ][ yPovlp ][ xPovlp ] = strengthControlScaledValelseinputTensor[ 0 ][ i][ yPovlp ][ xPovlp ][ 1 ] = strengthControlScaledVal}else if( nnpfc_inp_order_idc = = 1 )(84)for( yP = −nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)for( xP = −nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap;xP++ ) {inpCbVal = InpC( InpSampleVal( cTop + yP, cLeft + xP, CroppedHeight / SubHeightC,CroppedWidth / SubWidthC, CroppedCbPic[ i ] ) )inpCrVal = InpC(InpSampleVal( cTop + yP, cLeft + xP, CroppedHeight / SubHeightC,CroppedWidth / SubWidthC, CroppedCrPic[ i ] ) )yPovlp = yP + nnpfc_overlapxPovlp = xP +nnpfc_overlapif( !nnpfc_component_last_flag ) {inputTensor[ 0 ][ i ][ 0 ][yPovlp ][ xPovlp ] = inpCbValinputTensor[ 0 ][ i ][ 1 ][ yPovlp ][ xPovlp ] =inpCrVal} else {inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 0 ] =inpCbValinputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 1 ] = inpCrVal}if( nnpfc_auxiliary_inp_idc = = 1 )if( !nnpfc_component_last_flag )inputTensor[ 0 ][ i][ 2 ][ yPovlp ][ xPovlp ] = strengthControlScaledValelseinputTensor[ 0 ][ i][ yPovlp ][ xPovlp ][ 2 ] = strengthControlScaledVal}else if( nnpfc_inp_order_idc = = 2 )for( yP = −nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)for( xP = −nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap;xP++ ) {yY = cTop + yPxY = cLeft + xPyC = yY / SubHeightCxC = xY / SubWidthCinpYVal = InpY( InpSampleVal( yY, xY,CroppedHeight,CroppedWidth,CroppedYPic[ i ] ) )inpCbVal = InpC( InpSampleVal( yC, xC, CroppedHeight / SubHeightC,CroppedWidth / SubWidthC, CroppedCbPic[ i ] ) )inpCrVal = InpC(InpSampleVal( yC, xC, CroppedHeight / SubHeightC,CroppedWidth / SubWidthC,CroppedCrPic[ i ] ) )yPovlp = yP + nnpfc_overlapxPovlp = xP + nnpfc_overlapif( !nnpfc_component_last_flag ) {inputTensor[ 0 ][ i ][ 0 ][ yPovlp ][ xPovlp] = inpYValinputTensor[ 0 ][ i ][ 1 ][ yPovlp ][ xPovlp ] =inpCbValinputTensor[ 0 ][ i ][ 2 ][ yPovlp ][ xPovlp ] = inpCrVal} else{inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 0 ] = inpYValinputTensor[ 0 ][ i][ yPovlp ][ xPovlp ][ 1 ] = inpCbValinputTensor[ 0 ][ i ][ yPovlp ][ xPovlp][ 2 ] = inpCrVal}if( nnpfc_auxiliary_inp_idc = = 1 )if( !nnpfc_component_last_flag )inputTensor[ 0 ][ i ][ 3 ][ yPovlp ][ xPovlp ] = strengthControlScaledValelseinputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 3 ] =strengthControlScaledVal}else if( nnpfc_inp_order_idc = = 3 )for( yP = −nnpfc_overlap; yP < inpPatchHeight+ nnpfc_overlap; yP++)for( xP = −nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {yTL = cTop + yP * 2xTL =cLeft + xP * 2yBR = yTL + 1xBR = xTL + 1yC = cTop / 2 + yPxC = cLeft / 2 +xPinpTLVal = InpY( InpSampleVal( yTL, xTL, CroppedHeight,CroppedWidth,CroppedYPic[ i ] ) )inpTRVal = InpY( InpSampleVal( yTL, xBR, CroppedHeight,CroppedWidth, CroppedYPic[ i ] ) )inpBLVal = InpY( InpSampleVal( yBR, xTL,CroppedHeight,CroppedWidth, CroppedYPic[ i ] ) )inpBRVal = InpY( InpSampleVal( yBR, xBR, CroppedHeight,CroppedWidth, CroppedYPic[ i ] ) )inpCbVal = InpC(InpSampleVal( yC, xC, CroppedHeight / 2,CroppedWidth / 2, CroppedCbPic[ i ] ))inpCrVal = InpC( InpSampleVal( yC, xC, CroppedHeight / 2,CroppedWidth / 2,CroppedCrPic[ i ] ) )yPovlp = yP + nnpfc_overlapxPovlp = xP + nnpfc_overlapif( !nnpfc_component_last_flag ) {inputTensor[ 0 ][ i ][ 0 ][ yPovlp ][ xPovlp] = inpTLValinputTensor[ 0 ][ i ][ 1 ][ yPovlp ][ xPovlp ] =inpTRValinputTensor[ 0 ][ i ][ 2 ][ yPovlp ][ xPovlp ] =inpBLValinputTensor[ 0 ][ i ][ 3 ][ yPovlp ][ xPovlp ] = inpBRValinputTensor[ 0 ][ i ][ 4 ][yPovlp ][ xPovlp ] = inpCbValinputTensor[ 0 ][ i ][ 5 ][ yPovlp ][ xPovlp ] =inpCrVal} else {inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 0 ] =inpTLValinputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 1 ] = inpTRValinputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 2 ] = inpBLValinputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 3 ] = inpBRValinputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 4 ] =inpCbValinputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 5 ] = inpCrVal}if( nnpfc_auxiliary_inp_idc = = 1 )if( !nnpfc_component_last_flag )inputTensor[ 0 ][i ][ 6 ][ yPovlp ][ xPovlp ] = strengthControlScaledValelseinputTensor[ 0 ][i ][ yPovlp ][ xPovlp ][ 6 ] = strengthControlScaledVal}}
[0160] `nnpfc_separate_colour_description_present_flag` equal to 1 indicates that the SEI message syntax structure specifies different combinations of color primaries, transmission characteristics, and matrix coefficients for the image generated by NNPF. `nnpfc_separate_colour_description_present_flag` equal to 0 indicates that the combination of color primaries, transmission characteristics, and matrix coefficients for the image generated by NNPF is the same as indicated in the CLVS VUI parameters.
[0161] nnpfc_colour_primaries has the same semantics as the vui_colour_primaries syntax element specified in sub-entry 7.3, except as follows:
[0162] – nnpfc_colour_primaries specifies the primary color of the image generated by NNPF as defined in the application SEI message, rather than the primary color used for CLVS.
[0163] – When nnpfc_colour_primaries does not exist in the NNPFC SEI message, the value of nnpfc_colour_primaries is presumed to be equal to vui_colour_primaries.
[0164] nnpfc_transfer_characteristics has the same semantics as the vui_transfer_characteristics syntax element specified in sub-entry 7.3, except as follows:
[0165] – nnpfc_transfer_characteristics specifies the transfer characteristics of images generated by NNPF as defined in the SEI message, rather than the transfer characteristics used for CLVS.
[0166] – When nnpfc_transfer_characteristics does not exist in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is presumed to be equal to vui_transfer_characteristics.
[0167] nnpfc_matrix_coeffs has the same semantics as the vui_matrix_coeffs syntax element specified in sub-entry 7.3, except as follows:
[0168] – nnpfc_matrix_coeffs specifies the matrix coefficients of the image generated by the NNPF specified in the application SEI message, instead of the matrix coefficients used for CLVS.
[0169] – When nnpfc_matrix_coeffs does not exist in the NNPFC SEI message, the value of nnpfc_matrix_coeffs is presumed to be equal to vui_matrix_coeffs.
[0170] The allowed values for – nnpfc_matrix_coeffs are not constrained by the chroma format of the decoded video image as indicated by the value of ChromaFormatIdc, which is the semantics of the VUI parameter.
[0171] – When nnpfc_matrix_coeffs equals 0, nnpfc_out_order_idc should not equal 1 or 3.
[0172] An nnpfc_out_format_idc value of 0 indicates that the sample values output by NNPF are real numbers, where the range of values from 0 to 1 (inclusive) is linearly mapped to the range of unsigned integer values from 0 to (1 << bitDepth) – 1 (inclusive) for any desired bit depth bitDepth in subsequent post-processing or display.
[0173] An nnpfc_out_format_idc value of 1 indicates that the luminance sample values output by NNPF are unsigned integers in the range of 0 to (1 << (nnpfc_out_tensor_luma_bitdepth_minus8 + 8)) − 1 (inclusive), and the chrominance sample values output by NNPF are unsigned integers in the range of 0 to (1 << (nnpfc_out_tensor_chroma_bitdepth_minus8 + 8)) − 1 (inclusive).
[0174] Values greater than 1 for nnpfc_out_format_idc are reserved for future ITU-T | ISO / IEC specifications and should not be present in bitstreams conforming to this version of the document. Decoders conforming to this version of the document must ignore NNPFC SEI messages containing reserved values for nnpfc_out_format_idc.
[0175] The value of nnpfc_out_tensor_luma_bitdepth_minus8 plus 8 specifies the bit depth of the luminance sample values in the output integer tensor. The value of nnpfc_out_tensor_luma_bitdepth_minus8 must be in the range of 0 to 24 (inclusive).
[0176] The value of nnpfc_out_tensor_chroma_bitdepth_minus8 plus 8 specifies the bit depth of the chroma sample values in the output integer tensor. The value of nnpfc_out_tensor_chroma_bitdepth_minus8 must be in the range of 0 to 24 (inclusive).
[0177] When nnpfc_purpose & 0x10 is not equal to 0, the value of nnpfc_out_format_idc must be equal to 1, and at least one of the following conditions must be true:
[0178] – nnpfc_out_tensor_luma_bitdepth_minus8 + 8 is greater than BitDepth Y .
[0179] – nnpfc_out_tensor_chroma_bitdepth_minus8 + 8 is greater than BitDepth C .
[0180] nnpfc_out_order_idc indicates the output order of samples generated by NNPF.
[0181] The value of nnpfc_out_order_idc must be in the range of 0 to 3 (inclusive) in the bitstream conforming to this version of the document. Values of nnpfc_out_order_idc from 4 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and should not exist in the bitstream conforming to this version of the document. Decoders conforming to this version of the document must ignore NNPFC SEI messages with nnpfc_out_order_idc in the range of 4 to 255 (inclusive). Values of nnpfc_out_order_idc greater than 255 should not exist in the bitstream conforming to this version of the document and are not reserved for future use.
[0182] When nnpfc_purpose & 0x02 is not equal to 0, nnpfc_out_order_idc should not be equal to 3.
[0183] Table 22 contains informative descriptions of the nnpfc_out_order_idc values.
[0184] Table 22 – Description of nnpfc_out_order_idc values
[0185] nnpfc_out_order_idc describe 0 The output tensor contains only the brightness matrix, so the number of channels is 1. 1 The output tensor contains only the chroma matrix, so the number of channels is 2. 2 The output tensor contains luminance and chrominance matrices, therefore the number of channels is 3. 3 The output tensor contains four luminance matrices and two chrominance matrices, therefore the number of channels is 6. The nnpfc_out_order_idc can only be used when the output chrominance format is 4:2:0. 4..255 reserve
[0186] The procedure StoreOutputTensors() is used to derive sample values from the output tensor outputTensor, which is the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic. The output tensor outputTensor is given vertical sample coordinates cTop and horizontal sample coordinates cLeft, which specify the top-left sample position of a small block of samples included in the input tensor. The procedure StoreOutputTensors() is defined as follows:
[0187] for( i = 0; i < numOutputPics; i++ ) {if( nnpfc_out_order_idc = = 0)for( yP = 0; yP < outPatchHeight; yP++)for( xP = 0; xP < outPatchWidth; xP++) {yY = cTop * outPatchHeight / inpPatchHeight + yPxY = cLeft * outPatchWidth / inpPatchWidth + xPif ( yY < nnpfc_pic_height_in_luma_samples && xY <nnpfc_pic_width_in_luma_samples )if( !nnpfc_component_last_flag )FilteredYPic[ i ][ xY ][yY ] = outputTensor[ 0 ][ i ][ 0 ][ yP ][ xP ]elseFilteredYPic[ i][ xY ][ yY ] = outputTensor[ 0 ][ i ][ yP ][ xP ][ 0 ]}else if( nnpfc_out_order_idc = = 1 )(85)for( yP = 0; yP < outPatchCHeight; yP++)for( xP = 0;xP < outPatchCWidth; xP++ ) {xSrc = cLeft * horCScaling + xPySrc = cTop *verCScaling + yPif ( ySrc < nnpfc_pic_height_in_luma_samples / outSubHeightC&&xSrc < nnpfc_pic_width_in_luma_samples / outSubWidthC )if( !nnpfc_component_last_flag ) {FilteredCbPic[ i ][ xSrc ][ ySrc ] = outputTensor[ 0 ][ i ][ 0 ][ yP ][ xP ]FilteredCrPic[ i ][ xSrc ][ ySrc ] = outputTensor[ 0 ][ i ][ 1 ][ yP ][ xP ]} else{FilteredCbPic[ i ][ xSrc ][ ySrc ] =outputTensor[ 0 ][ i ][ yP ][ xP ][ 0 ]FilteredCrPic[ i ][ xSrc ][ ySrc ] =outputTensor[ 0 ][ i ][ yP ][ xP ][ 1 ]}}else if( nnpfc_out_order_idc = = 2)for( yP = 0; yP < outPatchHeight; yP++)for( xP = 0; xP < outPatchWidth; xP++) {yY = cTop * outPatchHeight / inpPatchHeight + yPxY = cLeft * outPatchWidth / inpPatchWidth + xPyC = yY / outSubHeightC xC = xY / outSubWidthC yPc = ( yP / outSubHeightC ) * outSubHeightCxPc = ( xP / outSubWidthC ) * outSubWidthCif( yY < nnpfc_pic_height_in_luma_samples && xY < nnpfc_pic_width_in_luma_samples)if( !nnpfc_component_last_flag ) {FilteredYPic[ i ][ xY ][ yY ] =outputTensor[ 0 ][ i ][ 0 ][ yP ][ xP ]FilteredCbPic[ i ][ xC ][ yC ] =outputTensor[ 0 ][ i ][ 1 ][ yPc ][ xPc ]FilteredCrPic[ i ][ xC ][ yC ] =outputTensor[ 0 ][ i ][ 2 ][ yPc ][ xPc ]} else {FilteredYPic[ i ][ xY ][ yY] = outputTensor[ 0 ][ i ][ yP ][ xP ][ 0 ]FilteredCbPic[ i ][ xC ][ yC ] =outputTensor[ 0 ][ i ][ yPc ][ xPc ][ 1 ]FilteredCrPic[ i ][ xC][ yC ] =outputTensor[ 0 ][ i ][ yPc ][ xPc ][ 2 ]}}else if( nnpfc_out_order_idc = =3 )for( yP = 0; yP < outPatchHeight; yP++ )for( xP = 0; xP < outPatchWidth;xP++ ) {ySrc = cTop / 2 * outPatchHeight / inpPatchHeight + yPxSrc = cLeft / 2 * outPatchWidth / inpPatchWidth + xPif ( ySrc < nnpfc_pic_height_in_luma_samples / 2 &&xSrc < nnpfc_pic_width_in_luma_samples / 2 )if( !nnpfc_component_last_flag ) {FilteredYPic[ i ][ xSrc * 2 ][ ySrc * 2 ] =outputTensor[ 0 ][ i ][ 0 ][ yP ][ xP ]FilteredYPic[ i ][ xSrc * 2 + 1 ][ySrc * 2 ] = outputTensor[ 0 ][ i ][ 1 ][ yP ][ xP ]FilteredYPic[ i ][ xSrc *2 ][ ySrc * 2 + 1 ] = outputTensor[ 0 ][ i ][ 2 ][ yP ][ xP ]FilteredYPic[ i][ xSrc * 2 + 1][ ySrc * 2 + 1 ] = outputTensor[ 0 ][ i ][ 3 ][ yP ][ xP ]FilteredCbPic[ i ][ xSrc ][ ySrc ] = outputTensor[ 0 ][ i ][ 4 ][ yP ][ xP ]FilteredCrPic[ i ][ xSrc ][ ySrc ] = outputTensor[ 0 ][ i ][ 5 ][ yP ][ xP ]}else {FilteredYPic[ i ][ xSrc * 2 ][ ySrc * 2 ] = outputTensor[ 0 ][ i ][ yP][ xP ][ 0 ]FilteredYPic[i ][ xSrc * 2 + 1 ][ ySrc * 2 ] = outputTensor[ 0 ][ i ][ yP ][ xP ][ 1 ]FilteredYPic[ i ][ xSrc * 2 ][ ySrc * 2 + 1 ] =outputTensor[ 0 ][ i ][ yP ][ xP ][ 2 ]FilteredYPic[ i ][ xSrc * 2 + 1][ ySrc* 2 + 1 ] = outputTensor[ 0 ][ i ][ yP ][ xP ][ 3 ]FilteredCbPic[ i ][ xSrc ][ ySrc ] = outputTensor[ 0 ][ i ][ yP ][ xP ][ 4 ]FilteredCrPic[ i ][ xSrc ][ ySrc ] = outputTensor[ 0 ][ i ][ yP ][ xP ][ 5 ]}}}
[0188] nnpfc_overlap indicates the horizontal and vertical sample counts of overlap between adjacent input tensors in NNPF. The value of nnpfc_overlap must be in the range of 0 to 16383 (inclusive).
[0189] The nnpfc_constant_patch_size_flag setting being equal to 1 indicates that NNPF accepts the precise patch size as input, as indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. The flag nnpfc_constant_patch_size_flag being 0 indicates that NNPF accepts any patch size with a width of inpPatchWidth and a height of inpPatchHeight as input, such that the width of the extended patch (i.e., the patch plus the overlapping area) (which is equal to inpPatchWidth + 2 * nnpfc_overlap) is a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, and the height of the extended patch (which is equal to inpPatchHeight + 2 * nnpfc_overlap) is a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap.
[0190] Incrementing 1 by nnpfc_patch_width_minus1 (when nnpfc_constant_patch_size_flag equals 1) indicates the horizontal sample count for the required patch size of the NNPF input. The value of nnpfc_patch_width_minus1 must be in the range of 0 to Min(32766, CroppedWidth − 1) (inclusive).
[0191] Incrementing 1 by nnpfc_patch_height_minus1 (when nnpfc_constant_patch_size_flag equals 1) indicates the vertical sample count for the patch size required for NNPF inputs. The value of nnpfc_patch_height_minus1 must be in the range of 0 to Min(32766, CroppedHeight − 1) (inclusive).
[0192] `nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap` (when `nnpfc_constant_patch_size_flag` equals 0) indicates the common divisor of all allowed values for the width of the extended patch required for the NNPF input. The value of `nnpfc_extended_patch_width_cd_delta_minus1` must be in the range of 0 to Min(32766, CroppedWidth − 1) (inclusive).
[0193] `nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap` (when `nnpfc_constant_patch_size_flag` equals 0) indicates the common divisor of all allowed values for the height of the extended patch required for the NNPF input. The value of `nnpfc_extended_patch_height_cd_delta_minus1` must be in the range of 0 to Min(32766, CroppedHeight − 1) (inclusive).
[0194] This makes the variables inpPatchWidth and inpPatchHeight the width and height of the small block, respectively.
[0195] If nnpfc_constant_patch_size_flag equals 0, then the following applies:
[0196] The values of inpPatchWidth and inpPatchHeight are provided externally by means not specified in this document, or are set by the post-processing itself.
[0197] The value of `inpPatchWidth + 2 * nnpfc_overlap` must be a positive integer multiple of `nnpfc_extended_patch_width_cd_delta_minus` (1 + 1 + 2 * nnpfc_overlap), and `inpPatchWidth` must be less than or equal to `CroppedWidth`. The value of `inpPatchHeight + 2 * nnpfc_overlap` must be a positive integer multiple of `nnpfc_extended_patch_height_cd_delta_minus` (1 + 1 + 2 * nnpfc_overlap), and `inpPatchHeight` must be less than or equal to `CroppedHeight`.
[0198] Otherwise (nnpfc_constant_patch_size_flag equals 1), the value of inpPatchWidth is set to equal to nnpfc_patch_width_minus1 + 1, and the value of inpPatchHeight is set to equal to nnpfc_patch_height_minus1 + 1.
[0199] The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, and outPatchCHeight are derived as follows:
[0200] outPatchWidth = ( nnpfc_pic_width_in_luma_samples * inpPatchWidth ) / CroppedWidth(86)
[0201] outPatchHeight = (nnpfc_pic_height_in_luma_samples * inpPatchHeight) / CroppedHeight(87)
[0202] horCScaling = SubWidthC / outSubWidthC(88)
[0203] verCScaling = SubHeightC / outSubHeightC(89)
[0204] outPatchCWidth = outPatchWidth * horCScaling(90)
[0205] outPatchCHeight = outPatchHeight * verCScaling(91)
[0206] outPatchWidth * CroppedWidth must be equal to nnpfc_pic_width_in_luma_samples * inpPatchWidth, and outPatchHeight * CroppedHeight must be equal to nnpfc_pic_height_in_luma_samples * inpPatchHeight. This is a requirement for bitstream consistency.
[0207] The nnpfc_padding_type indicates the padding process when referencing sample locations outside the boundaries of the cropped decoded output image, as described in Table 23. The value of nnpfc_padding_type must be in the range of 0 to 15 (inclusive).
[0208] Table 23 – Informative description of nnpfc_padding_type values
[0209] nnpfc_padding_type describe 0 Zero fill 1 Copy fill 2 Reflection fill 3 Surround fill 4 Fixed fill 5..15 reserve
[0210] nnpfc_luma_padding_val indicates the luminance value to be used for padding when nnpfc_padding_type is equal to 4.
[0211] nnpfc_cb_padding_val indicates the Cb value to be used for padding when nnpfc_padding_type is equal to 4.
[0212] nnpfc_cr_padding_val indicates the Cr value to be used for padding when nnpfc_padding_type is equal to 4.
[0213] The function InpSampleVal( y, x, picHeight, picWidth, croppedPic ) (where the inputs are the vertical sample position y, the horizontal sample position x, the image height picHeight, the image width picWidth, and the sample array croppedPic) returns the value of sampleVal, which is derived as follows:
[0214] Note 6 – For the input of the function InpSampleVal(), the vertical position is listed before the horizontal position to be compatible with the input tensor conventions of some inference engines.
[0215] if( nnpfc_padding_type = = 0 )if( y < 0 | | x < 0 | | y >=picHeight | | 1, x ) ][ Clip3( 0, picHeight − 1, y ) ]else if( nnpfc_padding_type = = 2 ) sampleVal = croppedPic[ Reflect( picWidth − 1, x ) ][ Reflect( picHeight − 1, y ) ]else if( nnpfc_padding_type = = 3 ) if( y >= 0 && y< picHeight )sampleVal = croppedPic[ Wrap( picWidth − 1, x ) ][ y ] else if(nnpfc_padding_type = = 4 ) if( y < 0 | | =nnpfc_cb_padding_valsampleVal[ 2 ] = nnpfc_cr_padding_valelsesampleVal =croppedPic[ x ][ y ]
[0216] The following example procedure can be used with NNPF PostProcessingFilter() to generate (multiple) filtered and / or interpolated images in a block-by-block manner, containing Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, as indicated by nnpfc_out_order_idc:
[0217] if( nnpfc_inp_order_idc = = 0 | = PostProcessingFilter( inputTensor )StoreOutputTensors( )}elseif( nnpfc_inp_order_idc = = 1 )for( cTop = 0; cTop < CroppedHeight / SubHeightC; cTop += inpPatchHeight )for( cLeft = 0; cLeft < CroppedWidth / SubWidthC; cLeft += inpPatchWidth ) {(93)DeriveInputTensors( )outputTensor= PostProcessingFilter( inputTensor )StoreOutputTensors( )}else if( nnpfc_inp_order_idc = = 3 )for( cTop = 0; cTop < CroppedHeight; cTop +=inpPatchHeight * 2 )for( cLeft = 0; cLeft < CroppedWidth; cLeft +=inpPatchWidth * 2 ) {DeriveInputTensors( )outputTensor = PostProcessingFilter( inputTensor )StoreOutputTensors( )}
[0218] The order of the images in the stored output tensor is the output order, and the output order generated by applying NNPF in the output order is interpreted as the output order (and does not conflict with the output order of the input images).
[0219] A value of 1 for nnpfc_complexity_info_present_flag indicates the existence of one or more syntax elements that indicate the complexity of the NNPF associated with nnpfc_id. A value of 0 for nnpfc_complexity_info_present_flag indicates the absence of a syntax element that indicates the complexity of the NNPF associated with nnpfc_id.
[0220] `nnpfc_parameter_type_idc` equal to 0 indicates that the neural network uses only integer parameters. `nnpfc_parameter_type_flag` equal to 1 indicates that the neural network can use floating-point or integer parameters. `nnpfc_parameter_type_idc` equal to 2 indicates that the neural network uses only binary parameters. `nnpfc_parameter_type_idc` equal to 3 is reserved for future use by ITU-T | ISO / IEC and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document must ignore NNPFC SEI messages with `nnpfc_parameter_type_idc` equal to 3.
[0221] The values 0, 1, 2, and 3 for nnpfc_log2_parameter_bit_length_minus3 indicate that the neural network does not use parameters with bit lengths greater than 8, 16, 32, and 64, respectively. When nnpfc_parameter_type_idc exists and nnpfc_log2_parameter_bit_length_minus3 does not exist, the neural network does not use parameters with a bit length greater than 1.
[0222] `nnpfc_num_parameters_idc` indicates the maximum number of neural network parameters for NNPF, in powers of 2048. `nnpfc_num_parameters_idc` equal to 0 indicates that the maximum number of neural network parameters is unknown. The value of `nnpfc_num_parameters_idc` must be in the range of 0 to 52 (inclusive). Values of `nnpfc_num_parameters_idc` greater than 52 are reserved for future use by ITU-T | ISO / IEC and should not exist in bitstreams conforming to this version of this document. Decoders conforming to this version of this document must ignore NNPFCSEI messages with `nnpfc_num_parameters_idc` greater than 52.
[0223] If the value of nnpfc_num_parameters_idc is greater than 0, then the variable maxNumParameters is deduced as follows:
[0224] maxNumParameters = ( 2 048 << nnpfc_num_parameters_idc ) − 1(94)
[0225] The requirement that the number of neural network parameters in NNPF must be less than or equal to maxNumParameters is a requirement for bitstream consistency.
[0226] A value greater than 0 for nnpfc_num_kmac_operations_idc indicates that the maximum number of multiply-accumulate operations per sample in NNPF is less than or equal to nnpfc_num_kmac_operations_idc * 1000. A value of 0 for nnpfc_num_kmac_operations_idc indicates that the maximum number of multiply-accumulate operations in the network is unknown. The value of nnpfc_num_kmac_operations_idc must be in the range of 0 to 2^32 − 2 (inclusive).
[0227] `nnpfc_total_kilobyte_size` greater than 0 indicates the total size in kilobytes required to store the uncompressed parameters of the neural network. The total size in bits is equal to or greater than the sum of the bits used to store each parameter. `nnpfc_total_kilobyte_size` is the total size in bits divided by 8000 and rounded down. `nnpfc_total_kilobyte_size` equal to 0 indicates that the total size required to store the parameters of the neural network is unknown. The value of `nnpfc_total_kilobyte_size` must be in the range of 0 to 2^32 − 2 (inclusive).
[0228] nnpfc_reserved_zero_bit_b must be equal to 0 in the bitstream conforming to this version of the document. The decoder must ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_b is not equal to 0.
[0229] nnpfc_payload_byte[i] contains the i-th byte of a bitstream conforming to ISO / IEC 15938-17. The byte sequence nnpfc_payload_byte[i] for all existing values must be a complete bitstream conforming to ISO / IEC 15938-17.
[0230] 8.29 Neural Network Post-Processing Filter Activation SEI Message
[0231] 8.29.1 Neural Network Post-Processing Filter Activation SEI Message Syntax
[0232] nn_post_filter_activation( payloadSize ) { descriptor nnpfa_target_id ue(v) nnpfa_cancel_flag u(1) if (!nnpfa_cancel_flag) nnpfa_persistence_flag u(1) }
[0233] 8.29.2 Neural Network Post-Processing Filter Activation of SEI Message Semantics
[0234] The Neural Network Post-Processing Filter Activation (NNPFA) SEI message activates or deactivates a target Neural Network Post-Processing Filter (NNPF) identified by nnpfa_target_id, which can be used for post-processing filtering of a set of images. For a specific image where the NNPF is activated, the target NNPF is the NNPF defined by the last NNPFC SEI message whose nnpfc_id is equal to nnpfa_target_id, which is in the decoding order before the first VCL NAL unit of the current image and is not a repetition of the NNPFC SEI message containing the underlying NNPF.
[0235] Note 1 – For example, multiple NNPFA SEI messages may exist for the same image when NNPF is used for different purposes or for filtering different color components.
[0236] nnpfa_target_id indicates the target NNPF, which is specified by one or more NNPFC SEI messages that are related to the current image and whose nnpfc_id is equal to nnfpa_target_id.
[0237] The value of nnpfa_target_id must be in the range of 0 to 2^32 − 2 (inclusive). Values of nnpfa_target_id from 256 to 511 (inclusive) and from 231 to 232 − 2 (inclusive) are reserved for future use by ITU-T | ISO / IEC. Decoders conforming to this document of this version must ignore NNPFA SEI messages when they encounter nnpfa_target_id in the range of 256 to 511 (inclusive) or from 231 to 232 − 2 (inclusive).
[0238] NNPFA SEI messages with a specific value of nnpfa_target_id should not exist in the current PU unless one or both of the following conditions are true:
[0239] – Within the current CLVS, there exists an NNPFC SEI message whose nnpfc_id is equal to a specific value of nnpfa_target_id that exists in the PU preceding the current PU in the decoding order.
[0240] – There is an NNPFC SEI message in the current PU with a specific value of nnpfc_id equal to nnpfa_target_id.
[0241] When a PU contains both an NNPFC SEI message with a specific value of nnpfc_id and an NNPFA SEI message with nnpfa_target_id equal to a specific value of nnpfc_id, the NNPFC SEI message must precede the NNPFA SEI message in the decoding order.
[0242] A `nnpfa_cancel_flag` of 1 indicates that the persistence of a target NNPF established by any previous NNPFA SEI message with the same `nnpfa_target_id` as the current SEI message is cancelled; that is, the target NNPF is no longer used unless it is activated by another NNPFA SEI message with the same `nnpfa_target_id` as the current SEI message and `nnpfa_cancel_flag` equal to 0. A `nnpfa_cancel_flag` of 0 indicates that `nnpfa_persistence_flag` follows.
[0243] The nnpfa_persistence_flag specifies the persistence of the target NNPF for the current layer.
[0244] Setting nnpfa_persistence_flag to 0 specifies that the target NNPF is used only for post-processing filtering of the current image.
[0245] Setting nnpfa_persistence_flag to 1 specifies that the target NNPF can be used for post-processing filtering of the current image and all subsequent images of the current layer (in output order) until one or more of the following conditions are true:
[0246] – A new CLVS begins for the current layer.
[0247] – End of bitstream.
[0248] - Outputs the images in the current layer that are associated with the NNPFA SEI message that has the same nnpfa_target_id as the current SEI message and whose nnpfa_cancel_flag is equal to 1, and are in the order of output following the current image.
[0249] Note 2 – Do not apply the target NNPF to the subsequent image in the current layer that is associated with an NNPFA SEI message that has the same nnpfa_target_id as the current SEI message and whose nnpfa_cancel_flag is equal to 1.
[0250] Let nnpfcTargetPictures be the set of pictures associated with the last NNPFC SEI message whose nnpfc_id is equal to nnpfa_target_id in the decoding order preceding the current NNPFA SEI message. Let nnpfaTargetPictures be the set of pictures for which the target NNPF is activated by the current NNPFA SEI message. The bitstream consistency requirement is that any pictures included in nnpfaTargetPictures must also be included in nnpfcTargetPictures.
[0251] 4. Technical problems solved by the disclosed embodiments
[0252] The example designs for Neural Network Post-Processing Filter Feature (NNPFC) SEI messages and Neural Network Post-Processing Filter Activation (NNPFA) SEI messages have the following problems:
[0253] First, (multiple) NNPF methods can modify VUI-related information in the output image. Various VUI-related information is missing from the NNPF syntax table.
[0254] Second, NNPF processing may require a StrengthControlVal set by the decoder or system. The current design may send an invalid StrengthControlVal to NNPF processing.
[0255] Third, when the purpose of NNPF is to indicate coloring (and possibly other types of format changes or upsampling), nnpfc_out_order_idc should be constrained to avoid or reduce meaningless cases.
[0256] Fourth, in some cases, the fill value is transmitted via signaling when not in use.
[0257] Fifth, in the VSEI v3 text, the constraints of nnpfc_matrix_coeffs may be unreasonable. For example, some constraints depend on nnpfc_out_tensor_chroma_bitdepth_minus8 and nnpfc_out_tensor_luma_bitdepth_minus8, while nnpfc_out_tensor_chroma_bitdepth_minus8 and / or nnpfc_out_tensor_luma_bitdepth_minus8 may not have specified values.
[0258] Sixth, in the VSEI v3 text, the inference of nnpfc_full_range_flag can be unreasonable. For example, when nnpfc_separate_colour_description_present_flag equals 0, the estimated value of nnpfc_full_range_flag can be different from vui_full_range_flag.
[0259] Seventh, inference for nnpfc_chroma_sample_loc_type_frame is missing in the VSEI v3 text.
[0260] Eighth, in the VSEI v3 text, the signaling for nnpfc_matrix_coeffs and nnpfc_full_range_flag depends on the output tensor format. This dependency may be unreasonable, as matrix_coeffs and full_range_flag are video attributes.
[0261] 5. List of solutions and implementation examples
[0262] To address the aforementioned problems, methods outlined below are disclosed. These aspects should be considered as examples for interpreting general concepts, and not interpreted in a narrow sense. Furthermore, these examples can be applied individually or in any combination.
[0263] 1) To solve problem 1, one or more of the following aspects are specified:
[0264] a. When the NNPF output image is in 4:2:0 color format, the position of the chroma sample points can be indicated by one or more syntax elements transmitted via signal transmission.
[0265] i. In one example, when the NNPF purpose indicates colorization (and possibly other types of format changes or upsampling) and (multiple) NNPF output images are in 4:2:0 color format, the position of the chroma samples of the NNPF output images is specified by the signaling syntax element.
[0266] b. When the color space of the NNPF output image differs from that of the decoded image or the cropped decoded output image, the signal transmission syntax element can be used to indicate whether the NNPF output image is in the full range.
[0267] c. Different sets of amplitude ratio related parameters can be transmitted via signaling in the NNPF SEI message.
[0268] i. In one example, a first flag is transmitted via signaling to indicate whether the NNPF process changes the amplitude ratio.
[0269] ii. In one example, different sets of amplitude ratio parameters, such as aspect_ratio_idc, sar_width, and sar_height, can be transmitted via signal.
[0270] 1. In one example, whether these parameters are transmitted via signal depends on the value of the first flag.
[0271] 2. In one example, when these parameters are not present, the aspect ratio attribute is assumed to be the same as the decoded image or the cropped decoded output image in CLVS.
[0272] d. The source scan type of the NNPF procedure's output can be indicated by signaling one or more syntax elements.
[0273] i. When not present, the source scan type is presumed to be the same as the decoded image or the cropped decoded output image in CLVS.
[0274] ii. In one example, a first flag is transmitted via signal transmission to indicate whether a preferred scan type is transmitted via signal transmission.
[0275] 1. In one example, a second flag is transmitted via signaling to indicate whether overscanning is preferred.
[0276] iii. In one example, the preferred scan type is only transmitted via signal transmission when the image resolution changes.
[0277] iv. When one or more related syntax elements are not present, the scan type is presumed to be the same as in VUI.
[0278] e. A preferred method of displaying the output of an NNPF procedure can be indicated by signaling one or more syntax elements.
[0279] i. When not present, the preferred display method is presumed to be the same as the decoded image or the cropped decoded output image in CLVS.
[0280] 2) To address problem 2, one or more of the following aspects are specified:
[0281] a. When using NNPF, StrengthControlVal is set to equal to (SliceQp). Y The value of SliceQp is calculated as follows: (63 + QpBdOffset) ÷ (63 + QpBdOffset) Y It is the first slice of currCodedPic's SliceQp Y .
[0282] b. Alternatively, when using NNPF, StrengthControlVal is set to equal to (SliceQp). Y The value of SliceQp is given by (+A)÷B, where SliceQp Y It is the first slice of currCodedPic's SliceQp Y A is the maximum possible value of QpBdOffset, and B is the value of SliceQp. Y The maximum possible value.
[0283] i. In one example, A equals 48 and B equals 111.
[0284] c. Alternatively, StrengthControlVal can be transmitted via signaling in the NNPF SEI message.
[0285] i. In one example, StrengthControlVal can be signaled in an NNPFC SEI message.
[0286] ii. In one example, NNPF can be activated by signaling StrengthControlVal in an NNPFA SEI message.
[0287] 1. In one example, StrengthControlVal can be signaled within a specific range and converted into a range of input values for NNPF.
[0288] 3) To address problem 3, one or more of the following aspects are specified:
[0289] a. When the NNPF purpose indicates colorization (and possibly other types of formatting changes or upsampling), the chroma matrix must exist in the output tensor.
[0290] i. In one example, it is specified that nnpfc_out_order_idc should not be equal to 0 when nnpfc_purpose & 0x20 is not equal to 0.
[0291] 4) To address problem 4, one or more of the following aspects are specified:
[0292] a. When the input tensor does not contain a brightness matrix, the signaling for brightness padding values can be skipped.
[0293] b. When the input tensor does not contain a chroma matrix, the signaling for the chroma fill value can be skipped.
[0294] 5) To address issue 5, the constraints of nnpfc_matrix_coeffs depend on the existence of nnpfc_out_tensor_chroma_bitdepth_minus8 and / or nnpfc_out_tensor_luma_bitdepth_minus8:
[0295] a. In one example, when both nnpfc_out_tensor_chroma_bitdepth_minus8 and nnpfc_out_tensor_luma_bitdepth_minus8 exist, and nnpfc_out_tensor_chroma_bitdepth_minus8 equals nnpfc_out_tensor_luma_bitdepth_minus8, and nnpfc_out_order_idc equals 2, outSubHeightC equals 1, and outSubWidthC equals 1, the value of nnpfc_matrix_coeffs should not be equal to 0.
[0296] b. In one example, when both nnpfc_out_tensor_chroma_bitdepth_minus8 and nnpfc_out_tensor_luma_bitdepth_minus8 exist; and nnpfc_out_tensor_chroma_bitdepth_minus8 equals nnpfc_out_tensor_luma_bitdepth_minus8, or when both nnpfc_out_tensor_chroma_bitdepth_minus8 and nnpfc_out_tensor_luma_bitdepth_minus8 exist; and nnpfc_out_tensor_chroma_bitdepth_minus8 equals nnpfc_out_tensor_luma_bitdepth_minus8 + 1, nnpfc_out_order_idc equals 2, outSubHeightC equals 1, and outSubWidthC equals 1, the value of nnpfc_matrix_coeffs should not be equal to 8.
[0297] 6) To solve problem 6, it is stipulated that when it does not exist, the value of nnpfc_full_range_flag must be presumed to be equal to vui_full_range_flag.
[0298] a. Alternatively, if not present, the value of nnpfc_full_range_flag must be presumed to be 0.
[0299] 7) To solve problem 7, it is stipulated that when nnpfc_chroma_sample_loc_type_frame does not exist, the value must be presumed to be equal to 6.
[0300] a. In addition, when nnpfc_out_colour_format_idc is absent and equal to 1, the value of nnpfc_chroma_sample_loc_type_frame is presumed to be 6, which indicates that the location of the chroma sample is unknown or unspecified or specified in any other way not specified in this document.
[0301] b. Alternatively, when not present, the value of nnpfc_chroma_sample_loc_type_frame shall be presumed to be equal to vui_chroma_sample_loc_type_frame.
[0302] i. Furthermore, when nnpfc_out_colour_format_idc is absent and equal to 1, the value of nnpfc_chroma_sample_loc_type_frame is presumed to be vui_chroma_sample_loc_type_frame.
[0303] ii. Furthermore, when nnpfc_out_colour_format_idc is absent and nnpfc_out_colour_format_idc is equal to 1 and vui_chroma_sample_loc_type_frame exists, the value of nnpfc_chroma_sample_loc_type_frame is presumed to be vui_chroma_sample_loc_type_frame.
[0304] 8) To address problem 8, it is proposed that the signaling and / or inference of nnpfc_matrix_coeffs and / or nnpfc_full_range_flag be independent of whether the sample values of the NNPF output are real numbers or integers.
[0305] a. In one example, the signaling for nnpfc_matrix_coeffs and nnpfc_full_range_flag is not conditional on nnpfc_out_format_idc.
[0306] 6. Examples
[0307] The following are some example embodiments of the aspects outlined in Section 5. Most of the relevant parts that have been added or modified are shown in bold, and some of the deleted parts are shown in italic bold. There may be other edited changes that are not highlighted.
[0308] 6.1 Example 1
[0309] This embodiment covers aspects of items 1, 1a, and 1b, and all their sub-items, as outlined in Section 5. The text modifications are based on JVET-AC2032-v2.
[0310] 8.28.1 Characteristics of Neural Network Post-Processing Filters and SEI Message Syntax
[0311] nn_post_filter_characteristics( payloadSize ) { descriptor ... ue(v) if( ( nnpfc_purpose & 0x02 ) != 0 ) nnpfc_out_sub_c_flag u(1) if( ( nnpfc_purpose & 0x20 ) != 0 ) { nnpfc_out_colour_format_idc u(2) if(nnpfc_out_colour_format_idc = 1) nnpfc_chroma_sample_loc_type_frame ue(v) } ... nnpfc_separate_colour_description_present_flag u(1) if( nnpfc_separate_colour_description_present_flag ) { nnpfc_colour_primaries u(8) nnpfc_transfer_characteristics u(8) nnpfc_matrix_coeffs u(8) nnpfc_full_range_flag u(1) } ...
[0312] 8.28.2 Characteristics of Neural Network Post-Processing Filters and Semantics of SEI Messages ...
[0314] When both nnpfc_purpose & 0x02 and nnpfc_purpose & 0x20 are equal to 0, outSubWidthC and outSubHeightC are presumed to be equal to SubWidthC and SubHeightC, respectively.
[0315] `nnpfc_chroma_sample_loc_type_frame` specifies the position of the chroma samples in the output image when it is not equal to 6 and `nnpfc_out_colour_format_idc` is equal to 1 (4:2:0 color format). Figure 1 As shown. `nnpfc_chroma_sample_loc_type_frame` equal to 6 and `nnpfc_out_colour_format_idc` equal to 1 (4:2:0 color format) indicate that the location of the chroma sample points is unknown, unspecified, or specified by any other means not specified in this specification. The value of `nnpfc_chroma_sample_loc_type_frame` must be in the range of 0 to 6 (inclusive). ...
[0317] `nnpfc_separate_colour_description_present_flag` equal to 1 indicates that the SEI message syntax structure specifies different combinations of the primary colors, transport characteristics, and matrix coefficients of the image generated by NNPF, as well as the scaling and offset values applied to the matrix coefficients. `nnpfc_separate_colour_description_present_flag` equal to 0 indicates that the combination of the primary colors, transport characteristics, and matrix coefficients of the image generated by NNPF, as well as the scaling and offset values applied to the matrix coefficients, is the same as indicated in the CLVS VUI parameters.
[0318] nnpfc_colour_primaries has the same semantics as the vui_colour_primaries syntax element specified in sub-entry 7.3, except as follows:
[0319] – nnpfc_colour_primaries specifies the primary color of the image generated by NNPF as defined in the application SEI message, rather than the primary color used for CLVS.
[0320] – When nnpfc_colour_primaries does not exist in the NNPFC SEI message, the value of nnpfc_colour_primaries is presumed to be equal to vui_colour_primaries.
[0321] nnpfc_transfer_characteristics has the same semantics as the vui_transfer_characteristics syntax element specified in sub-entry 7.3, except as follows:
[0322] – nnpfc_transfer_characteristics specifies the transfer characteristics of images generated by NNPF as defined in the SEI message, rather than the transfer characteristics used for CLVS.
[0323] – When nnpfc_transfer_characteristics does not exist in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is presumed to be equal to vui_transfer_characteristics.
[0324] nnpfc_matrix_coeffs has the same semantics as the vui_matrix_coeffs syntax element specified in sub-entry 7.3, except as follows:
[0325] – nnpfc_matrix_coeffs specifies the matrix coefficients of the image generated by the NNPF specified in the application SEI message, instead of the matrix coefficients used for CLVS.
[0326] – When nnpfc_matrix_coeffs does not exist in the NNPFC SEI message, the value of nnpfc_matrix_coeffs is presumed to be equal to vui_matrix_coeffs.
[0327] The allowed values for – nnpfc_matrix_coeffs are not constrained by the chroma format of the decoded video image as indicated by the value of ChromaFormatIdc, which is the semantics of the VUI parameter.
[0328] – When nnpfc_matrix_coeffs equals 0, nnpfc_out_order_idc should not equal 1 or 3.
[0329] nnpfc_full_range_flag has the same semantics as the vui_full_range_flag syntax element specified in sub-entry 7.3, except as follows:
[0330] – nnpfc_full_range_flag specifies the scaling and offset values applied in association with the matrix coefficients of the image generated by NNPF as specified in the application SEI message, rather than the color primary colors used for CLVS.
[0331] – When nnpfc_colour_primaries does not exist in the NNPFC SEI message, the value of nnpfc_full_range_flag is presumed to be equal to vui_full_range_flag. ...
[0333] 6.2 Example 2
[0334] This embodiment covers aspects of items 2 and 2.b. and all their sub-items as outlined in Section 5. The text modifications are based on JVET-AC2005-v1.
[0335] D.12.11 Use of SEI Message in Neural Network Post-Processing Filter Characteristics ...
[0337] To interpret NNPFC SEI messages, the following variables are defined:
[0338] – If pictureRateUpsamplingFlag equals 1, and there is a second NNPF defined by at least one NNPFC SEI message, activated by an NNPFA SEI message for currCodedPic, and nnpfc_purpose equals 4, then the following applies:
[0339] – CroppedWidth is set to be equal to nnpfc_pic_width_in_luma_samples as defined for the second NNPF.
[0340] – CroppedHeight is set to be equal to nnpfc_pic_height_in_luma_samples as defined for the second NNPF.
[0341] – Otherwise, the following applies:
[0342] – CroppedWidth is set to the value of pps_pic_width_in_luma_samples−SubWidthC * (pps_conf_win_left_offset + pps_conf_win_right_offset) equal to currCodedPic.
[0343] – CroppedHeight is set to the value of pps_pic_height_in_luma_samples − SubHeightC * (pps_conf_win_top_offset + pps_conf_win_bottom_offset) of currCodedPic.
[0344] – The luminance sample array CroppedYPic[i] and the chrominance sample arrays CroppedCbPic[i] and CroppedCrPic[i] (if present) are derived for each value of i in the range of 0 to numInputPics − 1 (inclusive):
[0345] – Make sourcePic a cropped decoded output image of PicOrderCntVal with inputPicPoc[i] in the CLVS containing currCodedPic.
[0346] – If pictureRateUpsamplingFlag equals 0, the following applies:
[0347] – The luminance sample array CroppedYPic[i] and the chrominance sample arrays CroppedCbPic[i] and CroppedCrPic[i] (if present) are set as two-dimensional arrays of the decoded sample values of the Y, Cb and Cr components of sourcePic, respectively.
[0348] – Otherwise (pictureRateUpsamplingFlag equals 1), the following applies:
[0349] – The variable sourceWidth is set to the value of pps_pic_width_in_luma_samples−SubWidthC * (pps_conf_win_left_offset + pps_conf_win_right_offset) of sourcePic.
[0350] – The variable sourceHeight is set to the value of sourcePic's pps_pic_height_in_luma_samples − SubHeightC * (pps_conf_win_top_offset + pps_conf_win_bottom_offset).
[0351] – If sourceWidth equals CroppedWidth and sourceHeight equals CroppedHeight, then inputPic is set to be the same as sourcePic.
[0352] – Otherwise (sourceWidth is not equal to CroppedWidth or sourceHeight is not equal to CroppedHeight), the following applies:
[0353] – There must be an NNPF defined by at least one NNPFC SEI message, activated by an NNPFA SEI message for the sourcePic, and with nnpfc_purpose equal to 4, nnpfc_pic_width_in_luma_samples equal to CroppedWidth, and nnpfc_pic_height_in_luma_samples equal to CroppedHeight. This is hereby referred to as the super-resolution NNPF.
[0354] – inputPic is set as the output of the neural network inference of the super-resolution NNPF with sourcePic as input.
[0355] – The luminance sample array CroppedYPic[i] and the chrominance sample arrays CroppedCbPic[i] and CroppedCrPic[i] (if present) are set as two-dimensional arrays of the decoded sample values of the Y, Cb and Cr components of the inputPic, respectively.
[0356] – BitDepth Y and BitDepth C They are all set to equal BitDepth.
[0357] – ChromaFormatIdc is set to equal sps_chroma_format_idc.
[0358] – StrengthControlVal is set to equal the first slice of currCodedPic (SliceQp) Y The value of + 48 ) ÷ 63 111. ...
[0360] 6.3 Example 3
[0361] This embodiment covers aspects of items 4, 4a, and 4b, and all their sub-items, as outlined in Section 5. The text modifications are based on JVET-AC2032-v2.
[0362] 8.28.1 Characteristics of Neural Network Post-Processing Filters and SEI Message Syntax
[0363] nn_post_filter_characteristics( payloadSize ) { Descriptor ... if( nnpfc_padding_type = 4 ) { if( nnpfc_inp_order_idc != 1 ) nnpfc_luma_padding_val ue(v) if( nnpfc_inp_order_idc != 0 ) { nnpfc_cb_padding_val ue(v) nnpfc_cr_padding_val ue(v) } } ...
[0364] 6.4 Example 4
[0365] This embodiment covers aspects of items 5.5.a, 5.b, 6, and 7. The text modifications are based on JVET-AE2006-v2.
[0366] 8.28.2.2 Characteristics of Neural Network Post-Processing Filters and Semantics of SEI Messages ......
[0368] `nnpfc_separate_colour_description_present_flag` equal to 1 indicates that the SEI message syntax structure specifies different combinations of the primary colors, transfer characteristics, matrix coefficients, and scaling and offset values associated with the matrix coefficients of the image generated by NNPF. `nnpfc_separate_colour_description_present_flag` equal to 0 indicates that the combination of the primary colors, transfer characteristics, matrix coefficients, and scaling and offset values associated with the matrix coefficients of the image generated by NNPF is the same as implied by the VUI parameters `vui_colour_primaries`, `vui_transfer_characteristics`, `vui_matrix_coeffs`, and `vui_full_range_flag` indicated or presumed by CLVS.
[0369] nnpfc_colour_primaries has the same semantics as the vui_colour_primaries syntax element specified in sub-entry 7.3, except as follows:
[0370] – nnpfc_colour_primaries specifies the primary color of the image generated by NNPF as defined in the application SEI message, rather than the primary color used for CLVS.
[0371] – When nnpfc_colour_primaries does not exist in the NNPFC SEI message, the value of nnpfc_colour_primaries is presumed to be equal to vui_colour_primaries.
[0372] nnpfc_transfer_characteristics has the same semantics as the vui_transfer_characteristics syntax element specified in sub-entry 7.3, except as follows:
[0373] – nnpfc_transfer_characteristics specifies the transfer characteristics of images generated by NNPF as defined in the SEI message, rather than the transfer characteristics used for CLVS.
[0374] – When nnpfc_transfer_characteristics does not exist in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is presumed to be equal to vui_transfer_characteristics.
[0375] nnpfc_matrix_coeffs describes the equations used to derive luminance and chrominance signals from green, blue, and red, or Y, Z, and X primary colors. Its semantics apply to images generated by applying the NNPF specified in this SEI message and are as defined for MatrixCoefficients in Recommendation ITU-T H.273 | ISO / IEC 23091-2, where BitDepth... Y and BitDepth C They are respectively equal to outTensorBitDepth Y and outTensorBitDepth C .
[0376] When nnpfc_matrix_coeffs does not exist in the NNPFC SEI message, the value of nnpfc_matrix_coeffs is presumed to be equal to vui_matrix_coeffs.
[0377] nnpfc_matrix_coeffs should not be equal to 0 unless both of the following conditions are true:
[0378] Both nnpfc_out_tensor_chroma_bitdepth_minus8 and nnpfc_out_tensor_luma_bitdepth_minus8 exist.
[0379] – nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8.
[0380] – nnpfc_out_order_idc equals 2, outSubHeightC equals 1, and outSubWidthC equals 1.
[0381] Note: nnpfc_matrix_coeffs must exist when all three conditions above are true and vui_matrix_coeffs equals 0.
[0382] nnpfc_matrix_coeffs should not be equal to 8 unless one of the following conditions is true:
[0383] Both nnpfc_out_tensor_chroma_bitdepth_minus8 and nnpfc_out_tensor_luma_bitdepth_minus8 exist; and nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8.
[0384] Both nnpfc_out_tensor_chroma_bitdepth_minus8 and nnpfc_out_tensor_luma_bitdepth_minus8 exist; and nnpfc_out_tensor_chroma_bitdepth_minus8 equals nnpfc_out_tensor_luma_bitdepth_minus8 + 1, nnpfc_out_order_idc equals 2, outSubHeightC equals 1, and outSubWidthC equals 1.
[0385] Note: nnpfc_matrix_coeffs must exist when one of the above two conditions is true and vui_matrix_coeffs equals 8.
[0386] `nnpfc_full_range_flag` indicates the scaling and offset values applied in association with the matrix coefficients specified by `nnpfc_matrix_coeffs`. Its semantics are identical to those specified for the `VideoFullRangeFlag` parameter in Recommendation ITU-T H.273 | ISO / IEC 23091-2. When not present, the value of `nnpfc_full_range_flag` is presumed to be equal to `vui_full_range_flag0`.
[0387] A value of 1 for `nnpfc_chroma_loc_info_present_flag` indicates the presence of the `nnpfc_chroma_sample_loc_type_frame` syntax element in the NNPFC SEI message. A value of 0 for `nnpfc_chroma_loc_info_present_flag` indicates the absence of the `nnpfc_chroma_sample_loc_type_frame` syntax element in the NNPFC SEI message. When `nnpfc_chroma_loc_info_present_flag` does not exist, its value is presumed to be 0. The value of `nnpfc_chroma_loc_info_present_flag` must be 0 when `ColourizationFlag` is 0 or when `nnpfc_out_colour_format_idc` is not equal to 1.
[0388] `nnpfc_chroma_sample_loc_type_frame` specifies the position of the chroma samples in the output image when it is not equal to 6 and `nnpfc_out_colour_format_idc` is equal to 1. Figure 1 As shown. `nnpfc_chroma_sample_loc_type_frame` equal to 6 and `nnpfc_out_colour_format_idc` equal to 1 indicates that the location of the chroma sample point is unknown, unspecified, or specified in any other way not specified in this document. The value of `nnpfc_chroma_sample_loc_type_frame` must be in the range of 0 to 6 (inclusive). When it does not exist and `nnpfc_out_colour_format_idc` equals 1, the value of `nnpfc_chroma_sample_loc_type_frame` is presumed to be 6, which indicates that the location of the chroma sample point is unknown, unspecified, or specified in any other way not specified in this document.
[0389] When nnpfc_out_colour_format_idc equals 1 (4:2:0 chroma format) and the output image is intended to be interpreted according to Recommendation ITU-R BT.2020 or Recommendation ITU-R BT.2100, nnpfc_chroma_loc_info_present_flag should be equal to 1, and nnpfc_chroma_sample_loc_type_frame should be equal to 2. ......
[0391] 6.5 Example 5
[0392] This embodiment covers aspects of Project 8. Text modifications are based on JVET-AE2006-v2.
[0393] 8.28.2.1 Characteristics of Neural Network Post-Processing Filters and SEI Message Syntax
[0394] nn_post_filter_characteristics( payloadSize ) { Descriptor ... nnpfc_separate_colour_description_present_flag u(1) if( nnpfc_separate_colour_description_present_flag ) { nnpfc_colour_primaries u(8) nnpfc_transfer_characteristics u(8) if( nnpfc_out_format_idc == 1 ) { nnpfc_matrix_coeffs u(8) nnpfc_full_range_flag u(1) } } ...
[0395] 6.6 Example 6
[0396] This embodiment covers aspects of projects 1.d.ii and 1.d.iii. Text modifications are based on JVET-AE2006-v2.
[0397] 8.28.2.1 Characteristics of Neural Network Post-Processing Filters and SEI Message Syntax
[0398] nn_post_filter_characteristics( payloadSize ) { Descriptor ... if( ResolutionResamplingFlag ) nnpfc_overscan_info_present_flag u(1) if( nnpfc_overscan_info_present_flag ) nnpfc_overscan_appropriate_flag u(1) ...
[0399] 8.28.2.2 Characteristics of Neural Network Post-Processing Filters and Semantics of SEI Messages ......
[0401] An nnpfc_overscan_info_present_flag value of 1 indicates that nnpfc_overscan_appropriate_flag exists. When nnpfc_overscan_info_present_flag is equal to 0 or does not exist, the preferred display method for the image output by NNPF is unknown, unspecified, or specified externally.
[0402] `nnpfc_overscan_appropriate_flag` equal to 1 indicates that the image output by NNPF is suitable for display using overscan. `nnpfc_overscan_appropriate_flag` equal to 0 indicates that the image output by NNPF contains visually important information throughout the entire area leading to the image's edges, making overscanning unsuitable. Instead, it should be displayed using an exact match between the display area and the edges, or underscanning. As used in this paragraph, the term "overscan" refers to a display process where some portions near the image's boundaries are not visible in the display area. The term "underscan" describes a display process where the entire image is visible in the display area, but does not cover the entire display area. For display processes that use neither overscan nor underscan, the display area perfectly matches the area of the image. ...
[0404] 6.7 Example 7
[0405] This embodiment covers aspects of projects 1.d.ii, 1.d.iii, and 1.d.iv. Text modifications are based on JVET-AE2006-v2.
[0406] 8.28.2.1 Characteristics of Neural Network Post-Processing Filters and SEI Message Syntax
[0407] nn_post_filter_characteristics( payloadSize ) { Descriptor ... if( ResolutionResamplingFlag ) nnpfc_scan_type_info_present_flag u(1) if(nnpfc_scan_type_info_present_flag) nnpfc_scan_type_idc u(2) ...
[0408] 8.28.2.2 Characteristics of Neural Network Post-Processing Filters and Semantics of SEI Messages ......
[0410] A value of 1 for nnpfc_scan_type_info_present_flag indicates that nnpfc_scan_type_idc exists. A value of 0 for nnpfc_scan_type_info_present_flag indicates that nnpfc_scan_type_idc does not exist. When nnpfc_scan_type_info_present_flag does not exist, it is presumed that nnpfc_scan_type_info_present_flag is equal to 0.
[0411] `nnpfc_scan_type_idc` equal to 0 indicates that the preferred display method for images output by NNPF is unknown, unspecified, or externally specified. `nnpfc_scan_type_idc` equal to 1 indicates that images output by NNPF are suitable for display using overscan. `nnpfc_scan_type_idc` equal to 2 indicates that images output by NNPF contain visually important information across the entire area to the edges of the image, such that overscan should not be used. Instead, they should be displayed using an exact match between the display area and the edges, or using underscan. As used in this paragraph, the term "overscan" refers to a display process where some portions near the image boundaries are not visible in the display area. The term "underscan" describes a display process where the entire image is visible in the display area, but does not cover the entire display area. For display processes that use neither overscan nor underscan, the display area perfectly matches the area of the image. The value of `nnpfc_scan_type_idc` should not be equal to 2. When it does not exist, the value of nnpfc_scan_type_idc is presumed to be equal to 0. ...
[0413] 7. References
[0414] [1] ITU-T and ISO / IEC, “Efficient video coding and decoding”, Recommendation ITU-T H.265 | ISO / IEC23008-2 (current version).
[0415] [2] J. Chen, E. Alshina, GJ Sullivan, J.-R. Ohm, J. Boyce, “Algorithmic description of Joint Exploration Test Model 7 (JEM7)”, JVET-G1001, August 2017.
[0416] [3] Recommendation ITU-T H.266 | ISO / IEC 23090-3, “Multi-functional video coding and decoding”, 2022.
[0417] [4] ITU-T Recommendation H.274 | ISO / IEC 23002-7, “Multifunctional supplementary enhancement information messages for encoding and decoding video bitstreams”, 2022.
[0418] [5] S. McCarthy, S. Deshpande, M. Hannuksela, Hendry, G. Sullivan and Y.-K. Wang (eds.), “Improvement considerations for SEI messages in neural network post-processing filters”, JVET output document JVET-AC2032, is publicly available online at: https: / / jvet-experts.org / doc_end_user / current_document.php?id=12585.
[0419] Figure 2 This is a block diagram illustrating an example video processing system 4000 in which various embodiments disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8-bit or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0420] System 4000 may include a codec component 4004 capable of implementing the various codec or encoding methods described in this disclosure. Codec component 4004 may reduce the average bit rate from the video input 4002 to the output of codec component 4004 to produce a codec representation of the video. Codec techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of codec component 4004 may be stored or transmitted via a communication connection such as that represented by component 4006. The bitstream (or codec) representation of the video received at input 4002, whether stored or communicated, may be used by component 4008 to generate pixel values or displayable video to be transmitted to display interface 4010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it is understood that a codec tool or operation is used at the encoder, and a corresponding decoding tool or operation that reverses the codec result will be performed by the decoder.
[0421] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE), etc. The embodiments described in this disclosure can be embodied in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0422] Figure 3 This is a block diagram of an example video processing apparatus 4100. Apparatus 4100 can be used to implement one or more methods described herein. Apparatus 4100 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processors(multiple) 4102 can be configured to implement one or more methods described in this disclosure. The memories(multiple) 4104 can be used to store data and code for implementing the methods and embodiments described herein. The video processing circuitry 4106 can be used to implement some embodiments described in this disclosure in hardware circuitry. In some embodiments, the video processing circuitry 4106 may be at least partially included in the processor 4102, such as a graphics coprocessor.
[0423] Figure 4This is a flowchart of an example method 4200 for video processing. In step 4202, method 4200 determines whether the NNPF output image is within the full range when the color space of the neural network post-processing filter (NNPF) output image differs from the color space of the decoded image or the color space of the cropped decoded output image. In step 4204, a conversion between visual media data and a bitstream is performed based on the NNPF. This conversion may include encoding at the encoder, decoding at the decoder, or a combination thereof.
[0424] It should be noted that method 4200 can be implemented in an apparatus for processing video data, including a processor and a non-transitory memory having instructions thereon, such as a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In this case, the instructions cause the processor to perform method 4200 when executed by the processor. Furthermore, method 4200 can be executed by a non-transitory computer-readable medium including a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec device to perform method 4200 when executed by a processor.
[0425] Figure 5 This is a block diagram illustrating an example video encoding / decoding system 4300 that can utilize embodiments of the present disclosure. The video encoding / decoding system 4300 may include a source device 4310 and a destination device 4320. The source device 4310 generates encoded video data, and this source device 4310 may be referred to as a video encoding device. The destination device 4320 can decode the encoded video data generated by the source device 4310, and this destination device 4320 may be referred to as a video decoding device.
[0426] Source device 4310 may include video source 4312, video encoder 4314, and input / output (I / O) interface 4316. Video source 4312 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations of such sources. Video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and associated data. Codec pictures are codec representations of pictures. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or transmitter. Encoded video data may be transmitted directly to destination device 4320 via network 4330 through I / O interface 4316. Encoded video data may also be stored on storage medium / server 4340 for access by destination device 4320.
[0427] The destination device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may acquire encoded video data from the source device 4310 or the storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the destination device 4320 or may be external to the destination device 4320, wherein the destination device 4320 may be configured to interface with an external display device.
[0428] The video encoder 4314 and the video decoder 4324 can operate according to video compression standards, such as HEVC, VVC and other existing and / or further standards.
[0429] Figure 6 This is a block diagram illustrating an example of a video encoder 4400, which can be... Figure 5 The system 4300 shown includes a video encoder 4314. The video encoder 4400 can be configured to perform any or all embodiments of this disclosure. The video encoder 4400 includes multiple functional components. The embodiments described in this disclosure can be shared among the various components of the video encoder 4400. In some examples, a processor can be configured to perform any or all embodiments described in this disclosure.
[0430] The functional components of the video encoder 4400 may include a segmentation unit 4401; a prediction unit 4402, which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra-frame prediction unit 4406; a residual generation unit 4407; a transform processing unit 4408; a quantization unit 4409; an inverse quantization unit 4410; an inverse transform unit 4411; a reconstruction unit 4412; a buffer 4413; and an entropy coding unit 4414.
[0431] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0432] Furthermore, some components such as the motion estimation unit 4404 and the motion compensation unit 4405 can be highly integrated, but for illustrative purposes, they are shown separately in the example of the video encoder 4400.
[0433] The segmentation unit 4401 can segment an image into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.
[0434] The mode selection unit 4403 can select one of several encoding / decoding modes (intra-frame encoding / decoding or inter-frame encoding / decoding), for example, based on error results, and provide the resulting intra-frame or inter-frame encoded / decoded block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 4403 can select an intra-frame / inter-frame joint prediction (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 4403 can also select a resolution for the block based on motion vectors (e.g., sub-pixel precision or integer pixel precision).
[0435] To perform inter-frame prediction on the current video block, motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. Motion compensation unit 4405 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 4413 other than the image associated with the current video block.
[0436] The motion estimation unit 4404 and the motion compensation unit 4405 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.
[0437] In some examples, motion estimation unit 4404 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images in list 0 or list 1. Motion estimation unit 4404 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0438] In other examples, motion estimation unit 4404 can perform bidirectional prediction on the current video block. Motion estimation unit 4404 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 1. Motion estimation unit 4404 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 4404 can output the reference indices and the motion vector of the current video block as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0439] In some examples, the motion estimation unit 4404 can output a complete set of motion information for use in the decoder's decoding process. In some examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 4404 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0440] In one example, the motion estimation unit 4404 may indicate a value to the video decoder 4500 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.
[0441] In another example, motion estimation unit 4404 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0442] As discussed above, the video encoder 4400 can transmit motion vectors via signaling in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge Mode Signaling.
[0443] Intra-prediction unit 4406 can perform intra-prediction on the current video block. When intra-prediction unit 4406 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples of other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0444] The residual generation unit 4407 can generate residual data for the current video block by subtracting (or more) predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0445] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform subtraction operations.
[0446] The transform processing unit 4408 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0447] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0448] The inverse quantization unit 4410 and the inverse transform unit 4411 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block, which is stored in the buffer 4413.
[0449] After the video block is reconstructed by reconstruction unit 4412, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0450] Entropy encoding unit 4414 can receive data from other functional components of video encoder 4400. When entropy encoding unit 4414 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0451] Figure 7 This is a block diagram illustrating an example of a video decoder 4500, which can be... Figure 5 The system 4300 shown includes a video decoder 4324. The video decoder 4500 can be configured to perform any or all embodiments of this disclosure. In the example shown, the video decoder 4500 includes multiple functional components. The embodiments described in this disclosure can be shared among the various components of the video decoder 4500. In some examples, a processor can be configured to perform any or all embodiments described in this disclosure.
[0452] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-frame prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 4400.
[0453] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 4501 can decode the entropy-encoded video data, and based on the entropy-decoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 4502 can determine this information, for example, by executing AMVP and Merge modes.
[0454] The motion compensation unit 4502 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. The identifier of the interpolation filter to be used, with sub-pixel accuracy, can be included in the syntax element.
[0455] The motion compensation unit 4502 can use interpolation filters, such as those used by the video encoder 4400 during the encoding of a video block, to calculate interpolations for sub-integer pixels of a reference block. The motion compensation unit 4502 can determine the interpolation filter used by the video encoder 4400 based on the received syntax information, and the motion compensation unit 4502 can use the interpolation filter to generate a prediction block.
[0456] The motion compensation unit 4502 may use some syntax information to determine the size of the blocks of (multiple) frames and / or (multiple) stripes used to encode the encoded video sequence, segmentation information describing how each macroblock of the picture of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame codec block, and other information for decoding the encoded video sequence.
[0457] Intra-prediction unit 4503 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 4504 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501. Inverse transform unit 4505 applies the inverse transform.
[0458] The reconstruction unit 4506 can add the residual block to the corresponding predicted block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be used to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 4507, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0459] Figure 8 This is a schematic diagram of an example encoder 4600. Encoder 4600 is suitable for implementing VVC techniques. Encoder 4600 includes three loop filters: a deblocking filter (DF) 4602, a sample adaptive compensation (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike DF 4602, which uses predefined filters, SAO 4604 and ALF 4606 utilize the original samples of the current image, reducing the mean square error between the original and reconstructed samples by adding offsets and applying finite impulse response (FIR) filters respectively, and utilizing the encoded / decoded side information through signal transmission offsets and filter coefficients. ALF 4606 is located in the last processing stage of each image and can be considered as a tool for attempting to capture and repair artifacts caused by previous stages.
[0460] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using a reference image obtained from a reference image buffer 4612. Residual blocks from inter-frame or intra-frame prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are then fed into an entropy codec component 4618. The entropy codec component 4618 entropy codes and decodes the prediction results and quantized transform coefficients and transmits them toward a video decoder (not shown). The quantization component output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. REC component 4624 is able to output images to DF 4602, SAO 4604 and ALF 4606 for filtering before these images are stored in reference image buffer 4612.
[0461] The following provides a list of preferred solutions as examples.
[0462] The following solutions illustrate examples of the embodiments discussed herein.
[0463] 1. A method for processing media data, comprising: determining whether the NNPF output image is in the full range when the color space of a neural network post-processing filter (NNPF) output image differs from the color space of a decoded image or the color space of a cropped decoded output image; and performing a conversion between visual media data and a bitstream based on the NNPF output image.
[0464] 2. The method according to Solution 1, wherein when the NNPF output image is in 4:2:0 color format, one or more syntax elements are transmitted via signal transmission to indicate the position of chroma sample points.
[0465] 3. The method according to any one of solutions 1-2, wherein when the NNPF purpose indicates coloring and the NNPF output image is in 4:2:0 color format, the position of the chroma sample points of the NNPF output image is specified by a signal transmission syntax element.
[0466] 4. The method according to any one of solutions 1-3, wherein a set of amplitude ratio related parameters is transmitted via signal in the NNPF Supplemental Enhancement Information (SEI) message.
[0467] 5. The method according to any one of solutions 1-4, wherein a signal transmission flag is used to indicate whether the NNPF processing changes the amplitude ratio.
[0468] 6. The method according to any one of solutions 1-5, wherein aspect_ratio_idc, sar_width, sar_height or a combination thereof are included in the NNPF SEI message according to the value of the flag.
[0469] 7. The method according to any one of solutions 1-6, wherein when aspect_ratio_idc, sar_width, and sar_height are not present in the bitstream, the aspect ratio attribute for the NNPF output image is presumed to be the same as the aspect ratio attribute for the decoded image or the aspect ratio attribute for the cropped decoded output image in the codec layer video sequence (CLVS).
[0470] 8. The method according to any one of solutions 1-7, wherein one or more syntax elements indicating the source scan type of the output of the NNPF process are conditionally transmitted via signal transmission, and wherein when the syntax element indicating the source scan type is not present, the source scan type for the NNPF output picture is presumed to be the same as the source scan type for the decoded picture or the source scan type for the cropped decoded output picture in the CLVS.
[0471] 9. The method according to any one of solutions 1-8, wherein one or more syntax elements indicating a preferred display mechanism for the output of the NNPF process are conditionally transmitted via signal transmission, and wherein when the syntax element indicating the preferred display mechanism is absent, the preferred display mechanism for the NNPF output picture is presumed to be the same as the preferred display mechanism for the decoded picture or the preferred display mechanism for the cropped decoded output picture in the CLVS.
[0472] 10. The method according to any one of solutions 1-9, wherein when using NNPF, StrengthControlVal is set to a value equal to (SliceQpY + QpBdOffset) ÷ (63 + QpBdOffset), where SliceQpY is the SliceQpY of the first slice of currCodedPic.
[0473] 11. The method according to any one of solutions 1-10, wherein when using NNPF, StrengthControlVal is set to a value equal to (SliceQpY + A) ÷ B, where SliceQpY is the SliceQpY of the first slice of currCodedPic, A is the maximum possible value of QpBdOffset, and B is the maximum possible value of SliceQpY.
[0474] 12. The method according to any one of solutions 1-11, wherein StrengthControlVal is transmitted via signaling in the NNPF SEI message or the Neural Network Post-Processing Filter Activation (NNPFA) SEI message.
[0475] 13. The method according to any one of solutions 1-12, wherein StrengthControlVal is transmitted via signal within a specific range and converted into the input value range of the NNPF.
[0476] 14. The method according to any one of solutions 1-13, wherein the chroma matrix must exist in the output tensor when the NNPF purpose indicates colorization, format change or upsampling.
[0477] 15. The method according to any one of solutions 1-14, wherein nnpfc_out_order_idc should not be equal to 0 when nnpfc_purpose & 0x20 is not equal to 0.
[0478] 16. The method according to any one of solutions 1-15, wherein when the input tensor does not contain a brightness matrix, the signaling of brightness fill values is skipped.
[0479] 17. The method according to any one of solutions 1-16, wherein when the input tensor does not contain a chroma matrix, signaling for chroma fill values is skipped.
[0480] 18. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of solutions 1-17.
[0481] 19. A non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec apparatus to perform the method according to any one of solutions 1-17 when executed by a processor.
[0482] 20. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining, when the color space of a neural network post-processing filter (NNPF) output image differs from the color space of a decoded image or a cropped decoded output image, whether the NNPF output image is in the full range by signal transmission syntax elements; and generating a bitstream based on the determination.
[0483] 21. A method for storing a bitstream of video, comprising: determining, when the color space of a neural network post-processing filter (NNPF) output image differs from the color space of a decoded image or the color space of a cropped decoded output image, whether the NNPF output image is in the full range by signaling a syntax element to indicate that the NNPF output image is in the full range; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0484] 22. A method, apparatus or system described in this disclosure.
[0485] In the described solution, the encoder conforms to the format rules by generating a codec representation based on those rules. In the described solution, the decoder parses the syntax elements in the codec representation using known information about their presence or absence, based on the format rules, to produce the decoded video.
[0486] In this disclosure, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, and vice versa. For example, the bitstream representation of the current video block can correspond to bits that are co-located or distributed at different locations in a bitstream defined by a syntax. For example, a macroblock can be encoded based on the error residual values after transformation and encoding / decoding, and also using bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can, based on this determination, parse the bitstream knowing whether some fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude specific syntax fields and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.
[0487] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of substances influencing machine-readable propagation signals, or one or more combinations thereof. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for an associated computer program, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof. Propagation signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.
[0488] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.
[0489] The processing and logic flows described in this disclosure can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processing and logic flows can also be executed by special-purpose logic circuitry, and the apparatus can also be implemented as special-purpose logic circuitry, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs).
[0490] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors in any type of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor storage devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0491] While this disclosure contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features specific to particular embodiments of this disclosure. Certain features described in the context of individual embodiments in this disclosure may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. Furthermore, although features may function in certain combinations as described above, and even were originally claimed in this manner, in some cases one or more features in a claimed combination may be removed from that combination, and the claimed combination may be for sub-combinations or variations thereof.
[0492] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed sequentially in the particular order or sequence shown, or requiring all shown operations to be performed in order to achieve the desired result. Furthermore, the partitioning of various system components in the embodiments described in this disclosure should not be construed as requiring such partitioning in all embodiments.
[0493] Only a few implementations and examples are described, and other implementations, improvements and variations may be made based on what is described and shown in this disclosure.
[0494] When there is no intermediary component (other than a line, trace, or other medium between the first and second components), the first component is directly coupled to the second component. When there is an intermediary component between the first and second components other than a line, trace, or other medium, the first component is indirectly coupled to the second component. The term "coupled" and its variations include direct coupling and indirect coupling. The use of the term "about" means including a range of ±10% of the following figures, unless otherwise specified.
[0495] While several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. The present examples are intended to be illustrative rather than limiting and are not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.
[0496] Furthermore, the technologies, systems, subsystems, and methods described and illustrated as discrete or separate in the various embodiments can be combined or integrated with other systems, modules, technologies, or methods without departing from the scope of this disclosure. Other items shown or discussed as couplings can be directly connected or indirectly coupled or communicated through some interface, device, or intermediate component (whether electrical, mechanical, or otherwise). Other examples of changes, substitutions, and modifications will be apparent to those skilled in the art upon reference to this disclosure and can be made without departing from the spirit and scope of this disclosure.
Claims
1. A method for processing media data, comprising: When the color space of the output image of the Neural Network Post-Processing Filter (NNPF) differs from the color space of the decoded image or the color space of the cropped decoded output image, it is determined to use a signal transmission syntax element to indicate whether the NNPF output image is in the full range. as well as The conversion between visual media data and bitstream is performed based on the NNPF output image.
2. The method of claim 1, wherein one or more syntax elements indicating the source scan type of the output of the NNPF process are conditionally transmitted via signal transmission, and wherein a first flag is transmitted via signal transmission to indicate whether a preferred scan type is transmitted via signal transmission.
3. The method according to any one of claims 1-2, wherein a second flag is transmitted via signal to indicate whether overscanning is preferred.
4. The method according to any one of claims 1-3, wherein the preferred scan type is transmitted via signal only when the image resolution changes.
5. The method according to any one of claims 1-4, wherein the scan type is presumed to be the same as the scan type indicated in the Video Availability Information (VUI) parameter.
6. The method according to any one of claims 1-5, wherein the constraint on the first syntax element (nnpfc_matrix_coeffs) depends on the presence of the second syntax element (nnpfc_out_tensor_chroma_bitdepth_minus8) and / or the third syntax element (nnpfc_out_tensor_luma_bitdepth_minus8) in the bitstream.
7. The method according to any one of claims 1-6, wherein the value of the first syntax element (nnpfc_matrix_coeffs) must not be equal to 0 if the following condition is met: a) Both the second syntax element (nnpfc_out_tensor_chroma_bitdepth_minus8) and the third syntax element (nnpfc_out_tensor_luma_bitdepth_minus8) exist in the bitstream; b) The value of the second syntax element (nnpfc_out_tensor_chroma_bitdepth_minus8) is equal to the value of the third syntax element (nnpfc_out_tensor_luma_bitdepth_minus8); c) The value of the fourth syntax element (nnpfc_out_order_idc) is equal to 2; d) The value of the first parameter (outSubHeightC) is equal to 1; and e) The value of the second parameter (outSubWidthC) is equal to 1.
8. The method according to any one of claims 1-7, wherein the value of the first syntax element (nnpfc_matrix_coeffs) must not be equal to 8 if the following condition is met: a) Both the second syntax element (nnpfc_out_tensor_chroma_bitdepth_minus8) and the third syntax element (nnpfc_out_tensor_luma_bitdepth_minus8) exist in the bitstream; b) Or: i) The value of the second syntax element (nnpfc_out_tensor_chroma_bitdepth_minus8) is equal to the value of the third syntax element (nnpfc_out_tensor_luma_bitdepth_minus8); or ii) The value of the second syntax element (nnpfc_out_tensor_chroma_bitdepth_minus8) is equal to the value of the third syntax element (nnpfc_out_tensor_luma_bitdepth_minus8) plus 1; c) The value of the fourth syntax element (nnpfc_out_order_idc) is equal to 2; d) The value of the first parameter (outSubHeightC) is equal to 1; and e) The value of the second parameter (outSubWidthC) is equal to 1.
9. The method according to any one of claims 1-8, wherein when the fifth syntax element (nnpfc_full_range_flag) is not present in the bitstream, the value of the fifth syntax element is presumed to be equal to the value of the sixth syntax element (vui_full_range_flag) or equal to 0.
10. The method according to any one of claims 1-9, wherein when the seventh syntax element (nnpfc_chroma_sample_loc_type_frame) is not present in the bitstream, the value of the seventh syntax element is presumed to be equal to 6.
11. The method according to any one of claims 1-10, wherein when the value of the eighth syntax element (nnpfc_out_colour_format_idc) is equal to 1 and the seventh syntax element (nnpfc_chroma_sample_loc_type_frame) is not present in the bitstream, the value of the seventh syntax element equal to 6 indicates that the location of the chroma sample is unknown or unspecified.
12. The method according to any one of claims 1-9, wherein the value of the seventh syntax element (nnpfc_chroma_sample_loc_type_frame) is presumed to be equal to the value of the ninth syntax element when the seventh syntax element (nnpfc_chroma_sample_loc_type_frame) is not present in the bitstream and satisfies either: a) when the value of the eighth syntax element (nnpfc_out_colour_format_idc) is equal to 1, or b) when the value of the eighth syntax element is equal to 1 and the ninth syntax element (vui_chroma_sample_loc_type_frame) is present in the bitstream.
13. The method according to any one of claims 1-12, wherein the signaling and / or presumption of the first syntax element (nnpfc_matrix_coeffs) and / or the fifth syntax element (nnpfc_full_range_flag) are independent of whether the sample values of the NNPF output are real numbers or integers.
14. The method according to any one of claims 1-13, wherein the signaling of the first syntax element (nnpfc_matrix_coeffs) and the fifth syntax element (nnpfc_full_range_flag) is not conditional on the tenth syntax element (nnpfc_out_format_idc).
15. An apparatus for processing video data, comprising: A processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-14.
16. A non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec apparatus to perform the method according to any one of claims 1-14 when executed by a processor.
17. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: When the color space of the output image of the Neural Network Post-Processing Filter (NNPF) differs from the color space of the decoded image or the color space of the cropped decoded output image, it is determined to use a signal transmission syntax element to indicate whether the NNPF output image is in the full range. as well as Based on the determination, a bit stream is generated.
18. A method for storing a bitstream of video, comprising: When the color space of the output image of the Neural Network Post-Processing Filter (NNPF) differs from the color space of the decoded image or the color space of the cropped decoded output image, it is determined to use a signal transmission syntax element to indicate whether the NNPF output image is in the full range. Based on the determination, a bit stream is generated; as well as The bit stream is stored in a non-transitory computer-readable recording medium.