Value range of neural network post-processing filter about syntax element and encoding and decoding method

By using an encoding and decoding method for the input format indicator of the neural network post-processing filter characteristics, the problem of low video encoding and decoding efficiency in existing technologies is solved, achieving efficient video data processing and meeting the network transmission requirements with high bandwidth demands.

CN121464633APending Publication Date: 2026-02-03DOUYIN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480046078.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-07
Filing Date
2024-07-05
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies lack effective syntax elements and encoding/decoding methods when dealing with the characteristics of neural network post-processing filters, resulting in low efficiency in video data processing and failing to meet the needs of high-bandwidth Internet and digital communication networks.

Method used

The encoding and decoding method using Neural Network Post-Processing Filter Characteristics (NNPFC) input format indicator optimizes the video encoding and decoding process by determining the value of the NNPFC input format indicator as the syntax element of the encoding and decoding u(N). Based on this, the conversion between video data and bitstream is performed.

Benefits of technology

It improves the efficiency and quality of video data processing, meets the video transmission requirements of the high-bandwidth Internet and digital communication networks, and enhances the performance of video encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121464633A_ABST
    Figure CN121464633A_ABST
Patent Text Reader

Abstract

A mechanism for processing video data is disclosed. The mechanism includes determining a value of a neural network post-processing filter characteristic (NNPFC) input format indicator (nnpfcinformatinc) to be encoded as a syntax element for u (N) encoding, where N is an integer greater than 0. Conversion between the visual media data and the bitstream is performed based on the NNPFC input format indicator.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This patent application claims the benefit of U.S. Provisional Patent Application No. 63 / 512,346, filed July 7, 2023, which is incorporated herein by reference. Technical Field

[0003] This disclosure relates to the generation, storage, and use of digital audio and video media information in file formats. Background Technology

[0004] Digital video consumes the largest share of bandwidth on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is likely to continue to grow. Summary of the Invention

[0005] The first aspect relates to a method for processing media data, comprising: determining that the value of a neural network post-processing filter characteristic (NNPFC) input format indicator (nnpfc_inp_format_inc) is encoded and decoded into a syntax element of u(N) encoding and decoding, wherein N is an integer greater than 0; and performing a conversion between visual media data and a bitstream based on the NNPFC input format indicator.

[0006] The second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform any of the aspects described above.

[0007] The third aspect relates to a non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, the video codec device performs the methods of any of the preceding aspects.

[0008] The fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining that the value of a neural network post-processing filter characteristic (NNPFC) input format indicator (nnpfc_inp_format_inc) is encoded and decoded into a syntax element of u(N) encoding and decoding, wherein N is an integer greater than 0; and generating the bitstream based on the NNPFC input format indicator.

[0009] The fifth aspect relates to a method for storing a bitstream of video, comprising: determining that the value of a neural network post-processing filter characteristic (NNPFC) input format indicator (nnpfc_inp_format_inc) is encoded and decoded into a syntax element of u(N) encoding and decoding, wherein N is an integer greater than 0; generating the bitstream based on the NNPFC input format indicator; and storing the bitstream in a non-transitory computer-readable recording medium.

[0010] The sixth aspect relates to the methods, apparatus, or systems described in this disclosure.

[0011] For clarity, any of the embodiments described above may be combined with any one or more other foregoing embodiments to create new embodiments within the scope of this disclosure.

[0012] These and other features will be more clearly understood through the following detailed description in conjunction with the accompanying drawings and claims. Attached Figure Description

[0013] To gain a more complete understanding of this disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals denote like parts.

[0014] Figure 1 An example of deriving the luminance channel from the luminance component is shown.

[0015] Figure 2 This is a block diagram illustrating an example video processing system.

[0016] Figure 3 This is a block diagram of an example video processing device.

[0017] Figure 4 This is a flowchart of an example method for video processing according to embodiments of the present disclosure.

[0018] Figure 5 This is a block diagram illustrating an example video codec system.

[0019] Figure 6 This is a block diagram showing an example encoder.

[0020] Figure 7 This is a block diagram showing an example decoder.

[0021] Figure 8 This is a schematic diagram of an example encoder. Detailed Implementation

[0022] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or yet to be developed. This disclosure should not be limited in any way to the illustrative implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but can be modified within the scope of the appended claims, together with their entire equivalents.

[0023] The use of chapter headings in this disclosure is for ease of understanding and not to limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter only. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs. In this disclosure, editorial changes relative to the Multi-Functional Video Codec (VVC) specification are indicated in the text by bold italics (indicating deleted text) and bold (indicating added text).

[0024] 1. Overview

[0025] This disclosure relates to image / video encoding and decoding techniques. Specifically, it relates to the value range and encoding / decoding methods of Neural Network Post-Processing Filter (NNPF) syntax elements, and to encoding / decoding methods of indicator syntax elements in NNPF Supplemental Enhancement Information (SEI) messages. These ideas can be applied individually or in various combinations for video bitstreams encoded and decoded by any codec, such as the Multi-Functional Video Codec (VVC) standard and / or the Multi-Functional SEI Message (VSEI) standard for encoding and decoding video bitstreams.

[0026] 2. Abbreviations

[0027] The following abbreviations may be used in this document: Adaptive Parameter Set (APS), Access Unit (AU), Codec Layer Video Sequence (CLVS), Codec Layer Video Sequence Start (CLVSS), Cyclic Redundancy Check (CRC), Codec Video Sequence (CVS), Finite Impulse Response (FIR), Intra-Frame Random Access Point (IRAP), Internet Engineering Task Force (IETF), Network Abstraction Layer (NAL), Picture Parameter Set (PPS), Picture Unit (PU), Random Access Skip Before (RASL) Picture, Supplemental Enhancement Information (SEI), Stepped Temporal Sublayer Access (STSA), Uniform Resource Identifier (URI), Video Codec Layer (VCL), Multifunctional Supplemental Enhancement Information (VSEI) and Video Availability Information (VUI) described in Recommendation ITU-T H.274 | ISO / IEC 23002-7, and Multifunctional Video Codec (VVC) described in Recommendation ITU-T H.266 | ISO / IEC 23090-3.

[0028] 3. Further discussion

[0029] 3.1 Video Coding and Decoding Standards

[0030] Video coding standards have evolved primarily through the development of standards by the well-known International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) and the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). ITU-T developed the H.261 and H.263 standards, ISO / IEC developed the Moving Picture Experts Group (MPEG)-1 and MPEG-4 Vision, and the two organizations jointly developed the H.262 / MPEG-2 video standard and the H.264 / MPEG-4 Advanced Video Coding (AVC) standard and the H.265 / HEVC[1]. Starting with H.262, video coding standards are based on a hybrid video coding architecture, which utilizes temporal prediction plus transform coding. In order to explore future video coding technologies beyond High Efficiency Video Coding (HEVC), the Joint Video Exploration Team (JVET) was jointly established by the Video Coding Experts Group (VCEG) and MPEG in 2015. Since then, JVET has adopted many new approaches and incorporated them into a reference software called the Joint Exploration Model (JEM)[2]. Later, when the Multifunctional Video Coding (VVC) project was officially launched, JVET was renamed the Joint Video Experts Group (JVET). VVC[3] is a new coding standard that aims to reduce the bit rate by 50% compared to HEVC. The standard was finalized by JVET at its 19th meeting, which ended on July 1, 2020.

[0031] The Multi-Functional Video Coding (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) [3] and the associated Multi-Functional Supplemental Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) [4] for codec video bitstreams have been designed for the widest range of applications, including traditional uses such as television broadcasting, video conferencing or playback from storage media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, synthesis and merging of content from multiple codec video bitstreams, multi-view video, scalable layered coding and decoding and viewport-adaptive 360° immersive media.

[0032] The Basic Video Codec (EVC) standard (ISO / IEC 23094-1) is another video codec standard recently developed by MPEG.

[0033] 3.2 General SEI messages and SEI messages in VVC and VSEI

[0034] SEI messages assist in processes related to decoding, display, or other purposes. However, SEI messages are not essential for constructing luma or chroma samples during the decoding process. Standard-compliant decoders do not need to process this information to achieve output order consistency. Some SEI messages are necessary for checking bitstream consistency and output timing decoder consistency. Other SEI messages are not necessary for checking bitstream consistency.

[0035] Appendix D of VVC specifies the syntax and semantics of SEI message payloads for some SEI messages, and specifies the use of SEI messages and VUI parameters with the syntax and semantics specified in ITU-TH.274 | ISO / IEC 23002-7.

[0036] 3.3 Signaling of Neural Network Post-processing Filters

[0037] JVET-AD2006[5] includes specifications for two SEI messages used for signaling of neural network post-processing filters: the Neural Network Post-Processing Filter Feature (NNPFC) SEI message and the Neural Network Post-Processing Filter Activation (NNPFA) SEI message. JVET-AD2005[6] includes specifications for the use of NNPFC SEI messages in VVC bitstreams.

[0038] The specifications for NNPFC and NNPFA SEI messages in JVET-AD2006 and the usage specifications for NNPFC SEI messages in VVC bitstreams are as follows. In addition, Working Group (WG) 05 output document N0158[5] and JVET-AC2032[6] include the specifications for two SEI messages for signaling used in neural network post-processing filters, as shown below.

[0039] 8.28 Neural Network Post-Processing Filter SEI Message

[0040] 8.28.1 General Post-Processing Filtering Procedure Using NNPF

[0041] 8.28.1.1 Overview

[0042] The input to this process is a bitstream, BitstreamToFilter. The output is a list of NNPF output images, ListNnpfOutputPics.

[0043] First, BitstreamToFilter is decoded, and the list CroppedDecodedPictures is set to a list of cropped decoded images generated by decoding BitstreamToFilter in output order.

[0044] Second, for each cropped decoded picture in CroppedDecodedPictures and with one or more NNPFs activated, the filtering process for a picture as specified in sub-entry 8.28.1.2 is called repeatedly in the output order.

[0045] The order of the images in ListNnpfOutputPics is the output order.

[0046] Within ListNnpfOutputPics, there should be no more than one image associated with any given output time instance. When multiple NNPFs are activated for any given image in CroppedDecodedPictures, and only one of the multiple NNPFs can be selected for application (although any NNPF can be selected), the above constraints apply regardless of which NNPF is selected for that particular image.

[0047] 8.28.1.2 Filtering process for an image using NNPF

[0048] The filtering process specified in this sub-entry is applied to each cropped decoded picture (referred to as the current picture) in CroppedDecodedPictures that has one or more NNPF activated.

[0049] When NNPF is applied to the current image, the filtered and / or interpolated image is generated by NNPF by applying the NNPF procedure specified in the semantics of the NNPFC SEI message to the current image in a block-by-block manner.

[0050] When NNPF is applied to the current image, the order in which the images generated by NNPF through the NNPF application process are stored in the output tensor of NNPF is the output order.

[0051] When the applied NNPF is the last NNPF applied to the current image, the images generated by the NNPF and output by the NNPF process are included in ListNnpfOutputPics in the same order as when the images were stored in the NNPF output tensor.

[0052] 8.28.2 Characteristics of Neural Network Post-Processing Filters (SEI Messages)

[0053] 8.28.2.1 Characteristics of Neural Network Post-Processing Filters and SEI Message Syntax

[0054]

[0055]

[0056]

[0057]

[0058] 8.28.2.2 Characteristics of Neural Network Post-Processing Filters and Semantics of SEI Messages

[0059] The Neural Network Post-Processing Filter Characteristics (NNPFC) SEI message specifies the neural networks that can be used as post-processing filters. The Neural Network Post-Processing Filter Activation (NNPFA) SEI message indicates the use of a specified Neural Network Post-Processing Filter (NNPF) for a particular image.

[0060] To use this SEI message, the following variables need to be defined:

[0061] - Input the image width and height, in units of brightness samples, which are referred to as CroppedWidth and CroppedHeight in this article.

[0062] - The luminance sample array CroppedYPic[idx] and chrominance sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] (if present) of the input image, where the index idx ranges from 0 to numInputPics-1 (inclusive of boundary values), which are used as input to NNPF.

[0063] - BitDepth of the luminance sample array used for the input image Y .

[0064] - Bit depth of the chroma sample array (if any) used for the input image C .

[0065] - Chroma format indicator, referred to as ChromaFormatIdc in this document, as described in sub-entry 7.3.

[0066] - When nnpfc_auxiliary_inp_idc equals 1, the filter strength control value array StrengthControlVal[idx] must contain real numbers in the range of 0 to 1 (inclusive) for input images with index idx in the range of 0 to numInputPics-1.

[0067] The input image at index 0 corresponds to the image whose NNPF is activated by the NNPFA SEI message defined by that NNPFC SEI message. Input images with indices i in the range of 1 to numInputPics-1 (inclusive) precede the input image at index i-1 in the output order.

[0068] The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc, as specified in Table 1.

[0069] Note 1 – More than one NNPFC SEI message can exist for the same image. When more than one NNPFC SEI message with different nnpfc_id values ​​exists or is activated for the same image, their nnpfc_purpose and nnpfc_mode_idc values ​​can be the same or different.

[0070] `nnpfc_purpose` indicates the purpose of the NNPF, as specified in Table 1, where `(nnpfc_purpose & bitMask)` not equal to 0 indicates that the NNPF has a purpose associated with the `bitMask` value in Table 1. When `nnpfc_purpose` is greater than 0 and `(nnpfc_purpose & bitMask)` is equal to 0, the purpose associated with the `bitMask` value is not applicable to the NNPF. When `nnpfc_purpose` is equal to 0, the NNPF can be used based on the determination of the application.

[0071] In bitstreams conforming to this version of this document, the value of nnpfc_purpose must be in the range of 0 to 63 (inclusive). Values ​​of nnpfc_purpose from 64 to 65,535 (inclusive) are reserved for future use by ITU-T | ISO / IEC and should not exist in bitstreams conforming to this version of this document. Decoders conforming to this version of this document must ignore NNPFC SEI messages with nnpfc_purpose in the range of 64 to 65,535 (inclusive).

[0072] Table 1 – Definition of nnpfc_purpose

[0073]

[0074] The variables chromaUpsamplingFlag, resolutionResamplingFlag, pictureRateUpsamplingFlag, bitDepthUpsamplingFlag, and colourizationFlag respectively specify whether nnpfc_purpose indicates the purpose of NNPF, including chroma upsampling, resolution resampling, picture rate upsampling, bit depth upsampling, and colourization. These variables are derived as follows:

[0075] chromaUpsamplingFlag = ( ( nnpfc_purpose & 0x02 ) > 0 ) ? 1 : 0

[0076] resolutionResamplingFlag = ( ( nnpfc_purpose & 0x04 ) > 0 ) ? 1 : 0

[0077] pictureRateUpsamplingFlag = ( ( nnpfc_purpose & 0x08 ) > 0 ) ? 1 : 0(76)

[0078] bitDepthUpsamplingFlag = ( ( nnpfc_purpose & 0x10 ) > 0 ) ? 1 : 0

[0079] colourizationFlag = ( ( nnpfc_purpose & 0x20 ) > 0 ) ? 1 : 0

[0080] Note 2 – When the reserved value of nnpfc_purpose is used by ITU-T | ISO / IEC in the future, the syntax of this SEI message can be extended using the following syntax elements, provided that nnpfc_purpose is equal to that value.

[0081] When ChromaFormatIdc equals 3, chromaUpsamplingFlag must equal 0.

[0082] When ChromaFormatIdc or chromaUpsamplingFlag is not equal to 0, colourizationFlag must be equal to 0.

[0083] When an input image with pictureRateUpsamplingFlag equal to 1 and index 0 is associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5, all input images are associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5 and the same fp_current_frame_is_frame0_flag value.

[0084] nnpfc_id contains an identifier that can be used to identify NNPF. The value of nnpfc_id must be between 0 and 2. 32 The range is -2 (inclusive). The value of nnpfc_id is from 256 to 511 (inclusive) and from 2... 31 Up to 2 32 -2 (including boundary values) is reserved for future use by ITU-T | ISO / IEC. Decoders conforming to this document for this version will respond when encountering nnpfc_id in the range of 256 to 511 (including boundary values) or 2. 31 Up to 2 32 When an NNPFC SEI message is received within the range of -2 (including boundary values), the SEI message must be ignored.

[0085] The following applies when the NNPFC SEI message is the first NNPFC SEI message in the current CLVS with a specific nnpfc_id value and in decoding order:

[0086] - This SEI message specifies the underlying NNPF.

[0087] - This SEI message applies to the currently decoded image and all subsequent decoded images of the current layer in output order, until the end of the current CLVS.

[0088] An nnpfc_base_flag value of 1 indicates that the SEI message specifies the underlying NNPF. An nnpf_base_flag value of 0 indicates that the SEI message specifies an update relative to the underlying NNPF.

[0089] The following constraints apply to the value of nnpfc_base_flag:

[0090] - When the NNPFC SEI message is the first NNPFC SEI message in the current CLVS with a specific nnpfc_id value in decoding order, the value of nnpfc_base_flag must be equal to 1.

[0091] - When NNPFC SEI message nnpfcB is not the first NNPFC SEI message in the current CLVS with a specific nnpfc_id value in decoding order, and the value of nnpfc_base_flag is equal to 1, the NNPFC SEI message must be a duplicate of the first NNPFC SEI message nnpfcA with the same nnpfc_id value in decoding order, that is, the payload content of nnpfcB must be the same as the payload content of nnpfcA.

[0092] When nnpfc_base_flag equals 0, the following applies:

[0093] This SEI message defines an update to the base NNPF relative to the first decoded NNPF with the same nnpfc_id value. Updates are not cumulative; rather, each update is applied to the base NNPF, which is defined by the first NNPFC SEI message in the current CLVS with a specific nnpfc_id value in decoded order. The NNPF defined by this SEI message is obtained by applying the update defined by this SEI message relative to the base NNPF with the same nnpfc_id value.

[0094] - This SEI message applies to the current decoded image and all subsequent decoded images of the current layer in output order, up to the end of the current CLVS, or up to but not including the decoded image within the current CLVS that is in output order after the current decoded image and is associated with a subsequent NNPFC SEI message within the current CLVS in decoding order, whichever has an nnpfc_base_flag equal to 0 and the specific nnpfc_id value, whichever is earlier.

[0095] An nnpfc_mode_idc value of 0 indicates that the SEI message contains an ISO / IEC 15938-17 bitstream that specifies the underlying NNPF (when nnpfc_base_flag equals 1) or an update relative to the underlying NNPF with the same nnpfc_id value (when nnpfc_base_flag equals 0).

[0096] When nnpfc_base_flag equals 1, nnpfc_mode_idc equals 1, which specifies that the underlying NNPF associated with the nnpfc_id value is a neural network identified by the URI indicated by nnpfc_uri, which has a format identified by the tag URI nnpfc_tag_uri.

[0097] When nnpfc_base_flag equals 0, nnpfc_mode_idc equals 1, indicating that the update relative to the base NNPF with the same nnpfc_id value is defined by the URI indicated by nnpfc_uri, which has a format identified by the tag URI nnpfc_tag_uri.

[0098] In bitstreams conforming to this version of this document, the value of nnpfc_mode_idc must be in the range of 0 to 1 (inclusive). Values ​​of nnpfc_mode_idc from 2 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and should not exist in bitstreams conforming to this version of this document. Decoders conforming to this version of this document must ignore NNPFC SEI messages with nnpfc_mode_idc in the range of 2 to 255 (inclusive). Values ​​of nnpfc_mode_idc greater than 255 should not exist in bitstreams conforming to this version of this document and are not reserved for future use.

[0099] nnpfc_reserved_zero_bit_a must be equal to 0 in the bitstream conforming to this version of the document. The decoder must ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_a is not equal to 0.

[0100] The nnpfc_tag_uri contains a tag URI that has the syntax and semantics as specified in IETF RFC 4151, identifying the format and related information of a neural network used as the underlying NNPF or an updated NNPF with the same nnpfc_id value as specified by the nnpfc_uri.

[0101] Note 3 – nnpfc_tag_uri can uniquely identify the format of neural network data as specified by nnrpf_uri, without the need for a central registry.

[0102] The nnpfc_tag_uri equal to “tag:iso.org,2023:15938-17” indicates that the neural network data identified by nnpfc_uri conforms to ISO / IEC 15938-17.

[0103] nnpfc_uri contains a URI that has the syntax and semantics as specified in IETF Internet Standard 66, identifying a neural network used as the underlying NNPF or an updated NNPF relative to the underlying NNPF with the same nnpfc_id value.

[0104] The value of nnpfc_property_present_flag equal to 1 indicates the existence of syntax elements related to the filter's purpose, input format, output format, and complexity. The value of nnpfc_property_present_flag equal to 0 indicates the absence of syntax elements related to the filter's purpose, input format, output format, and complexity.

[0105] When nnpfc_base_flag is equal to 1, nnpfc_property_present_flag must be equal to 1.

[0106] When nnpfc_property_present_flag equals 0, the values ​​of all syntax elements that can only exist when nnpfc_property_present_flag equals 1 are presumed to be equal to the corresponding syntax elements in the NNPFC SEI message containing the underlying NNPF that the SEI message updates.

[0107] The following constraints apply when the NNPFC SEI message nnpfcCurr is not the first NNPFC SEI message with a specific nnpfc_id value in the current CLVS in decoding order, is not a duplicate of the first NNPFC SEI message with that specific nnpfc_id (i.e., the value of nnpfc_base_flag is equal to 0), and the value of nnpfc_property_present_flag is equal to 1:

[0108] The value of nnpfc_purpose in the NNPFC SEI message must be the same as the value of nnpfc_purpose in the first NNPFC SEI message in the current CLVS that has that particular nnpfc_id value in the decoding order.

[0109] - The value of the syntax element in the NNPFC SEI message that is after nnpfc_property_present_flag and before nnpfc_complexity_info_present_flag in the decoding order must be the same as the value of the corresponding syntax element in the first NNPFC SEI message in the current CLVS that has that particular nnpfc_id value in the decoding order.

[0110] - The nnpfc_complexity_info_present_flag in the first NNPFC SEI message (hereinafter referred to as nnpfcBase) in the current CLVS with this specific nnpfc_id value in decoding order must be equal to 0 or both nnpfc_complexity_info_present_flag must be equal to 1, and all of the following must apply:

[0111] – The nnpfc_parameter_type_idc in nnpfcCurr must be equal to the nnpfc_parameter_type_idc in nnpfcBase.

[0112] – nnpfc_log2_parameter_bit_length_minus3 in nnpfcCurr (if it exists) must be less than or equal to nnpfc_log2_parameter_bit_length_minus3 in nnpfcBase.

[0113] – If nnpfc_num_parameters_idc in nnpfcBase is equal to 0, then nnpfc_num_parameters_idc in nnpfcCurr must be equal to 0.

[0114] Otherwise (nnpfc_num_parameters_idc in nnpfcBase is greater than 0), nnpfc_num_parameters_idc in nnpfcCurr must be greater than 0 and less than or equal to nnpfc_num_parameters_idc in nnpfcBase.

[0115] – If nnpfc_num_kmac_operations_idc in nnpfcBase is equal to 0, then nnpfc_num_kmac_operations_idc in nnpfcCurr must be equal to 0.

[0116] – Otherwise (nnpfc_num_kmac_operations_idc in nnpfcBase is greater than 0), nnpfc_num_kmac_operations_idc in nnpfcCurr must be greater than 0 and less than or equal to nnpfc_num_kmac_operations_idc in nnpfcBase.

[0117] – If nnpfc_total_kilobyte_size in nnpfcBase is equal to 0, then nnpfc_total_kilobyte_size in nnpfcCurr must be equal to 0.

[0118] – Otherwise (nnpfc_total_kilobyte_size in nnpfcBase is greater than 0), nnpfc_total_kilobyte_size in nnpfcCurr must be greater than 0 and less than or equal to nnpfc_total_kilobyte_size in nnpfcBase.

[0119] The increment of 1 in `nnpfc_num_input_pics_minus1` specifies the number of images used as input to NNPF. The value of `nnpfc_num_input_pics_minus1` must be in the range of 0 to 63 (inclusive). When `pictureRateUpsamplingFlag` equals 1, the value of `nnpfc_num_input_pics_minus1` must be greater than 0.

[0120] The variable numInputPics, which specifies the number of images used as input to NNPF, is derived as follows:

[0121] numInputPics = nnpfc_num_input_pics_minus1 + 1 (77)

[0122] `nnpfc_input_pic_output_flag[i]` equal to 1 indicates that NNPF generates the corresponding output image for the i-th input image. `nnpfc_input_pic_output_flag[i]` equal to 0 indicates that NNPF does not generate the corresponding output image for the i-th input image. When `nnpfc_num_input_pics_minus1` equals 0, `nnpfc_input_pic_output_flag[0]` is presumed to be equal to 1. When `pictureRateUpsamplingFlag` equals 0 and `nnpfc_num_input_pics_minus1` is greater than 0, `nnpfc_input_pic_output_flag[i]` must be equal to 1 for at least one value of `i` within the range of 0 to `nnpfc_num_input_pics_minus1` (inclusive).

[0123] `nnpfc_absent_input_pic_zero_flag` equal to 1 indicates that NNPF expects input images not present in the bitstream to be represented by an array of samples with sample values ​​equal to 0. `nnpfc_absent_input_pic_zero_flag` equal to 0 indicates that NNPF expects input images not present in the bitstream to be represented by the closest input image in the bitstream in output order.

[0124] `nnpfc_out_sub_c_flag` specifies the values ​​of variables `outSubWidthC` and `outSubHeightC` when `chromaUpsamplingFlag` equals 1. `nnpfc_out_sub_c_flag` equal to 1 specifies that both `outSubWidthC` and `outSubHeightC` equal to 1. `nnpfc_out_sub_c_flag` equal to 0 specifies that `outSubWidthC` equals 2 and `outSubHeightC` equals 1. When `ChromaFormatIdc` equals 2 and `nnpfc_out_sub_c_flag` exists, the value of `nnpfc_out_sub_c_flag` must be equal to 1.

[0125] When `colourizationFlag` equals 1, `nnpfc_out_colour_format_idc` specifies the color format of the NNPF output, thus defining the values ​​of the variables `outSubWidthC` and `outSubHeightC`. `nnpfc_out_colour_format_idc` equals 1 specifies that the NNPF output color format is 4:2:0, and both `outSubWidthC` and `outSubHeightC` equal 2. `nnpfc_out_colour_format_idc` equals 2 specifies that the NNPF output color format is 4:2:2, with `outSubWidthC` equal to 2 and `outSubHeightC` equal to 1. `nnpfc_out_colour_format_idc` equals 3 specifies that the NNPF output color format is 4:4:4, with both `outSubWidthC` and `outSubHeightC` equal to 1. The value of `nnpfc_out_colour_format_idc` should not be 0.

[0126] When both chromaUpsamplingFlag and colourizationFlag are equal to 0, outSubWidthC and outSubHeightC are presumed to be equal to SubWidthC and SubHeightC, respectively.

[0127] `nnpfc_pic_width_num_minus1 + 1` and `nnpfc_pic_width_denom_minus1 + 1` define the numerator and denominator of the resampling ratio of the NNPF output image width relative to `CroppedWidth`, respectively. The value of `(nnpfc_pic_width_num_minus1 + 1) ÷ (nnpfc_pic_width_denom_minus1 + 1)` must be in the range of 1 ÷ 16 to 16 (inclusive). When `nnpfc_pic_width_num_minus1` and `nnpfc_pic_width_denom_minus1` do not exist, their values ​​are presumed to be 0.

[0128] The variable nnpfcOutputPicWidth, representing the width of the luminance sample array of the (multiple) images generated by applying the NNPF identified by nnpfc_id to (multiple) input images, is derived as follows:

[0129] nnpfcOutputPicWidth = Ceil( CroppedWidth ( nnpfc_pic_width_num_minus1 + 1 ) ÷ ( nnpfc_pic_width_denom_minus1 + 1 ) ) (78)

[0130] The value of nnpfcOutputPicWidth % outSubWidthC must be equal to 0 as a requirement for bitstream consistency.

[0131] `nnpfc_pic_height_num_minus1 + 1` and `nnpfc_pic_height_denom_minus1 + 1` define the numerator and denominator of the resampling ratio of the NNPF output image height relative to `CroppedHeight`, respectively. The value of `(nnpfc_pic_height_num_minus1 + 1) ÷ (nnpfc_pic_height_denom_minus1 + 1)` must be in the range of 1 ÷ 16 to 16 (inclusive). When `nnpfc_pic_height_num_minus1` and `nnpfc_pic_height_denom_minus1` do not exist, their values ​​are presumed to be equal to 0.

[0132] The variable nnpfcOutputPicHeight, representing the height of the luminance sample array of the resulting image(s) generated by applying the NNPF identified by nnpfc_id to (multiple) input images, is derived as follows:

[0133] nnpfcOutputPicHeight = Ceil( CroppedHeight ( nnpfc_pic_height_num_minus1 + 1 ) ÷ ( nnpfc_pic_height_denom_minus1 + 1 ) ) (79)

[0134] The requirement for bitstream consistency is that the value of nnpfcOutputPicHeight % outSubHeightC must be equal to 0.

[0135] When nnpfc_pic_width_num_minus1, nnpfc_pic_width_denom_minus1, nnpfc_pic_height_num_minus1, and nnpfc_pic_height_denom_minus1 exist, at least one of the following must be true:

[0136] The value of nnpfcOutputPicWidth is not equal to CroppedWidth.

[0137] The value of nnpfcOutputPicHeight is not equal to CroppedHeight.

[0138] `nnpfc_interpolated_pics[i]` specifies the number of interpolated pictures generated by NNPF between the i-th picture and the (i+1)-th picture used as input to NNPF. The value of `nnpfc_interpolated_pics[i]` must be in the range of 0 to 63 (inclusive). For at least one value of `i` in the range of 0 to `nnpfc_num_input_pics_minus1-1` (inclusive), the value of `nnpfc_interpolated_pics[i]` must be greater than 0.

[0139] The variables NumInpPicsInOutputTensor, which specify the number of images with corresponding input images and existing in the NNPF output tensor, InpIdx[idx], which specifies the input image index of the idx-th image existing in the NNPF output tensor and having corresponding input images, and numOutputPics, which specifies the total number of images existing in the NNPF output tensor, are derived as follows:

[0140] for( i = 0, numOutputPics = 0; i < numInputPics; i++ )

[0141] if( nnpfc_input_pic_output_flag[ i ] ) {

[0142] InpIdx[ numOutputPics ] = i

[0143] numOutputPics++

[0144] } (80)

[0145] NumInpPicsInOutputTensor = numOutputPics

[0146] if( pictureRateUpsamplingFlag )

[0147] for( i = 0; i <= numInputPics - 2; i++ )

[0148] numOutputPics += nnpfc_interpolated_pics[ i ]

[0149] A value of 1 for nnpfc_component_last_flag indicates that the last dimension of both the input tensor (inputTensor) and the output tensor (outputTensor) generated by NNPF is used for the current channel. A value of 0 for nnpfc_component_last_flag indicates that the third dimension of both the input tensor (inputTensor) and the output tensor (outputTensor) generated by NNPF is used for the current channel.

[0150] Note 4 – The first dimension in both the input and output tensors is used for batch indexing, a practice in some neural network frameworks. While the formula in the semantics of this SEI message uses a batch size corresponding to a batch index equal to 0, the batch size used as input for neural network inference is determined by the post-processing implementation.

[0151] Note 5 – For example, when nnpfc_inp_order_idc equals 3 and nnpfc_auxiliary_inp_idc equals 1, the input tensor has 7 channels, including four luminance matrices, two chrominance matrices, and one auxiliary input matrix. In this case, the procedure DeriveInputTensors() will derive each of these 7 channels of the input tensor one by one, and when a particular channel of these channels is processed, that channel is referred to as the current channel during the procedure.

[0152] `nnpfc_inp_format_idc` specifies the method for converting the sample values ​​of the input image into NNPF input values. When `nnpfc_inp_format_idc` equals 0, the NNPF input values ​​are real numbers, and the functions `InpY()` and `InpC()` are defined as follows:

[0153] InpY( x ) = x ÷ ( ( 1 << BitDepth Y ) - 1 ) (81)

[0154] InpC( x )= x ÷ ( ( 1 << BitDepth C ) - 1 ) (82)

[0155] When nnpfc_inp_format_idc equals 1, the input values ​​of NNPF are unsigned integers, and the functions InpY() and InpC() are defined as follows:

[0156] shiftY = BitDepth Y - inpTensorBitDepth Y

[0157] if (inpTensorBitDepth) Y >= BitDepth Y )

[0158] InpY( x ) = x << ( inpTensorBitDepth Y - BitDepth Y (83)

[0159] else

[0160] InpY( x ) = Clip3(0, ( 1 << inpTensorBitDepth Y ) - 1, ( x + ( 1 << (shiftY - 1 ) ) ) >> shiftY )

[0161] shiftC = BitDepth C - inpTensorBitDepth C

[0162] if (inpTensorBitDepth) C >= BitDepth C )

[0163] InpC( x ) = x << ( inpTensorBitDepth C - BitDepth C (84)

[0164] else

[0165] InpC( x ) = Clip3(0, ( 1 << inpTensorBitDepth C ) - 1, ( x + ( 1 << (shiftC - 1 ) ) ) >> shiftC )

[0166] The variable inpTensorBitDepthY is deduced from the syntax element nnpfc_inp_tensor_luma_bitdepth_minus8 as specified below. The variable inpTensorBitDepthC is deduced from the syntax element nnpfc_inp_tensor_chroma_bitdepth_minus8 as specified below.

[0167] Values ​​greater than 1 for nnpfc_inp_format_idc are reserved for future ITU-T | ISO / IEC specifications and should not be present in bitstreams conforming to this version of the document. Decoders conforming to this version of the document must ignore NNPFC SEI messages containing reserved values ​​for nnpfc_inp_format_idc.

[0168] A value greater than 0 for nnpfc_auxiliary_inp_idc indicates that the auxiliary input data exists in the input tensor of NNPF. A value equal to 0 for nnpfc_auxiliary_inp_idc indicates that the auxiliary input data does not exist in the input tensor. A value equal to 1 for nnpfc_auxiliary_inp_idc indicates that the auxiliary input data is derived according to Equation 85.

[0169] In bitstreams conforming to this version of this document, the value of nnpfc_auxiliary_inp_idc must be in the range of 0 to 1 (inclusive). Values ​​of nnpfc_auxiliary_inp_idc from 2 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and should not exist in bitstreams conforming to this version of this document. Decoders conforming to this version of this document must ignore NNPFC SEI messages with nnpfc_auxiliary_inp_idc in the range of 2 to 255 (inclusive). Values ​​of nnpfc_auxiliary_inp_idc greater than 255 should not exist in bitstreams conforming to this version of this document and are not reserved for future use.

[0170] nnpfc_inp_order_idc indicates the method for sorting the sample array of the input image to form the input tensor of NNPF.

[0171] In bitstreams conforming to this version of this document, the value of nnpfc_inp_order_idc must be in the range of 0 to 3 (inclusive). Values ​​of nnpfc_inp_order_idc from 4 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and should not exist in bitstreams conforming to this version of this document. Decoders conforming to this version of this document must ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 4 to 255 (inclusive). Values ​​of nnpfc_inp_order_idc greater than 255 should not exist in bitstreams conforming to this version of this document and are not reserved for future use.

[0172] When ChromaFormatIdc is not equal to 1, nnpfc_inp_order_idc should not be equal to 3.

[0173] When ChromaFormatIdc equals 0, nnpfc_inp_order_idc must equal 0.

[0174] When chromaUpsamplingFlag equals 1, nnpfc_inp_order_idc should not equal 0.

[0175] Table 2 contains informative descriptions of the nnpfc_inp_order_idc values.

[0176] Table 2 – Description of nnpfc_inp_order_idc values

[0177]

[0178] Figure 1 An example of deriving luminance channels from luminance components is shown. For instance, when nnpfc_inp_order_idc equals 3, the four luminance channels (right) are derived from the luminance components (left).

[0179] `nnpfc_inp_tensor_luma_bitdepth_minus8` plus 8 specifies the bit depth of the luminance sample values ​​in the input integer tensor. The value of `inpTensorBitDepthY` is derived as follows:

[0180] inpTensorBitDepth Y = nnpfc_inp_tensor_luma_bitdepth_minus8 + 8 (85)

[0181] The requirement for bitstream consistency is that the value of nnpfc_inp_tensor_luma_bitdepth_minus8 must be in the range of 0 to 24 (inclusive).

[0182] `nnpfc_inp_tensor_chroma_bitdepth_minus8` plus 8 specifies the bit depth of the chroma sample values ​​in the input integer tensor. The value of `inpTensorBitDepthC` is derived as follows:

[0183] inpTensorBitDepthC = nnpfc_inp_tensor_chroma_bitdepth_minus8 + 8 (86)

[0184] The requirement for bitstream consistency is that the value of nnpfc_inp_tensor_chroma_bitdepth_minus8 must be in the range of 0 to 24 (inclusive).

[0185] When nnpfc_auxiliary_inp_idc equals 1, the variable strengthControlScaledVal is derived as follows:

[0186] for( i = 0; i < numInputPics; i++ )

[0187] if( nnpfc_inp_format_idc = = 1 ) (87)

[0188] if( nnpfc_inp_order_idc == 0 || nnpfc_inp_order_idc == 2||

[0189] nnpfc_inp_order_idc == 3)

[0190] strengthControlScaledVal[ i ] =

[0191] Floor (StrengthControlVal[i] ( ( 1 < <inpTensorBitDepth Y ) - 1 ) )

[0192] else if( nnpfc_inp_order_idc == 1 )

[0193] strengthControlScaledVal[ i ] =

[0194] Floor (StrengthControlVal[i] ( ( 1 < <inpTensorBitDepth C ) - 1 ) )

[0195] else

[0196] strengthControlScaledVal[ i ] = StrengthControlVal[ i ]

[0197] Small blocks are rectangular arrays of samples from the components of an image (e.g., luminance or chrominance components).

[0198] The process of deriving the input tensor `inputTensors()` for a given vertical sample coordinate `cTop` and horizontal sample coordinate `cLeft` (which specifies the top-left sample position of a small block of samples included in the input tensor) is defined as follows:

[0199] for( i = 0; i < numInputPics; i++ ) {

[0200] if( nnpfc_inp_order_idc == 0 )

[0201] for( yP = -nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)

[0202] for( xP = -nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {

[0203] inpVal = InpY( InpSampleVal( cTop + yP, cLeft +xP, CroppedHeight,

[0204] CroppedWidth, CroppedYPic[ i ], 0 ) )

[0205] yPovlp = yP + nnpfc_overlap

[0206] xPovlp = xP + nnpfc_overlap

[0207] if( !nnpfc_component_last_flag )

[0208] inputTensor[ 0 ][ i ][ 0 ][ yPovlp ][ xPovlp] = inpVal

[0209] else

[0210] inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 0] = inpVal

[0211] if( nnpfc_auxiliary_inp_idc = = 1 )

[0212] if( !nnpfc_component_last_flag )

[0213] inputTensor[ 0 ][ i ][ 1 ][ yPovlp ][xPovlp ] = strengthControlScaledVal[ i ]

[0214] else

[0215] inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp][ 1 ] = strengthControlScaledVal[ i ]

[0216] }

[0217] else if( nnpfc_inp_order_idc = = 1 ) (88)

[0218] for( yP = -nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)

[0219] for( xP = -nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {

[0220] inpCbVal = InpC( InpSampleVal( cTop + yP, cLeft +xP, CroppedHeight / SubHeightC,

[0221] CroppedWidth / SubWidthC, CroppedCbPic[ i ], 1 ) )

[0222] inpCrVal = InpC( InpSampleVal( cTop + yP, cLeft +xP, CroppedHeight / SubHeightC,

[0223] CroppedWidth / SubWidthC, CroppedCrPic[ i ], 2 ) )

[0224] yPovlp = yP + nnpfc_overlap

[0225] xPovlp = xP + nnpfc_overlap

[0226] if( !nnpfc_component_last_flag ) {

[0227] inputTensor[ 0 ][ i ][ 0 ][ yPovlp ][ xPovlp] = inpCbVal

[0228] inputTensor[ 0 ][ i ][ 1 ][ yPovlp ][ xPovlp] = inpCrVal

[0229] } else {

[0230] inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 0] = inpCbVal

[0231] inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 1] = inpCrVal

[0232] }

[0233] if( nnpfc_auxiliary_inp_idc = = 1 )

[0234] if( !nnpfc_component_last_flag )

[0235] inputTensor[ 0 ][ i ][ 2 ][ yPovlp ][xPovlp ] = strengthControlScaledVal[ i ]

[0236] else

[0237] inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp][ 2 ] = strengthControlScaledVal[ i ]

[0238] }

[0239] else if( nnpfc_inp_order_idc = = 2 )

[0240] for( yP = -nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)

[0241] for( xP = -nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {

[0242] yY = cTop + yP

[0243] xY = cLeft + xP

[0244] yC = yY / SubHeightC

[0245] xC = xY / SubWidthC

[0246] inpYVal = InpY( InpSampleVal( yY, xY,CroppedHeight,

[0247] CroppedWidth, CroppedYPic[ i ], 0 ) )

[0248] inpCbVal = InpC( InpSampleVal( yC, xC,CroppedHeight / SubHeightC,

[0249] CroppedWidth / SubWidthC, CroppedCbPic[ i ], 1 ) )

[0250] inpCrVal = InpC( InpSampleVal( yC, xC,CroppedHeight / SubHeightC,

[0251] CroppedWidth / SubWidthC, CroppedCrPic[ i ], 2 ) )

[0252] yPovlp = yP + nnpfc_overlap

[0253] xPovlp = xP + nnpfc_overlap

[0254] if( !nnpfc_component_last_flag ) {

[0255] inputTensor[ 0 ][ i ][ 0 ][ yPovlp ][ xPovlp] = inpYVal

[0256] inputTensor[ 0 ][ i ][ 1 ][ yPovlp ][ xPovlp] = inpCbVal

[0257] inputTensor[ 0 ][ i ][ 2 ][ yPovlp ][ xPovlp] = inpCrVal

[0258] } else {

[0259] inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 0] = inpYVal

[0260] inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 1] = inpCbVal

[0261] inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 2] = inpCrVal

[0262] }

[0263] if( nnpfc_auxiliary_inp_idc = = 1 )

[0264] if( !nnpfc_component_last_flag )

[0265] inputTensor[ 0 ][ i ][ 3 ][ yPovlp ][xPovlp ] = strengthControlScaledVal[ i ]

[0266] else

[0267] inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp][ 3 ] = strengthControlScaledVal[ i ]

[0268] }

[0269] else if( nnpfc_inp_order_idc = = 3 )

[0270] for( yP = -nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)

[0271] for( xP = -nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {

[0272] yTL = cTop + yP 2

[0273] xTL = cLeft + xP 2

[0274] yBR = yTL + 1

[0275] xBR = xTL + 1

[0276] yC = cTop / 2 + yP

[0277] xC = cLeft / 2 + xP

[0278] inpTLVal = InpY( InpSampleVal( yTL, xTL,CroppedHeight,

[0279] CroppedWidth, CroppedYPic[ i ], 0 ) )

[0280] inpTRVal = InpY( InpSampleVal( yTL, xBR,CroppedHeight,

[0281] CroppedWidth, CroppedYPic[ i ], 0 ) )

[0282] inpBLVal = InpY( InpSampleVal( yBR, xTL,CroppedHeight,

[0283] CroppedWidth, CroppedYPic[ i ], 0 ) )

[0284] inpBRVal = InpY( InpSampleVal( yBR, xBR,CroppedHeight,

[0285] CroppedWidth, CroppedYPic[ i ], 0 ) )

[0286] inpCbVal = InpC( InpSampleVal( yC, xC,CroppedHeight / 2,

[0287] CroppedWidth / 2, CroppedCbPic[ i ], 1) )

[0288] inpCrVal = InpC( InpSampleVal( yC, xC,CroppedHeight / 2,

[0289] CroppedWidth / 2, CroppedCrPic[ i ], 2) )

[0290] yPovlp = yP + nnpfc_overlap

[0291] xPovlp = xP + nnpfc_overlap

[0292] if( !nnpfc_component_last_flag ) {

[0293] inputTensor[ 0 ][ i ][ 0 ][ yPovlp ][ xPovlp] = inpTLVal

[0294] inputTensor[ 0 ][ i ][ 1 ][ yPovlp ][ xPovlp] = inpTRVal

[0295] inputTensor[ 0 ][ i ][ 2 ][ yPovlp ][ xPovlp] = inpBLVal

[0296] inputTensor[ 0 ][ i ][ 3 ][ yPovlp ][ xPovlp] = inpBRVal

[0297] inputTensor[ 0 ][ i ][ 4 ][ yPovlp ][ xPovlp] = inpCbVal

[0298] inputTensor[ 0 ][ i ][ 5 ][ yPovlp ][ xPovlp] = inpCrVal

[0299] } else {

[0300] inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 0] = inpTLVal

[0301] inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 1] = inpTRVal

[0302] inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 2] = inpBLVal

[0303] inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 3] = inpBRVal

[0304] inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 4] = inpCbVal

[0305] inputTensor[ 0 ][ i ][ yPovlp ][ xPovlp ][ 5] = inpCrVal

[0306] }

[0307] if( nnpfc_auxiliary_inp_idc = = 1 )

[0308] if( !nnpfc_component_last_flag )

[0309] inputTensor[ 0 ][ i ][ 6 ][ yPovlp ][xPovlp ] = strengthControlScaledVal[ i ]

[0310] else

[0311] inputTensor[ 0 ][ i ][ yPovlp ][xPovlp ][ 6 ] = strengthControlScaledVal[ i ]

[0312] }

[0313] }

[0314] When nnpfc_out_format_idc is equal to 0, it indicates that the sample values output by NNPF are real numbers, and the value range from 0 to 1 (including the boundary values) is linearly mapped to the unsigned integer value range from 0 to (1 << bitDepth) - 1 (including the boundary values), where bitDepth is any bit depth required for subsequent post-processing or display.

[0315] When nnpfc_out_format_idc is equal to 1, it indicates that the luminance sample values output by NNPF are unsigned integers within the range from 0 to (1 << outTensorBitDepthY) - 1 (including the boundary values), and the chrominance sample values output by NNPF are unsigned integers within the range from 0 to (1 << outTensorBitDepthC) - 1 (including the boundary values).

[0316] Values of nnpfc_out_format_idc greater than 1 are reserved for future ITU-T | ISO / IEC specifications and shall not be present in the bitstream conforming to this version of this document. The decoder conforming to this version of this document shall ignore the NNPFC SEI messages containing the reserved values of nnpfc_out_format_idc.

[0317] nnpfc_out_order_idc indicates the output order of the samples generated by NNPF.

[0318] In bitstreams conforming to this version of this document, the value of nnpfc_out_order_idc must be in the range of 0 to 3 (inclusive). Values ​​of nnpfc_out_order_idc from 4 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and should not exist in bitstreams conforming to this version of this document. Decoders conforming to this version of this document must ignore NNPFC SEI messages with nnpfc_out_order_idc in the range of 4 to 255 (inclusive). Values ​​of nnpfc_out_order_idc greater than 255 should not exist in bitstreams conforming to this version of this document and are not reserved for future use.

[0319] When chromaUpsamplingFlag equals 1, nnpfc_out_order_idc should not equal 0 or 3.

[0320] When colourizationFlag equals 1, nnpfc_out_order_idc should not equal 0.

[0321] Table 3 contains informative descriptions of the nnpfc_out_order_idc values.

[0322] Table 3 – Description of nnpfc_out_order_idc values

[0323]

[0324] `nnpfc_out_tensor_luma_bitdepth_minus8` plus 8 specifies the bit depth of the luminance sample values ​​in the output integer tensor. The value of `nnpfc_out_tensor_luma_bitdepth_minus8` must be in the range of 0 to 24 (inclusive). The value of `outTensorBitDepthY` is derived as follows:

[0325] outTensorBitDepth Y = nnpfc_out_tensor_luma_bitdepth_minus8 + 8 (89)

[0326] `nnpfc_out_tensor_chroma_bitdepth_minus8` plus 8 specifies the bit depth of the chroma sample values ​​in the output integer tensor. The value of `nnpfc_out_tensor_chroma_bitdepth_minus8` must be in the range of 0 to 24 (inclusive). The value of `outTensorBitDepthC` is derived as follows:

[0327] outTensorBitDepthC = nnpfc_out_tensor_chroma_bitdepth_minus8 + 8 (90)

[0328] When bitDepthUpsamplingFlag equals 1, the value of nnpfc_out_format_idc must be equal to 1, and at least one of the following conditions must be true:

[0329] - nnpfc_out_tensor_luma_bitdepth_minus8 exists, and outTensorBitDepth Y Greater than BitDepth Y .

[0330] - nnpfc_out_tensor_chroma_bitdepth_minus8 exists, and outTensorBitDepth C Greater than BitDepth C .

[0331] When nnpfc_inp_tensor_luma_bitdepth_minus8, nnpfc_inp_tensor_chroma_bitdepth_minus8, nnpfc_out_tensor_luma_bitdepth_minus8, and nnpfc_out_tensor_chroma_bitdepth_minus8 exist, and outTensorBitDepthY is greater than inpTensorBitDepthY, outTensorBitDepthC should not be less than inpTensorBitDepthC. When nnpfc_inp_tensor_luma_bitdepth_minus8, nnpfc_inp_tensor_chroma_bitdepth_minus8, nnpfc_out_tensor_luma_bitdepth_minus8, and nnpfc_out_tensor_chroma_bitdepth_minus8 exist, and outTensorBitDepthC is greater than inpTensorBitDepthC, outTensorBitDepthY should not be less than inpTensorBitDepthY.

[0332] The procedure StoreOutputTensors(), used to derive sample values ​​from the output tensor outputTensor for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft (which specifies the top-left sample position of a small block of samples included in the input tensor), is defined as follows:

[0333] for( i = 0; i < numOutputPics; i++ ) {

[0334] if( nnpfc_out_order_idc == 0 )

[0335] for( yP = 0; yP < outPatchHeight; yP++)

[0336] for( xP = 0; xP < outPatchWidth; xP++ ) {

[0337] yY = cTop outPatchHeight / inpPatchHeight + yP

[0338] xY = cLeft outPatchWidth / inpPatchWidth + xP

[0339] if ( yY < nnpfcOutputPicHeight && xY <nnpfcOutputPicWidth )

[0340] if( !nnpfc_component_last_flag )

[0341] FilteredYPic[ i ][ xY ][yY ] =outputTensor[ 0 ][ i ][ 0 ][ yP ][ xP ]

[0342] else

[0343] FilteredYPic[ i ][ xY ][ yY ] =outputTensor[ 0 ][ i ][ yP ][ xP ][ 0 ]

[0344] }

[0345] else if( nnpfc_out_order_idc = = 1 ) (91)

[0346] for( yP = 0; yP < outPatchCHeight; yP++)

[0347] for( xP = 0; xP < outPatchCWidth; xP++ ) {

[0348] xSrc = cLeft horCScaling + xP

[0349] ySrc = cTop verCScaling + yP

[0350] if ( ySrc < nnpfcOutputPicHeight / outSubHeightC&&

[0351] xSrc < nnpfcOutputPicWidth / outSubWidthC )

[0352] if( !nnpfc_component_last_flag ) {

[0353] FilteredCbPic[ i ][ xSrc ][ ySrc ] =outputTensor[ 0 ][ i ][ 0 ][ yP ][ xP ]

[0354] FilteredCrPic[ i ][ xSrc ][ ySrc ] =outputTensor[ 0 ][ i ][ 1 ][ yP ][ xP ]

[0355] } else {

[0356] FilteredCbPic[ i ][ xSrc ][ ySrc ] =outputTensor[ 0 ][ i ][ yP ][ xP ][ 0 ]

[0357] FilteredCrPic[ i ][ xSrc ][ ySrc ] =outputTensor[ 0 ][ i ][ yP ][ xP ][ 1 ]

[0358] }

[0359] }

[0360] else if( nnpfc_out_order_idc = = 2 )

[0361] for( yP = 0; yP < outPatchHeight; yP++)

[0362] for( xP = 0; xP < outPatchWidth; xP++ ) {

[0363] yY = cTop outPatchHeight / inpPatchHeight + yP

[0364] xY = cLeft outPatchWidth / inpPatchWidth + xP

[0365] yC = yY / outSubHeightC

[0366] xC = xY / outSubWidthC

[0367] yPc = ( yP / outSubHeightC ) outSubHeightC

[0368] xPc = ( xP / outSubWidthC ) outSubWidthC

[0369] if ( yY < nnpfcOutputPicHeight && xY <nnpfcOutputPicWidth )

[0370] if( !nnpfc_component_last_flag ) {

[0371] FilteredYPic[ i ][ xY ][ yY ] =outputTensor[ 0 ][ i ][ 0 ][ yP ][ xP ]

[0372] FilteredCbPic[ i ][ xC ][ yC ] =outputTensor[ 0 ][ i ][ 1 ][ yPc ][ xPc ]

[0373] FilteredCrPic[ i ][ xC ][ yC ] =outputTensor[ 0 ][ i ][ 2 ][ yPc ][ xPc ]

[0374] } else {

[0375] FilteredYPic[ i ][ xY ][ yY ] =outputTensor[ 0 ][ i ][ yP ][ xP ][ 0 ]

[0376] FilteredCbPic[ i ][ xC ][ yC ] =outputTensor[ 0 ][ i ][ yPc ][ xPc ][ 1 ]

[0377] FilteredCrPic[ i ][ xC ][ yC ] =outputTensor[ 0 ][ i ][ yPc ][ xPc ][ 2 ]

[0378] }

[0379] }

[0380] else if( nnpfc_out_order_idc = = 3 )

[0381] for( yP = 0; yP < outPatchHeight; yP++ )

[0382] for( xP = 0; xP < outPatchWidth; xP++ ) {

[0383] ySrc = cTop / 2 outPatchHeight / inpPatchHeight+ yP

[0384] xSrc = cLeft / 2 outPatchWidth / inpPatchWidth+ xP

[0385] if ( ySrc < nnpfcOutputPicHeight / 2 &&

[0386] xSrc < nnpfcOutputPicWidth / 2 )

[0387] if( !nnpfc_component_last_flag ) {

[0388] FilteredYPic[ i ][ xSrc 2 ][ ySrc 2 ] = outputTensor[ 0 ][ i ][ 0 ][ yP ][ xP ]

[0389] FilteredYPic[ i ][ xSrc 2 + 1 ][ ySrc 2 ] = outputTensor[ 0 ][ i ][ 1 ][ yP ][ xP ]

[0390] FilteredYPic[ i ][ xSrc 2 ][ ySrc 2 + 1 ] = outputTensor[ 0 ][ i ][ 2 ][ yP ][ xP ]

[0391] FilteredYPic[ i ][ xSrc 2 + 1][ ySrc 2 + 1 ] = outputTensor[ 0 ][ i ][ 3 ][ yP ][ xP ]

[0392] FilteredCbPic[ i ][ xSrc ][ ySrc ] =outputTensor[ 0 ][ i ][ 4 ][ yP ][ xP ]

[0393] FilteredCrPic[ i ][ xSrc ][ ySrc ] =outputTensor[ 0 ][ i ][ 5 ][ yP ][ xP ]

[0394] } else {

[0395] FilteredYPic[ i ][ xSrc 2 ][ ySrc 2 ] = outputTensor[ 0 ][ i ][ yP ][ xP ][ 0 ]

[0396] FilteredYPic[ i ][ xSrc 2 + 1 ][ ySrc 2 ] = outputTensor[ 0 ][ i ][ yP ][ xP ][ 1 ]

[0397] FilteredYPic[ i ][ xSrc 2 ][ ySrc 2 + 1 ] = outputTensor[ 0 ][ i ][ yP ][ xP ][ 2 ]

[0398] FilteredYPic[ i ][ xSrc 2 + 1][ ySrc 2 + 1 ] = outputTensor[ 0 ][ i ][ yP ][ xP ][ 3 ]

[0399] FilteredCbPic[ i ][ xSrc ][ ySrc ] =outputTensor[ 0 ][ i ][ yP ][ xP ][ 4 ]

[0400] FilteredCrPic[ i ][ xSrc ][ ySrc ] =outputTensor[ 0 ][ i ][ yP ][ xP ][ 5 ]

[0401] }

[0402] }

[0403] }

[0404] `nnpfc_separate_colour_description_present_flag` equal to 1 indicates that the SEI message syntax structure specifies different combinations of the primary color, transport characteristics, matrix coefficients, and scaling and offset values ​​associated with the matrix coefficients for the image generated by NNPF. `nnpfc_separate_colour_description_present_flag` equal to 0 indicates that the combination of the primary color, transport characteristics, matrix coefficients, and scaling and offset values ​​associated with the matrix coefficients for the image generated by NNPF is the same as the combination indicated in the CLVS VUI parameters.

[0405] nnpfc_colour_primaries has the same semantics as the vui_colour_primaries syntax element specified in sub-entry 7.3, except as follows:

[0406] - nnpfc_colour_primaries specifies the primary color of the image generated by NNPF as defined in the application SEI message, instead of the primary color used for CLVS.

[0407] - When nnpfc_colour_primaries does not exist in the NNPFC SEI message, the value of nnpfc_colour_primaries is presumed to be equal to vui_colour_primaries.

[0408] nnpfc_transfer_characteristics has the same semantics as the vui_transfer_characteristics syntax element specified in sub-entry 7.3, except as follows:

[0409] - nnpfc_transfer_characteristics specifies the transfer characteristics of images generated by NNPF as defined in the application SEI message, rather than the transfer characteristics used for CLVS.

[0410] - When nnpfc_transfer_characteristics does not exist in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is presumed to be equal to vui_transfer_characteristics.

[0411] The nnpfc_matrix_coeffs describes the equations used to derive luminance and chrominance signals from the green, blue, and red, or Y, Z, and X primary colors. Its semantics apply to images generated by applying the NNPF specified in this SEI message and are identical to the semantics specified for MatrixCoefficients in Recommendation ITU-T H.273 | ISO / IEC 23091-2, where BitDepthY and BitDepthC are equal to outTensorBitDepthY and outTensorBitDepthC, respectively.

[0412] When nnpfc_matrix_coeffs does not exist in the NNPFC SEI message, the value of nnpfc_matrix_coeffs is presumed to be equal to vui_matrix_coeffs.

[0413] nnpfc_matrix_coeffs should not be equal to 0 unless both of the following conditions are true:

[0414] - nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8.

[0415] - nnpfc_out_order_idc equals 2, outSubHeightC equals 1, and outSubWidthC equals 1.

[0416] nnpfc_matrix_coeffs should not be equal to 8 unless one of the following conditions is true:

[0417] - nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8.

[0418] - nnpfc_out_tensor_chroma_bitdepth_minus8 equals nnpfc_out_tensor_luma_bitdepth_minus8 + 1, nnpfc_out_order_idc equals 2, outSubHeightC equals 1, and outSubWidthC equals 1.

[0419] The `nnpfc_full_range_flag` indicates the scaling and offset values ​​applied in association with the matrix coefficients specified by `nnpfc_matrix_coeffs`. Its semantics are identical to those specified for the `VideoFullRangeFlag` parameter in Recommendation ITU-T H.273 | ISO / IEC 23091-2. When not present, the value of `nnpfc_full_range_flag` is presumed to be 0.

[0420] A value of 1 for nnpfc_chroma_loc_info_present_flag indicates the presence of the nnpfc_chroma_sample_loc_type_frame syntax element in the NNPFC SEI message. A value of 0 for nnpfc_chroma_loc_info_present_flag indicates the absence of the nnpfc_chroma_sample_loc_type_frame syntax element in the NNPFC SEI message. The value of nnpfc_chroma_loc_info_present_flag must be 0 when colourizationFlag is 0 or nnpfc_out_colour_format_idc is not equal to 1.

[0421] When `nnpfc_chroma_sample_loc_type_frame` is not equal to 6 and `nnpfc_out_colour_format_idc` is equal to 1, it specifies the position of the chroma samples in the output image. Figure 1 As shown. `nnpfc_chroma_sample_loc_type_frame` equal to 6 and `nnpfc_out_colour_format_idc` equal to 1 indicate that the location of the chroma sample point is unknown, unspecified, or specified in any other way not specified in this document. The value of `nnpfc_chroma_sample_loc_type_frame` must be in the range of 0 to 6 (inclusive).

[0422] nnpfc_overlap indicates the horizontal and vertical sample counts of overlap between adjacent input tensors in NNPF. The value of nnpfc_overlap must be in the range of 0 to 16383 (inclusive).

[0423] `nnpfc_constant_patch_size_flag` equal to 1 indicates that NNPF accepts the exact patch size as input, indicated by `nnpfc_patch_width_minus1` and `nnpfc_patch_height_minus1`. `nnpfc_constant_patch_size_flag` equal to 0 indicates that NNPF accepts any patch size with a width of `inpPatchWidth` and a height of `inpPatchHeight`, such that the width of the expanded patch (i.e., the patch plus the overlapping area) is equal to `inpPatchWidth + 2`. nnpfc_overlap) is nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 The height of the expanded patch is a positive integer multiple of nnpfc_overlap, and it is equal to inpPatchHeight + 2. nnpfc_overlap) is nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 A positive integer multiple of nnpfc_overlap.

[0424] Incrementing 1 to nnpfc_patch_width_minus1 indicates the horizontal sample count for the required patch size of the NNPF input when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_width_minus1 must be in the range of 0 to Min(32766, CroppedWidth - 1) (inclusive).

[0425] Incrementing 1 to nnpfc_patch_height_minus1 indicates the vertical sample count for the required patch size of the NNPF input when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_height_minus1 must be in the range of 0 to Min(32766, CroppedHeight - 1) (inclusive).

[0426] nnpfc_extended_patch_width_cd_delta_minus1+1+2 `nnpfc_overlap` indicates the common divisor of all allowed values ​​for the width of the extended patch required for the NNPF input, when `nnpfc_constant_patch_size_flag` is equal to 0. The value of `nnpfc_extended_patch_width_cd_delta_minus1` must be in the range of 0 to Min(32766, CroppedWidth - 1) (inclusive).

[0427] nnpfc_extended_patch_height_cd_delta_minus1+1+2 `nnpfc_overlap` indicates the common divisor of all allowed values ​​for the height of the extended patch required for the NNPF input, when `nnpfc_constant_patch_size_flag` is equal to 0. The value of `nnpfc_extended_patch_height_cd_delta_minus1` must be in the range of 0 to Min(32766, CroppedHeight - 1) (inclusive).

[0428] This makes the variables inpPatchWidth and inpPatchHeight the width and height of the small block, respectively.

[0429] If nnpfc_constant_patch_size_flag equals 0, then the following applies:

[0430] The values ​​of inpPatchWidth and inpPatchHeight are provided externally by means not specified in this document or set by the post-processing itself.

[0431] - inpPatchWidth + 2 The value of nnpfc_overlap must be nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 The value must be a positive integer multiple of nnpfc_overlap, and inpPatchWidth must be less than or equal to CroppedWidth. inpPatchHeight + 2 The value of nnpfc_overlap must be nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 The value must be a positive integer multiple of nnpfc_overlap, and inpPatchHeight must be less than or equal to CroppedHeight.

[0432] Otherwise (nnpfc_constant_patch_size_flag equals 1), the value of inpPatchWidth is set to equal to nnpfc_patch_width_minus1 + 1, and the value of inpPatchHeight is set to equal to nnpfc_patch_height_minus1 + 1.

[0433] The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, and outPatchCHeight are derived as follows:

[0434] outPatchWidth = (nnpfcOutputPicWidth inpPatchWidth ) / CroppedWidth (92)

[0435] outPatchHeight = (nnpfcOutputPicHeight inpPatchHeight ) / CroppedHeight (93)

[0436] horCScaling = SubWidthC / outSubWidthC (94)

[0437] verCScaling = SubHeightC / outSubHeightC (95)

[0438] outPatchCWidth = outPatchWidth horCScaling (96)

[0439] outPatchCHeight = outPatchHeight verCScaling (97)

[0440] The requirement for bitstream consistency is outPatchWidth CroppedWidth must be equal to nnpfcOutputPicWidth inpPatchWidth and outPatchHeight CroppedHeight must be equal to nnpfcOutputPicHeight inpPatchHeight.

[0441] `nnpfc_padding_type` indicates the padding process when referencing sample locations outside the boundaries of the input image, as described in Table 4. In the bitstream conforming to this version of this document, the value of `nnpfc_padding_type` must be in the range of 0 to 4 (inclusive of boundary values). Values ​​of `nnpfc_padding_type` from 5 to 15 (inclusive of boundary values) are reserved for future use by ITU-T | ISO / IEC and should not exist in the bitstream conforming to this version of this document. Decoders conforming to this version of this document must ignore NNPFC SEI messages with `nnpfc_padding_type` in the range of 5 to 15 (inclusive of boundary values). Values ​​of `nnpfc_padding_type` greater than 15 should not exist in the bitstream conforming to this version of this document and are not reserved for future use.

[0442] Table 4 – Informative description of nnpfc_padding_type values

[0443]

[0444] nnpfc_luma_padding_val indicates the luminance value to be used for padding when nnpfc_padding_type is equal to 4. The value of nnpfc_luma_padding_val must be in the range of 0 to (1 << BitDepthY) - 1 (inclusive).

[0445] nnpfc_cb_padding_val indicates the Cb value to be used for padding when nnpfc_padding_type is equal to 4. The value of nnpfc_cb_padding_val must be in the range of 0 to (1 << BitDepthC) - 1 (inclusive).

[0446] nnpfc_cr_padding_val indicates the Cr value to be used for padding when nnpfc_padding_type is equal to 4. The value of nnpfc_cr_padding_val must be in the range of 0 to (1 << BitDepthC) - 1 (inclusive).

[0447] The function InpSampleVal( y,x, picHeight, picWidth, croppedPic, cIdx ) takes as input the vertical sample position y, the horizontal sample position x, the image height picHeight, the image width picWidth, the sample array croppedPic, and the component index cIdx (luminance is 0, Cb is 1, Cr is 2). It returns the value of sampleVal as derived below:

[0448] Note 6 – For the input of the function InpSampleVal(), the vertical position is listed before the horizontal position to be compatible with the input tensor conventions of some inference engines.

[0449] if( nnpfc_padding_type == 0 )

[0450] if( y < 0 || x < 0 || y >= picHeight || x >= picWidth )

[0451] sampleVal = 0

[0452] else

[0453] sampleVal = croppedPic[ x ][ y ](98)

[0454] else if( nnpfc_padding_type == 1 )

[0455] sampleVal = croppedPic[ Clip3( 0, picWidth - 1, x ) ][ Clip3( 0,picHeight - 1, y ) ]

[0456] else if( nnpfc_padding_type == 2 )

[0457] sampleVal = croppedPic[ Reflect( picWidth - 1, x ) ][ Reflect(picHeight - 1, y ) ]

[0458] else if( nnpfc_padding_type == 3 )

[0459] if ( y >= 0 && y < picHeight )

[0460] sampleVal = croppedPic[ Wrap( picWidth - 1, x ) ][ y ]

[0461] else if( nnpfc_padding_type == 4 )

[0462] if( y < 0 || x < 0 || y >= picHeight || x >= picWidth )

[0463] sampleVal = ( cIdx = = 0 ? nnpfc_luma_padding_val :

[0464] ( cIdx = = 1 ? nnpfc_cb_padding_val : nnpfc_cr_padding_val ) )

[0465] else

[0466] sampleVal = croppedPic[x][y]

[0467] The NNPF PostProcessingFilter() is a target NNPF derived from the semantics of the NNPFA SEI message. The following example procedure can be used with NNPF PostProcessingFilter() to generate (multiple) filtered and / or interpolated images in a block-by-block manner, containing Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, as indicated by nnpfc_out_order_idc:

[0468] if( nnpfc_inp_order_idc == 0 || nnpfc_inp_order_idc == 2 )

[0469] for( cTop = 0; cTop < CroppedHeight; cTop += inpPatchHeight )

[0470] for( cLeft = 0; cLeft < CroppedWidth; cLeft +=inpPatchWidth ) {

[0471] DeriveInputTensors( )

[0472] outputTensor = PostProcessingFilter( inputTensor )

[0473] StoreOutputTensors( )

[0474] }

[0475] else if( nnpfc_inp_order_idc = = 1 )

[0476] for( cTop = 0; cTop < CroppedHeight / SubHeightC; cTop +=inpPatchHeight )

[0477] for( cLeft = 0; cLeft < CroppedWidth / SubWidthC; cLeft +=inpPatchWidth ) { (99)

[0478] DeriveInputTensors( )

[0479] outputTensor = PostProcessingFilter( inputTensor )

[0480] StoreOutputTensors( )

[0481] }

[0482] else if( nnpfc_inp_order_idc = = 3 )

[0483] for( cTop = 0; cTop < CroppedHeight; cTop += inpPatchHeight 2)

[0484] for( cLeft = 0; cLeft < CroppedWidth; cLeft +=inpPatchWidth 2) {

[0485] DeriveInputTensors()

[0486] outputTensor = PostProcessingFilter( inputTensor )

[0487] StoreOutputTensors()

[0488] }

[0489] The image generated by NNPF at index i contains sample arrays FilteredYPic[i], FilteredCbPic[i], and FilteredCrPic[i] (if present), derived by Equation 99. The image generated by NNPF does not include overlapping regions.

[0490] The NNPF process consists of the process defined in Equation 99, and then outputs the images generated by NNPF in ascending index order, wherein all images generated by NNPF through NNPF interpolation are output, and those images generated by NNPF corresponding to any input image of NNPF are output according to the semantics specified in the NNPFA SEI message.

[0491] A value of 1 for nnpfc_complexity_info_present_flag indicates the existence of one or more syntax elements that indicate the complexity of the NNPF associated with nnpfc_id. A value of 0 for nnpfc_complexity_info_present_flag indicates the absence of a syntax element that indicates the complexity of the NNPF associated with nnpfc_id.

[0492] `nnpfc_parameter_type_idc` equal to 0 indicates that the neural network uses only integer parameters. `nnpfc_parameter_type_flag` equal to 1 indicates that the neural network can use floating-point or integer parameters. `nnpfc_parameter_type_idc` equal to 2 indicates that the neural network uses only binary parameters. `nnpfc_parameter_type_idc` equal to 3 is reserved for future use by ITU-T | ISO / IEC and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document must ignore NNPFC SEI messages with `nnpfc_parameter_type_idc` equal to 3.

[0493] The values ​​0, 1, 2, and 3 for nnpfc_log2_parameter_bit_length_minus3 indicate that the neural network does not use parameters with bit lengths greater than 8, 16, 32, and 64, respectively. When nnpfc_parameter_type_idc exists but nnpfc_log2_parameter_bit_length_minus3 does not exist, the neural network does not use parameters with a bit length greater than 1.

[0494] `nnpfc_num_parameters_idc` indicates the maximum number of neural network parameters for NNPF, in powers of 2048. `nnpfc_num_parameters_idc` equal to 0 indicates that the maximum number of neural network parameters is unknown. The value of `nnpfc_num_parameters_idc` must be in the range of 0 to 52 (inclusive). Values ​​of `nnpfc_num_parameters_idc` greater than 52 are reserved for future use by ITU-T | ISO / IEC and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document must ignore NNPFC SEI messages with `nnpfc_num_parameters_idc` greater than 52.

[0495] If the value of nnpfc_num_parameters_idc is greater than 0, then the variable maxNumParameters is deduced as follows:

[0496] maxNumParameters = ( 2 048 << nnpfc_num_parameters_idc ) – 1 (100)

[0497] The requirement for bitstream consistency is that the number of neural network parameters in NNPF must be less than or equal to maxNumParameters.

[0498] A value greater than 0 for nnpfc_num_kmac_operations_idc indicates that the maximum number of multiply-accumulate operations per sample in NNPF is less than or equal to nnpfc_num_kmac_operations_idc. 1000. An nnpfc_num_kmac_operations_idc value of 0 indicates that the maximum number of multiply-accumulate operations in the network is unknown. The value of nnpfc_num_kmac_operations_idc must be between 0 and 2. 32 - 2 (including boundary values).

[0499] `nnpfc_total_kilobyte_size` greater than 0 indicates the total size, in kilobytes, required to store the uncompressed parameters of the neural network. The total size, in bits, is equal to or greater than the sum of the bits used to store each parameter. `nnpfc_total_kilobyte_size` is the total size in bits divided by 8000 and rounded down. `nnpfc_total_kilobyte_size` equal to 0 indicates that the total size required to store the neural network parameters is unknown. The value of `nnpfc_total_kilobyte_size` must be between 0 and 2. 32 - 2 (including boundary values).

[0500] A value of 0 for nnpfc_metadata_extension_num_bits indicates that nnpfc_reserved_metadata_extension does not exist. A value greater than 0 for nnpfc_metadata_extension_num_bits indicates the bit length of nnpfc_reserved_metadata_extension. In this version of the document, nnpfc_metadata_extension_num_bits must be equal to 0. Values ​​of nnpfc_metadata_extension_num_bits from 1 to 2048 (inclusive) are reserved for future use by ITU-T | ISO / IEC and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document must allow any value of nnpfc_metadata_extension_num_bits in the range of 0 to 2048 (inclusive). Values ​​greater than 2048 for nnpfc_metadata_extension_num_bits should not exist in the bitstream conforming to this version of the document and are not reserved for future use.

[0501] The nnpfc_reserved_metadata_extension should not exist in the bitstream conforming to this version of the document. However, decoders conforming to this version of the document must ignore the presence and value of nnpfc_reserved_metadata_extension. When present, the bit length of nnpfc_reserved_metadata_extension is equal to nnpfc_metadata_extension_num_bits.

[0502] nnpfc_reserved_zero_bit_b must be equal to 0 in the bitstream conforming to this version of the document. The decoder must ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_b is not equal to 0.

[0503] nnpfc_payload_byte[i] contains the i-th byte of a bitstream conforming to ISO / IEC 15938-17. For all existing values ​​of i, the byte sequence nnpfc_payload_byte[i] must be a complete bitstream conforming to ISO / IEC 15938-17.

[0504] 8.28.3 Activation of SEI Messages in Neural Network Post-Processing Filters

[0505] 8.28.3.1 Syntax for SEI Message Activation of Neural Network Post-processing Filters

[0506]

[0507] 8.23.3.2 Activating SEI message semantics using neural network post-processing filters

[0508] The Neural Network Post-Processing Filter Activation (NNPFA) SEI message activates or deactivates the target Neural Network Post-Processing Filter (NNPF), identified by nnpfa_target_id and nnpfa_target_base_flag, for possible post-processing filtering of a set of images. For a specific image where the NNPF is activated, the target NNPF is derived as follows:

[0509] - If nnpfa_target_base_flag equals 1, then the target NNPF is the base NNPF whose nnpfc_id equals nnpfa_target_id.

[0510] - Otherwise (nnpfa_target_base_flag equals 0), the target NNPF is the NNPF defined by the last NNPFC SEI message with an nnpfc_id equal to nnpfa_target_id, which is in the decoding order before the first VCL NAL unit of the current image and is not a duplicate of the NNPFC SEI message containing the base NNPF.

[0511] Note – Multiple NNPFA SEI messages can exist for the same image, for example, when NNPF is used for different purposes or for filtering different color components.

[0512] `nnpfa_target_id` indicates the target NNPF, which is specified by one or more NNPFC SEI messages associated with the current image and having an `nnpfc_id` equal to `nnpfa_target_id`. The value of `nnpfa_target_id` must be between 0 and 2. 32 The range is -2 (including boundary values).

[0513] NNPFA SEI messages with a specific value of nnpfa_target_id should not exist in the current PU unless one or both of the following conditions are true:

[0514] - Within the current CLVS, there exists an NNPFC SEI message whose nnpfc_id is equal to the specific value of nnpfa_target_id in the PU preceding the current PU in the decoding order.

[0515] - There is an NNPFC SEI message in the current PU with a specific value of nnpfc_id equal to nnpfa_target_id.

[0516] When a PU contains both an NNPFC SEI message with a specific value of nnpfc_id and an NNPFA SEI message with nnpfa_target_id equal to a specific value of nnpfc_id, the NNPFC SEI message must precede the NNPFA SEI message in the decoding order.

[0517] A `nnpfa_cancel_flag` equal to 1 indicates that the persistence of a target NNPF established by any previous NNPFA SEI message with the same `nnpfa_target_id` as the current SEI message is cancelled; that is, the target NNPF is no longer used unless it is activated by another NNPFASEI message with the same `nnpfa_target_id` as the current SEI message and a `nnpfa_cancel_flag` equal to 0. A `nnpfa_cancel_flag` equal to 0 indicates that `nnpfa_target_base_flag`, `nnpfa_persistence_flag`, and `nnpfa_num_output_entries` follow.

[0518] An nnpfa_target_base_flag value of 1 indicates that the target NNPF is the base NNPF whose nnpfc_id is equal to nnpfa_target_id. An nnpfa_target_base_flag value of 0 indicates that the target NNPF is the NNPF defined by the last NNPFC SEI message with an nnpfc_id equal to nnpfa_target_id, which is in decoding order before the first VCL NAL unit of the current image and is not a duplicate of the NNPFC SEI message containing the base NNPF.

[0519] The nnpfa_persistence_flag specifies the persistence of the target NNPF for the current layer.

[0520] Setting nnpfa_persistence_flag to 0 specifies that the target NNPF can only be used for post-processing filtering of the current image.

[0521] Setting nnpfa_persistence_flag to 1 specifies that the target NNPF can be used for post-processing filtering of the current image and all subsequent images in the current layer in output order, until one or more of the following conditions are true:

[0522] - A new CLVS begins for the current layer.

[0523] - End of bitstream.

[0524] - Output the images in the current layer that are associated with the NNPFA SEI message that has the same nnpfa_target_id as the current SEI message and an nnpfa_cancel_flag equal to 1, in the order of output.

[0525] Note – The target NNPF should not be applied to the subsequent picture in the current layer that is associated with an NNPFA SEI message that has the same nnpfa_target_id as the current SEI message and an nnpfa_cancel_flag equal to 1.

[0526] Let nnpfcTargetPictures be the set of pictures associated with the last NNPFA SEI message whose nnpfc_id equals nnpfa_target_id, preceding the current NNPFA SEI message in decoding order. Let nnpfaTargetPictures be the set of pictures that activate the target NNPF through the current NNPFA SEI message. Bitstream consistency requires that any pictures included in nnpfaTargetPictures must also be included in nnpfcTargetPictures.

[0527] nnpfa_num_output_entries specifies the number of nnpfa_output_flag[i] syntax elements present in the NNPFA SEI message. The value of nnpfa_num_output_entries must be in the range of 0 to NumInpPicsInOutputTensor (inclusive).

[0528] An NNpfa_output_flag[i] equal to 1 indicates that the NNPF-generated image corresponding to the input image with index InpIdx[i] is output by the NNPF procedure activated by the NNPFA SEI message, where the NNPF procedure is defined in the semantics of the NNPFA SEI message. An NNpfa_output_flag[i] equal to 0 indicates that the NNPF-generated image corresponding to the input image with index InpIdx[i] is not output by the NNPF procedure activated by the NNPFA SEI message. When nnpfa_num_output_entries is less than NumInpPicsInOutputTensor, for each i value in the range from nnpfa_num_output_entries to NumInpPicsInOutputTensor-1 (inclusive of boundary values), nnpfa_output_flag[i] is presumed to be equal to 1.

[0529] D.12.11 Use of SEI Message in Neural Network Post-Processing Filter Characteristics

[0530] Let currPic be the cropped decoded output image. For this cropped decoded output image, the Neural Network Post-Processing Filter (NNPF) defined by the Neural Network Post-Processing Filter Feature (NNPFC) SEI message is activated by the Neural Network Post-Processing Filter Activation (NNPFA) SEI message, and currLayerId is the nuh_layer_id value of currPic.

[0531] The requirement for bitstream consistency is that when a picture unit contains an NNPFA SEI message, the value of ph_pic_output_flag in the picture header of that picture unit must be equal to 1.

[0532] Note – Since only the cropped decoded output image is used as the input image for NNPF, the value of ph_pic_output_flag in the image header of the encoded / decoded image corresponding to each input image of NNPF is equal to 1.

[0533] The variable pictureRateUpsamplingFlag is set to equal to (nnpfc_purpose & 0x08) != 0.

[0534] The variable numInputPics is set to equal nnpfc_num_input_pics_minus1 + 1.

[0535] The variable numInferences is derived as follows:

[0536] - If pictureRateUpsamplingFlag equals 1, the NNPFA SEI message activating this NNPF has nnpfa_persistence_flag equal to 1, nnpfc_interpolated_pics[i] is greater than 0 only for a single i value greater than 0, and currPic is the last picture in the output-order bitstream with nuh_layer_id equal to currLayerId, then the variable numPostRoll is set to the value of i that makes nnpfc_interpolated_pics[i] greater than 0, and the variable numInferences is set to 1 + numPostRoll.

[0537] Otherwise, the variable numInferences is set to 1.

[0538] For each value of j in the range of 0 to numInferences-1 (inclusive), the following applies:

[0539] - The arrays inputPic[i] and inputPresentFlag[i], representing all input images and the existence of input images within the bitstream, respectively, with i ranging from 0 to numInputPics-1 (inclusive of boundary values), are defined as follows:

[0540] - When j is greater than 0, for each k value in the range of 0 to j-1 (inclusive of boundary values), inputPic[k] is set to currPic, and inputPresentFlag[k] is set to equal to 0.

[0541] - The j-th input image inputPic[j] is set to equal currPic, and inputPresentFlag[j] is set to equal 1.

[0542] - When numInputPics is greater than 1, for each value of i in ascending order of i within the range of j+1 to numInputPics-j-1 (inclusive of boundary values), the following applies:

[0543] - If pictureRateUpsamplingFlag equals 1, currPic is associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5 and a specific fp_current_frame_is_frame0_flag value, and there exists a cropped decoded output image prevPic, which is the last image in output order among all cropped decoded output images associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5 and the same fp_current_frame_is_frame0_flag value, then inputPic[i] is set to prevPic, and inputPresentFlag[i] is set to equal to 1.

[0544] Otherwise, if pictureRateUpsamplingFlag equals 0 and there exists a picture prevPic that is the last picture in output order among all cropped decoded output pictures before inputPic[i-1] with nuh_layer_id equal to currLayerId, then inputPic[i] is set to prevPic and inputPresentFlag[i] is set to 1.

[0545] - Otherwise, if currPic is not associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5, and there exists a cropped decoded output image prevPic that is the last cropped decoded output image in output order among all cropped decoded output images before inputPic[i-1] with nuh_layer_id equal to currLayerId, then inputPic[i] is set to prevPic, and inputPresentFlag[i] is set to equal to 1.

[0546] - Otherwise, the following applies:

[0547] – inputPic[i] is set to the same image as inputPic[i-1], and inputPresentFlag[i] is set to 0.

[0548] – The requirement for bitstream consistency is that num_interpolated_pics[i-1] should not be greater than 0.

[0549] - To interpret NNPFC SEI messages, the following variables are specified:

[0550] - If numInputPics is greater than 1, and there exists a second NNPF defined by at least one NNPFC SEI message, activated by an NNPFA SEI message for currPic, and nnpfc_purpose equals 4, then the following applies:

[0551] – CroppedWidth is set to be equal to nnpfc_pic_width_in_luma_samples as defined for the second NNPF.

[0552] – CroppedHeight is set to be equal to nnpfc_pic_height_in_luma_samples as defined for the second NNPF.

[0553] - Otherwise, the following applies:

[0554] - CroppedWidth is set to equal pps_pic_width_in_luma_samples of curPic -SubWidthC (pps_conf_win_left_offset + pps_conf_win_right_offset) value.

[0555] - CroppedHeight is set to equal to currPic's pps_pic_height_in_luma_samples -SubHeightC The value of (pps_conf_win_top_offset + pps_conf_win_bottom_offset).

[0556] - For each value i in the range of 0 to numInputPics-1 (inclusive), the luminance sample array CroppedYPic[i] and the chrominance sample arrays CroppedCbPic[i] and CroppedCrPic[i] (if they exist) are derived as follows:

[0557] - The variable sourcePic is derived as follows:

[0558] - If inputPresentFlag[i] equals 1 or nnpfc_absent_input_pic_zero_flag equals 0, then sourcePic is set to inputPic[i].

[0559] - Otherwise (inputPresentFlag[i] equals 0 and nnpfc_absent_input_pic_zero_flag equals 1), sourcePic is set to an image of 0 CroppedWidth × CroppedHeight samples in the luminance sample array and 0 (CroppedWidth / SubWidthC) × (CroppedHeight / SubHeightC) samples in the Cb and Cr sample arrays.

[0560] - If numInputPics equals 1, then the following applies:

[0561] - The luminance sample array CroppedYPic[i] and the chrominance sample arrays CroppedCbPic[i] and CroppedCrPic[i] (if present) are set as two-dimensional arrays of the decoded sample values ​​of the Y, Cb and Cr components of sourcePic, respectively.

[0562] - Otherwise (numInputPics is greater than 1), the following applies:

[0563] - The variable sourceWidth is set to equal pps_pic_width_in_luma_samples of sourcePic - SubWidthC (pps_conf_win_left_offset + pps_conf_win_right_offset) value.

[0564] - The variable sourceHeight is set to equal pps_pic_height_in_luma_samples of sourcePic - SubHeightC The value of (pps_conf_win_top_offset + pps_conf_win_bottom_offset).

[0565] - If sourceWidth equals CroppedWidth and sourceHeight equals CroppedHeight, then inputPic is set to be the same as sourcePic.

[0566] - Otherwise (sourceWidth is not equal to CroppedWidth or sourceHeight is not equal to CroppedHeight), the following applies:

[0567] - There must be an NNPF defined by at least one NNPFC SEI message, activated by an NNPFA SEI message for sourcePic, and with nnpfc_purpose equal to 4, nnpfc_pic_width_in_luma_samples equal to CroppedWidth, and nnpfc_pic_height_in_luma_samples equal to CroppedHeight. This is referred to as the super-resolution NNPF.

[0568] - resampledPic is set as the output of the neural network inference of the super-resolution NNPF, where sourcePic is the input.

[0569] - The luminance sample array CroppedYPic[i] and the chrominance sample arrays CroppedCbPic[i] and CroppedCrPic[i] (if present) are set as two-dimensional arrays of the decoded sample values ​​of the Y, Cb and Cr components of the resampledPic, respectively.

[0570] - BitDepth Yand BitDepth C All of them are set to equal BitDepth.

[0571] - ChromaFormatIdc is set to equal sps_chroma_format_idc.

[0572] - The array StrengthControlVal[i], which specifies the filter strength control value of the input image for NNPF, and contains all values ​​of i in the range of 0 to numInputPics-1 (inclusive), is derived as follows:

[0573] - StrengthControlVal[i] is set to equal to (firstSliceQp) Y The value of (63 + QpBdOffset) ÷ (63 + QpBdOffset), where firstSliceQp Y The SliceQp of the first slice equal to inputPic[i] Y .

[0574] Within a picture unit, there should be no more than two NNPFC SEI messages with the same nnpfc_id. When two NNPFC SEI messages with the same nnpfc_id exist within a picture unit, these SEI messages must have different content. When two NNPFC SEI messages with the same nnpfc_id but different content exist within the same picture unit, both NNPFC SEI messages must be within the same SEI NAL unit.

[0575] 4. The technical problem solved by the disclosed technical solution

[0576] The current design of the Neural Network Post-Processing Filter (NNPFC) SEI message has the following problems:

[0577] Syntax elements whose names end with "_idc" are also called indicator syntax elements. Some indicator syntax elements are ue(v) encoded, such as nnpfc_inp_format_idc, nnpfc_out_format_idc, nnpfc_mode_idc, nnpfc_inp_order_idc, nnpfc_out_order_idc, and / or nnpfc_auxiliary_inp_idc. However, to save bit rate, it may be necessary to reduce the maximum value of these syntax elements and replace the ue(v) encoded syntax elements with u(N) encoded syntax elements (where N is a positive integer). Furthermore, since the number of possible values ​​for each of these indicator syntax elements is very limited—for example, some indicator syntax elements have no more than four possible values—it is simpler to encode one or more of these indicator syntax elements with fixed-length encoding (i.e., u(N) encoding, where N is a positive integer) rather than encoding them as ue(v) encoded syntax elements.

[0578] 5. List of solutions and implementation examples

[0579] To address the aforementioned problems, methods outlined below are disclosed. This invention should be considered as an example of interpreting general concepts and not interpreted in a narrow sense. Furthermore, these inventions can be applied individually or in any combination.

[0580] 1) To solve problem 2, instead of using ue(v) encoding or decoding syntax elements, some or all of the following syntax elements are encoded or decoded as u(N) encoding syntax elements, where N is an integer greater than 0.

[0581] a. In one example, for one or more of these syntax elements, the value of N is equal to 2, 3, or 4.

[0582] b. In one example, different syntax elements can use different values ​​of N.

[0583] i. In one example, nnpfc_inp_format_idc and nnpfc_out_format_idc can be encoded and decoded as syntax elements of u(3), while nnpfc_mode_idc, nnpfc_inp_order_idc, nnpfc_out_order_idc and nnpfc_auxiliary_inp_idc can be encoded and decoded as syntax elements of u(4).

[0584] c. In one example, nnpfc_mode_idc is encoded as a syntax element of u(N) encoding, not ue(v) encoding. Therefore, the value of nnpfc_mode_idc is between 0 and 2. N -1 (including boundary values).

[0585] d. In one example, nnpfc_inp_format_idc is encoded as a syntax element of u(N) encoding, not ue(v) encoding. Therefore, the value of nnpfc_inp_format_idc is between 0 and 2. N -1 (including boundary values).

[0586] e. In one example, nnpfc_inp_order_idc is encoded and decoded as a syntax element of u(N) encoding, rather than a syntax element of ue(v) encoding. Therefore, the value of nnpfc_inp_order_idc is between 0 and 2. N -1 (including boundary values).

[0587] f. In one example, nnpfc_auxiliary_inp_idc is encoded as a syntax element of u(N) encoding, rather than a syntax element of ue(v) encoding. Therefore, the value of nnpfc_auxiliary_inp_idc is between 0 and 2. N -1 (including boundary values).

[0588] g. In one example, nnpfc_out_format_idc is encoded and decoded as a syntax element of u(N) encoding, rather than a syntax element of ue(v) encoding. Therefore, the value of nnpfc_out_format_idc is between 0 and 2. N -1 (including boundary values).

[0589] h. In one example, nnpfc_out_order_idc is encoded and decoded as a syntax element of u(N) encoding, rather than a syntax element of ue(v) encoding. Therefore, the value of nnpfc_out_order_idc is between 0 and 2. N -1 (including boundary values).

[0590] i. In one example, any combination of the above syntax elements can be encoded as a syntax element encoded by u(N) rather than a syntax element encoded by ue(v).

[0591] j. In one example, all the above syntax elements are encoded as u(N) encoded syntax elements, rather than ue(v) encoded syntax elements.

[0592] 6. Example.

[0593] The following are some example embodiments of the invention aspects outlined in Section 5 of the previous article.

[0594] Most of the relevant sections that have been added or modified are shown in bold, and some of the deleted sections are shown in bold italic. There may be other changes that are essentially editable, and therefore are not indicated.

[0595] 6.1 Example 1.

[0596] This embodiment refers to items 1 and 2 and all their sub-items as outlined in Section 5 of the previous article.

[0597] 8.28.1 Characteristics of Neural Network Post-Processing Filters and SEI Message Syntax

[0598]

[0599] 8.28.2 Characteristics of Neural Network Post-Processing Filters and Semantics of SEI Messages

[0600] This specifies the method for converting sample values ​​from an input image into NNPF input values.

[0601] When nnpfc_inp_format_idc equals 0, the input values ​​of NNPF are real numbers, and the functions InpY() and InpC() are defined as follows:

[0602] InpY( x ) = x ÷ ( ( 1 << BitDepth Y ) - 1 ) (81)

[0603] InpC( x )= x ÷ ( ( 1 << BitDepth C ) - 1 ) (82)

[0604] When nnpfc_inp_format_idc equals 1, the input values ​​of NNPF are unsigned integers, and the functions InpY() and InpC() are defined as follows:

[0605] shiftY = BitDepth Y - inpTensorBitDepth Y

[0606] if (inpTensorBitDepth) Y >= BitDepth Y )

[0607] InpY( x ) = x << ( inpTensorBitDepth Y - BitDepth Y (83)

[0608] else

[0609] InpY( x ) = Clip3(0, ( 1 << inpTensorBitDepth Y ) - 1, ( x + ( 1 << (shiftY - 1 ) ) ) >> shiftY )

[0610] shiftC = BitDepth C - inpTensorBitDepth C

[0611] if (inpTensorBitDepth) C >= BitDepth C )

[0612] InpC( x ) = x << ( inpTensorBitDepth C - BitDepth C (84)

[0613] else

[0614] InpC( x ) = Clip3(0, ( 1 << inpTensorBitDepth C) - 1, ( x + ( 1 << (shiftC - 1 ) ) ) >> shiftC )

[0615] variable inpTensorBitDepth Y It is derived from the syntax element nnpfc_inp_tensor_luma_bitdepth_minus8 as specified below. The variable inpTensorBitDepth C It is derived from the syntax element nnpfc_inp_tensor_chroma_bitdepth_minus8 as specified below.

[0616]

[0617] An equal-0 value indicates that the sample value output by NNPF is a real number, where the range of values ​​from 0 to 1 (inclusive) is linearly mapped to the range of unsigned integer values ​​from 0 to (1 << bitDepth) - 1 (inclusive), where bitDepth is any bit depth required for subsequent post-processing or display.

[0618] An nnpfc_out_format_idc value equal to 1 indicates that the luminance sample values ​​output by NNPF are between 0 and (1 < 1). <outTensorBitDepth Y An unsigned integer in the range of 0 to 1 (inclusive), and the chroma sample values ​​output by NNPF are in the range of 0 to (1 << outTensorBitDepth). C An unsigned integer within the range of -1 (including boundary values).

[0619] nnpfc_out_format_idc Reserved for future use by ITU-T | ISO / IEC Furthermore, it should not exist in the bitstream conforming to this version of the document. Decoders conforming to this version of the document must ignore it. nnpfc_out_format_idc NNPFC SEI message.

[0620] 6.2 Example 2.

[0621] This embodiment refers to item 2 and all its sub-items as outlined in section 5 of the previous article.

[0622] 8.28.1 Characteristics of Neural Network Post-Processing Filters and SEI Message Syntax

[0623]

[0624] 8.28.2 Characteristics of Neural Network Post-Processing Filters and Semantics of SEI Messages

[0625] An equal value of 0 indicates that the SEI message contains an ISO / IEC 15938-17 bitstream that specifies the base NNPF (when nnpfc_base_flag equals 1) or an update relative to the base NNPF with the same nnpfc_id value (when nnpfc_base_flag equals 0).

[0626] When nnpfc_base_flag equals 1, nnpfc_mode_idc equals 1, which specifies that the underlying NNPF associated with the nnpfc_id value is a neural network identified by the URI indicated by nnpfc_uri, which has a format identified by the tag URI nnpfc_tag_uri.

[0627] When nnpfc_base_flag equals 0, nnpfc_mode_idc equals 1, indicating that the update relative to the base NNPF with the same nnpfc_id value is defined by the URI indicated by nnpfc_uri, which has a format identified by the tag URI nnpfc_tag_uri.

[0628] In the bitstream conforming to this version of the document, the value of nnpfc_mode_idc must be in the range of 0 to 1 (inclusive). The value of nnpfc_mode_idc is 2 to... (Including boundary values) are reserved for future use by ITU-T | ISO / IEC and should not exist in the bitstream conforming to this version of the document. Decoders conforming to this version of the document must ignore nnpfc_mode_idc values ​​between 2 and... NNPFC SEI messages within the range (including boundary values).

[0629] A value greater than 0 indicates that the auxiliary input data exists in the input tensor of NNPF. nnpfc_auxiliary_inp_idc equal to 0 indicates that the auxiliary input data does not exist in the input tensor. nnpfc_auxiliary_inp_idc equal to 1 indicates that the auxiliary input data is derived according to Equation 85.

[0630] In the bitstream conforming to this version of the document, the value of nnpfc_auxiliary_inp_idc must be in the range of 0 to 1 (inclusive). The value of nnpfc_auxiliary_inp_idc is 2 to... (Including boundary values) are reserved for future use by ITU-T | ISO / IEC and should not exist in the bitstream conforming to this version of the document. Decoders conforming to this version of the document must ignore nnpfc_auxiliary_inp_idc in ranges 2 to 1. NNPFC SEI messages within the range (including boundary values).

[0631] Instructs on a method for sorting the sample array of the input image to form the input tensor of NNPF.

[0632] In the bitstream conforming to this version of the document, the value of nnpfc_inp_order_idc must be in the range of 0 to 3 (inclusive). The value of nnpfc_inp_order_idc must be 4 to... (Including boundary values) are reserved for future use by ITU-T | ISO / IEC and should not exist in the bitstream conforming to this version of the document. Decoders conforming to this version of the document must ignore nnpfc_inp_order_idc in ranges 4 to 5. NNPFC SEI messages within the range (including boundary values).

[0633] When ChromaFormatIdc is not equal to 1, nnpfc_inp_order_idc should not be equal to 3.

[0634] When ChromaFormatIdc equals 0, nnpfc_inp_order_idc must equal 0.

[0635] When chromaUpsamplingFlag equals 1, nnpfc_inp_order_idc should not equal 0.

[0636] Table 2 contains informative descriptions of the nnpfc_inp_order_idc values.

[0637]

[0638] Indicates the output order of the samples generated by NNPF.

[0639] In the bitstream conforming to this version of the document, the value of nnpfc_out_order_idc must be in the range of 0 to 3 (inclusive). The value of nnpfc_out_order_idc is 4 to... (Including boundary values) are reserved for future use by ITU-T | ISO / IEC and should not exist in the bitstream conforming to this version of the document. Decoders conforming to this version of the document must ignore nnpfc_out_order_idc in ranges 4 to 5. NNPFC SEI messages within the range (including boundary values).

[0640] When chromaUpsamplingFlag equals 1, nnpfc_out_order_idc should not equal 0 or 3.

[0641] When colourizationFlag equals 1, nnpfc_out_order_idc should not equal 0.

[0642] Table 3 contains informative descriptions of the nnpfc_out_order_idc values.

[0643]

[0644] 7. References

[0645] [1] [ITU-T and ISO / IEC, “High efficiency video coding,” Rec. ITU-TH.265 | ISO / IEC 23008-2 (in force edition).

[0646] [2] J. Chen, E. Alshina, GJ Sullivan, J.-R. Ohm, J. Boyce, "Algorithm description of Joint Exploration Test Model 7 (JEM7)," JVET-G1001, Aug. 2017.

[0647] [3] Rec. ITU-T H.266 | ISO / IEC 23090-3, “Versatile Video Coding,” 2022.

[0648] [4] Rec. ITU-T Rec. H.274 | ISO / IEC 23002-7, “Versatile SupplementalEnhancement Information Messages for Coded Video Bitstreams,” 2022.

[0649] [5] S. McCarthy, S. Deshpande, M. Hannuksela, Hendry, G. Sullivan,and Y.-K. Wang (editors), "Improvements under consideration for neuralnetwork post filter SEI messages," JVET output document JVET-AC2032, publiclyavailable online herein: https: / / jvet-experts.org / doc_end_user / current_document.php?id=12585.

[0650] [6] E. François, B. Bross, M. M. Hannuksela, A. Tourapis, and Y.-K.Wang (editors), “New level and systems-related supplemental enhancementinformation for VVC (Draft 5)”, JVET output document JVET-AD2005, publiclyavailable online herein: https: / / www.jvet-experts.org / doc_end_user / current_document.php?id=12975.

[0651] Figure 2This is a block diagram illustrating an example video processing system 4000 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0652] System 4000 may include an encoding component 4004 capable of implementing various encoding / decoding or coding methods described in this disclosure. Encoding component 4004 can reduce the average bit rate from the video input 4002 to the output of encoding component 4004 to produce an encoded representation of the video. Encoding techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of encoding component 4004 may be stored or transmitted via a communication connection such as that represented by component 4006. The stored or communicated bitstream (or encoded) representation of the video received at input 4002 may be used by component 4008 to generate pixel values ​​or displayable video that is sent to display interface 4010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “encoding / decoding” operations or tools, it is understood that encoding tools or operations are used by the encoder, and corresponding decoding tools or operations that reverse the encoded result will be performed by the decoder.

[0653] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE), etc. The technologies described in this disclosure can be embodied in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0654] Figure 3This is a block diagram of an example video processing apparatus 4100. Apparatus 4100 can be used to implement one or more methods described herein. Apparatus 4100 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processors(multiple) 4102 can be configured to implement one or more methods described in this disclosure. The memories(multiple) 4104 can be used to store data and code for implementing the methods and techniques described herein. The video processing circuitry 4106 can be used to implement some of the techniques described in this disclosure in hardware circuitry. In some embodiments, the video processing circuitry 4106 may be at least partially included in the processor 4102, such as a graphics coprocessor.

[0655] Figure 4 This is a flowchart of an example method 4200 for video processing according to an embodiment of the present disclosure. Method 4200 can be performed by a codec device (e.g., an encoder, decoder, etc.) configured to apply NNPF to a video or a portion thereof. That is, method 4200 can be implemented in an apparatus for processing video data, the apparatus including a processor and a non-transitory memory having instructions thereon, such as a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform method 4200. Furthermore, method 4200 can be performed by a non-transitory computer-readable medium comprising a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec device to perform method 4200 when executed by a processor. Method 4200 can be applied during the encoding / decoding process in which the codec device applies one or more filters to a video or a portion thereof.

[0656] In box 4202, the encoding / decoding device determines the value of the neural network post-processing filter characteristic (NNPFC) input format indicator (nnpfc_inp_format_inc) to be encoded / decoded as a syntax element of u(N), where N is an integer greater than 0.

[0657] In box 4204, the codec device performs conversion between visual media data and bitstream based on the NNPFC input format indicator. The conversion may include encoding at the encoder, decoding at the decoder, or a combination thereof.

[0658] Figure 5This is a block diagram illustrating an example video encoding / decoding system 4300 from which the techniques of this disclosure can be utilized. The video encoding / decoding system 4300 may include a source device 4310 and a target device 4320. The source device 4310 generates encoded video data, and this source device 4310 may be referred to as a video encoding device. The target device 4320 can decode the encoded video data generated by the source device 4310, and this target device 4320 may be referred to as a video decoding device.

[0659] Source device 4310 may include video source 4312, video encoder 4314, and input / output (I / O) interface 4316. Video source 4312 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or sources of computer graphics systems used to generate video data, or combinations of these sources. Video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and associated data. Codec pictures are codec representations of pictures. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to target device 4320 via network 4330 through I / O interface 4316. Encoded video data may also be stored on storage medium / server 4340 for access by target device 4320.

[0660] Target device 4320 may include I / O interface 4326, video decoder 4324, and display device 4322. I / O interface 4326 may include a receiver and / or a modem. I / O interface 4326 may acquire encoded video data from source device 4310 or storage medium / server 4340. Video decoder 4324 may decode the encoded video data. Display device 4322 may display the decoded video data to a user. Display device 4322 may be integrated with target device 4320 or may be external to target device 4320, wherein target device 4320 may be configured to interface with an external display device.

[0661] The video encoder 4314 and the video decoder 4324 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or further standards.

[0662] Figure 6 This is a block diagram illustrating an example of a video encoder 4400, which can be... Figure 5The system 4300 shown includes a video encoder 4314. The video encoder 4400 can be configured to perform any or all of the techniques disclosed herein. The video encoder 4400 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 4400. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0663] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402 (which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, an intra-frame prediction unit 4406, a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy coding unit 4414).

[0664] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0665] Furthermore, some components such as the motion estimation unit 4404 and the motion compensation unit 4405 can be highly integrated, but for illustrative purposes, they are shown separately in the example of the video encoder 4400.

[0666] The segmentation unit 4401 can segment an image into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.

[0667] The mode selection unit 4403 can select one of several encoding / decoding modes (intra-frame encoding / decoding or inter-frame encoding / decoding), for example, based on error results, and provide the resulting intra-frame or inter-frame encoded / decoded block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 4403 can select an intra-frame / inter-frame joint prediction (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 4403 can also select a resolution for the block based on motion vectors (e.g., sub-pixel precision or integer pixel precision).

[0668] To perform inter-frame prediction on the current video block, motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. Motion compensation unit 4405 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 4413 other than the image associated with the current video block.

[0669] The motion estimation unit 4404 and the motion compensation unit 4405 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.

[0670] In some examples, motion estimation unit 4404 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 4404 can then generate a reference index indicating the reference image in list 0 or list 1 (which contains the reference video block) and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0671] In other examples, motion estimation unit 4404 can perform bidirectional prediction on the current video block. Motion estimation unit 4404 can search for a reference video block for the current video block in the reference images of list 0, and can also search for another reference video block for the current video block in the reference images of list 1. Motion estimation unit 4404 can then generate reference indices indicating the reference images (containing reference video blocks) in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 4404 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0672] In some examples, the motion estimation unit 4404 can output a complete set of motion information for use in the decoder's decoding process. In some examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 4404 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0673] In one example, the motion estimation unit 4404 may indicate a value to the video decoder 4500 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.

[0674] In another example, motion estimation unit 4404 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0675] As discussed above, the video encoder 4400 can transmit motion vectors via signaling in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge Mode Signaling.

[0676] Intra-prediction unit 4406 can perform intra-prediction on the current video block. When intra-prediction unit 4406 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples of other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0677] The residual generation unit 4407 can generate residual data for the current video block by subtracting (or more) predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0678] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform subtraction operations.

[0679] The transform processing unit 4408 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0680] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0681] The inverse quantization unit 4410 and the inverse transform unit 4411 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block and store it in the buffer 4413.

[0682] After the video block is reconstructed by reconstruction unit 4412, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0683] The entropy coding unit 4414 can receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives data, it can perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[0684] Figure 7 This is a block diagram illustrating an example of a video decoder 4500, which can be... Figure 5 The system 4300 shown includes a video decoder 4324. The video decoder 4500 can be configured to perform any or all of the techniques disclosed herein. In the example shown, the video decoder 4500 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 4500. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0685] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-frame prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 4400.

[0686] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 4501 can decode the entropy-encoded video data, and the motion compensation unit 4502 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 4502 can determine such information, for example, by executing AMVP and Merge modes.

[0687] The motion compensation unit 4502 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used, with sub-pixel accuracy, can be included in the syntax element.

[0688] The motion compensation unit 4502 can use the interpolation filter used by the video encoder 4400 during the encoding of a video block to calculate the interpolation for sub-integer pixels of the reference block. The motion compensation unit 4502 can determine the interpolation filter used by the video encoder 4400 based on the received syntax information, and the motion compensation unit 4502 can use the interpolation filter to generate a prediction block.

[0689] The motion compensation unit 4502 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame codec block, and other information for decoding the encoded video sequence.

[0690] Intra-prediction unit 4503 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 4504 dequantizes (i.e., inverse quantization) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501. Inverse transform unit 4505 applies the inverse transform.

[0691] The reconstruction unit 4506 can add the residual block to the corresponding predicted block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be used to filter the decoded block to eliminate block artifacts. The decoded video block is then stored in a buffer 4507 to provide a reference block for subsequent motion compensation / intra-frame prediction, and also generates decoded video for presentation on a display device.

[0692] Figure 8This is a schematic diagram of the example encoder 4600. Encoder 4600 is suitable for implementing VVC techniques. Encoder 4600 includes three loop filters: a deblocking filter (DF) 4602, a sample adaptive compensation (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike DF 4602, which uses predefined filters, SAO 4604 and ALF 4606 utilize the original samples of the current image, respectively, by adding an offset and by applying a finite impulse response (FIR) filter, and by utilizing the side information from the encoding and decoding through signal transmission offset and filter coefficients to reduce the mean square error between the original and reconstructed samples. ALF 4606 is located in the last processing stage of each image and can be thought of as a tool to attempt to capture and repair artifacts caused by previous stages.

[0693] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using a reference image obtained from a reference image buffer 4612. Residual blocks from inter-frame or intra-frame prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are then fed into an entropy encoding / decoding component 4618. The entropy encoding / decoding component 4618 entropy-encodes and decodes the prediction results and quantized transform coefficients and transmits them to a video decoder (not shown). The quantization component output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. REC component 4624 is able to output images to DF 4602, SAO 4604 and ALF 4606 for filtering before these images are stored in reference image buffer 4612.

[0694] The following is a list of some preferred solutions.

[0695] The following solutions illustrate examples of the techniques discussed in this article.

[0696] 1. A method for processing media data, comprising: determining to apply a metadata extension mechanism to a neural network post-processing filter (NNPF) when the neural network post-processing filter feature (NNPFC) mode identification code (nnpfc_mode_idc) is equal to 0, and determining not to apply the metadata extension mechanism when nnpfc_mode_idc is not equal to 0, the metadata extension mechanism including the syntax element NNPF metadata extension bit count (nnpfc_metadata_extension_num_bits) and NNPFC reserved metadata extension (nnpfc_reserved_metadata_extension); and performing a conversion between visual media data and a bitstream based on the NNPF.

[0697] 2. The method according to Solution 1, wherein when nnpfc_mode_idc is not equal to 0, the signaling of the syntax element nnpfc_metadata_extension_num_bits is skipped.

[0698] 3. The method according to any one of solutions 1-2, wherein when nnpfc_metadata_extension_num_bits does not exist, the value of nnpfc_metadata_extension_num_bits is presumed to be equal to 0.

[0699] 4. The method according to any one of solutions 1-3, wherein when the NNPFC Supplemental Enhancement Information (SEI) message nnpfcB is not the first NNPFC SEI message in the current codec layer video sequence (CLVS) with an NNPFC identifier (nnpfc_id) value in the decoding order, the current CLVS contains nnpfcB, and the value of the NNPFC base flag (nnpfc_base_flag) is equal to 1, the NNPFC SEI message must be a repetition of the first NNPFC SEI message nnpfcA in the current CLVS with the same nnpfc_id value in the decoding order.

[0700] 5. The method according to any one of solutions 1-4, wherein the load content of nnpfcB must be the same as the load content of nnpfcA.

[0701] 6. The method according to any one of solutions 1-5, wherein the current CLVS is a CLVS containing the NNPFC SEI message nnpfcB, and the first NNPFC SEI message nnpfcA is within the current CLVS.

[0702] 7. The method according to any one of solutions 1-6, wherein the NNPFC missing input image zero flag (nnpfc_absent_input_pic_zero_flag) equal to 0 indicates that the NNPF expects the input image inputPicA, which is not present in the bitstream, to be represented by the input image that is closest to inputPicA in output order among all images in the bitstream after inputPicA in output order.

[0703] 8. The method according to any one of solutions 1-7, wherein nnpfc_absent_input_pic_zero_flag equal to 0 indicates that the NNPF expects an input image inputPicA that does not exist in the bitstream to be represented by the input image that is closest to inputPicA in output order among all images in the bitstream that precede inputPicA in output order.

[0704] 9. The method according to any one of solutions 1-8, wherein nnpfc_absent_input_pic_zero_flag equal to 0 indicates that the NNPF expects an input image inputPicA that does not exist in the bitstream to be represented by an input image inputPicB.

[0705] 10. The method according to any one of solutions 1-9, wherein when inputPicA precedes the first image in the bitstream in the output order, inputPicB is the image in the bitstream that is closest to inputPicA in the output order among all images in the bitstream that follow inputPicA in the output order.

[0706] 11. The method according to any one of solutions 1-10, wherein when inputPicA follows the last image in the bitstream in output order, inputPicB is the image in the bitstream that is closest to inputPicA in output order among all images in the bitstream that precede inputPicA in output order.

[0707] 12. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of solutions 1-11.

[0708] 13. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec device to perform any one of the methods in solutions 1-11 when executed by a processor.

[0709] 14. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method comprises: determining, when a Neural Network Post-Processing Filter Feature (NNPFC) mode identification code (nnpfc_mode_idc) is equal to 0, to apply a metadata extension mechanism to a Neural Network Post-Processing Filter (NNPF), and when nnpfc_mode_idc is not equal to 0, to determine not to apply the metadata extension mechanism, the metadata extension mechanism comprising the syntax element NNPF metadata extension bit count (nnpfc_metadata_extension_num_bits) and NNPFC reserved metadata extension (nnpfc_reserved_metadata_extension); and generating a bitstream based on the determination.

[0710] 15. A method for storing a bitstream of video, comprising: determining, when a Neural Network Post-Processing Filter Feature (NNPFC) mode identification code (nnpfc_mode_idc) is equal to 0, to apply a metadata extension mechanism to a Neural Network Post-Processing Filter (NNPF), and when nnpfc_mode_idc is not equal to 0, to determine not to apply the metadata extension mechanism, the metadata extension mechanism including a syntax element NNPF metadata extension bit count (nnpfc_metadata_extension_num_bits) and an NNPFC reserved metadata extension (nnpfc_reserved_metadata_extension); generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0711] 16. The methods, apparatus or systems described in this disclosure.

[0712] The following solutions illustrate further examples of the techniques discussed in this article.

[0713] 1. A method for processing media data, comprising: determining that the value of a neural network post-processing filter characteristic (NNPFC) input format indicator (nnpfc_inp_format_inc) is encoded and decoded into a syntax element of u(N) encoding and decoding, wherein N is an integer greater than 0; and performing a conversion between visual media data and a bitstream based on the NNPFC input format indicator.

[0714] 2. The method according to Solution 1, wherein the NNPFC output format indicator (nnpfc_out_format_idc) is encoded and decoded into a syntax element of u(N) encoding and decoding, where N is an integer greater than 0.

[0715] 3. The method according to any one of solutions 1-2, wherein the NNPFC mode indicator (nnpfc_mode_idc) is encoded and decoded into a syntax element of u(N) encoding and decoding, where N is an integer greater than 0.

[0716] 4. The method according to any one of solutions 1-3, wherein the NNPFC input order indicator (nnpfc_inp_order_idc) is encoded and decoded into a syntax element of u(N) encoding and decoding, where N is an integer greater than 0.

[0717] 5. The method according to any one of solutions 1-4, wherein the NNPFC output order indicator (nnpfc_out_order_idc) is encoded and decoded into a syntax element of u(N) encoding and decoding, where N is an integer greater than 0.

[0718] 6. The method according to any one of solutions 1-5, wherein the NNPFC auxiliary input indicator (nnpfc_auxiliary_inp_idc) is encoded and decoded into a syntax element of u(N) encoding and decoding, where N is an integer greater than 0.

[0719] 7. The method according to any one of solutions 1-6, wherein N is a value equal to 2, 3 or 4.

[0720] 8. The method according to any one of solutions 1-7, wherein nnpfc_inp_format_idc and nnpfc_out_format_idc are encoded and decoded as syntax elements of u(3), and wherein nnpfc_mode_idc, nnpfc_inp_order_idc, nnpfc_out_order_idc and nnpfc_auxiliary_inp_idc are encoded and decoded as syntax elements of u(4).

[0721] 9. The method according to any one of solutions 1-8, wherein nnpfc_mode_idc is encoded / decoded as a syntax element of u(N) encoding / decoding, and wherein the value of nnpfc_mode_idc is constrained to between 0 and 2. N -1 (including boundary values).

[0722] 10. The method according to any one of solutions 1-9, wherein nnpfc_inp_format_idc is encoded / decoded into a syntax element of u(N) encoding / decoding, and wherein the value of nnpfc_inp_format_idc is constrained to between 0 and 2. N -1 (including boundary values).

[0723] 11. The method according to any one of solutions 1-10, wherein nnpfc_inp_order_idc is encoded / decoded into a syntax element of u(N) encoding / decoding, and wherein the value of nnpfc_inp_order_idc is constrained to between 0 and 2. N -1 (including boundary values).

[0724] 12. The method according to any one of solutions 1-11, wherein nnpfc_auxiliary_inp_idc is encoded / decoded into a syntax element of u(N) encoding / decoding, and wherein the value of nnpfc_auxiliary_inp_idc is constrained to between 0 and 2. N -1 (including boundary values).

[0725] 13. The method according to any one of solutions 1-12, wherein nnpfc_out_format_idc is encoded / decoded into a syntax element of u(N) encoding / decoding, and wherein the value of nnpfc_out_format_idc is constrained to between 0 and 2. N -1 (including boundary values).

[0726] 14. The method according to any one of solutions 1-13, wherein nnpfc_out_order_idc is encoded / decoded into a syntax element of u(N) encoding / decoding, and wherein the value of nnpfc_out_order_idc is constrained to between 0 and 2. N -1 (including boundary values).

[0727] 15. The method according to any one of solutions 1-14, wherein nnpfc_inp_format_idc indicates a mechanism for converting sample values ​​of an input image into input values ​​of a neural network post-processing filter (NNPF).

[0728] 16. The method according to any one of solutions 1-15, wherein nnpfc_inp_order_idc indicates a mechanism for sorting the sample array of the input image to form the input tensor of NNPF.

[0729] 17. The method according to any one of solutions 1-16, wherein nnpfc_out_order_idc indicates the output order of samples generated by NNPF.

[0730] 18. The method according to any one of solutions 1-17, wherein nnpfc_out_format_idc indicates whether the sample value output by NNPF is a real number or an unsigned integer.

[0731] 19. The method according to any one of solutions 1-18, wherein nnpfc_mode_idc indicates whether the supplementary enhancement information (SEI) message specifies the underlying NNPF or an update to the underlying NNPF.

[0732] 20. The method according to any one of solutions 1-19, wherein nnpfc_auxiliary_inp_idc indicates whether auxiliary input data exists in the input tensor of NNPF.

[0733] 21. The method according to any one of solutions 1-20, wherein the conversion includes encoding the visual media data into a bitstream.

[0734] 22. The method according to any one of solutions 1-20, wherein the conversion includes decoding the visual media data from the bitstream.

[0735] 23. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of solutions 1-22.

[0736] 24. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec device to perform any one of solutions 1-22 when executed by a processor.

[0737] 25. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method comprises: determining that the value of a neural network post-processing filter characteristic (NNPFC) input format indicator (nnpfc_inp_format_inc) is encoded and decoded into a syntax element of u(N) encoding and decoding, wherein N is an integer greater than 0; and generating the bitstream based on the NNPFC input format indicator.

[0738] 26. A method for storing a bitstream of video, comprising: determining that the value of a neural network post-processing filter characteristic (NNPFC) input format indicator (nnpfc_inp_format_inc) is encoded and decoded into a syntax element of u(N) encoding and decoding, wherein N is an integer greater than 0; generating the bitstream based on the NNPFC input format indicator; and storing the bitstream in a non-transitory computer-readable recording medium.

[0739] In the described solution, the encoder conforms to the format rules by generating a codec representation based on those rules. In the described solution, the decoder parses the syntax elements in the codec representation using known information about their presence or absence, based on the format rules, to produce the decoded video.

[0740] In this disclosure, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, and vice versa. For example, the bitstream representation of a current video block can correspond to bits at co-positions or propagated at different positions in a bitstream defined by a syntax. For example, a macroblock can be encoded based on the error residual values ​​after transformation and encoding / decoding, and can also use bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can parse the bitstream based on the determination, knowing whether certain fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude specific syntax fields, and generate the encoding / decoding representation accordingly by including or excluding syntax fields from the encoding / decoding representation.

[0741] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in one or more combinations thereof. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more computer program instruction modules encoded on a computer-readable medium for execution by a data processing apparatus or for controlling the operation of the data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of substances influencing machine-readable propagation signals, or a combination thereof. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for an associated computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. Propagation signals are artificially generated signals, such as machine-generated electrical signals, optical signals, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.

[0742] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.

[0743] The processing and logic flows described in this disclosure can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processing and logic flows can also be executed by dedicated logic circuitry, and the apparatus can be implemented as dedicated logic circuitry, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs).

[0744] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors in any type of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are the processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor storage devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0745] While this disclosure contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features specific to particular embodiments of a particular art. Certain features described in the context of individual embodiments in this disclosure may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. Furthermore, although features may function in certain combinations as described above, and even were originally claimed in this manner, in some cases one or more features in the claimed combination may be removed from that combination, and the claimed combination may be for sub-combinations or variations thereof.

[0746] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed sequentially in the particular order or sequence shown, or requiring all shown operations to be performed in order to achieve the desired result. Furthermore, the partitioning of various system components in the embodiments described in this disclosure should not be construed as requiring such partitioning in all embodiments.

[0747] Only a few implementations and examples are described, and other implementations, improvements and variations may be made based on what is described and shown in this disclosure.

[0748] When there are no intermediate components other than a line, trace, or other medium between the first and second components, the first component is directly coupled to the second component. When there are intermediate components other than a line, trace, or other medium between the first and second components, the first component is indirectly coupled to the second component. The term "coupled" and its variations include direct coupling and indirect coupling. The use of the term "approximately" means including a range of ±10% of the following figures, unless otherwise specified.

[0749] While several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. The examples presented are intended to be illustrative rather than restrictive and are not intended to be limited to the details set forth herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.

[0750] Furthermore, the technologies, systems, subsystems, and methods described and illustrated as discrete or separate in the various embodiments can be combined or integrated with other systems, modules, technologies, or methods without departing from the scope of this disclosure. Other items shown or discussed as couplings can be directly connected or indirectly coupled or communicated through some interface, device, or intermediate component (whether electrical, mechanical, or otherwise). Examples of other changes, substitutions, and modifications can be determined by those skilled in the art and can be made without departing from the spirit and scope of this disclosure.

Claims

1. A method for processing media data, comprising: The value of the input format indicator (nnpfc_inp_format_inc) for determining the characteristics of the Neural Network Post-Processing Filter (NNPFC) is encoded and decoded into u(N) encoding / decoding syntax elements, where N is an integer greater than 0; and The conversion between visual media data and bitstream is performed based on the NNPFC input format indicator.

2. The method according to claim 1, wherein, The NNPFC output format indicator (nnpfc_out_format_idc) is encoded and decoded into u(N) encoding and decoding syntax elements, where N is an integer greater than 0.

3. The method according to any one of claims 1-2, wherein, The NNPFC mode indicator (nnpfc_mode_idc) is encoded and decoded as a syntax element of u(N), where N is an integer greater than 0.

4. The method according to any one of claims 1-3, wherein, The NNPFC input order indicator (nnpfc_inp_order_idc) is encoded and decoded into a syntax element of u(N), where N is an integer greater than 0.

5. The method according to any one of claims 1-4, wherein, The NNPFC output order indicator (nnpfc_out_order_idc) is encoded and decoded into a syntax element of u(N), where N is an integer greater than 0.

6. The method according to any one of claims 1-5, wherein, The NNPFC auxiliary input indicator (nnpfc_auxiliary_inp_idc) is encoded and decoded as a syntax element of u(N), where N is an integer greater than 0.

7. The method according to any one of claims 1-6, wherein, N is a value equal to 2, 3, or 4.

8. The method according to any one of claims 1-7, wherein, nnpfc_inp_format_idc and nnpfc_out_format_idc are encoded and decoded as syntax elements of u(3), and among them, nnpfc_mode_idc, nnpfc_inp_order_idc, nnpfc_out_order_idc and nnpfc_auxiliary_inp_idc are encoded and decoded as syntax elements of u(4).

9. The method according to any one of claims 1-8, wherein, nnpfc_mode_idc is encoded and decoded into a syntax element of u(N) encoding and decoding, and the value of nnpfc_mode_idc is constrained to between 0 and 2. N -1 (including boundary values).

10. The method according to any one of claims 1-9, wherein, nnpfc_inp_format_idc is encoded and decoded into a syntax element of u(N) encoding and decoding, and the value of nnpfc_inp_format_idc is constrained to between 0 and 2. N -1 (including boundary values).

11. The method according to any one of claims 1-10, wherein, nnpfc_inp_order_idc is encoded and decoded into a syntax element of u(N) encoding and decoding, and the value of nnpfc_inp_order_idc is constrained to between 0 and 2. N -1 (including boundary values).

12. The method according to any one of claims 1-11, wherein, nnpfc_auxiliary_inp_idc is encoded and decoded into a syntax element of u(N) encoding and decoding, and the value of nnpfc_auxiliary_inp_idc is constrained to between 0 and 2. N The range is -1 (including boundary values).

13. The method according to any one of claims 1-12, wherein, nnpfc_out_format_idc is encoded and decoded into syntax elements of u(N) encoding and decoding, and the value of nnpfc_out_format_idc is constrained to between 0 and 2. N -1 (including boundary values).

14. The method according to any one of claims 1-13, wherein, nnpfc_out_order_idc is encoded and decoded into a syntax element of u(N) encoding and decoding, and the value of nnpfc_out_order_idc is constrained to between 0 and 2. N -1 (including boundary values).

15. The method according to any one of claims 1-14, wherein, nnpfc_inp_format_idc indicates the mechanism for converting sample values ​​of an input image into input values ​​for a neural network post-processing filter (NNPF).

16. The method according to any one of claims 1-15, wherein, nnpfc_inp_order_idc indicates the mechanism for sorting the sample array of the input image to form the input tensor of NNPF.

17. The method according to any one of claims 1-16, wherein, nnpfc_out_order_idc indicates the output order of samples generated by NNPF.

18. The method according to any one of claims 1-17, wherein, nnpfc_out_format_idc indicates whether the sample values ​​output by NNPF are real numbers or unsigned integers.

19. The method according to any one of claims 1-18, wherein, The nnpfc_mode_idc indicates whether the Supplemental Enhancement Information (SEI) message specifies the underlying NNPF or an update to the underlying NNPF.

20. The method according to any one of claims 1-19, wherein, nnpfc_auxiliary_inp_idc indicates whether auxiliary input data exists in the input tensor of NNPF.

21. The method according to any one of claims 1-20, wherein, The conversion includes encoding the visual media data into a bitstream.

22. The method according to any one of claims 1-20, wherein, The conversion includes decoding the visual media data from the bitstream.

23. An apparatus for processing video data, comprising: processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1-22.

24. A non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec apparatus to perform the method of any one of claims 1-22 when executed by a processor.

25. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein, The method includes: The value of the input format indicator (nnpfc_inp_format_inc) for determining the characteristics of the Neural Network Post-Processing Filter (NNPFC) is encoded and decoded into u(N) encoding / decoding syntax elements, where N is an integer greater than 0; and The bitstream is generated based on the NNPFC input format indicator.

26. A method for storing a video bitstream, comprising: The value of the input format indicator (nnpfc_inp_format_inc) for determining the characteristics of the neural network post-processing filter (NNPFC) is encoded and decoded into u(N) encoding and decoding syntax elements, where N is an integer greater than 0; The bitstream is generated based on the NNPFC input format indicator; and The bit stream is stored in a non-transitory computer-readable recording medium.