Neural network post-processing filter output picture and destination combination

By ensuring consistent output time instances in the list of output images from the neural network post-processing filter, the problem of inconsistent output in video encoding and decoding is solved, improving processing efficiency and consistency. It is applicable to a variety of video encoding and decoding standards and bitstream formats.

CN120982092APending Publication Date: 2025-11-18DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480024097.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-06
Filing Date
2024-04-03
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies suffer from inconsistent output time instances when processing images output by neural network post-processing filters, resulting in low efficiency in video data processing.

Method used

By determining that there should be no more than one instance in the list of output images from the neural network post-processing filter that relates to any particular output time, and based on this, the video data is converted to a bitstream, and the corresponding instructions are executed using the processor and non-transitory memory to achieve consistent output processing.

Benefits of technology

It improves the efficiency and consistency of video data processing, optimizes the video encoding and decoding process, and is applicable to a variety of video encoding and decoding standards and bitstream formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120982092A_ABST
    Figure CN120982092A_ABST
Patent Text Reader

Abstract

A mechanism for processing video data is disclosed. The mechanism includes determining that there is no more than one picture involving any particular output time instance within a list of neural network post-processing filter (NNPF) output pictures. A conversion between the visual media data and the bitstream is then performed based on the NNPF output picture.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross Reference to Related Patent Applications

[0002] This patent application claims the benefit of International Patent Application No. PCT / CN2023 / 086728, filed April 6, 2023, the teachings and disclosure of which are incorporated by reference in their entirety. TECHNICAL FIELD

[0003] This patent document relates to the generation, storage, and use of digital audio-visual media information in file formats. BACKGROUND

[0004] Digital video accounts for the largest bandwidth use on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video usage can continue to grow. SUMMARY

[0005] A first aspect relates to a method of processing video data, comprising: determining that there should not be more than one picture referring to any particular output time instance within a list of neural network post filter (NNPF) output pictures; and performing a conversion between a visual media data and a bitstream based on the NNPF output pictures.

[0006] A second aspect relates to an apparatus of processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform any of the above aspects.

[0007] A third aspect relates to a non-transitory computer readable medium comprising a computer program product for use by a video coding device, the computer program product including instructions stored on the non-transitory computer readable medium that when executed by a processor cause the video coding device to perform the method of any of the above aspects.

[0008] A fourth aspect relates to a non-transitory computer readable recording medium storing a bitstream of a video generated by a method performed by a video processing apparatus, wherein the method comprises: determining that there should not be more than one picture referring to any particular output time instance within a list of neural network post filter (NNPF) output pictures; and generating the bitstream based on the determination.

[0009] A fifth aspect relates to a method for storing a bitstream of a video, comprising: determining that there should not be more than one picture referring to any particular output time instance within a list of neural network post filter (NNPF) output pictures; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer readable recording medium.

[0010] For clarity, any of the foregoing embodiments may be combined with one or more other foregoing embodiments to create new embodiments within the scope of this disclosure.

[0011] These and other features will be more clearly understood through the following detailed description with reference to the accompanying drawings and claims. Attached Figure Description

[0012] For a more complete understanding of this disclosure, reference is now made to the following brief description, along with the accompanying drawings and detailed description, wherein like reference numerals denote like parts.

[0013] Figure 1 An example of deriving the luminance channel from the luminance component is shown.

[0014] Figure 2 This is a block diagram illustrating an example video processing system.

[0015] Figure 3 This is a block diagram of an example video processing device.

[0016] Figure 4 This is a flowchart of an example method for video processing.

[0017] Figure 5 This is a block diagram illustrating an example video codec system.

[0018] Figure 6 This is a block diagram showing an example encoder.

[0019] Figure 7 This is a block diagram showing an example decoder.

[0020] Figure 8 This is a schematic diagram of an example encoder.

[0021] Figure 9 This is a flowchart of an example method for video processing. Detailed Implementation

[0022] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or yet to be developed. This disclosure should not be limited in any way to the illustrative implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but can be modified within the scope of the appended claims and all their equivalents.

[0023] Chapter headings are used in this document for ease of understanding, not to limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter only. Furthermore, H.266 terminology is used in some descriptions merely for ease of understanding, not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs. In this document, edited text changes are indicated by bold italics to indicate deleted text and by bold text to indicate added text, relative to the Multi-Functional Video Codec (VVC) specification.

[0024] 1. Preliminary Discussion

[0025] This document relates to image / video codec techniques. Specifically, this disclosure relates to the output image of a signal transmission and a specified neural network post-processing filter (NNPF), particularly when different NNPF purposes are combined. This idea can be applied alone or in various combinations to video bitstreams encoded or decoded by any codec (e.g., the Multi-Functional Video Codec (VVC) standard and / or the Multi-Functional Supplemental Enhancement Information (SEI) Message (VSEI) standard for encoding and decoding video bitstreams).

[0026] 2. Abbreviations

[0027] Adaptive Parameter Set (APS), Access Unit (AU), Codec Layer Video Sequence (CLVS), Codec Layer Video Sequence Start (CLVSS), Cyclic Redundancy Check (CRC), Codec Video Sequence (CVS), Finite Impulse Response (FIR), Intra-Frame Random Access Point (IRAP), Network Abstraction Layer (NAL), Neural Network Post-Processing Filter (NNPF), Neural Network Post-Processing Filter Activation (NNPFA), Neural Network Post-Processing Filter Characteristics (NNPFC), Picture Parameter Set (PPS), Picture Unit (PU), Random Access Skip Preamble (RASL) Picture, Supplemental Enhancement Information (SEI), Stepped Temporal Sublayer Access (STSA), Uniform Resource Identifier (URI), Video Codec Layer (VCL), Multifunctional Supplemental Enhancement Information (VSEI) described in Recommendation ITU-T H.274|ISO / IEC 23002-7, Video Availability Information (VUI), and Multifunctional Video Codec (VVC) described in Recommendation ITU-T H.266|ISO / IEC 23090-3.

[0028] 3. Further discussion

[0029] 3.1 Video codec standards

[0030] Video coding standards have evolved primarily through the development of standards by the Telecommunication Standardization Sector of the International Telecommunication Union (ITU-T) and the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). ITU-T developed the H.261 and H.263 standards, ISO / IEC developed the Moving Picture Experts Group (MPEG)-1 and MPEG-4 Vision, and the two organizations jointly developed the H.262 / MPEG-2 video standard and the H.264 / MPEG-4 Advanced Video Coding (AVC) standard and the H.265 / High Efficiency Video Coding (HEVC) standard[1]. Starting with H.262, video coding standards are based on a hybrid video coding architecture, which utilizes temporal prediction plus transform coding. In order to explore video coding technologies other than High Efficiency Video Coding (HEVC), the Joint Video Exploration Team (JVET) was established by the Video Coding Experts Group (VCEG) and the Moving Picture Experts Group (MPEG). In addition, JVET has adopted some methods and incorporated them into a reference software called the Joint Exploration Model (JEM)[2]. When the Multi-Functional Video Codec (VVC) project was officially launched, JVET was later renamed the Joint Video Experts Group (JVET). VVC[3] is a codec standard that aims to reduce the bit rate by 50% compared to HEVC.

[0031] The Multi-Functional Video Coding (VVC) standard (ITU-T H.266|ISO / IEC 23090-3)[3] and the related Multi-Functional Supplemental Enhancement Information (VSEI) standard (ITU-T H.274|ISO / IEC 23002-7)[4] are designed for use in the widest range of applications, including simple uses such as television broadcasting, video conferencing or playback from storage media, as well as more advanced use cases such as adaptive bitrate streaming, video region extraction, synthesis and merging of content from multiple codec video bitstreams, multi-view video, scalable layered coding and decoding, and viewport-adaptive 360° immersive media.

[0032] The Basic Video Codec (EVC) standard (ISO / IEC 23094-1) is another video codec standard developed by MPEG.

[0033] 3.2 General SEI messages in VVC and VSEI

[0034] SEI messages assist in processes related to decoding, display, or other purposes. However, SEI messages are not essential for constructing luma or chroma samples during the decoding process. Standard-compliant decoders do not need to process this information to achieve output order consistency. Some SEI messages are necessary for checking bitstream consistency and output timing decoder consistency. Other SEI messages are not necessary for checking bitstream consistency.

[0035] Appendix D of VVC specifies the syntax and semantics of SEI message payloads for some SEI messages, and specifies the use of SEI messages and VUI parameters whose syntax and semantics are specified in ITU-T H.274|ISO / IEC 23002-7.

[0036] 3.3 Signaling of Neural Network Post-Processing Filters

[0037] JVET-AC2032[5] includes specifications for two SEI messages used for signaling of neural network post-filters, as shown below.

[0038] 8.28 Neural Network Post-Processing Filter Characteristics SEI Message

[0039] 8.28.1 Characteristics of Neural Network Post-Processing Filters and SEI Message Syntax

[0040]

[0041]

[0042]

[0043] 8.28.2 Characteristics of Neural Network Post-Processing Filters and Semantics of SEI Messages

[0044] The Neural Network Post-Processing Filter Characteristics (NNPFC) SEI message specifies the neural networks that can be used as post-processing filters. The Neural Network Post-Processing Filter Activation (NNPFA) SEI message indicates the use of a specified Neural Network Post-Processing Filter (NNPF) for a particular image.

[0045] To use this SEI message, the following variables need to be defined:

[0046] - Input the width and height of the image, in units of brightness samples, which are denoted as CroppedWidth and CroppedHeight in this paper, respectively.

[0047] - Input images containing luminance sample arrays CroppedYPic[idx] and chrominance sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] (if they exist), where the index idx ranges from 0 to numInputPics-1 (inclusive), which are used as inputs to NNPF.

[0048] - BitDepth of the brightness sample array of the input image Y .

[0049] - The bit depth (BitDepth) of the input image's chroma sample array (if any). C .

[0050] - Chroma format indicator, referred to herein as ChromaFormatIdc, as described in subclause 7.3.

[0051] - When nnpfc_auxiliary_inp_idc equals 1, the filter strength control value StrengthControlVal should be a real number in the range of 0 to 1 (inclusive).

[0052] The input image at index 0 corresponds to the image for which the NNPF defined by the NNPFC SEI message is activated via the NNPFA SEI message. Input images with indices i in the range of 1 to numInputPics-1 (inclusive) precede the input image at index i-1 in the output order.

[0053] When an input image with index 0 and nnpfc_purpose&0x08 is not equal to 0 is associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5, all input images are associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5 and fp_current_frame_is_frame0_flag equal to the same value.

[0054] The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc, as specified in Table 2.

[0055] Note 1 – More than one NNPFC SEI message can exist for the same image. When more than one NNPFC SEI message with different values ​​of nnpfc_id exists or is activated for the same image, they can have the same or different values ​​of nnpfc_purpose and nnpfc_mode_idc.

[0056] nnpfc_purpose indicates the purpose of NNPF, as specified in Table 20.

[0057] The value of nnpfc_purpose should be in the range of 0 to 63 (inclusive) in the bitstream conforming to this version of this document. Values ​​of nnpfc_purpose from 64 to 65535 (inclusive) are reserved for future use by ITU-T|ISO / IEC and should not exist in the bitstream conforming to this version of this document. Decoders conforming to this version of this document should ignore NNPFC SEI messages with nnpfc_purpose in the range of 64 to 65535 (inclusive).

[0058] Table 20 - Definition of nnpfc_purpose

[0059]

[0060]

[0061] Note 2 – When the reserved value of nnpfc_purpose is used by ITU-T|ISO / IEC in the future, the syntax of this SEI message can be extended using the following syntax elements, provided that nnpfc_purpose is equal to that value.

[0062] When ChromaFormatIdc equals 3, nnpfc_purpose&0x02 should equal 0.

[0063] When ChromaFormatIdc or nnpfc_purpose&0x02 is not equal to 0, nnpfc_purpose&0x20 should be equal to 0.

[0064] The nnpfc_id contains an identifier that can be used to identify NNPF. The value of nnpfc_id should be in the range of 0 to 2^32-2 (inclusive). Values ​​of nnpfc_id from 256 to 511 (inclusive) and from 231 to 232-2 (inclusive) are reserved for future use by ITU-T|ISO / IEC. Decoders conforming to this version of this document should ignore NNPFC SEI messages when they encounter nnpfc_id in the range of 256 to 511 (inclusive) or in the range of 231 to 232-2 (inclusive).

[0065] When the NNPFC SEI message is the first NNPFC SEI message in the current CLVS with a specific nnpfc_id value in decoding order, the following applies:

[0066] - This SEI message specifies the basic NNPF.

[0067] - This SEI message applies to the currently decoded image and all subsequent decoded images of the current layer in the order of output, until the current CLVS ends.

[0068] An nnpfc_mode_idc value of 0 indicates that the SEI message contains an ISO / IEC 15938-17 bitstream, which specifies the basic NNPF or an update of the basic NNPF with the same nnpfc_id value.

[0069] When the NNPFC SEI message is the first NNPFC SEI message in the current CLVS with a specific nnpfc_id value in the decoding order, nnpfc_mode_idc equals 1, indicating that the basic NNPF associated with the nnpfc_id value is a neural network identified by the URI indicated by nnpfc_uri, where the URI has a format identified by the tag URI nnpfc_tag_uri.

[0070] When an NNPFC SEI message is neither the first NNPFC SEI message in the current CLVS with a specific nnpfc_id value in decoding order, nor a duplicate of the first NNPFC SEI message with that specific nnpfc_id value in decoding order, an nnpfc_mode_idc equal to 1 indicates that the update relative to the underlying NNPF with the same nnpfc_id value is defined by the URI indicated by nnpfc_uri, where the URI has a format identified by the tag URI nnpfc_tag_uri.

[0071] The value of nnpfc_mode_idc should be in the range of 0 to 1 (inclusive) in the bitstream conforming to this version of the document. Values ​​of nnpfc_mode_idc from 2 to 255 (inclusive) are reserved for future use by ITU-T|ISO / IEC and should not exist in the bitstream conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages with nnpfc_mode_idc in the range of 2 to 255 (inclusive). Values ​​of nnpfc_mode_idc greater than 255 should not exist in the bitstream conforming to this version of the document and are not reserved for future use.

[0072] When the SEI message is the first NNPFCSEI message in the current CLVS with a specific nnpfc_id value in the decoding order, NNPF PostProcessingFilter() is assigned the same as the base NNPF.

[0073] When the SEI message is neither the first NNPFC SEI message in the current CLVS with a specific nnpfc_id value in the decoding order, nor a duplicate of the first NNPFC SEI message with that specific nnpfc_id value in the decoding order, NNPFPostProcessingFilter() is obtained by applying the update of the SEI message definition to the base NNPF.

[0074] Updates are not cumulative; instead, each update is applied to the base NNPF, which is defined by the first NNPFC SEI message in the current CLVS with a specific nnpfc_id value in the decoding order.

[0075] nnpfc_reserved_zero_bit_a should be equal to 0 in the bitstream conforming to this version of the document. The decoder should ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_a is not equal to 0.

[0076] The nnpfc_tag_uri contains a tag URI, which has the syntax and semantics as specified in Internet Engineering Task Force (IETF) Request for Comment (RFC) 4151. The tag URI identifies the format and associated information about the neural network used as the base NNPF or relative to the base NNPF having the same nnpfc_id value specified by the nnpfc_uri.

[0077] Note 3 – nnpfc_tag_uri enables the unique identification of neural network data in the format specified by nnrpf_uri without the need for a central registry.

[0078] The nnpfc_tag_uri equal to “tag:iso.org,2023:15938-17” indicates that the neural network data identified by nnpfc_uri conforms to ISO / IEC 15938-17.

[0079] nnpfc_uri contains a URI with syntax and semantics as specified in IETF Internet Standard 66, which identifies a neural network used as a base NNPF or an update relative to a base NNPF with the same nnpfc_id value.

[0080] The value of nnpfc_property_present_flag equal to 1 indicates the existence of syntax elements related to the filter's purpose, input format, output format, and complexity. The value of nnpfc_property_present_flag equal to 0 indicates the absence of syntax elements related to the filter's purpose, input format, output format, and complexity.

[0081] When the SEI message is the first NNPFCSEI message in the current CLVS with a specific nnpfc_id value in the decoding order, nnpfc_property_present_flag should be equal to 1.

[0082] When nnpfc_property_present_flag equals 0, the values ​​of all syntax elements that can only exist when nnpfc_property_present_flag equals 1 and for which no presumed value is specified are presumed to be equal to the corresponding syntax element in the NNPFC SEI message containing the base NNPF (for which the SEI provides updates).

[0083] An nnpfc_base_flag value of 1 indicates that the SEI message specifies the basic NNPF. An nnpf_base_flag value of 0 indicates that the SEI message specifies an update relative to the basic NNPF. When it does not exist, the value of nnpfc_base_flag is presumed to be 0.

[0084] The following constraints apply to the value of nnpfc_base_flag:

[0085] - When the NNPFC SEI message is the first NNPFC SEI message in the current CLVS with a specific nnpfc_id value in the decoding order, the value of nnpfc_base_flag should be equal to 1.

[0086] - When NNPFC SEI message nnpfcB is not the first NNPFC SEI message in the current CLVS with a specific nnpfc_id value in the decoding order, and the value nnpfc_base_flag is equal to 1, the NNPFC SEI message should be a duplicate of the first NNPFC SEI message nnpfcA with the same nnpfc_id in the decoding order, that is, the payload content of nnpfcB should be the same as the payload content of nnpfcA.

[0087] When an NNPFC SEI message is not the first NNPFC SEI message with a specific nnpfc_id value in the current CLVS in decoding order, and is not a duplicate of the first NNPFC SEI message with that specific nnpfc_id, the following applies:

[0088] - This SEI message defines an update relative to the base NNPF with the same nnpfc_id value, ordered by the decoding sequence.

[0089] - This SEI message applies to the current decoded picture and all subsequent decoded pictures of the current layer in output order, until the end of the current CLVS, or until, but not including, the decoded picture that is in the current CLVS after the current decoded picture in output order and is associated with the subsequent NNPFC SEI message in the current CLVS with that particular nnpfc_id value in decoding order, whichever is earlier.

[0090] When the NNPFC SEI message nnpfcCurr is not the first NNPFC SEI message in the current CLVS with a specific nnpfc_id value in decoding order, is not a duplicate of the first NNPFC SEI message with that specific nnpfc_id (i.e., the value of nnpfc_base_flag is equal to 0), and the value of nnpfc_property_present_flag is equal to 1, the following constraints apply:

[0091] The value of nnpfc_purpose in the -NNPFC SEI message should be the same as the value of nnpfc_purpose in the first NNPFC SEI message with a specific nnpfc_id value in the current CLVS in the decoding order.

[0092] The value of the syntax element in the NNPFC SEI message that follows nnpfc_base_flag and precedes nnpfc_complexity_info_present_flag in the decoding order should be the same as the value of the corresponding syntax element in the first NNPFC SEI message in the current CLVS that has that particular nnpfc_id value in the decoding order.

[0093] - Either nnpfc_complexity_info_present_flag should be equal to 0, or nnpfc_complexity_info_present_flag should be equal to 1 in the first NNPFC SEI message (hereinafter referred to as nnpfcBase) in the current CLVS with that particular nnpfc_id value in the decoding order, and all of the following apply:

[0094] The nnpfc_parameter_parameter_type_idc in -nnpfcCurr should be equal to the nnpfc_parameter_parameter_type_idc in nnpfcBase.

[0095] The nnpfc_log2_parameter_bit_length_minus3 in -nnpfcCurr (if it exists) should be less than or equal to the nnpfc_log2_parameter_bit_length_minus3 in nnpfcBase.

[0096] - If nnpfc_num_parameters_idc in nnpfcBase is equal to 0, then nnpfc_num_parameters_idc in nnpfcCurr should be equal to 0.

[0097] Otherwise (nnpfc_num_parameters_idc in nnpfcBase is greater than 0), nnpfc_num_parameters_idc in nnpfcCurr should be greater than 0 and less than or equal to nnpfc_num_parameters_idc in nnpfcBase.

[0098] - If nnpfc_num_kmac_operations_idc in nnpfcBase is equal to 0, then nnpfc_num_kmac_operations_idc in nnpfcCurr should be equal to 0.

[0099] Otherwise (nnpfc_num_kmac_operations_idc in nnpfcBase is greater than 0), nnpfc_num_kmac_operations_idc in nnpfcCurr should be greater than 0 and less than or equal to nnpfc_num_kmac_operations_idc in nnpfcBase.

[0100] - If nnpfc_total_kilobyte_size in nnpfcBase is equal to 0, then nnpfc_total_kilobyte_size in nnpfcCurr should be equal to 0.

[0101] Otherwise (nnpfc_total_kilobyte_size in nnpfcBase is greater than 0), nnpfc_total_kilobyte_size in nnpfcCurr should be greater than 0 and less than or equal to nnpfc_total_kilobyte_size in nnpfcBase.

[0102] When `nnpfc_purpose&0x02` is not equal to 0, `nnpfc_out_sub_c_flag` specifies the values ​​of variables `outSubWidthC` and `outSubHeightC`. `nnpfc_out_sub_c_flag` equal to 1 specifies that `outSubWidthC` and `outSubHeightC` are both equal to 1. `nnpfc_out_sub_c_flag` equal to 0 specifies that `outSubWidthC` is equal to 2 and `outSubHeightC` is equal to 1. When `ChromaFormatIdc` is equal to 2 and `nnpfc_out_sub_c_flag` exists, the value of `nnpfc_out_sub_c_flag` should be equal to 1.

[0103] When `nnpfc_purpose&0x20` is not equal to 0, `nnpfc_out_colour_format_idc` specifies the color format of the NNPF output, thus defining the values ​​of the variables `outSubWidthC` and `outSubHeightC`. `nnpfc_out_colour_format_idc` equal to 1 specifies the NNPF output color format as 4:2:0, and both `outSubWidthC` and `outSubHeightC` are equal to 2. `nnpfc_out_colour_format_idc` equal to 2 specifies the NNPF output color format as 4:2:2, and both `outSubWidthC` and `outSubHeightC` are equal to 1. `nnpfc_out_colour_format_idc` equal to 3 specifies the NNPF output color format as 4:2:4, and both `outSubWidthC` and `outSubHeightC` are equal to 1. The value of `nnpfc_out_colour_format_idc` should not be equal to 0.

[0104] When both nnpfc_purpose&0x02 and nnpfc_purpose&0x20 are equal to 0, outSubWidthC and outSubHeightC are presumed to be equal to SubWidthC and SubHeightC, respectively.

[0105] `nnpfc_pic_width_in_luma_samples` and `nnpfc_pic_height_in_luma_samples` specify the width and height of the luminance sample array for the image, which is produced by applying the NNPF identifier `nnpfc_id` to the cropped decoded output image. When `nnpfc_pic_width_in_luma_samples` and `nnpfc_pic_height_in_luma_samples` are not present, they are presumed to be equal to `CroppedWidth` and `CroppedHeight`, respectively. The value of `nnpfc_pic_width_in_luma_samples` should be in the range from `CroppedWidth` to `CroppedWidth*16-1` (inclusive). The value of `nnpfc_pic_height_in_luma_samples` should be in the range from `CroppedHeight` to `CroppedHeight*16-1` (inclusive).

[0106] The increment of 1 in `nnpfc_num_input_pics_minus1` specifies the number of decoded output images used as input to NNPF. The value of `nnpfc_num_input_pics_minus1` should be in the range of 0 to 63 (inclusive). When `nnpfc_purpose&0x08` is not equal to 0, the value of `nnpfc_num_input_pics_minus1` should be greater than 0.

[0107] `nnpfc_interpolated_pics[i]` specifies the number of interpolated pictures generated by NNPF between the i-th picture and the (i+1)-th picture used as input to NNPF. The value of `nnpfc_interpolated_pics[i]` should be in the range of 0 to 63 (inclusive). For at least one i in the range of 0 to `nnpfc_num_input_pics_minus1-1` (inclusive), the value of `nnpfc_interpolated_pics[i]` should be greater than 0.

[0108] `nnpfc_input_pic_output_flag[i]` equal to 1 indicates that NNPF generates the corresponding output image for the i-th input image. `nnpfc_input_pic_output_flag[i]` equal to 0 indicates that NNPF does not generate the corresponding output image for the i-th input image.

[0109] The variables numInputPics, which specify the number of images used as input to NNPF, and numOutputPics, which specify the total number of images generated by NNPF, are derived as follows:

[0110]

[0111]

[0112] `nnpfc_component_last_flag` equal to 1 indicates that the last dimension of the NNPF input tensor and the output tensor generated by NNPF is used for the current channel. `nnpfc_component_last_flag` equal to 0 indicates that the third dimension of the NNPF input tensor and the output tensor generated by NNPF is used for the current channel.

[0113] Note 4 – The first dimension in both the input and output tensors is used for the batch index, which is a practice in some neural network frameworks. Although the formula in the semantics of this SEI message uses the batch size corresponding to a batch index of 0, the batch size used as the input for neural network inference is determined by the post-processing implementation.

[0114] Note 5 – For example, when nnpfc_inp_order_idc equals 3 and nnpfc_auxiliary_inp_idc equals 1, the input tensor has 7 channels, including four luminance matrices, two chrominance matrices, and one auxiliary input matrix. In this case, the procedure DeriveInputTensors() will derive each of these 7 channels of the input tensor one by one, and when processing a particular channel, that channel is referred to as the current channel during the procedure.

[0115] `nnpfc_inp_format_idc` specifies the method for converting the sample values ​​of the cropped decoded output image into NNPF input values. When `nnpfc_inp_format_idc` equals 0, the NNPF input values ​​are real numbers, and the functions `InpY()` and `InpC()` are defined as follows:

[0116] InpY(x)=x÷((1< <BitDepth Y )-1) (77)

[0117] InpC(x)=x÷((1< <BitDepth C )-1) (78)

[0118] When nnpfc_inp_format_idc equals 1, the input values ​​for NNPF are unsigned integers, and the functions InpY() and InpC() are defined as follows:

[0119] shiftY=BitDepthY-inpTensorBitDepthY

[0120] if(inpTensorBitDepthY>=BitDepthY)

[0121] InpY(x)=x<<(inpTensorBitDepthY-BitDepthY)(79)

[0122] else

[0123] InpY(x) = Clip3(0, (1 < <inpTensorBitDepthY)-1,(x+(1<<(shiftY-1)))> >shiftY)

[0124] shiftC=BitDepthC-inpTensorBitDepthC

[0125] if(inpTensorBitDepthC>=BitDepthC)

[0126] InpC(x)=x<<(inpTensorBitDepthC-BitDepthC)(80)

[0127] else

[0128] InpC(x) = Clip3(0, (1 < <inpTensorBitDepthC)-1,(x+(1<<(shiftC-1)))> >shiftC)

[0129] The variable inpTensorBitDepthY is deduced from the syntax element nnpfc_inp_tensor_luma_bitdepth_minus8 as specified below. The variable inpTensorBitDepthC is deduced from the syntax element nnpfc_inp_tensor_chroma_bitdepth_minus8 as specified below.

[0130] Values ​​greater than 1 for nnpfc_inp_format_idc are reserved for future ITU-T|ISO / IEC specifications and should not be present in bitstreams conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages containing reserved values ​​for nnpfc_inp_format_idc.

[0131] `nnpfc_inp_tensor_luma_bitdepth_minus8` plus 8 specifies the bit depth of the luminance sample values ​​in the input integer tensor. The value of `inpTensorBitDepthY` is derived as follows:

[0132] inpTensorBitDepth Y =nnpfc_inp_tensor_luma_bitdepth_minus8+8 (81)

[0133] The requirement for bitstream consistency is that the value of nnpfc_inp_tensor_luma_bitdepth_minus8 should be in the range of 0 to 24 (inclusive).

[0134] `nnpfc_inp_tensor_chroma_bitdepth_minus8` plus 8 specifies the bit depth of the chroma sample values ​​in the input integer tensor. The value of `inpTensorBitDepthC` is derived as follows:

[0135] inpTensorBitDepth C =nnpfc_inp_tensor_chroma_bitdepth_minus8+8 (82)

[0136] The requirement for bitstream consistency is that the value of nnpfc_inp_tensor_chroma_bitdepth_minus8 should be in the range of 0 to 24 (inclusive).

[0137] nnpfc_inp_order_idc indicates a method for sorting the sample arrays of the cropped decoded output image into one of the input images for NNPF.

[0138] The value of nnpfc_inp_order_idc should be in the range of 0 to 3 (inclusive) in the bitstream conforming to this version of the document. Values ​​of 4 to 255 (inclusive) for nnpfc_inp_order_idc are reserved for future use by ITU-T|ISO / IEC and should not exist in the bitstream conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 4 to 255 (inclusive). Values ​​of nnpfc_inp_order_idc greater than 255 should not exist in the bitstream conforming to this version of the document and are not reserved for future use.

[0139] When ChromaFormatIdc is not equal to 1, nnpfc_inp_order_idc should not be equal to 3.

[0140] Table 21 contains an informative description of the nnpfc_inp_order_idc value.

[0141] Table 21 - Description of nnpfc_inp_order_idc values

[0142]

[0143]

[0144] Figure 1 An example of deriving the luminance channel from the luminance component is shown.

[0145] Small blocks are rectangular arrays of samples from the components of an image (e.g., luminance or chrominance components).

[0146] A value greater than 0 for nnpfc_auxiliary_inp_idc indicates that the auxiliary input data exists in the input tensor of NNPF. A value equal to 0 for nnpfc_auxiliary_inp_idc indicates that the auxiliary input data does not exist in the input tensor. A value equal to 1 for nnpfc_auxiliary_inp_idc indicates that the auxiliary input data is derived according to Equation 84.

[0147] The value of nnpfc_auxiliary_inp_idc should be in the range of 0 to 1 (inclusive) in the bitstream conforming to this version of the document. For nnpfc_inp_order_idc, values ​​from 2 to 255 (inclusive) are reserved for future use by ITU-T|ISO / IEC and should not exist in the bitstream conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 2 to 255 (inclusive). Values ​​of nnpfc_inp_order_idc greater than 255 should not exist in the bitstream conforming to this version of the document and are not reserved for future use.

[0148] When nnpfc_auxiliary_inp_idc equals 1, the variable strengthControlScaledVal is derived as follows:

[0149] if(nnpfc_inp_format_idc==1)

[0150] strengthControlScaledVal=Floor(StrengthControlVal*((1< <inpTensorBitDepthY)-1))(83)

[0151] else

[0152] strengthControlScaledVal=StrengthControlVal

[0153] The procedure DeriveInputTensors() is used to derive an input tensor inputTensor given vertical sample coordinates cTop and horizontal sample coordinates cLeft. These vertical and horizontal coordinates specify the top-left position of the small block of samples included in the input tensor. The procedure DeriveInputTensors() is defined as follows:

[0154]

[0155]

[0156]

[0157]

[0158]

[0159] `nnpfc_separate_colour_description_present_flag` equal to 1 indicates that the SEI message syntax structure specifies different combinations of color primaries, transmission characteristics, and matrix coefficients for the image generated by NNPF. `nnpfc_separate_colour_description_present_flag` equal to 0 indicates that the combination of color primaries, transmission characteristics, and matrix coefficients for the image generated by NNPF is the same as the combination indicated in the CLVS VUI parameters.

[0160] nnpfc_colour_primaries has the same semantics as the vui_colour_primaries syntax element specified in sub-clause 7.3, except as follows:

[0161] –nnpfc_colour_primaries specifies the primary color of the image generated by NNPF as defined in the application SEI message, instead of the primary color used for CLVS.

[0162] – When nnpfc_colour_primaries does not exist in the NNPFC SEI message, the value of nnpfc_colour_primaries is presumed to be equal to vui_colour_primaries.

[0163] nnpfc_transfer_characteristics has the same semantics as the vui_transfer_characteristics syntax element specified in sub-clause 7.3, except as follows:

[0164] –nnpfc_transfer_characteristics specifies the transfer characteristics of images generated by NNPF as defined in the SEI message, rather than the transfer characteristics used for CLVS.

[0165] – When nnpfc_transfer_characteristics does not exist in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is presumed to be equal to vui_transfer_characteristics.

[0166] nnpfc_matrix_coeffs has the same semantics as the vui_matrix_coeffs syntax element specified in sub-clause 7.3, except as follows:

[0167] – The nnpfc_matrix_coeffs specifies the matrix coefficients of the picture generated by the NNPF specified in the applied SEI message, rather than the matrix coefficients for CLVS.

[0168] – When the nnpfc_matrix_coeffs does not exist in the NNPF SEI message, the value of the nnpfc_matrix_coeffs is presumed to be equal to the vui_matrix_coeffs.

[0169] – The allowed values of the nnpfc_matrix_coeffs are not constrained by the chroma format of the decoded video picture indicated by the value of ChromaFormatIdc of the semantics of the VUI parameters.

[0170] – When the nnpfc_matrix_coeffs is equal to 0, the nnpfc_out_order_idc shall not be equal to 1 or 3.

[0171] The nnpfc_out_format_idc equal to 0 indicates that the sample values output by the NNPF are real numbers, where the value range from 0 to 1 (including the end values) is linearly mapped to the unsigned integer value range from 0 to (1<<bitDepth)–1 (including the end values) for any desired bit depth bitDepth for subsequent post-processing or display.

[0172] The nnpfc_out_format_idc equal to 1 indicates that the luma sample values output by the NNPF are unsigned integers within the range from 0 to (1<<(nnpfc_out_tensor_luma_bitdepth_minus8+8))-1 (including the end values), and the chroma sample values output by the NNPF are unsigned integers within the range from 0 to (1<<(nnpfc_out_tensor_chroma_bitdepth_minus8+8))-1 (including the end values).

[0173] Values of nnpfc_out_format_idc greater than 1 are reserved for future ITU-T|ISO / IEC specifications and shall not be present in the bitstream conforming to this version of this document. A decoder conforming to this version of this document shall ignore the NNPF SEI message containing a reserved value of nnpfc_out_format_idc.

[0174] The value of nnpfc_out_tensor_luma_bitdepth_minus8 plus 8 specifies the bit depth of the luminance sample values ​​in the output integer tensor. The value of nnpfc_out_tensor_luma_bitdepth_minus8 should be in the range of 0 to 24 (inclusive).

[0175] The value of nnpfc_out_tensor_chroma_bitdepth_minus8 plus 8 specifies the bit depth of the chroma sample values ​​in the output integer tensor. The value of nnpfc_out_tensor_chroma_bitdepth_minus8 should be in the range of 0 to 24 (inclusive).

[0176] When nnpfc_purpose&0x10 is not equal to 0, the value of nnpfc_out_format_idc should be equal to 1, and at least one of the following conditions should be true:

[0177] –nnpfc_out_tensor_luma_bitdepth_minus8+8 is greater than BitDepth Y .

[0178] –nnpfc_out_tensor_chroma_bitdepth_minus8+8 is greater than BitDepth C .

[0179] nnpfc_out_order_idc indicates the output order of samples generated by NNPF.

[0180] The value of nnpfc_out_order_idc should be in the range of 0 to 3 (inclusive) in the bitstream conforming to this version of the document. Values ​​of 4 to 255 (inclusive) for nnpfc_out_order_idc are reserved for future use by ITU-T|ISO / IEC and should not exist in the bitstream conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages with nnpfc_out_order_idc in the range of 4 to 255 (inclusive). Values ​​of nnpfc_out_order_idc greater than 255 should not exist in the bitstream conforming to this version of the document and are not reserved for future use.

[0181] When nnpfc_purpose&0x02 is not equal to 0, nnpfc_out_order_idc should not be equal to 3.

[0182] Table 22 contains an informative description of the nnpfc_out_order_idc value.

[0183] Table 22 - Description of nnpfc_out_order_idc values

[0184]

[0185]

[0186] The procedure StoreOutputTensors() is used to derive the sample values ​​from the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the given vertical sample coordinates cTop and horizontal sample coordinates cLeft of the output tensor outputTensor. The vertical sample coordinates cTop and horizontal sample coordinates cLeft specify the top-left sample position of the small block of samples included in the input tensor. The procedure StoreOutputTensors() is specified as follows:

[0187]

[0188]

[0189]

[0190] nnpfc_overlap indicates the horizontal and vertical sample counts of the overlap between adjacent input tensors in NNPF. The value of nnpfc_overlap should be in the range of 0 to 16,383 (inclusive).

[0191] The nnpfc_constant_patch_size_flag being equal to 1 indicates that NNPF fully accepts the patch sizes indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1 as input. The nnpfc_constant_patch_size_flag setting being 0 indicates that NNPF accepts any patch size with a width inpPatchWidth and a height inpPatchHeight as input, such that the width of the extended patch (i.e., the patch plus the overlapping area) (which is equal to inpPatchWidth + 2 * nnpfc_overlap) is a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, and the height of the extended patch (which is equal to inpPatchHeight + 2 * nnpfc_overlap) is a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap.

[0192] Incrementing 1 by nnpfc_patch_width_minus1 (when nnpfc_constant_patch_size_flag equals 1) indicates the horizontal sample count for the required patch size of the NNPF input. The value of nnpfc_patch_width_minus1 should be in the range of 0 to Min(32766, CroppedWidth-1) (inclusive).

[0193] Incrementing 1 by 1 (when nnpfc_patch_height_minus1 equals 1) indicates the vertical sample count for the patch size required for NNPF inputs. The value of nnpfc_patch_height_minus1 should be in the range of 0 to Min(32,766, CroppedHeight-1) (inclusive).

[0194] `nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap` (when `nnpfc_constant_patch_size_flag` equals 0) indicates the common divisor of all allowed values ​​for the width of the extended patch required for the NNPF input. The value of `nnpfc_extended_patch_width_cd_delta_minus1` should be in the range of 0 to Min(32766, CroppedWidth-1) (inclusive).

[0195] `nnpfc_extended_patch_height_cd_delta_minus1` plus 1 plus 2 * `nnpfc_overlap` (when `nnpfc_constant_patch_size_flag` equals 0) indicates the common divisor of all allowed values ​​for the height of the extended patch required for the NNPF input. The value of `nnpfc_extended_patch_height_cd_delta_minus1` should be in the range of 0 to Min(32766, CroppedHeight-1) (inclusive).

[0196] Let the variables inpPatchWidth and inpPatchHeight be the width and height of the small block, respectively.

[0197] If nnpfc_constant_patch_size_flag equals 0, then the following applies:

[0198] The values ​​of -inpPatchWidth and inpPatchHeight are provided by external means not specified in this document or set by the post-processing itself.

[0199] The value of -inpPatchWidth+2*nnpfc_overlap should be a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1+1+2*nnpfc_overlap, and inpPatchWidth should be less than or equal to CroppedWidth. The value of inpPatchHeight+2*nnpfc_overlap should be a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1+1+2*nnpfc_overlap, and inpPatchHeight should be less than or equal to CroppedHeight.

[0200] Otherwise (nnpfc_constant_patch_size_flag equals 1), the value of inpPatchWidth is set to equal to nnpfc_patch_width_minus1+1, and the value of inpPatchHeight is set to equal to nnpfc_patch_height_minus1+1.

[0201] The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, and outPatchCHeight are derived as follows:

[0202] outPatchWidth=(nnpfc_pic_width_in_luma_samples*inpPatchWidth) / CroppedWidth (86)

[0203] outPatchHeight=(nnpfc_pic_height_in_luma_samples*inpPatchHeight) / CroppedHeight (87)

[0204] horCScaling=SubWidthC / outSubWidthC (88)

[0205] verCScaling=SubHeightC / outSubHeightC (89)

[0206] outPatchCWidth=outPatchWidth*horCScaling (90)

[0207] outPatchCHeight=outPatchHeight*verCScaling (91)

[0208] The requirement for bitstream consistency is that outPatchWidth * CroppedWidth should be equal to nnpfc_pic_width_in_luma_samples * inpPatchWidth, and outPatchHeight * CroppedHeight should be equal to nnpfc_pic_height_in_luma_samples * inpPatchHeight.

[0209] The nnpfc_padding_type indicates the padding process when referencing sample locations outside the boundaries of the cropped decoded output image, as described in Table 23. The value of nnpfc_padding_type should be in the range of 0 to 15 (inclusive).

[0210] Table 23 - Informative Description of nnpfc_padding_type Values

[0211] nnpfc_padding_type Description 0 Zero padding 1 Copy padding 2 Reflective padding 3 Wrap padding 4 Fixed padding 5..15 Preserve

[0212] nnpfc_luma_padding_val indicates the luminance value to be used for padding when nnpfc_padding_type is equal to 4.

[0213] nnpfc_cb_padding_val indicates the Cb value to be used for padding when nnpfc_padding_type is equal to 4.

[0214] nnpfc_cr_padding_val indicates the Cr value to be used for padding when nnpfc_padding_type is equal to 4.

[0215] The function InpSampleVal(y,x,picHeight,picWidth,croppedPic) (where the inputs are the vertical sample position y, the horizontal sample position x, the image height picHeight, the image width picWidth, and the sample array croppedPic) returns the value of sampleVal derived as follows:

[0216] Note 6 – For the input of the function InpSampleVal(), the vertical position is listed before the horizontal position to be compatible with the input tensor conventions of some inference engines.

[0217]

[0218]

[0219] The following example procedure can be used with NNPF PostProcessingFilter() to generate (multiple) filtered and / or interpolated images in small chunks, containing Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated by nnpfc_out_order_idc:

[0220]

[0221]

[0222] The order of the images in the stored output tensor is the output order, and the output order generated by applying NNPF in the output order is interpreted as the output order (and does not conflict with the output order of the input images).

[0223] A value of 1 for nnpfc_complexity_info_present_flag indicates the existence of one or more syntax elements that indicate the complexity of the NNPF associated with nnpfc_id. A value of 0 for nnpfc_complexity_info_present_flag indicates the absence of a syntax element that indicates the complexity of the NNPF associated with nnpfc_id.

[0224] An `nnpfc_parameter_type_idc` value of 0 indicates that the neural network uses only integer parameters. An `nnpfc_parameter_type_flag` value of 1 indicates that the neural network can use either floating-point or integer parameters. An `nnpfc_parameter_type_idc` value of 2 indicates that the neural network uses only binary parameters. An `nnpfc_parameter_type_idc` value of 3 is reserved for future use by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages with `nnpfc_parameter_type_idc` equal to 3.

[0225] The values ​​0, 1, 2, and 3 for nnpfc_log2_parameter_bit_length_minus3 indicate that the neural network does not use parameters with bit lengths greater than 8, 16, 32, and 64, respectively. When nnpfc_parameter_type_idc exists and nnpfc_log2_parameter_bit_length_minus3 does not exist, the neural network does not use parameters with a bit length greater than 1.

[0226] `nnpfc_num_parameters_idc` indicates the maximum number of neural network parameters for NNPF, in powers of 2048. `nnpfc_num_parameters_idc` equal to 0 indicates that the maximum number of neural network parameters is unknown. The value of `nnpfc_num_parameters_idc` should be in the range of 0 to 52 (inclusive). Values ​​of `nnpfc_num_parameters_idc` greater than 52 are reserved for future use by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this version of this document. Decoders conforming to this version of this document should ignore NNPFCSEI messages with `nnpfc_num_parameters_idc` greater than 52.

[0227] If the value of nnpfc_num_parameters_idc is greater than 0, then the variable maxNumParameters is deduced as follows:

[0228] maxNumParameters = (2 048 < <nnpfc_num_parameters_idc)-1 (94)

[0229] The requirement for bitstream consistency is that the number of neural network parameters in NNPF should be less than or equal to maxNumParameters.

[0230] A value greater than 0 for nnpfc_num_kmac_operations_idc indicates that the maximum number of multiply-accumulate operations per sample in NNPF is less than or equal to nnpfc_num_kmac_operations_idc * 1000. A value of 0 for nnpfc_num_kmac_operations_idc indicates that the maximum number of multiply-accumulate operations in the network is unknown. The value of nnpfc_num_kmac_operations_idc should be in the range of 0 to 2^32 - 2 (inclusive).

[0231] `nnpfc_total_kilobyte_size` greater than 0 indicates the total size in kilobytes required to store the uncompressed parameters of the neural network. The total size in bits is equal to or greater than the sum of the bits used to store each parameter. `nnpfc_total_kilobyte_size` is the total size in bits divided by 8000 and rounded down. `nnpfc_total_kilobyte_size` equal to 0 indicates that the total size required to store the parameters of the neural network is unknown. The value of `nnpfc_total_kilobyte_size` should be in the range of 0 to 2^32 - 2 (inclusive).

[0232] nnpfc_reserved_zero_bit_b should be equal to 0 in the bitstream conforming to this version of the document. The decoder should ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_b is not equal to 0.

[0233] nnpfc_payload_byte[i] contains the i-th byte of a bitstream conforming to ISO / IEC 15938-17. The sequence of bytes nnpfc_payload_byte[i] for all existing values ​​i should be a complete bitstream conforming to ISO / IEC 15938-17.

[0234] 8.29 Neural Network Post-Processing Filter Activation SEI Message

[0235] 8.29.1 Neural Network Post-Processing Filter Activation SEI Message Syntax

[0236]

[0237] 8.29.2 Neural Network Post-Processing Filter Activation of SEI Message Semantics

[0238] The Neural Network Post-Processing Filter Activation (NNPFA) SEI message activates or deactivates the target Neural Network Post-Processing Filter (NNPF), identified by nnpfa_target_id, for a set of images. For a specific image where the NNPF is activated, the target NNPF is defined by the last NNPFC SEI message whose nnpfc_id equals nnpfa_target_id. This last NNPFC SEI message precedes the first VCLNAL unit in the current image according to the decoding order and is not a repetition of the NNPFC SEI message containing the basic NNPF.

[0239] Note 1 – Multiple NNPFA SEI messages can exist for the same image, for example, when NNPF is used for different purposes or for filtering different color components.

[0240] nnpfa_target_id indicates the target NNPF, which is specified by one or more NNPFC SEI messages that involve the current picture and have an nnpfc_id equal to nnpfa_target_id.

[0241] The value of nnpfa_target_id should be in the range of 0 to 232-2 (inclusive). Values ​​of nnpfa_target_id from 256 to 511 (inclusive) and from 231 to 232-2 (inclusive) are reserved for future use by ITU-T|ISO / IEC. Decoders conforming to this document should ignore NNPFASEI messages when they encounter nnpfa_target_id in the range of 256 to 511 (inclusive) or 231 to 232-2 (inclusive).

[0242] NNPFA SEI messages with a specific value of nnpfa_target_id should not exist in the current PU unless one or both of the following conditions are true:

[0243] - Within the current CLVS, there are NNPFC SEI messages in the PUs preceding the current PU, in the decoding order, where nnpfc_id is equal to a specific value of nnpfa_target_id.

[0244] - There is an NNPFC SEI message in the current PU with a specific value of nnpfc_id equal to nnpfa_target_id.

[0245] When a PU contains both an NNPFC SEI message with a specific value of nnpfc_id and an NNPFA SEI message with nnpfa_target_id equal to a specific value of nnpfc_id, the NNPFC SEI message should precede the NNPFA SEI message in the decoding order.

[0246] A `nnpfa_cancel_flag` of 1 indicates that the persistence of the target NNPF established by any previous NNPFA SEI message with the same `nnpfa_target_id` as the current SEI message is cancelled; that is, the target NNPF is no longer used unless it is activated by another NNPFASEI message with the same `nnpfa_target_id` as the current SEI message and `nnpfa_cancel_flag` equal to 0. A `nnpfa_cancel_flag` of 0 indicates that `nnpfa_persistence_flag` follows.

[0247] The nnpfa_persistence_flag specifies the persistence of the target NNPF for the current layer.

[0248] Setting nnpfa_persistence_flag to 0 specifies that the target NNPF is used only for post-processing filtering of the current image.

[0249] Setting nnpfa_persistence_flag to 1 specifies that the target NNPF can be used for post-processing filtering of the current image and all subsequent images of the current layer in output order, until one or more of the following conditions are true:

[0250] - A new CLVS begins for the current layer.

[0251] -End of bitstream.

[0252] - The images associated with the NNPFA SEI message in the current layer that have the same nnpfa_target_id as the current SEI message and an nnpfa_cancel_flag equal to 1 are output after the current image in the output order.

[0253] Note 2 – Do not apply the target NNPF to the subsequent picture in the current layer that is associated with an NNPFA SEI message that has the same nnpfa_target_id as the current SEI message and an nnpfa_cancel_flag equal to 1.

[0254] Make nnpfcTargetPictures the set of pictures involved in the last NNPFA SEI message that precedes the current NNPFA SEI message in decoding order, and whose nnpfc_id is equal to nnpfa_target_id. Make nnpfaTargetPictures the set of pictures that activated the target NNPF through the current NNPFA SEI message. For bitstream consistency, any pictures included in nnpfaTargetPictures should also be included in nnpfcTargetPictures.

[0255] 3.4 Use of NNPFC SEI messages in VVC bitstream

[0256] JVET-AC2005[6] includes the specification for the use of NNPFC SEI messages in VVC bitstreams, as follows:

[0257] D.12.11 Use of SEI Messages in Neural Network Post-Processing Filter Characteristics

[0258] Make currCodedPic a encoded image. For this encoded image, the Neural Network Post-Processing Filter (NNPF) defined by the Neural Network Post-Processing Filter Feature (NNPFC) SEI message is activated by the Neural Network Post-Processing Filter Activation (NNPFA) SEI message.

[0259] The variable pictureRateUpsamplingFlag is set to equal to (nnpfc_purpose&0x08)! = 0.

[0260] The variable numInputPics is set to equal nnpfc_num_input_pics_minus1+1.

[0261] For all values ​​of i in the range from 0 to numInputPics-1 (inclusive), the array inputPicPoc[i] (which specifies the image order count for the input images in NNPF) is derived as follows:

[0262] -inputPicPoc[0] is set to equal PicOrderCntVal of currCodedPic.

[0263] - When numInputPics is greater than 1, for each value of i in the range from 1 to numInputPics-1 (inclusive), the following applies in ascending order of i:

[0264] - If currCodedPic is associated with a Frame Encapsulation Arrangement SEI message having a specific value of fp_arrangement_type equal to 5 and fp_current_frame_is_frame0_flag equal to 5, then inputPicPoc[i] is set to equal to the PicOrderCntVal of the image that precedes the image associated with index i-1 in the output order and is associated with a Frame Encapsulation Arrangement SEI message having the same value of fp_arrangement_type equal to 5 and fp_current_frame_is_frame0_flag equal to 5.

[0265] - Otherwise (currCodedPic is not associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5), inputPicPoc[i] is set to equal to the PicOrderCntVal of the image preceding the image associated with index i-1 in the output order.

[0266] To interpret NNPFC SEI messages, the following variables are defined:

[0267] - If pictureRateUpsamplingFlag equals 1 and there exists a second NNPF defined by at least one NNPFC SEI message, activated for currCodedPic via an NNPFA SEI message, and nnpfc_purpose equals 4, then the following applies:

[0268] -CroppedWidth is set to be equal to nnpfc_pic_width_in_luma_samples as defined for the second NNPF.

[0269] -CroppedHeight is set to be equal to nnpfc_pic_height_in_luma_samples as defined for the second NNPF.

[0270] -Otherwise, the following applies:

[0271] - For currCodedPic, CroppedWidth is set to equal to pps_pic_width_in_luma_samples

[0272] The value of -SubWidthC*(pps_conf_win_left_offset+pps_conf_win_right_offset).

[0273] - For currCodedPic, CroppedHeight is set to equal to pps_pic_height_in_luma_samples

[0274] -SubHeightC*(pps_conf_win_top_offset+pps_conf_win_bottom_offset) value.

[0275] - For each value of i in the range of 0 to numInputPics-1 (inclusive), the luminance sample array CroppedYPic[i] and the chrominance sample arrays CroppedCbPic[i] and CroppedCrPic[i] (if they exist) are derived as follows:

[0276] - Make sourcePic the cropped decoded output image of PicOrderCntVal, which is equal to inputPicPoc[i], in the CLVS containing currCodedPic.

[0277] - If pictureRateUpsamplingFlag equals 0, the following applies:

[0278] The luminance sample array CroppedYPic[i] and the chrominance sample arrays CroppedCbPic[i] and CroppedCrPic[i] (if present) are set as two-dimensional arrays of the decoded sample values ​​of the Y, Cb, and Cr components of sourcePic, respectively.

[0279] - Otherwise (pictureRateUpsamplingFlag equals 1), the following applies:

[0280] For sourcePic, the variable sourceWidth is set to the value equal to pps_pic_width_in_luma_samples-SubWidthC*(pps_conf_win_left_offset+pps_conf_win_right_offset).

[0281] - For sourcePic, the variable sourceHeight is set to the value of pps_pic_height_in_luma_samples-SubHeightC*(pps_conf_win_top_offset+pps_conf_win_bottom_offset).

[0282] - If sourceWidth equals CroppedWidth and sourceHeight equals CroppedHeight, then inputPic is set to be the same as sourcePic.

[0283] - Otherwise (sourceWidth is not equal to CroppedWidth or sourceHeight is not equal to CroppedHeight), the following applies:

[0284] - There should exist an NNPF, hereinafter referred to as the super-resolution NNPF, which is defined by at least one NNPFC SEI message, activated for the sourcePic by an NNPFA SEI message, and has nnpfc_purpose equal to 4, nnpfc_pic_width_in_luma_samples equal to CroppedWidth, and nnpfc_pic_height_in_luma_samples equal to CroppedHeight.

[0285] -inputPic is set to the output of the neural network inference of the super-resolution NNPF with sourcePic as input.

[0286] - The luminance sample array CroppedYPic[i] and the chrominance sample arrays CroppedCbPic[i] and CroppedCrPic[i] (if present) are set as two-dimensional arrays of the decoded sample values ​​of the Y, Cb and Cr components of the inputPic, respectively.

[0287] Both BitDepthY and BitDepthC are set to equal to BitDepth.

[0288] -ChromaFormatIdc is set to equal sps_chroma_format_idc.

[0289] -StrengthControlVal is set to the value of SliceQpY ÷ 63 of the first stripe of currCodedPic.

[0290] No more than one NNPFC SEI message with the same nnpfc_id value should exist in a single image unit. If two NNPFC SEI messages with the same nnpfc_id value exist in a single image unit, these SEI messages should have different content. If two NNPFC SEI messages with the same nnpfc_id but different content exist in the same image unit, these two NNPFC SEI messages should be in the same SEI NAL unit.

[0291] 4. The technical problem solved by the disclosed technical solution

[0292] The example design of the Neural Network Post-Processing Filter Feature (NNPFC) SEI message and the Neural Network Post-Processing Filter Activation (NNPFA) SEI message, as well as the use of the NNPFC SEI message in the VVC bitstream, have the following problems:

[0293] First, when the purpose of NNPF is only image rate upsampling, in addition to some interpolated images, NNPF may output images corresponding to one or more input images.

[0294] Second, when the purpose of NNPF includes image rate upsampling, regardless of whether there are other NNPF purposes or other types of upsampling, NNPF may output images corresponding to a specific input image multiple times.

[0295] Third, when the NNPF objective includes image rate upsampling with one or more of the following NNPF objectives or upsampling types: 1) general visual quality improvement, 2) chroma upsampling, 3) resolution upsampling, 4) bit depth upsampling, 5) colorization, and 6) depth image generation, NNPF may not output an image corresponding to any of the input images. However, such a combination is meaningless.

[0296] Fourth, the purpose of “having general visible quality improvement” and / or “not having general visible quality improvement” is not clearly defined.

[0297] Fifth, when the NNPF objective includes image rate upsampling, regardless of whether other NNPF objectives or other types of upsampling are present, NNPF may output inconsistent images between two neighboring input images.

[0298] 5. List of solutions and implementation examples

[0299] To address the problems described above, methods outlined below are disclosed. These aspects should be considered as examples for interpreting general concepts, and not interpreted in a narrow sense. Furthermore, these examples can be applied individually or in any combination.

[0300] 1) To address problem 1, it is stipulated that when the purpose of NNPF is only image rate upsampling, NNPF should not output an image corresponding to any input image. In other words, when the purpose of NNPF is only image rate upsampling, all flags nnpfc_input_pic_output_flag[i] should be equal to 0.

[0301] a. In one example, when the purpose of NNPF includes image rate upsampling, the signaling syntax element nnpfc_input_pic_output_enable_flag is used to specify whether NNPF outputs an image corresponding to one or more input images, in addition to some interpolated images.

[0302] i. In one example, when nnpfc_input_pic_output_enable_flag equals 1, the flag nnpfc_input_pic_output_flag[i] is signaled.

[0303] ii. In one example, when the syntax element nnpfc_input_pic_output_enable_flag is equal to 0, the flags nnpfc_input_pic_output_flag[i] are not transmitted via signaling, and their values ​​are presumed to be equal to 0.

[0304] b. In one example, when the NNPF objective only involves image rate upsampling, the nnpfc_input_pic_output_enable_flag is not transmitted via semaphore, and its value is presumed to be equal to 0.

[0305] 2) To solve problem 2, one or more of the following aspects are specified:

[0306] a. In one example, it is specified that (multiple) NNPFs are allowed to output images multiple times for a specific input image.

[0307] i. In one example, additionally, when (multiple) NNPFs output images multiple times for a particular input image, the final NNPF output image for that particular input image is deduced to be the image last output by any NNPF for that particular input image.

[0308] ii. In one example, additionally, when (multiple) NNPFs output images multiple times for a particular input image, the final NNPF output image corresponding to the particular input image is deduced to be the image first output by any NNPF for that particular input image.

[0309] iii. In one example, additionally, when (multiple) NNPFs output images multiple times for a particular input image, the final NNPF output image for that particular input image is derived as the average of all images output by (multiple) NNPFs for that particular input image.

[0310] b. In one example, it is specified that (multiple) NNPFs are not allowed to output the image multiple times for any given input image. i. In one example, additionally, it is required that (multiple) NNPFs output the image corresponding to each input image exactly once.

[0311] ii. In one example, additionally, it is required that NNPF can output an image corresponding to a specific input image only if the corresponding output image was not generated by (multiple) previous NNPFs.

[0312] iii. In one example, additionally, it is required that NNPF can output an image corresponding to a specific input image only if the corresponding output image has not been previously generated by any NNPF.

[0313] iv. In one example, additionally, it is required that NNPF should not output an image corresponding to a specific input image when (multiple) previous NNPFs have already output images corresponding to a specific input image.

[0314] v. In one example, additionally, it is required that NNPF should not output the image corresponding to the specific input image if the image corresponding to the specific input image has already been output by any NNPF.

[0315] c. In one example, it is specified that (multiple) NNPFs are allowed to output images multiple times for a specific output order or output time instance.

[0316] i. In one example, additionally, when (multiple) NNPFs output images multiple times for a particular output order or output time instance, the final NNPF output image for that particular output order or output time instance is deduced to be the image last output by any NNPF for that particular output order or output time instance.

[0317] ii. In one example, additionally, when (multiple) NNPFs output images multiple times for a particular output order or output time instance, the final NNPF output image for the particular output order or output time instance is derived as the image first output by NNPF for the particular output order or output time instance.

[0318] iii. In one example, additionally, when (multiple) NNPFs output images multiple times for a particular output order or output time instance, the final NNPF output image for the particular output order or output time instance is derived as the average of all images output by (multiple) NNPFs for the particular output order or output time instance.

[0319] d. In one example, it is specified that (multiple) NNPFs are not allowed to output images multiple times for any particular output order or output time instance.

[0320] i. In one example, additionally, it is required that (multiple) NNPF instances output the image once and only once for any given output order or output time instance.

[0321] 3) To address problem 3, when the NNPF objective includes image rate upsampling with one or more of the following NNPF objectives or upsampling types: 1) general visual quality improvement, 2) chroma upsampling, 3) resolution upsampling, 4) bit depth upsampling, 5) colorization, and 6) depth image generation, one or more of the following aspects are specified:

[0322] a. NNPF should output an image corresponding to at least one input image. In other words, at least one of the flags nnpfc_input_pic_output_flag[i] should be equal to 1.

[0323] b. NNPF should output an image corresponding to all input images. In other words, all flags nnpfc_input_pic_output_flag[i] should be equal to 1.

[0324] 4) To address problem 4, one or more of the following aspects are specified:

[0325] a. In one example, it is specified that the purpose corresponding to "with general visual quality improvement" should not involve any format changes, including resolution, image rate, color format, bit depth, etc.

[0326] i. In one example, additionally, it is not permitted to combine the purpose of "having general visual quality improvement" with any other purpose that results in a change in format.

[0327] b. In one example, it is specified that for the purpose of “no general visual quality improvement”, there should be some format changes, including resolution, image rate, color format, bit depth, etc.

[0328] i. In one example, additionally, it is not permitted to use the purpose of "no general visual quality improvement" alone without combining it with any other purpose that results in a change in format.

[0329] 5) To address problem 5, one or more of the following aspects are specified:

[0330] a. When the purpose of NNPF includes image rate upsampling, it is required that for any two distinct activations of NNPF, at most one image should be used as the input image for both activations.

[0331] When one or more NNPFs have the purpose of including image rate upsampling, it is required that for any two distinct activations of these NNPFs, at most one image should be used as the input image for both activations.

[0332] 6. References

[0333] [1]ITU-T and ISO / IEC, “Efficient video coding and decoding”, Recommendation ITU-T H.265|ISO / IEC 23008-2 (current version).

[0334] [2] J. Chen, E. Alshina, G. J. Sullivan, J.-R. Ohm, J. Boyce, “Algorithmic description of Joint Exploration Test Model 7 (JEM7)”, JVET-G1001, August 2017.

[0335] [3] Recommendation ITU-T H.266|ISO / IEC 23090-3, “Multi-functional video coding and decoding”, 2022.

[0336] [4] ITU-T Recommendation H.274 | ISO / IEC 23002-7, “Multifunctional supplementary enhancement information messages for encoded and decoded video bitstreams”, 2022.

[0337] [5] S. McCarthy, S. Deshpande, M. Hannuksela, Hendry, G. Sullivan and Y.-K. Wang (eds.), “Improvement considerations for SEI messages in neural network post-processing filters”, JVET output document JVET-AC2032, is publicly available online at: https: / / jvet-experts.org / doc_end_user / current_document.php?id=12585.

[0338] [6] E. Francois, B. Bross, M.M. Hannuksela, A.M. Tourapis and Y.-K. Wang (eds.), “New Levels and System-Related Supplemental Enhancements for VVC (Draft 4),” JVET Output Document JVET-AC2005, available online at: https: / / www.jvet-experts.org / doc_end_user / current_document.php?id=12574.

[0339] Figure 2 This is a block diagram illustrating an example video processing system 4000 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0340] System 4000 may include an encoding / decoding component 4004 capable of implementing the various encoding / decoding or coding methods described in this document. Encoding / decoding component 4004 can reduce the average bit rate from the video input 4002 to the output of encoding / decoding component 4004 to produce an encoded / decoded representation of the video. Encoding / decoding techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of encoding / decoding component 4004 may be stored or transmitted via a communication connection such as that represented by component 4006. The bitstream (or encoded / decoded) representation of the video received at input 4002, whether stored or communicated, may be used by component 4008 to generate pixel values ​​or displayable video that is sent to display interface 4010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “encoding / decoding” operations or tools, it is understood that encoding / decoding tools or operations are used by the encoder, and the corresponding decoding tools or operations that inversely convert the encoding / decoding results will be performed by the decoder.

[0341] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE), and so on. The technologies described in this document can be embodied in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0342] Figure 3 This is a block diagram of an example video processing apparatus 4100. Apparatus 4100 can be used to implement one or more methods described herein. Apparatus 4100 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processors(multiple) 4102 may be configured to implement one or more methods described herein. The memories(multiple) 4104 may be used to store data and code for implementing the methods and techniques described herein. The video processing circuitry 4106 may be used to implement some of the techniques described herein in hardware circuitry. In some embodiments, the video processing circuitry 4106 may be at least partially included in the processor 4102, for example, a graphics coprocessor.

[0343] Figure 4This is a flowchart of example method 4200 for video processing. In step 4204, a conversion between visual media data and a bitstream is performed based on rules. The rules stipulate that when the purpose of the Neural Network Post-Processing Filter (NNPF) is only to upsample the image rate, the NNPF should not output an image corresponding to any input image. The conversion may include encoding at the encoder, decoding at the decoder, or a combination thereof.

[0344] It should be noted that method 4200 can be implemented in an apparatus for processing video data, including a processor and a non-transitory memory having instructions thereon, such as a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform method 4200. Furthermore, method 4200 can be executed by a non-transitory computer-readable medium including a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec device to perform method 4200 when executed by a processor.

[0345] Figure 5 This is a block diagram illustrating an example video encoding / decoding system 4300 that can utilize the techniques disclosed herein. The video encoding / decoding system 4300 may include a source device 4310 and a destination device 4320. The source device 4310 generates encoded video data, and this source device 4310 may be referred to as a video encoding device. The destination device 4320 can decode the encoded video data generated by the source device 4310, and this destination device 4320 may be referred to as a video decoding device.

[0346] Source device 4310 may include video source 4312, video encoder 4314, and input / output (I / O) interface 4316. Video source 4312 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations thereof. Video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits forming a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to destination device 4320 via network 4330 through I / O interface 4316. Encoded video data may also be stored on storage medium / server 4340 for access by destination device 4320.

[0347] The destination device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may acquire encoded video data from the source device 4310 or the storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the destination device 4320 or may be external to the destination device 4320, wherein the destination device 4320 may be configured to interface with an external display device.

[0348] The video encoder 4314 and the video decoder 4324 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or further standards.

[0349] Figure 6 This is a block diagram illustrating an example of a video encoder 4400, which can be... Figure 5 The system 4300 shown includes a video encoder 4314. The video encoder 4400 can be configured to perform any or all of the techniques disclosed herein. The video encoder 4400 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 4400. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0350] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402, a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy coding unit 4414. The prediction unit 4402 may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra-frame prediction unit 4406.

[0351] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0352] Furthermore, some components such as the motion estimation unit 4404 and the motion compensation unit 4405 can be highly integrated, but for illustrative purposes, they are shown separately in the example of the video encoder 4400.

[0353] The segmentation unit 4401 can segment an image into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.

[0354] The mode selection unit 4403 can select one of several codec modes (intra-frame codec or inter-frame codec) based, for example, on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 4403 can select an intra-frame / inter-frame joint prediction (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 4403 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).

[0355] To perform inter-frame prediction on the current video block, motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. Motion compensation unit 4405 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 4413 other than the image associated with the current video block.

[0356] The motion estimation unit 4404 and the motion compensation unit 4405 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.

[0357] In some examples, motion estimation unit 4404 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 4404 can then generate a reference index indicating the reference image in list 0 or list 1, where the reference image contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0358] In other examples, motion estimation unit 4404 can perform bidirectional prediction on the current video block. Motion estimation unit 4404 can search for reference images in list 0 to find a reference video block for the current video block, and can also search for reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 4404 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 4404 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0359] In some examples, the motion estimation unit 4404 can output a complete set of motion information for use in the decoder's decoding process. In some examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 4404 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0360] In one example, the motion estimation unit 4404 may indicate a value to the video decoder 4500 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.

[0361] In another example, motion estimation unit 4404 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 4500 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0362] As discussed above, the video encoder 4400 can transmit motion vectors via signaling in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.

[0363] Intra-prediction unit 4406 can perform intra-prediction on the current video block. When intra-prediction unit 4406 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples of other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0364] The residual generation unit 4407 can generate residual data for the current video block by subtracting (or more) predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0365] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform subtraction operations.

[0366] The transform processing unit 4408 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0367] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0368] The inverse quantization unit 4410 and the inverse transform unit 4411 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block, which is stored in the buffer 4413.

[0369] After the video block is reconstructed by reconstruction unit 4412, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0370] Entropy encoding unit 4414 can receive data from other functional components of video encoder 4400. When entropy encoding unit 4414 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0371] Figure 7 This is a block diagram illustrating an example of a video decoder 4500, which can be... Figure 5 The system 4300 shown includes a video decoder 4324. The video decoder 4500 can be configured to perform any or all of the techniques disclosed herein. In the example shown, the video decoder 4500 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 4500. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0372] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-frame prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 4400.

[0373] Entropy decoding unit 4501 can retrieve encoded bitstreams. The encoded bitstreams may include entropy-encoded video data (e.g., encoded video data blocks). Entropy decoding unit 4501 can decode the entropy-encoded video data, and motion compensation unit 4502 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference image list index, and other motion information. Motion compensation unit 4502 can determine this information, for example, by executing AMVP and Merge modes.

[0374] The motion compensation unit 4502 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used, with sub-pixel accuracy, can be included in the syntax element.

[0375] The motion compensation unit 4502 can use the interpolation filter used by the video encoder 4400 during the encoding of a video block to calculate the interpolation for sub-integer pixels of the reference block. The motion compensation unit 4502 can determine the interpolation filter used by the video encoder 4400 based on the received syntax information, and the motion compensation unit 4502 can use the interpolation filter to generate the prediction block.

[0376] The motion compensation unit 4502 may use some syntax information to determine the size of the blocks of (multiple) frames and / or (multiple) stripes used to encode the encoded video sequence, segmentation information describing how each macroblock of the picture of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame codec block, and other information for decoding the encoded video sequence.

[0377] Intra-prediction unit 4503 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 4504 dequantizes (i.e., de-quantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501. Inverse transform unit 4505 applies the inverse transform.

[0378] The reconstruction unit 4506 can sum the residual block with the corresponding predicted block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be used to filter the decoded block to eliminate block artifacts. The decoded video block is then stored in a buffer 4507, which provides a reference block for subsequent motion compensation / intra-frame prediction, and the buffer 4507 also generates decoded video for presentation on a display device.

[0379] Figure 8 This is a schematic diagram of the example encoder 4600. Encoder 4600 is suitable for implementing VVC techniques. Encoder 4600 includes three loop filters: a deblocking filter (DF) 4602, a sample adaptive compensation (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike DF 4602, which uses predefined filters, SAO 4604 and ALF 4606 utilize the original samples of the current image, reducing the mean square error between the original and reconstructed samples by adding offsets and applying finite impulse response (FIR) filters respectively, and utilizing the encoded / decoded side information through signal transmission offsets and filter coefficients. ALF 4606 is located in the last processing stage of each image and can be considered as a tool attempting to capture and repair artifacts caused by previous stages.

[0380] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610, configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using a reference image obtained from a reference image buffer 4612. Residual blocks from inter-frame or intra-frame prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are then fed into an entropy codec component 4618. The entropy codec component 4618 entropy codes and decodes the prediction results and quantized transform coefficients and transmits them toward a video decoder (not shown). The quantization component output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. REC component 4624 can output images to DF 4602, SAO 4604 and ALF 4606 for filtering before these images are stored in reference image buffer 4612.

[0381] Figure 9This is a flowchart of example method 4700 for video processing. In step 4702, it is determined that no more than one image in the list of NNPF output images should belong to any particular output time instance. In step 4704, a conversion between visual media data and a bitstream is performed based on the NNPF output images. The conversion may include encoding at the encoder, decoding at the decoder, or a combination thereof.

[0382] It should be noted that method 4700 can be implemented in an apparatus for processing video data, including a processor and a non-transitory memory having instructions thereon, such as a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform method 4700. Furthermore, method 4700 can be executed by a non-transitory computer-readable medium including a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec device to perform method 4700 when executed by a processor.

[0383] The following is a list of some preferred solutions.

[0384] The following solutions illustrate examples of the techniques discussed in this article.

[0385] 1. A method for processing media data, comprising: performing a conversion between visual media data and a bitstream based on rules, wherein the rules specify that when the purpose of a neural network post-processing filter (NNPF) is only to perform image rate upsampling, the NNPF should not output an image corresponding to any input image.

[0386] 2. According to the method described in Solution 1, where all flags nnpfc_input_pic_output_flag[i] should be equal to 0 when the NNPF objective only includes image rate upsampling.

[0387] 3. The method according to any one of solutions 1-2, wherein when the purpose of the NNPF includes image rate upsampling, the signal transmission syntax element nnpfc_input_pic_output_enable_flag is used to specify whether the NNPF outputs an image corresponding to one or more input images, in addition to some interpolated images.

[0388] 4. The method according to any one of solutions 1-3, wherein when nnpfc_input_pic_output_enable_flag is equal to 1, the flag nnpfc_input_pic_output_flag[i] is transmitted by signal.

[0389] 5. The method according to any one of solutions 1-4, wherein when the syntax element nnpfc_input_pic_output_enable_flag is equal to 0, the flag nnpfc_input_pic_output_flag[i] is not transmitted by signal, and their values ​​are presumed to be equal to 0.

[0390] 6. The method according to any one of solutions 1-5, wherein when the NNPF objective only includes image rate upsampling, nnpfc_input_pic_output_enable_flag is not transmitted via signaling, and the value is presumed to be equal to 0.

[0391] 7. The method according to any one of solutions 1-6, wherein NNPF is allowed to output the image multiple times for a specific input image.

[0392] 8. The method according to any one of solutions 1-7, wherein the rule specifies one or more of the following: when NNPF outputs images multiple times for a specific input image, the final NNPF output image corresponding to the specific input image is derived as the image last output by any NNPF for the specific input image; when NNPF outputs images multiple times for a specific input image, the final NNPF output image corresponding to the specific input image is derived as the image first output by any NNPF for the specific input image; or when NNPF outputs images multiple times for a specific input image, the final NNPF output image for the specific input image is derived as the average of all images output by NNPF for the specific input image.

[0393] 9. The method according to any one of solutions 1-8, wherein NNPF is not allowed to output the image multiple times for a specific input image.

[0394] 10. The method according to any one of solutions 1-9, wherein the rule specifies one or more of the following: NNPF shall output the image corresponding to each input image once and only once; NNPF may output the image corresponding to a specific input image only if the corresponding output image was not generated by a previous NNPF; NNPF may output the image corresponding to a specific input image only if the corresponding output image has not been previously generated by any NNPF; NNPF shall not output the image corresponding to a specific input image if a previous NNPF has already output the image corresponding to the specific input image; or NNPF shall not output the image corresponding to a specific input image if the image corresponding to the specific input image has already been output by any NNPF.

[0395] 11. The method according to any one of solutions 1-10, wherein NNPF is allowed to output images multiple times for a specific output order or output time instance.

[0396] 12. The method according to any one of solutions 1-11, wherein the rule specifies one or more of the following: when NNPF outputs images multiple times for a specific output order or output time instance, the final NNPF output image for the specific output order or output time instance is derived as the image last output by any NNPF for the specific output order or output time instance; when NNPF outputs images multiple times for a specific output order or output time instance, the final NNPF output image for the specific output order or output time instance is derived as the image first output by NNPF for the specific output order or output time instance; or when NNPF outputs images multiple times for a specific output order or output time instance, the final NNPF output image for the specific output order or output time instance is derived as the average of all images output by NNPF for the specific output order or output time instance.

[0397] 13. The method according to any one of solutions 1-12, wherein the rule specifies one or more of the following: for any particular output order or output time instance, NNPF is not allowed to output images multiple times; or for any particular output order or output time instance, NNPF shall output images once and only once.

[0398] 14. The method according to any one of solutions 1-13, wherein when the NNPF objective includes image rate upsampling having one or more of the following NNPF objectives or upsampling types: 1) having general visual quality improvement, 2) having chroma upsampling, 3) having resolution upsampling, 4) having bit depth upsampling, 5) having shading, and 6) having depth image generation, one or more of the following aspects are specified: the NNPF shall output an image corresponding to at least one input image, and at least one of the flags nnpfc_input_pic_output_flag[i] shall be equal to 1; or the NNPF shall output an image corresponding to all input images, and all flags nnpfc_input_pic_output_flag[i] shall be equal to 1.

[0399] 15. The method according to any one of solutions 1-14, wherein the rule specifies one or more of the following: a purpose corresponding to a general improvement in visual quality should not involve any format change, including resolution, image rate, chroma format, or bit depth; a purpose to improve general visual quality is not permitted to be combined with any other purpose that results in a format change; a purpose without a general improvement in visual quality is permitted to be used with a format change, including resolution, image rate, chroma format, or bit depth; or a purpose without a general improvement in visual quality is not permitted to be used alone without being combined with any other purpose that results in a format change.

[0400] 16. The method according to any one of solutions 1-15, wherein the rule specifies one or more of the following: when the purpose of the NNPF includes image rate upsampling, for any two different activations of the NNPF, at most one image should be used as the input image for the two activations; or when one or more NNPFs have a purpose including image rate upsampling, for any two different activations of these NNPFs, at most one image should be used as the input image for the two activations.

[0401] 17. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of solutions 1-16.

[0402] 18. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium such that, when executed by a processor, the video codec device performs the method according to any one of solutions 1-16.

[0403] 19. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method comprises: determining to perform a conversion between visual media data and the bitstream based on rules, wherein the rules specify that when the purpose of a neural network post-processing filter (NNPF) is only to perform image rate upsampling, the NNPF should not output an image corresponding to any input image; and generating the bitstream based on the determination.

[0404] 20. A method for storing a bitstream of video, comprising: determining to perform a conversion between visual media data and a bitstream based on rules, wherein the rules specify that when the purpose of a neural network post-processing filter (NNPF) is only to perform image rate upsampling, the NNPF should not output an image corresponding to any input image; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0405] 21. A method, apparatus or system described in this document.

[0406] The following solutions illustrate further examples of the techniques discussed in this article.

[0407] 1. A method for processing media data, comprising: determining that there should be no more than one image in a list of neural network post-processing filter (NNPF) output images that relates to any particular output time instance; and performing a conversion between visual media data and a bitstream based on the NNPF output images.

[0408] 2. The method described in Solution 1, wherein NNPF is not allowed to output the image multiple times for any given input image.

[0409] 3. The method according to any one of solutions 1-2, wherein NNPF is not allowed to output images multiple times for any particular output order.

[0410] 4. The method according to any one of solutions 1-3, wherein NNPF is not allowed to output images multiple times for any particular output time instance.

[0411] 5. The method according to any one of solutions 1-4, wherein when the purpose of NNPF is only image rate upsampling, NNPF should not output an image corresponding to any input image.

[0412] 6. The method according to any one of solutions 1-5, wherein when the NNPF objective only includes image rate upsampling, all flags in the NNPF input image output flag syntax element (nnpfc_input_pic_output_flag[i]) should be equal to 0.

[0413] 7. The method according to any one of solutions 1-6, wherein when the NNPF objective includes image rate upsampling, the NNPF input image output enable flag syntax element (nnpfc_input_pic_output_enable_flag) is signaled to specify whether the NNPF output image corresponds to one or more input images, in addition to some interpolated images.

[0414] 8. The method according to any one of solutions 1-7, wherein when nnpfc_input_pic_output_enable_flag is equal to 1, nnpfc_input_pic_output_flag[i] is transmitted by signal.

[0415] 9. The method according to any one of solutions 1-8, wherein when the syntax element nnpfc_input_pic_output_enable_flag is equal to 0, the flag in nnpfc_input_pic_output_flag[i] is not transmitted by signal, and the flag value in nnpfc_input_pic_output_flag[i] is presumed to be equal to 0.

[0416] 10. The method according to any one of solutions 1-9, wherein when the NNPF objective only includes image rate upsampling, nnpfc_input_pic_output_enable_flag is not transmitted via signaling, and the value is presumed to be equal to 0.

[0417] 11. The method according to any one of solutions 1-10, wherein NNPF is allowed to output the image multiple times for a particular input image.

[0418] 12. The method according to any one of solutions 1-11, wherein when NNPF outputs images multiple times for a specific input image, the final NNPF output image corresponding to the specific input image is derived as the image last output by any NNPF for the specific input image, wherein when NNPF outputs images multiple times for a specific input image, the final NNPF output image corresponding to the specific input image is derived as the image first output by any NNPF for the specific input image, or wherein when NNPF outputs images multiple times for a specific input image, the final NNPF output image for the specific input image is derived as the average of all images output by NNPF for the specific input image.

[0419] 13. The method according to any one of solutions 1-12, wherein the NNPF shall output the image corresponding to each input image once and only once, wherein the NNPF may output the image corresponding to a specific input image only if the corresponding output image was not generated by a previous NNPF, wherein the NNPF may output the image corresponding to a specific input image only if the corresponding output image has not been previously generated by any NNPF, wherein the NNPF shall not output the image corresponding to a specific input image if a previous NNPF has already output the image corresponding to a specific input image, or wherein the NNPF shall not output the image corresponding to a specific input image if the image corresponding to a specific input image has already been output by any NNPF.

[0420] 14. The method according to any one of solutions 1-13, wherein NNPF is allowed to output images multiple times for a specific output order or output time instance.

[0421] 15. The method according to any one of solutions 1-14, wherein when NNPF outputs images multiple times for a specific output order or output time instance, the final NNPF output image for the specific output order or output time instance is derived as the image last output by any NNPF for the specific output order or output time instance, wherein when NNPF outputs images multiple times for a specific output order or output time instance, the final NNPF output image for the specific output order or output time instance is derived as the image first output by NNPF for the specific output order or output time instance, or wherein when NNPF outputs images multiple times for a specific output order or output time instance, the final NNPF output image for the specific output order or output time instance is derived as the average of all images output by NNPF for the specific output order or output time instance.

[0422] 16. The method according to any one of solutions 1-15, wherein NNPF shall output the image once and only once for any particular output order or output time instance.

[0423] 17. The method according to any one of solutions 1-16, wherein when the NNPF objective includes image rate upsampling having one or more of the following NNPF objectives or upsampling types: 1) having general visual quality improvement, 2) having chroma upsampling, 3) having resolution upsampling, 4) having bit depth upsampling, 5) having shading, and 6) having depth image generation, specifies one or more of the following aspects: the NNPF shall output an image corresponding to at least one input image, and at least one of the flags nnpfc_input_pic_output_flag[i] shall be equal to 1; or the NNPF shall output an image corresponding to all input images, and all flags nnpfc_input_pic_output_flag[i] shall be equal to 1.

[0424] 18. The method according to any one of solutions 1-17, wherein corresponding to the purpose of having general visual quality improvement shall not involve any format change, including resolution, image rate, chroma format, or bit depth; wherein the purpose of having general visual quality improvement is not permitted to be combined with any other purpose that causes a format change; wherein the purpose without general visual quality improvement is permitted to be used with a format change, including resolution, image rate, chroma format, or bit depth; or wherein the purpose without general visual quality improvement is not permitted to be used alone without being combined with any other purpose that causes a format change.

[0425] 19. The method according to any one of solutions 1-18, wherein when the purpose of the NNPF includes image rate upsampling, for any two different activations of the NNPF, at most one image should be used as the input image for the two activations; or wherein when one or more NNPFs have a purpose including image rate upsampling, for any two different activations of these NNPFs, at most one image should be used as the input image for the two activations.

[0426] 20. The method according to any one of solutions 1-19, wherein the conversion includes encoding the visual media data into the bitstream.

[0427] 21. The method according to any one of solutions 1-19, wherein the conversion includes decoding the visual media data from the bitstream.

[0428] 22. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of solutions 1-21.

[0429] 23. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium such that, when executed by a processor, the video codec device performs the method according to any one of solutions 1-21.

[0430] 24. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method comprises: determining that there should be no more than one image relating to any particular output time instance within a list of output images of a neural network post-processing filter (NNPF); and generating a bitstream based on the determination.

[0431] 25. A method for storing a bitstream of video, comprising: determining that there should not be more than one image relating to any particular output time instance in a list of images output by a neural network post-processing filter (NNPF); generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0432] In the described solution, the encoder conforms to the format rules by generating a codec representation based on those rules. In the described solution, the decoder parses the syntax elements in the codec representation using known information about their presence or absence, based on the format rules, to produce the decoded video.

[0433] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. For example, the bitstream representation of the current video block can correspond to bits at the same position in the bitstream defined by the syntax or bits propagated at different positions. For example, a macroblock can be encoded based on the error residual value after transformation and encoding, and can also use bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can resolve the bitstream based on determination, knowing whether some fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude specific syntax fields and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.

[0434] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more computer program instruction modules encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of substances influencing machine-readable propagation signals, or a combination of one or more. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for a related computer program, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. Propagation signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information to be transmitted to a suitable receiver device.

[0435] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the related program, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.

[0436] The processing and logic flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processing and logic flows can also be executed by special-purpose logic circuitry, and the devices can be implemented as special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0437] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or receive data from or transfer data to one or more mass storage devices via operative coupling, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor storage devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. Processors and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0438] While this patent document contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features specific to particular embodiments of a particular technology. In this patent document, certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. Furthermore, although features may function in certain combinations as described above, and even were originally claimed in this manner, in some cases one or more features in the claimed combination may be removed from that combination, and the claimed combination may be for sub-combinations or variations thereof.

[0439] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed sequentially in the specific order or sequence shown, or requiring all shown operations to be performed in order to achieve the desired result. Furthermore, the division of various system components in the embodiments described in this patent document should not be construed as requiring such division in all embodiments.

[0440] Only a few implementations and examples are described, and other implementations, improvements and variations can be made based on what is described and shown in this patent document.

[0441] When there is no intermediary component other than a line, trace, or other medium between the first and second components, the first component is directly coupled to the second component. When there is an intermediary component other than a line, trace, or other medium between the first and second components, the first component is indirectly coupled to the second component. The term "coupled" and its variations include direct coupling and indirect coupling. The use of the term "about" means including a range of ±10% of the following figures, unless otherwise specified.

[0442] While several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. The present examples are to be considered illustrative rather than restrictive and are not intended to be limited to the details set forth herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.

[0443] Furthermore, the technologies, systems, subsystems, and methods described and illustrated as discrete or separate in the various embodiments can be combined or integrated with other systems, modules, technologies, or methods without departing from the scope of this disclosure. Other items shown or discussed as couplings can be directly connected or indirectly coupled or communicated through some interface, device, or intermediate component, whether electrical, mechanical, or otherwise. Other examples of variations, substitutions, and modifications will be apparent to those skilled in the art and can be made without departing from the spirit and scope of this disclosure.

Claims

1. A method for processing media data, comprising: The list of output images for the Neural Network Post-Processing Filter (NNPF) should not contain more than one image that relates to any particular output time instance. as well as The conversion between visual media data and bitstream is performed based on the NNPF output image.

2. The method of claim 1, wherein for any given input image, NNPF is not allowed to output the image multiple times.

3. The method according to any one of claims 1-2, wherein NNPF is not allowed to output images multiple times for any particular output order.

4. The method according to any one of claims 1-3, wherein NNPF is not allowed to output images multiple times for any particular output time instance.

5. The method according to any one of claims 1-4, wherein when the purpose of NNPF is only image rate upsampling, NNPF should not output an image corresponding to any input image.

6. The method according to any one of claims 1-5, wherein when the NNPF objective only includes image rate upsampling, all flags in the NNPF input image output flag syntax element (nnpfc_input_pic_output_flag[i]) should be equal to 0.

7. The method according to any one of claims 1-6, wherein when the NNPF objective includes image rate upsampling, the NNPF input image output enable flag syntax element (nnpfc_input_pic_output_enable_flag) is transmitted via signal transmission to specify whether the NNPF output image corresponds to one or more input images, in addition to some interpolated images.

8. The method according to any one of claims 1-7, wherein when nnpfc_input_pic_output_enable_flag is equal to 1, nnpfc_input_pic_output_flag[i] is transmitted via signal.

9. The method according to any one of claims 1-8, wherein when the syntax element nnpfc_input_pic_output_enable_flag is equal to 0, the flag in nnpfc_input_pic_output_flag[i] is not transmitted by signal, and the flag value in nnpfc_input_pic_output_flag[i] is presumed to be equal to 0.

10. The method according to any one of claims 1-9, wherein when the NNPF objective comprises only image rate upsampling, nnpfc_input_pic_output_enable_flag is not transmitted via signaling, and the value is presumed to be equal to 0.

11. The method according to any one of claims 1-10, wherein NNPF is allowed to output the image multiple times for a particular input image.

12. The method according to any one of claims 1-11, wherein when the NNPF outputs images multiple times for a specific input image, the final NNPF output image corresponding to the specific input image is derived as the image last output by any NNPF for the specific input image, wherein when the NNPF outputs images multiple times for a specific input image, the final NNPF output image corresponding to the specific input image is derived as the image first output by any NNPF for the specific input image, or wherein when the NNPF outputs images multiple times for a specific input image, the final NNPF output image for the specific input image is derived as the average of all images output by the NNPF for the specific input image.

13. The method according to any one of claims 1-12, wherein the NNPF shall output the image corresponding to each input image once and only once, wherein the NNPF may output the image corresponding to a particular input image only when the corresponding output image was not generated by a previous NNPF, wherein the NNPF may output the image corresponding to a particular input image only when the corresponding output image has not been previously generated by any NNPF, wherein the NNPF shall not output the image corresponding to a particular input image when a previous NNPF has already output the image corresponding to a particular input image, or wherein the NNPF shall not output the image corresponding to a particular input image when the image corresponding to a particular input image has already been output by any NNPF.

14. The method according to any one of claims 1-13, wherein NNPF is allowed to output images multiple times for a specific output order or output time instance.

15. The method according to any one of claims 1-14, wherein when the NNPF outputs images multiple times for a specific output order or output time instance, the final NNPF output image for the specific output order or output time instance is derived as the image last output by any NNPF for the specific output order or output time instance, wherein when the NNPF outputs images multiple times for a specific output order or output time instance, the final NNPF output image for the specific output order or output time instance is derived as the image first output by the NNPF for the specific output order or output time instance, or wherein when the NNPF outputs images multiple times for a specific output order or output time instance, the final NNPF output image for the specific output order or output time instance is derived as the average of all images output by the NNPF for the specific output order or output time instance.

16. The method according to any one of claims 1-15, wherein for any particular output order or output time instance, NNPF shall output the image once and only once.

17. The method according to any one of claims 1-16, wherein when the NNPF objective includes image rate upsampling having one or more of the following NNPF objectives or upsampling types: 1) having general visual quality improvement, 2) having chroma upsampling, 3) having resolution upsampling, 4) having bit depth upsampling, 5) having colorization, and 6) having depth image generation, one or more of the following aspects are specified: The NNPF should output an image corresponding to at least one of the input images, and at least one flag in the flag nnpfc_input_pic_output_flag[i] should be equal to 1; or The NNPF should output an image corresponding to all the input images, and all flags nnpfc_input_pic_output_flag[i] should be equal to 1.

18. The method of any one of claims 1-17, wherein the objective of having a general improvement in visual quality should not involve any format change, including resolution, image rate, chroma format, or bit depth; wherein the objective of having a general improvement in visual quality is not permitted to be combined with any other objective that results in a format change; wherein the objective of not having a general improvement in visual quality is permitted to be used with a format change, including resolution, image rate, chroma format, or bit depth; or wherein the objective of not having a general improvement in visual quality is not permitted to be used alone without being combined with any other objective that results in a format change.

19. The method according to any one of claims 1-18, wherein when the purpose of the NNPF includes image rate upsampling, for any two distinct activations of the NNPF, at most one image should be used as the input image for the two activations; or wherein when one or more NNPFs have the purpose of including image rate upsampling, for any two distinct activations of these NNPFs, at most one image should be used as the input image for the two activations.

20. The method according to any one of claims 1-19, wherein the conversion comprises encoding the visual media data into the bitstream.

21. The method according to any one of claims 1-19, wherein the conversion comprises decoding the visual media data from the bitstream.

22. An apparatus for processing video data, comprising: processor; and a non-transitory memory thereon having instructions, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-21.

23. A non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium such that, when executed by a processor, the video codec apparatus performs the method according to any one of claims 1-21.

24. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: The list of output images for the Neural Network Post-Processing Filter (NNPF) should not contain more than one image that relates to any particular output time instance. as well as Based on the determination, a bit stream is generated.

25. A method for storing a bitstream of video, comprising: The list of output images for the Neural Network Post-Processing Filter (NNPF) should not contain more than one image that relates to any particular output time instance. Based on the determination, a bit stream is generated; as well as The bit stream is stored in a non-transitory computer-readable recording medium.