Method and device for video processing and medium
By applying a neural network post-processing filter (NNPF) to the video bitstream to reduce the width and/or height of the image, the shortcomings of existing NNPF technologies in resolution upsampling and downsampling are addressed, achieving functional expansion and efficiency improvement.
Patent Information
- Application Number
- CN202480024314.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-15
- Filing Date
- 2024-04-07
- Publication Date
- 2025-11-25
AI Technical Summary
Existing video encoding and decoding technologies have room for improvement in terms of functionality, especially the unmet need for expansion and enhancement of neural network post-processing filters (NNPF) for video processing in terms of resolution upsampling and downsampling.
By applying a neural network post-processing filter (NNPF) to transform images in the video bitstream, the width and/or height of the images are reduced to achieve resolution downsampling, thus expanding and enhancing the functionality of NNPF.
It enables NNPF to downsample at higher resolutions, expands its functionality, and improves the efficiency and effectiveness of video processing.
Smart Images

Figure CN121014201A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to neural network post-processing filters (NNPF). Background Technology
[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Multi-Functional Video Codec (VVC) standard. However, the functionality of video codec technologies is often expected to be further improved. Summary of the Invention
[0003] Embodiments of this disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: performing a conversion between video and video bitstreams, wherein a neural network post-processing filter (NNPF) is applied to at least one image associated with the video, the bitstream including a first indication indicating the purpose of the NNPF, and a first candidate among a plurality of candidates for that purpose indicating at least one of the following is applicable: reducing the width of at least one image, or reducing the height of at least one image.
[0005] Based on the method according to the first aspect of this disclosure, one of the candidates for NNPF is to indicate that reducing the width and / or height of at least one image is applicable. Compared to conventional solutions that only support resolution upsampling, the proposed method advantageously enables NNPF to support resolution downsampling. Thus, the functionality of NNPF can be extended and enhanced.
[0006] In a second aspect, an apparatus for video processing is provided. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.
[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of this disclosure.
[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by means of a device for video processing. The method includes: performing a conversion between video and video bitstreams, wherein a neural network post-processing filter (NNPF) is applied to at least one image associated with the video, the bitstream including a first indication indicating the purpose of the NNPF, and the first candidate among a plurality of candidates for that purpose indicating at least one of the following is applicable: reducing the width of at least one image, or reducing the height of at least one image.
[0009] In a fifth aspect, a method for storing a bitstream of video is proposed. The method includes: performing a conversion between video and video bitstreams, wherein a neural network post-processing filter (NNPF) is applied to at least one image associated with the video, the bitstream including a first indication indicating the purpose of the NNPF, and a first candidate among a plurality of candidates for that purpose indicating at least one of the following is applicable: reducing the width of at least one image, or reducing the height of at least one image; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] This summary aims to present, in a simplified form, the selected concepts further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0011] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0012] Figure 1 A block diagram of an example video codec system according to some embodiments of the present disclosure is shown;
[0013] Figure 2 A block diagram of a first example video encoder according to some embodiments of the present disclosure is shown;
[0014] Figure 3 A block diagram of an example video decoder according to some embodiments of the present disclosure is shown;
[0015] Figure 4 A schematic diagram of the brightness data channel is shown;
[0016] Figure 5 A flowchart of a method for video processing according to embodiments of the present disclosure is shown; and
[0017] Figure 6 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0018] In all the accompanying drawings, the same or similar reference numerals usually refer to the same or similar elements. Detailed Implementation
[0019] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0020] In the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0021] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, whether explicitly described or not, it is believed that such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0022] It should be understood that although the terms “first” and “second”, etc., can be used to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0023] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” and / or “having” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Example Environment
[0024] Figure 1This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0025] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.
[0026] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded and decoded representation of the video data. The bitstream may include encoded images and associated data. The decoded images are encoded representations of the images. The associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.
[0027] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.
[0028] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or further standards.
[0029] Figure 2This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.
[0030] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0031] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0032] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0033] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2 The examples are shown separately.
[0034] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0035] The mode selection unit 203 can, for example, select one of several coding modes (intra-coding or inter-coding) based on the error result, and provide the resulting intra-coded or inter-coded block to the residual generation unit 207 to generate residual block data, and provide it to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-inter-prediction joint prediction (CIIP) mode, in which prediction is based on inter-prediction signals and intra-prediction signals. In the case of inter-prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).
[0036] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.
[0037] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks that do not depend on macroblocks within the same image.
[0038] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0039] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search reference images in list 0 to find a reference video block for the current video block, and can also search reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference images containing reference video blocks in lists 0 and 1, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and multiple motion vectors of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.
[0040] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0041] In one example, the motion estimation unit 204 may indicate a value to the video decoder 300 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.
[0042] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0043] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0044] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0045] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0046] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform subtraction operations.
[0047] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0048] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0049] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0050] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0051] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0052] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.
[0053] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0054] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.
[0055] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and based on the entropy-encoded video data, motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine this information, for example, by performing AMVP and Merge mode. AMVP is used, which includes deriving several most likely candidates based on data from adjacent PBs and reference pictures. Motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially or temporally adjacent blocks.
[0056] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. Identifiers for interpolation filters used at sub-pixel precision can be included in the syntax elements.
[0057] The motion compensation unit 302 can use interpolation filters, such as those used by the video encoder 200 during the encoding of a video block, to calculate interpolated values for sub-integer pixels of a reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.
[0058] Motion compensation unit 302 may use at least some of the syntax information to determine the block size of the frames(multiple) and / or stripes(multiple) used to encode the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be the entire image or a region of the image.
[0059] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.
[0060] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding predicted block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0061] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, while some embodiments are described with reference to multi-function video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Additionally, although some embodiments describe video encoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Furthermore, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates. 1. Preliminary Discussion
[0062] This document relates to image / video codec techniques. Specifically, this disclosure relates to the definition and signaling of a neural network post-processing filter with downsampling capability transmitted as a signal in a video bitstream. Downsampling in this document can be bit-depth downsampling, for example, from 10 bits to 2 bits. Downsampling in this document can also be reducing one or both of the image width and image height, or both. For video bitstreams encoded by any codec (e.g., the Multi-Function Video Codec (VVC) standard and / or the Multi-Function Supplemental Enhancement Information (SEI) Message (VSEI) standard used for encoding and decoding video bitstreams), these ideas can be applied individually or in various combinations. 2. Abbreviation
[0063] Adaptive Parameter Set (APS), Access Unit (AU), Codec Layer Video Sequence (CLVS), Codec Layer Video Sequence Start (CLVSS), Cyclic Redundancy Check (CRC), Codec Video Sequence (CVS), Finite Impulse Response (FIR), Intra-Frame Random Access Point (IRAP), Network Abstraction Layer (NAL), Picture Parameter Set (PPS), Picture Unit (PU), Random Access Skip Before (RASL) Picture, Supplemental Enhancement Information (SEI), Stepped Temporal Sublayer Access (STSA), Video Codec Layer (VCL), Multifunctional Supplemental Enhancement Information (VSEI) described in Recommendation ITU-T H.274|ISO / IEC 23002-7, Video Availability Information (VUI), and Multifunctional Video Codec (VVC) described in Recommendation ITU-T H.266|ISO / IEC 23090-3. 3. Further discussion 3.1 Video codec standards
[0064] Video codec standards have primarily evolved through the development of standards by the International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) and the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed the Moving Picture Experts Group (MPEG)-1 and MPEG-4 Vision standards. The two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / High-Efficiency Video Codec (HEVC) standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was established by the Video Codecs Expert Group (VCEG) and the Moving Picture Experts Group (MPEG). Furthermore, JVET has adopted several methodologies and incorporated them into reference software called the Joint Exploration Model (JEM). When the Multi-Functional Video Codec (VVC) project was officially launched, JVET was later renamed the Joint Video Experts Group (JVET). VVC is a codec standard designed to reduce the bit rate by 50% compared to HEVC.
[0065] The Multi-Functional Video Coding (VVC) standard (ITU-T H.266|ISO / IEC 23090-3) and the associated Multi-Functional Supplemental Enhancement Information (VSEI) standard (ITU-T H.274|ISO / IEC 23002-7) for encoding and decoding video bitstreams are designed for the widest range of applications, including simple uses such as television broadcasting, video conferencing, or playback from storage media, as well as more advanced use cases such as adaptive bitrate streaming, video region extraction, synthesis and merging of content from multiple encoded video bitstreams, multi-view video, scalable layered coding and decoding, and viewport-adaptive 360° immersive media.
[0066] The Basic Video Codec (EVC) standard (ISO / IEC 23094-1) is another video codec standard developed by MPEG. 3.2 Common SEI messages and those in VVC and VSEI
[0067] SEI messages assist in processes related to decoding, display, or other purposes. However, SEI messages are not necessary for constructing luma or chroma samples during the decoding process. Standard-compliant decoders are not required to process this information for output order consistency. Some SEI messages are necessary for checking bitstream consistency and output timing decoder consistency. Other SEI messages are not necessary for checking bitstream consistency.
[0068] Appendix D of VVC specifies the syntax and semantics of SEI message payloads for some SEI messages, and specifies the use of SEI messages and VUI parameters with the syntax and semantics specified in ITU-T H.274|ISO / IEC 23002-7. 3.3 Signaling of Neural Network Post-Processing Filters
[0069] The following is an excerpt from the specifications of two SEI messages used for signaling in neural network post-processing filters. 8.28 Neural Network Post-Processing Filter Characteristics SEI Message 8.28.1 Characteristics of Neural Network Post-Processing Filters and SEI Message Syntax 8.28.2 Characteristics of Neural Network Post-Processing Filters and Semantics of SEI Messages
[0070] The Neural Network Post-Processing Filter Characteristics (NNPFC) SEI message specifies the neural networks that can be used as post-processing filters. The use of the specified Neural Network Post-Processing Filter (NNPF) for a specific image is indicated by the Neural Network Post-Processing Filter Activation (NNPFA) SEI message.
[0071] To use this SEI message, the following variables need to be specified: - Input the image width and height, in units of brightness samples, denoted as CroppedWidth and [missing information] respectively in this article. Cropped Height. - An array of luminance samples of the input image with index idx ranging from 0 to numInputPics-1 (inclusive). CroppedYPic[idx] and chromaticity sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] (if present) are used as inputs for NNPF. - BitDepth for the luminance sample array of the input image Y . - Bit depth of the chroma sample array (if any) for the input image. C . - Chroma format indicator, referred to herein as ChromaFormatIdc, as described in sub-entry 7.3. - When nnpfc_auxiliary_inp_idc equals 1, the filter strength control value StrengthControlVal should be a real number in the range of 0 to 1 (including boundary values).
[0072] The input image with index 0 corresponds to the image for which the NNPF defined by the NNPFC SEI message is activated via the NNPFA SEI message. The input images with index i (in the range of 1 to numInputPics-1 (inclusive)) precede the input images with index i-1 in the output order.
[0073] When an input image with index 0 and nnpfc_purpose&0x08 is not equal to 0 is associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5, all input images are associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5 and the same value as fp_current_frame_is_frame0_flag.
[0074] The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc, as specified in Table 2.
[0075] Note 1 – More than one NNPFC SEI message can exist for the same image. When more than one NNPFC SEI message with different values of nnpfc_id exists or is activated for the same image, they can have the same or different values of nnpfc_purpose and nnpfc_mode_idc.
[0076] nnpfc_purpose indicates the purpose of NNPF, as specified in Table 20.
[0077] In bitstreams conforming to this version of this document, the value of nnpfc_purpose should be in the range of 0 to 63 (inclusive). Values of nnpfc_purpose from 64 to 65,535 (inclusive) are reserved for future use by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this version of this document. Decoders conforming to this version of this document should ignore NNPFC SEI messages with nnpfc_purpose in the range of 64 to 65,535 (inclusive). Table 20 - Definition of nnpfc_purpose
[0078] Note 2 – When the reserved value of nnpfc_purpose is used by ITU-T|ISO / IEC in the future, the syntax of this SEI message can be extended using the following syntax elements, provided that nnpfc_purpose is equal to that value.
[0079] When ChromaFormatIdc equals 3, nnpfc_purpose&0x02 should equal 0.
[0080] When ChromaFormatIdc or nnpfc_purpose&0x02 is not equal to 0, nnpfc_purpose&0x20 should be equal to 0.
[0081] The nnpfc_id contains an identifier that can be used to identify NNPF. The value of nnpfc_id should be in the range of 0 to 2^32-2 (inclusive). Values of nnpfc_id from 2^56 to 5^11 (inclusive) and from 2^31 to 2^32-2 (inclusive) are reserved for future use by ITU-T|ISO / IEC. Decoders conforming to this version of this document should ignore NNPFC SEI messages when they encounter nnpfc_id values in the range of 2^56 to 5^11 (inclusive) or 2^31 to 2^32-2 (inclusive).
[0082] The following applies when an NNPFC SEI message has a specific nnpfc_id value within the current CLVS and is the first NNPFC SEI message in decoding order: - This SEI message specifies the basic NNPF. - This SEI message applies to the currently decoded image and all subsequent decoded images of the current layer (in output order) until the end of the current CLVS.
[0083] An nnpfc_mode_idc value of 0 indicates that the SEI message contains an ISO / IEC 15938-17 bitstream that specifies the basic NNPF or is updated relative to the basic NNPF with the same nnpfc_id value.
[0084] When an NNPFC SEI message has a specific nnpfc_id value within the current CLVS and is the first NNPFC SEI message in decoding order, nnpfc_mode_idc equals 1, indicating that the underlying NNPF associated with the nnpfc_id value is a neural network identified by the URI indicated by nnpfc_uri using a format identified by the tag URI nnpfc_tag_uri.
[0085] When an NNPFC SEI message is neither the first NNPFC SEI message with a specific nnpfc_id value within the current CLVS and in decoding order, nor a duplicate of the first NNPFC SEI message with a specific nnpfc_id value within the current CLVS and in decoding order, an nnpfc_mode_idc equal to 1 indicates that the update relative to the basic NNPF with the same nnpfc_id value is defined by the URI indicated by nnpfc_uri using the format identified by the tag URI nnpfc_tag_uri.
[0086] In bitstreams conforming to this version of this document, the value of nnpfc_mode_idc should be in the range of 0 to 1 (inclusive). Values of nnpfc_mode_idc from 2 to 255 (inclusive) are reserved for future use by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this version of this document. Decoders conforming to this version of this document should ignore NNPFC SEI messages with nnpfc_mode_idc in the range of 2 to 255 (inclusive). Values of nnpfc_mode_idc greater than 255 should not exist in bitstreams conforming to this version of this document and are not reserved for future use.
[0087] When the SEI message is the first NNPFCSEI message in decoding order and has a specific nnpfc_id value within the current CLVS, NNPF PostProcessingFilter() is assigned the same as the base NNPF.
[0088] When the SEI message is neither the first NNPFC SEI message with a specific nnpfc_id value in the current CLVS and in decoding order, nor a duplicate of the first NNPFC SEI message with a specific nnpfc_id value in the current CLVS and in decoding order, NNPF PostProcessingFilter() is obtained by applying the update defined by the SEI message to the base NNPF.
[0089] The updates are not cumulative; instead, each update is applied to a base NNPF, which is defined by the first NNPFC SEI message with a specific nnpfc_id value within the current CLVS and in decoding order.
[0090] nnpfc_reserved_zero_bit_a should be equal to 0 in the bitstream conforming to this version of the document. The decoder should ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_a is not equal to 0.
[0091] The nnpfc_tag_uri contains a tag URI with the syntax and semantics as specified in IETF RFC 4151, identifying the format and associated information about the neural network used as the base NNPF or an updated base NNPF with the same nnpfc_id value as specified by the nnpfc_uri.
[0092] Note 3 – nnpfc_tag_uri enables the unique identification of neural network data in the format specified by nnrpf_uri without requiring a central registry.
[0093] The nnpfc_tag_uri equal to “tag:iso.org,2023:15938-17” indicates that the neural network data identified by nnpfc_uri conforms to ISO / IEC 15938-17.
[0094] nnpfc_uri contains a URI with syntax and semantics as specified in IETF Internet Standard 66, identifying a neural network used as a base NNPF or an updated NNPF relative to a base NNPF with the same nnpfc_id value.
[0095] A value of 1 for nnpfc_property_present_flag indicates the presence of syntax elements related to the filter's purpose, input format, output format, and complexity. A value of 0 for nnpfc_property_present_flag indicates the absence of any syntax elements related to the filter's purpose, input format, output format, and complexity.
[0096] When the SEI message is the first NNPFCSEI message in decoding order and has a specific nnpfc_id value within the current CLVS, nnpfc_property_present_flag should be equal to 1.
[0097] When nnpfc_property_present_flag equals 0, the value of all syntax elements that can exist only when nnpfc_property_present_flag equals 1 and for which no presumed value is specified is presumed to be equal to the corresponding syntax element in the NNPFC SEI message containing the basic NNPF for which the SEI provides an update.
[0098] An nnpfc_base_flag value of 1 indicates that the SEI message specifies the basic NNPF. An nnpf_base_flag value of 0 indicates that the SEI message specifies an update relative to the basic NNPF. When it does not exist, the value of nnpfc_base_flag is presumed to be 0.
[0099] The following constraints apply to the value of nnpfc_base_flag: - When the NNPFC SEI message is the first NNPFC SEI message in decoding order and has a specific nnpfc_id value within the current CLVS, the value of nnpfc_base_flag should be equal to 1. - When NNPFC SEI message nnpfcB is not the first NNPFC SEI message in decoding order and has a specific nnpfc_id value within the current CLVS, and the value of nnpfc_base_flag is equal to 1, the NNPFC SEI message should be a duplicate of the first NNPFC SEI message nnpfcA in decoding order with the same nnpfc_id, that is, the payload content of nnpfcB should be the same as the payload content of nnpfcA.
[0100] When an NNPFC SEI message is not the first NNPFC SEI message in decoding order with a specific nnpfc_id value within the current CLVS, and is not a duplicate of the first NNPFC SEI message with that specific nnpfc_id, the following applies: - This SEI message defines an update relative to the base NNPF with the same nnpfc_id value and in the order of decoding. - This SEI message applies to the currently decoded image and all subsequent decoded images of the current layer (in output order), up to the end of the current CLVS, or up to but not including the decoded image that is in the current CLVS after the currently decoded image in output order and associated with a subsequent NNPFC SEI message that has that specific nnpfc_id value in the current CLVS and in decoding order, whichever is earlier.
[0101] When the NNPFC SEI message nnpfcCurr is not the first NNPFC SEI message in decoding order with a specific nnpfc_id value within the current CLVS, nor a duplicate of the first NNPFC SEI message with that specific nnpfc_id (i.e., the value of nnpfc_base_flag is equal to 0), and the value of nnpfc_property_present_flag is equal to 1, the following constraints apply: The value of nnpfc_purpose in the -NNPFC SEI message should be the same as the value of nnpfc_purpose in the first NNPFC SEI message that has that specific nnpfc_id value within the current CLVS and is in decoding order. The value of the syntax element in the NNPFC SEI message that is after nnpfc_base_flag and before nnpfc_complexity_info_present_flag in the decoding order should be the same as the value of the corresponding syntax element in the first NNPFC SEI message that has that particular nnpfc_id value in the current CLVS and is in the decoding order. - Either nnpfc_complexity_info_present_flag should be equal to 0, or both nnpfc_complexity_info_present_flags in the first NNPFC SEI message (hereinafter referred to as nnpfcBase) with that specific nnpfc_id value in the current CLVS and in decoding order should be equal to 1, and all of the following apply: The nnpfc_parameter_parameter_type_idc in -nnpfcCurr should be equal to the nnpfc_parameter_parameter_type_idc in nnpfcBase. The nnpfc_log2_parameter_bit_length_minus3 in -nnpfcCurr (if it exists) should be less than or equal to the nnpfc_log2_parameter_bit_length_minus3 in nnpfcBase. - If nnpfc_num_parameters_idc in nnpfcBase is equal to 0, then nnpfc_num_parameters_idc in nnpfcCurr should be equal to 0. Otherwise (nnpfc_num_parameters_idc in nnpfcBase is greater than 0), nnpfc_num_parameters_idc in nnpfcCurr should be greater than 0 and less than or equal to nnpfc_num_parameters_idc in nnpfcBase. - If nnpfc_num_kmac_operations_idc in nnpfcBase is equal to 0, then nnpfc_num_kmac_operations_idc in nnpfcCurr should be equal to 0. Otherwise (nnpfc_num_kmac_operations_idc in nnpfcBase is greater than 0), nnpfc_num_kmac_operations_idc in nnpfcCurr should be greater than 0 and less than or equal to nnpfc_num_kmac_operations_idc in nnpfcBase. - If nnpfc_total_kilobyte_size in nnpfcBase is equal to 0, then nnpfc_total_kilobyte_size in nnpfcCurr should be equal to 0. Otherwise (nnpfc_total_kilobyte_size in nnpfcBase is greater than 0), nnpfc_total_kilobyte_size in nnpfcCurr should be greater than 0 and less than or equal to nnpfc_total_kilobyte_size in nnpfcBase.
[0102] When nnpfc_purpose&0x02 or nnpfc_purpose&0x40 is not equal to 0, nnpfc_out_sub_c_flag specifies the values of variables outSubWidthC and outSubHeightC. nnpfc_out_sub_c_flag equal to 1 specifies that outSubWidthC equals 1 and outSubHeightC equals 1. nnpfc_out_sub_c_flag equal to 0 specifies that outSubWidthC equals 2 and outSubHeightC equals 1. When ChromaFormatIdc equals 2, nnpfc_purpose&0x40 equals 0, and nnpfc_out_sub_c_flag exists, the value of nnpfc_out_sub_c_flag should be equal to 1.
[0103] When `nnpfc_purpose&0x20` is not equal to 0, `nnpfc_out_colour_format_idc` specifies the color format of the NNPF output, thus defining the values of the variables `outSubWidthC` and `outSubHeightC`. `nnpfc_out_colour_format_idc` equal to 1 specifies that the NNPF output color format is 4:2:0, and both `outSubWidthC` and `outSubHeightC` are equal to 2. `nnpfc_out_colour_format_idc` equal to 2 specifies that the NNPF output color format is 4:2:2, and both `outSubWidthC` and `outSubHeightC` are equal to 1. `nnpfc_out_colour_format_idc` equal to 3 specifies that the NNPF output color format is 4:2:4, and both `outSubWidthC` and `outSubHeightC` are equal to 1. The value of `nnpfc_out_colour_format_idc` should not be equal to 0.
[0104] When both nnpfc_purpose&0x02 and nnpfc_purpose&0x20 are equal to 0, outSubWidthC and outSubHeightC are presumed to be equal to SubWidthC and SubHeightC, respectively.
[0105] `nnpfc_pic_width_in_luma_samples` and `nnpfc_pic_height_in_luma_samples` specify the width and height of the luminance sample array of the image obtained by applying the NNPF identified by `nnpfc_id` to the cropped, decoded output image, respectively. When `nnpfc_pic_width_in_luma_samples` and `nnpfc_pic_height_in_luma_samples` are not present, they are presumed to be equal to `CroppedWidth` and `CroppedHeight`, respectively. The value of `nnpfc_pic_width_in_luma_samples` should be in the range from `CroppedWidth` to `CroppedWidth*16-1` (inclusive). The value of `nnpfc_pic_height_in_luma_samples` should be in the range from `CroppedHeight` to `CroppedHeight*16-1` (inclusive).
[0106] The increment of 1 in `nnpfc_num_input_pics_minus1` specifies the number of decoded output images used as input for NNPF. The value of `nnpfc_num_input_pics_minus1` should be in the range of 0 to 63 (inclusive). When `nnpfc_purpose&0x08` is not equal to 0, the value of `nnpfc_num_input_pics_minus1` should be greater than 0.
[0107] `nnpfc_interpolated_pics[i]` specifies the number of interpolated pictures generated by NNPF between the i-th picture and the (i+1)-th picture used as input to NNPF. The value of `nnpfc_interpolated_pics[i]` should be in the range of 0 to 63 (inclusive). The value of `nnpfc_interpolated_pics[i]` should be greater than 0 for at least one `i` in the range of 0 to `nnpfc_num_input_pics_minus1-1` (inclusive).
[0108] `nnpfc_input_pic_output_flag[i]` equal to 1 indicates that NNPF generates the corresponding output image for the i-th input image. `nnpfc_input_pic_output_flag[i]` equal to 0 indicates that NNPF does not generate the corresponding output image for the i-th input image.
[0109] The variables numInputPics, which specify the number of images used as input to NNPF, and numOutputPics, which specify the total number of images generated by NNPF, are derived as follows:
[0110] A value of 1 for nnpfc_component_last_flag indicates that the last dimension of both the input tensor and the output tensor generated by NNPF is used for the current channel. A value of 0 for nnpfc_component_last_flag indicates that the third dimension of both the input tensor and the output tensor generated by NNPF is used for the current channel.
[0111] Note 4 – The first dimension in both the input and output tensors is used for the batch index, which is a practice in some neural network frameworks. Although the formula in the semantics of this SEI message uses the batch size corresponding to a batch index of 0, the batch size used as input for neural network inference is determined by the post-processing implementation.
[0112] Note 5 – For example, when nnpfc_inp_order_idc equals 3 and nnpfc_auxiliary_inp_idc equals 1, the input tensor has 7 channels, including four luminance matrices, two chrominance matrices, and one auxiliary input matrix. In this case, the process DeriveInputTensors() will derive each of these 7 channels of the input tensor one by one, and when a particular channel of these channels is processed, that channel is referred to as the current channel during the process.
[0113] `nnpfc_inp_format_idc` specifies the method for converting the sample values of the cropped, decoded output image into the input values for NNPF. When `nnpfc_inp_format_idc` equals 0, the input values to NNPF are real numbers, and the functions `InpY()` and `InpC()` are defined as follows: InpY(x)=x÷((1< <BitDepth Y )-1) (77) InpC(x)=x÷((1< <BitDepth C )-1) (78)
[0114] When nnpfc_inp_format_idc equals 1, the input value to NNPF is an unsigned integer, and the functions InpY() and InpC() are defined as follows:
[0115] The variable inpTensorBitDepthY is inferred from the syntax element nnpfc_inp_tensor_luma_bitdepth_minus8, as specified below. The variable inpTensorBitDepthC is inferred from the syntax element nnpfc_inp_tensor_chroma_bitdepth_minus8, as specified below.
[0116] Values greater than 1 for nnpfc_inp_format_idc are reserved for future ITU-T|ISO / IEC specifications and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages containing reserved values for nnpfc_inp_format_idc.
[0117] `nnpfc_inp_tensor_luma_bitdepth_minus8` plus 8 specifies the bit depth of the luminance sample values in the input integer tensor. The value of `inpTensorBitDepthY` is derived as follows: inpTensorBitDepth Y =nnpfc_inp_tensor_luma_bitdepth_minus8+8 (81)
[0118] The requirement for bitstream consistency is that the value of nnpfc_inp_tensor_luma_bitdepth_minus8 should be in the range of 0 to 24 (inclusive).
[0119] `nnpfc_inp_tensor_chroma_bitdepth_minus8` plus 8 specifies the bit depth of the chroma sample values in the input integer tensor. The value of `inpTensorBitDepthC` is derived as follows: inpTensorBitDepth C =nnpfc_inp_tensor_chroma_bitdepth_minus8+8 (82)
[0120] The requirement for bitstream consistency is that the value of nnpfc_inp_tensor_chroma_bitdepth_minus8 should be in the range of 0 to 24 (inclusive).
[0121] nnpfc_inp_order_idc indicates a method for sorting the sample arrays of the cropped, decoded output image into one of the input images for NNPF.
[0122] In this version of the bitstream conforming to this document, the value of nnpfc_inp_order_idc should be in the range of 0 to 3 (inclusive). Values of nnpfc_inp_order_idc from 4 to 255 (inclusive) are reserved for future use by ITU-T|ISO / IEC and should not exist in this version of the bitstream conforming to this document. Decoders conforming to this version of this document should ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 4 to 255 (inclusive). Values of nnpfc_inp_order_idc greater than 255 should not exist in this version of the bitstream conforming to this document and are not reserved for future use.
[0123] When ChromaFormatIdc is not equal to 1, nnpfc_inp_order_idc should not be equal to 3.
[0124] Table 21 contains informative descriptions of the nnpfc_inp_order_idc values. Table 21 - Description of nnpfc_inp_order_idc values
[0125] Figure 4 This example shows how to derive four luminance channels (right) from the luminance component (left) when nnpfc_inp_order_idc equals 3.
[0126] A patch is a rectangular array of samples from the components of an image (e.g., luminance or chrominance components).
[0127] A value greater than 0 for nnpfc_auxiliary_inp_idc indicates that the auxiliary input data exists in the input tensor of NNPF. A value equal to 0 for nnpfc_auxiliary_inp_idc indicates that the auxiliary input data does not exist in the input tensor. A value equal to 1 for nnpfc_auxiliary_inp_idc indicates that the auxiliary input data is derived according to Equation 84.
[0128] In this version of the bitstream conforming to this document, the value of nnpfc_auxiliary_inp_idc should be in the range of 0 to 1 (inclusive). Values of nnpfc_inp_order_idc from 2 to 255 (inclusive) are reserved for future use by ITU-T|ISO / IEC and should not exist in this version of the bitstream conforming to this document. Decoders conforming to this version of this document should ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 2 to 255 (inclusive). Values of nnpfc_inp_order_idc greater than 255 should not exist in this version of the bitstream conforming to this document and are not reserved for future use.
[0129] When nnpfc_auxiliary_inp_idc equals 1, the variable strengthControlScaledVal is derived as follows:
[0130] The process DeriveInputTensors() is used to derive the input tensor inputTensor, which is given the vertical sample coordinates cTop and the horizontal sample coordinates cLeft for the top-left sample position of a block of samples included in the input tensor. The process DeriveInputTensors() is defined as follows:
[0131] `nnpfc_separate_colour_description_present_flag` equal to 1 indicates that the SEI message syntax structure specifies different combinations of color primaries, transmission characteristics, and matrix coefficients for the image generated by NNPF. `nnpfc_separate_colour_description_present_flag` equal to 0 indicates that the combination of color primaries, transmission characteristics, and matrix coefficients for the image generated by NNPF is the same as indicated in the CLVS VUI parameters.
[0132] nnpfc_colour_primaries has the same semantics as the vui_colour_primaries syntax element specified in sub-entry 7.3, except as follows: –nnpfc_colour_primaries specifies the primary color of the image generated by applying NNPF as defined in the SEI message, instead of the primary color used for CLVS. – When nnpfc_colour_primaries does not exist in the NNPFC SEI message, the value of nnpfc_colour_primaries is presumed to be equal to vui_colour_primaries.
[0133] nnpfc_transfer_characteristics has the same semantics as the vui_transfer_characteristics syntax element specified in sub-entry 7.3, except as follows: –nnpfc_transfer_characteristics specifies the transfer characteristics of images generated by applying NNPF as defined in the SEI message, rather than the transfer characteristics used for CLVS. – When nnpfc_transfer_characteristics does not exist in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is presumed to be equal to vui_transfer_characteristics.
[0134] nnpfc_matrix_coeffs has the same semantics as the vui_matrix_coeffs syntax element specified in sub-entry 7.3, except as follows: The –nnpfc_matrix_coeffs specification specifies the matrix coefficients of the image generated by applying NNPF in the SEI message, instead of the matrix coefficients used for CLVS. – When nnpfc_matrix_coeffs does not exist in the NNPFC SEI message, the value of nnpfc_matrix_coeffs is presumed to be equal to vui_matrix_coeffs. – The values allowed for nnpfc_matrix_coeffs are not constrained by the chroma format of the decoded video image as indicated by the value of ChromaFormatIdc, which is the semantics of the VUI parameter. – When nnpfc_matrix_coeffs is equal to 0, nnpfc_out_order_idc shall not be equal to 1 or 3.
[0135] nnpfc_out_format_idc being equal to 0 indicates that the sample values output by NNPF are real numbers, where the value range from 0 to 1 (including the boundary values) is linearly mapped to the unsigned integer value range from 0 to (1 << bitDepth) – 1 (including the boundary values) for any desired bit depth bitDepth for subsequent post - processing or display.
[0136] nnpfc_out_format_idc being equal to 1 indicates that the luminance sample values output by NNPF are unsigned integers within the range from 0 to (1 << (nnpfc_out_tensor_luma_bitdepth_minus8 + 8)) - 1 (including the boundary values), and the chrominance sample values output by NNPF are unsigned integers within the range from 0 to (1 << (nnpfc_out_tensor_chroma_bitdepth_minus8 + 8)) - 1 (including the boundary values).
[0137] Values of nnpfc_out_format_idc greater than 1 are reserved for future specifications of ITU - T|ISO / IEC and shall not be present in the bitstream compliant with this version of the document. The decoder compliant with this version of the document shall ignore the NNPFC SEI message containing the reserved value of nnpfc_out_format_idc.
[0138] nnpfc_out_tensor_luma_bitdepth_minus8 plus 8 specifies the bit depth of the luminance sample values in the output integer tensor. The value of nnpfc_out_tensor_luma_bitdepth_minus8 shall be within the range from 0 to 24 (including the boundary values).
[0139] nnpfc_out_tensor_chroma_bitdepth_minus8 plus 8 specifies the bit depth of the chrominance sample values in the output integer tensor. The value of nnpfc_out_tensor_chroma_bitdepth_minus8 shall be within the range from 0 to 24 (including the boundary values).
[0140] When nnpfc_purpose & 0x10 is not equal to 0, the value of nnpfc_out_format_idc shall be equal to 1, and at least one of the following conditions shall be true: -nnpfc_out_tensor_luma_bitdepth_minus8+8 is greater than BitDepth Y . -nnpfc_out_tensor_chroma_bitdepth_minus8+8 is greater than BitDepth C .
[0141] nnpfc_out_order_idc indicates the output order of samples generated by NNPF.
[0142] In this version of the bitstream conforming to this document, the value of nnpfc_out_order_idc should be in the range of 0 to 3 (inclusive). Values of nnpfc_out_order_idc from 4 to 255 (inclusive) are reserved for future use by ITU-T|ISO / IEC and should not exist in this version of the bitstream conforming to this document. Decoders conforming to this version of this document should ignore NNPFC SEI messages with nnpfc_out_order_idc in the range of 4 to 255 (inclusive). Values of nnpfc_out_order_idc greater than 255 should not exist in this version of the bitstream conforming to this document and are not reserved for future use.
[0143] When nnpfc_purpose&0x02 is not equal to 0, nnpfc_out_order_idc should not be equal to 3.
[0144] Table 22 contains informative descriptions of the nnpfc_out_order_idc values. Table 22 - Description of nnpfc_out_order_idc values
[0145] The process StoreOutputTensors() is used to derive sample values from the output tensor outputTensor, which is a filtered array of output samples FilteredYPic, FilteredCbPic, and FilteredCrPic. The output tensor outputTensor is defined with respect to the given vertical sample coordinates cTop and horizontal sample coordinates cLeft of the top-left sample position of a block of samples included in the input tensor. The process StoreOutputTensors() is defined as follows:
[0146] nnpfc_overlap indicates the horizontal and vertical sample counts of overlap between adjacent input tensors in NNPF. The value of nnpfc_overlap should be in the range of 0 to 16383 (inclusive).
[0147] The nnpfc_constant_patch_size_flag setting being equal to 1 indicates that NNPF accepts the precise patch size as input, as indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. The nnpfc_constant_patch_size_flag setting to 0 indicates that NNPF accepts any patch size with a width of inpPatchWidth and a height of inpPatchHeight as input, such that the width of the extended patch (i.e., the patch plus the overlapping area) (which is equal to inpPatchWidth + 2 * nnpfc_overlap) is a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, and the height of the extended patch (which is equal to inpPatchHeight + 2 * nnpfc_overlap) is a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap.
[0148] Incrementing nnpfc_patch_width_minus1 by 1 indicates the horizontal sample count for the required patch size for the NNPF input, when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_width_minus1 should be in the range of 0 to Min(32766, CroppedWidth-1) (inclusive).
[0149] Incrementing nnpfc_patch_height_minus1 by 1 indicates the vertical sample count for the required patch size for the NNPF input, when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_height_minus1 should be in the range of 0 to Min(32766, CroppedHeight-1) (inclusive).
[0150] `nnpfc_extended_patch_width_cd_delta_minus1` plus 1 plus 2 * `nnpfc_overlap` indicates the common divisor of all allowed values for the width of the extended patch required for the NNPF input, when `nnpfc_constant_patch_size_flag` is equal to 0. The value of `nnpfc_extended_patch_width_cd_delta_minus1` should be in the range of 0 to Min(32766, CroppedWidth-1) (inclusive).
[0151] `nnpfc_extended_patch_height_cd_delta_minus1` plus 1 plus 2 * `nnpfc_overlap` indicates the common divisor of all allowed values for the height of the extended patch required for the NNPF input, when `nnpfc_constant_patch_size_flag` is equal to 0. The value of `nnpfc_extended_patch_height_cd_delta_minus1` should be in the range of 0 to `Min(32766, CroppedHeight-1)` (inclusive of boundary values).
[0152] Let the variables inpPatchWidth and inpPatchHeight be the width and height of the small block, respectively.
[0153] If nnpfc_constant_patch_size_flag equals 0, then the following applies: The values of -inpPatchWidth and inpPatchHeight are provided by external means not specified in this document, or set by the post-processor itself. The value of -inpPatchWidth+2*nnpfc_overlap should be a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1+1+2*nnpfc_overlap, and inpPatchWidth should be less than or equal to CroppedWidth. The value of inpPatchHeight+2*nnpfc_overlap should be a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1+1+2*nnpfc_overlap, and inpPatchHeight should be less than or equal to CroppedHeight.
[0154] Otherwise (nnpfc_constant_patch_size_flag equals 1), the value of inpPatchWidth is set to equal to nnpfc_patch_width_minus1+1, and the value of inpPatchHeight is set to equal to nnpfc_patch_height_minus1+1.
[0155] The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, and outPatchCHeight are derived as follows: outPatchWidth=(nnpfc_pic_width_in_luma_samples*inpPatchWidth) / CroppedWidth (86) outPatchHeight=(nnpfc_pic_height_in_luma_samples*inpPatchHeight) / CroppedHeight (87) horCScaling=SubWidthC / outSubWidthC (88) verCScaling=SubHeightC / outSubHeightC (89) outPatchCWidth=outPatchWidth*horCScaling (90) outPatchCHeight=outPatchHeight*verCScaling (91)
[0156] The requirement for bitstream consistency is that outPatchWidth * CroppedWidth should be equal to nnpfc_pic_width_in_luma_samples * inpPatchWidth, and outPatchHeight * CroppedHeight should be equal to nnpfc_pic_height_in_luma_samples * inpPatchHeight.
[0157] The nnpfc_padding_type indicates the padding process when referencing sample locations outside the boundaries of the cropped, decoded output image, as described in Table 23. The value of nnpfc_padding_type should be in the range of 0 to 15 (inclusive). Table 23 - Informative Description of nnpfc_padding_type Values
[0158] nnpfc_luma_padding_val indicates the luminance value to be used for padding when nnpfc_padding_type is equal to 4.
[0159] nnpfc_cb_padding_val indicates the Cb value to be used for padding when nnpfc_padding_type is equal to 4.
[0160] nnpfc_cr_padding_val indicates the Cr value to be used for padding when nnpfc_padding_type is equal to 4.
[0161] The function InpSampleVal(y,x,picHeight,picWidth,croppedPic) takes the vertical sample position y, the horizontal sample position x, the image height picHeight, the image width picWidth, and the sample array croppedPic as input, and returns the value of sampleVal derived as follows:
[0162] Note 6 – For the input of the function InpSampleVal(), the vertical position is listed before the horizontal position to be compatible with the input tensor conventions of some inference engines.
[0163] The following example procedure can be used with NNPF PostProcessingFilter() to generate (multiple) filtered and / or interpolated images in small chunks, containing Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, as indicated by nnpfc_out_order_idc:
[0164] The order of the images in the stored output tensor is the output order, and the output order generated by applying NNPF in the output order is interpreted as the output order (and does not conflict with the output order of the input images).
[0165] A value of 1 for nnpfc_complexity_info_present_flag indicates the existence of one or more syntax elements that indicate the complexity of the NNPF associated with nnpfc_id. A value of 0 for nnpfc_complexity_info_present_flag indicates the absence of a syntax element that indicates the complexity of the NNPF associated with nnpfc_id.
[0166] An `nnpfc_parameter_type_idc` value of 0 indicates that the neural network uses only integer parameters. An `nnpfc_parameter_type_flag` value of 1 indicates that the neural network can use either floating-point or integer parameters. An `nnpfc_parameter_type_idc` value of 2 indicates that the neural network uses only binary parameters. An `nnpfc_parameter_type_idc` value of 3 is reserved for future use by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages with `nnpfc_parameter_type_idc` equal to 3.
[0167] The values 0, 1, 2, and 3 for nnpfc_log2_parameter_bit_length_minus3 indicate that the neural network does not use parameters with bit lengths greater than 8, 16, 32, and 64, respectively. When nnpfc_parameter_type_idc exists and nnpfc_log2_parameter_bit_length_minus3 does not exist, the neural network does not use parameters with a bit length greater than 1.
[0168] `nnpfc_num_parameters_idc` indicates the maximum number of neural network parameters for NNPF, in powers of 2048. `nnpfc_num_parameters_idc` equal to 0 indicates that the maximum number of neural network parameters is unknown. The value of `nnpfc_num_parameters_idc` should be in the range of 0 to 52 (inclusive). Values of `nnpfc_num_parameters_idc` greater than 52 are reserved for future use by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this version of this document. Decoders conforming to this version of this document should ignore NNPFCSEI messages with `nnpfc_num_parameters_idc` greater than 52.
[0169] If the value of nnpfc_num_parameters_idc is greater than zero, then the variable maxNumParameters is deduced as follows: maxNumParameters = (2048 < <nnpfc_num_parameters_idc)-1 (94)
[0170] The requirement for bitstream consistency is that the number of neural network parameters in NNPF should be less than or equal to maxNumParameters.
[0171] A value greater than 0 for nnpfc_num_kmac_operations_idc indicates that the maximum number of multiply-accumulate operations per sample in NNPF is less than or equal to nnpfc_num_kmac_operations_idc * 1000. A value of 0 for nnpfc_num_kmac_operations_idc indicates that the maximum number of multiply-accumulate operations in the network is unknown. The value of nnpfc_num_kmac_operations_idc should be in the range of 0 to 2^32 - 2 (inclusive).
[0172] `nnpfc_total_kilobyte_size` greater than 0 indicates the total size, in kilobytes, required to store the uncompressed parameters of the neural network. The total size, in bits, is equal to or greater than the sum of the bits used to store each parameter. `nnpfc_total_kilobyte_size` is the total size in bits divided by 8000 and rounded to the nearest integer. `nnpfc_total_kilobyte_size` equal to 0 indicates that the total size required to store the neural network parameters is unknown. The value of `nnpfc_total_kilobyte_size` should be in the range of 0 to 2^32 - 2 (inclusive).
[0173] nnpfc_reserved_zero_bit_b should be equal to 0 in the bitstream conforming to this version of the document. The decoder should ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_b is not equal to 0.
[0174] `nnpfc_payload_byte[i]` contains the i-th byte of a bitstream conforming to ISO / IEC 15938-17. The sequence of bytes for all current values of `i`, `nnpfc_payload_byte[i]`, should be a complete bitstream conforming to ISO / IEC 15938-17. 8.29 Neural Network Post-Processing Filter Activation SEI Message 8.29.1 Neural Network Post-Processing Filter Activation SEI Message Syntax 8.29.2 Neural Network Post-Processing Filter Activation of SEI Message Semantics
[0175] The Neural Network Post-Processing Filter Activation (NNPFA) SEI message activates or deactivates a target Neural Network Post-Processing Filter (NNPF) identified by nnpfa_target_id, which may be used for post-processing filtering of a set of images. For a specific image to which the NNPF is activated, the target NNPF is defined by the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id, preceding the first VCL NAL unit in the current image in decoding order, and is not a repetition of an NNPFC SEI message containing the basic NNPF.
[0176] Multiple NNPFA SEI messages can exist for the same image, for example, when NNPF is used for different purposes or for filtering different color components.
[0177] nnpfa_target_id indicates the target NNPF, which is specified by one or more NNPFC SEI messages that are related to the current image and whose nnpfc_id is equal to nnpfa_target_id.
[0178] The value of nnpfa_target_id should be in the range of 0 to 2^32-2 (inclusive). Values of nnpfa_target_id from 256 to 511 (inclusive) and from 231 to 232-2 (inclusive) are reserved for future use by ITU-T|ISO / IEC. Decoders conforming to this document should ignore NNPFA SEI messages when they encounter nnpfa_target_id values in the range of 256 to 511 (inclusive) or 231 to 232-2 (inclusive).
[0179] NNPFA SEI messages with a specific value of nnpfa_target_id should not exist in the current PU unless one or both of the following conditions are true: - Within the current CLVS, there are NNPFC SEI messages in the PUs preceding the current PU, in the decoding order, where the nnpfc-id is equal to a specific value of nnpfa_target_id. - There is an NNPFC SEI message in the current PU with a specific value of nnpfc_id equal to nnpfa_target_id.
[0180] When a PU contains both an NNPFC SEI message with a specific value of nnpfc_id and an NNPFA SEI message with nnpfa_target_id equal to a specific value of nnpfc_id, the NNPFC SEI message should precede the NNPFA SEI message in the decoding order.
[0181] A `nnpfa_cancel_flag` of 1 indicates that the persistence of the target NNPF established by any previous NNPFA SEI message with the same `nnpfa_target_id` as the current SEI message is cancelled; that is, the target NNPF is no longer used unless it is activated by another NNPFA SEI message with the same `nnpfa_target_id` as the current SEI message and `nnpfa_cancel_flag` equal to 0. A `nnpfa_cancel_flag` of 0 indicates that `nnpfa_persistence_flag` follows.
[0182] The nnpfa_persistence_flag specifies the persistence of the target NNPF for the current layer.
[0183] Setting nnpfa_persistence_flag to 0 specifies that the target NNPF is used only for post-processing filtering of the current image.
[0184] The nnpfa_persistence_flag setting to 1 specifies that the target NNPF is used for post-processing filtering of the current image and all subsequent images of the current layer (in output order) until one or more of the following conditions are true: - A new CLVS begins for the current layer. -End of bitstream. - Output the images in the current layer that are associated with the NNPFA SEI message that has the same nnpfa_target_id as the current SEI message and an nnpfa_cancel_flag equal to 1, in the order of output.
[0185] Note 2 – Do not apply the target NNPF to the subsequent picture in the current layer that is associated with an NNPFA SEI message that has the same nnpfa_target_id as the current SEI message and an nnpfa_cancel_flag equal to 1.
[0186] Let nnpfcTargetPictures be the set of pictures associated with the last NNPFC SEI message whose nnpfc_id equals nnpfa_target_id, preceding the current NNPFA SEI message in decoding order. Let nnpfaTargetPictures be the set of pictures of the target NNPF activated by the current NNPFA SEI message. The requirement for bitstream consistency is that any pictures included in nnpfaTargetPictures should also be included in nnpfcTargetPictures. 4. The technical problem solved by the disclosed technical solution
[0187] The example design for the Neural Network Post-Processing Filter (NNPFC) SEI message has the following problems:
[0188] First, the purpose of a neural network post-processing filter with resolution upsampling capability is defined. However, in video applications, resolution downsampling is often required, and similar to resolution upsampling, using a neural network filter for resolution downsampling can provide better performance than using traditional methods. Therefore, there is a need for a neural network filter capable of resolution downsampling through signal transmission.
[0189] Second, the purpose of a neural network post-processing filter with bit-depth upsampling capability is defined. However, in video applications, bit-depth downsampling is often required, and similar to bit-depth upsampling, using a neural network filter for bit-depth downsampling can provide better performance than traditional methods. Therefore, a neural network filter capable of bit-depth downsampling through signal transmission is needed.
[0190] Third, the purpose of a neural network post-processing filter with chroma upsampling capability is defined. However, in video applications, chroma downsampling is often required, and similar to chroma upsampling, using a neural network filter for chroma downsampling can provide better performance than traditional methods. Therefore, a neural network filter capable of chroma downsampling over signal transmission is needed.
[0191] Fourth, the purpose of a neural network post-processing filter with image rate upsampling capability is defined. However, in video applications, image rate downsampling is often required, and similar to image rate upsampling, using a neural network filter for image rate downsampling can provide better performance than traditional methods. Therefore, a neural network filter capable of image rate downsampling through signal transmission is needed. 5. List of solutions and implementation examples
[0192] To address the aforementioned issues, the methods outlined below are disclosed. These aspects should be considered as examples for interpreting general concepts, and not interpreted in a narrow sense. Furthermore, these examples can be applied individually or combined in any way. 1) To solve problem 1, one or more of the following new objectives are defined: a. In one example, a new purpose for resolution downsampling is defined, wherein one or both of the width and height of the cropped decoded output image are reduced. i. In one example, it is further required that the purpose should not include both resolution upsampling and resolution downsampling. ii. In one example, the downsampling rate is required to be limited to the range of minR to maxR. 1. In one example, minR equals 1, 1 / 2, 1 / 4, 1 / 8, and 1 / 16. 2. In one example, maxR equals 1. b. In one example, a new purpose for resolution resampling is defined, wherein one or both of the width and height of the cropped decoded output image are reduced or increased. i. In one example, the resampling / scaling ratio is required to be limited to the range of minR to maxR. 1. In one example, minR equals 1, 1 / 2, 1 / 4, 1 / 8, and 1 / 16. 2. In one example, maxR equals 1, 2, 4, 8, 16. 2) To solve problem 2, one or more of the following new objectives are defined: a. In one example, a new objective for bit depth downsampling is defined, wherein one or both of the luminance bit depth and chrominance bit depth of the cropped decoded output image are reduced. i. In one example, it is further required that the purpose should not include both bit depth upsampling and bit depth downsampling. b. In one example, a new purpose for bit depth resampling is defined, wherein one or both of the luminance bit depth and chrominance bit depth of the cropped decoded output image are reduced or increased. 3) To address problem 3, one or more of the following new objectives are defined: a. In one example, a new purpose for chroma format downsampling is defined, wherein the chroma format is changed from 4:4:4 chroma format to 4:2:2 or 4:2:0 chroma format, or from 4:2:2 chroma format to 4:2:0 chroma format. i. In one example, it is further required that the purpose should not include both chroma format upsampling and chroma format downsampling. b. In one example, a new purpose for chroma format downsampling is defined, wherein the chroma format is changed from 4:4:4 chroma format to 4:0:0 chroma format, or from 4:2:2 chroma format to 4:0:0 chroma format, or from 4:2:0 chroma format to 4:0:0 chroma format. c. In one example, a new purpose for chroma format resampling is defined, wherein the chroma format is changed from 4:4:4 chroma format to 4:2:2 or 4:2:0 chroma format, or from 4:2:2 chroma format to 4:2:0 chroma format, or from 4:2:0 chroma format to 4:2:2 or 4:4:4 chroma format, or from 4:2:2 chroma format to 4:4:4 chroma format. d. In one example, a new purpose for chroma format resampling is defined, wherein the chroma format is changed from 4:4:4 to 4:2:2, 4:2:0, or 4:0:0, or from 4:2:2 to 4:2:0 or 4:0:0, or from 4:0:0 to 4:2:0, 4:2:2, or 4:4:4, or from 4:2:0 to 4:2:2 or 4:4:4, or from 4:2:2 to 4:4:4. 4) To address problem 4, one or more of the following new objectives are defined: a. In one example, a new purpose for image rate downsampling is defined. i. In one example, it is further required that the objective should not include both image rate upsampling and image rate downsampling. b. In one example, a new purpose for image rate resampling is defined. 6. Example
[0193] The following are some example implementations of the aspects outlined in Section 5 of the previous article.
[0194] Most of the relevant sections that have been added or modified are shown in bold, and some of the deleted sections are shown in both bold and italic fonts. Other changes that may have been editable are not indicated. 6.1 Example 1
[0195] This embodiment refers to items 1, 2, 3, and 4 and all their sub-items as outlined in Section 5 of the previous article. 8.28.1 Characteristics of Neural Network Post-Processing Filters and SEI Message Syntax 8.28.2 Characteristics of Neural Network Post-Processing Filters and Semantics of SEI Messages The Neural Network Post-Processing Filter Characteristics (NNPFC) SEI message specifies the neural networks that can be used as post-processing filters. The use of the specified Neural Network Post-Processing Filter (NNPF) for a specific image is indicated by the Neural Network Post-Processing Filter Activation (NNPFA) SEI message. To use this SEI message, the following variables need to be specified: - Input the image width and height, in units of brightness samples, which are represented as CroppedWidth and CroppedHeight in this article. - The luminance sample array CroppedYPic[idx] and chrominance sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] (if present) of the input image, with index idx in the range of 0 to numInputPics-1 (inclusive), are used as input for NNPF. - BitDepth for the luminance sample array of the input image Y . - Bit depth of the chroma sample array (if any) for the input image. C . - Chroma format indicator, referred to herein as ChromaFormatIdc, as described in sub-entry 7.3. - When nnpfc_auxiliary_inp_idc equals 1, the filter strength control value StrengthControlVal should be a real number in the range of 0 to 1 (including boundary values). The input image with index 0 corresponds to the image for which the NNPF defined by the NNPFC SEI message is activated via the NNPFA SEI message. The input images with index i (in the range of 1 to numInputPics-1 (inclusive)) precede the input images with index i-1 in the output order. When an input image with index 0 and nnpfc_purpose&0x08 is not equal to 0 is associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5, all input images are associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5 and the same value as fp_current_frame_is_frame0_flag. The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc, as specified in Table 2. Note 1 – More than one NNPFC SEI message can exist for the same image. When more than one NNPFC SEI message with different values of nnpfc_id exists or is activated for the same image, they can have the same or different values of nnpfc_purpose and nnpfc_mode_idc. nnpfc_purpose indicates the purpose of NNPF, as specified in Table 20. In a bitstream conforming to this version of the document, the value of nnpfc_purpose should be between 0 and... The range (including boundary values). The value of nnpfc_purpose Values up to 65535 (including boundary values) are reserved for future use by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this document. Decoders conforming to this document should ignore nnpfc_purpose in this version. NNPFC SEI messages in the range of 65,535 (inclusive). Table 20 – Definition of nnpfc_purpose Note 2 – When the reserved value of nnpfc_purpose is used by ITU-T|ISO / IEC in the future, the syntax of this SEI message can be extended using the following syntax elements, provided that nnpfc_purpose is equal to that value. When ChromaFormatIdc equals 3, nnpfc_purpose&0x02 should equal 0. When ChromaFormatIdc or nnpfc_purpose&0x02 is not equal to 0, nnpfc_purpose&0x20 should be equal to 0. nnpfc_id contains an identifier that can be used to identify NNPF. The value of nnpfc_id should be between 0 and 2. 32 The range is -2 (inclusive of boundary values). The values of nnpfc_id are 256 to 511 (inclusive of boundary values) and 2. 31 to 2 32 -2 (including boundary values) is reserved for future use by ITU-T|ISO / IEC. Decoders conforming to this document for this version encounter nnpfc_id in the range of 256 to 511 (including boundary values) or 2. 31 to 2 32 When an NNPFC SEI message is received within the range of -2 (including boundary values), the SEI message should be ignored. The following applies when an NNPFC SEI message has a specific nnpfc_id value within the current CLVS and is the first NNPFC SEI message in decoding order: - This SEI message specifies the basic NNPF. - This SEI message applies to the currently decoded image and all subsequent decoded images of the current layer (in output order) until the end of the current CLVS. ... When `nnpfc_purpose&0x20` is not equal to 0, `nnpfc_out_colour_format_idc` specifies the color format of the NNPF output, thus defining the values of the variables `outSubWidthC` and `outSubHeightC`. `nnpfc_out_colour_format_idc` equal to 1 specifies that the NNPF output color format is 4:2:0, and both `outSubWidthC` and `outSubHeightC` are equal to 2. `nnpfc_out_colour_format_idc` equal to 2 specifies that the NNPF output color format is 4:2:2, and both `outSubWidthC` and `outSubHeightC` are equal to 1. `nnpfc_out_colour_format_idc` equal to 3 specifies that the NNPF output color format is 4:2:4, and both `outSubWidthC` and `outSubHeightC` are equal to 1. The value of `nnpfc_out_colour_format_idc` should not be equal to 0. When nnpfc_purpose&0x02 and nnpfc_purpose&0x20 When both are equal to 0, outSubWidthC and outSubHeightC are presumed to be equal to SubWidthC and SubHeightC, respectively. nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples respectively specify the width and height of the luma sample array of the picture generated by applying the NNPF identified by nnpfc_id to the cropped decoded output picture. When nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples do not exist, they are presumed to be equal to CroppedWidth and CroppedHeight, respectively. The value of nnpfc_pic_width_in_luma_samples shall be in the range of CroppedWidth>>4 to CroppedWidth*16 - 1 (including the boundary values). The value of nnpfc_pic_height_in_luma_samples shall be in the range of CroppedHeight>>4 to CroppedHeight*16 - 1 (including the boundary values). nnpfc_num_input_pics_minus1 plus 1 specifies the number of decoded output pictures used as inputs for the NNPF. The value of nnpfc_num_input_pics_minus1 shall be in the range of 0 to 63 (including the boundary values). When nnpfc_purpose&0x08 is not equal to 0, the value of nnpfc_num_input_pics_minus1 shall be greater than 0. ... nnpfc_out_format_idc being equal to 0 indicates that the sample values output by the NNPF are real numbers, where the value range of 0 to 1 (including the boundary values) is linearly mapped to the unsigned integer value range of 0 to (1<<bitDepth) – 1 (including the boundary values) for any desired bit depth bitDepth for subsequent post - processing or display. An nnpfc_out_format_idc value of 1 indicates that the luminance sample values output by NNPF are unsigned integers in the range of 0 to (1 << (nnpfc_out_tensor_luma_bitdepth_minus8+8)) - 1 (inclusive of boundary values), and the chrominance sample values output by NNPF are unsigned integers in the range of 0 to (1 << (nnpfc_out_tensor_chroma_bitdepth_minus8+8)) - 1 (inclusive of boundary values). Values of nnpfc_out_format_idc greater than 1 are reserved for future ITU-T|ISO / IEC specifications and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages containing reserved values of nnpfc_out_format_idc. The value of nnpfc_out_tensor_luma_bitdepth_minus8 plus 8 specifies the bit depth of the luminance sample values in the output integer tensor. The value of nnpfc_out_tensor_luma_bitdepth_minus8 should be in the range of 0 to 24 (inclusive). `nnpfc_out_tensor_chroma_bitdepth_minus8` plus 8 specifies the bit depth of the chroma sample values in the output integer tensor. The value of `nnpfc_out_tensor_chroma_bitdepth_minus8` should be in the range of 0 to 24 (inclusive). When `nnpfc_purpose&0x10` is not equal to 0, the value of `nnpfc_out_format_idc` should be equal to 1, and at least one of the following conditions should be true: -nnpfc_out_tensor_luma_bitdepth_minus8+8 is greater than BitDepth Y , -nnpfc_out_tensor_chroma_bitdepth_minus8+8 is greater than BitDepth C , nnpfc_out_order_idc indicates the output order of samples generated by NNPF. In this version of the bitstream conforming to this document, the value of nnpfc_out_order_idc should be in the range of 0 to 3 (inclusive). Values of nnpfc_out_order_idc from 4 to 255 (inclusive) are reserved for future use by ITU-T|ISO / IEC and should not exist in this version of the bitstream conforming to this document. Decoders conforming to this version of this document should ignore NNPFC SEI messages with nnpfc_out_order_idc in the range of 4 to 255 (inclusive). Values of nnpfc_out_order_idc greater than 255 should not exist in this version of the bitstream conforming to this document and are not reserved for future use. ... 6.2 Example 2
[0196] This embodiment refers to items 1, 2, 3, and 4 and all their sub-items as outlined in Section 5 of the previous article. 8.28.1 Characteristics of Neural Network Post-Processing Filters and SEI Message Syntax 8.28.2 Characteristics of Neural Network Post-Processing Filters and Semantics of SEI Messages The Neural Network Post-Processing Filter Characteristics (NNPFC) SEI message specifies the neural networks that can be used as post-processing filters. The use of the specified Neural Network Post-Processing Filter (NNPF) for a specific image is indicated by the Neural Network Post-Processing Filter Activation (NNPFA) SEI message. To use this SEI message, the following variables need to be specified: - Input the image width and height, in units of brightness samples, which are represented as CroppedWidth and CroppedHeight in this article. - The luminance sample array CroppedYPic[idx] and chrominance sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] (if present) of the input image, with index idx in the range of 0 to numInputPics-1 (inclusive), are used as input for NNPF. - BitDepth for the luminance sample array of the input image Y . - Bit depth of the chroma sample array (if any) for the input image. C . - Chroma format indicator, referred to herein as ChromaFormatIdc, as described in sub-entry 7.3. - When nnpfc_auxiliary_inp_idc equals 1, the filter strength control value StrengthControlVal should be a real number in the range of 0 to 1 (inclusive). The input image with index 0 corresponds to the image for which the NNPF defined by the NNPFC SEI message is activated via the NNPFA SEI message. The input images with index i (in the range of 1 to numInputPics-1 (inclusive)) precede the input images with index i-1 in the output order. When an input image with index 0 and nnpfc_purpose&0x08 is not equal to 0 is associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5, all input images are associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5 and the same value as fp_current_frame_is_frame0_flag. The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc, as specified in Table 2. Note 1 – More than one NNPFC SEI message can exist for the same image. When more than one NNPFC SEI message with different values of nnpfc_id exists or is activated for the same image, they can have the same or different values of nnpfc_purpose and nnpfc_mode_idc. nnpfc_purpose indicates the purpose of NNPF, as specified in Table 20. In bitstreams conforming to this version of the document, the value of nnpfc_purpose should be in the range of 0 to 63 (inclusive). Values of nnpfc_purpose from 64 to 65535 (inclusive) are reserved for future use by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages with nnpfc_purpose in the range of 64 to 65535 (inclusive). Table 20 – Definition of nnpfc_purpose Note 2 – When the reserved value of nnpfc_purpose is used by ITU-T|ISO / IEC in the future, the syntax of this SEI message can be extended using the following syntax elements, provided that nnpfc_purpose is equal to that value. When ChromaFormatIdc or nnpfc_purpose&0x02 is not equal to 0, nnpfc_purpose&0x20 should be equal to 0. nnpfc_id contains an identifier that can be used to identify NNPF. The value of nnpfc_id should be between 0 and 2. 32 The range is -2 (inclusive of boundary values). The values of nnpfc_id are 256 to 511 (inclusive of boundary values) and 2. 31 to 2 32 -2 (including boundary values) is reserved for future use by ITU-T|ISO / IEC. Decoders conforming to this document for this version encounter nnpfc_id in the range of 256 to 511 (including boundary values) or 2. 31 to 2 32 When an NNPFC SEI message is received within the range of -2 (including boundary values), the SEI message should be ignored. The following applies when the NNPFC SEI message is the first NNPFC SEI message in decoding order with a specific nnpfc_id value within the current CLVS: - This SEI message specifies the basic NNPF. - This SEI message applies to the currently decoded image and all subsequent decoded images of the current layer (in output order) until the current CLVS ends. ... When `nnpfc_purpose&0x20` is not equal to 0, `nnpfc_out_colour_format_idc` specifies the color format of the NNPF output, thus defining the values of the variables `outSubWidthC` and `outSubHeightC`. `nnpfc_out_colour_format_idc` equal to 1 specifies that the NNPF output color format is 4:2:0, and both `outSubWidthC` and `outSubHeightC` are equal to 2. `nnpfc_out_colour_format_idc` equal to 2 specifies that the NNPF output color format is 4:2:2, and both `outSubWidthC` and `outSubHeightC` are equal to 1. `nnpfc_out_colour_format_idc` equal to 3 specifies that the NNPF output color format is 4:2:4, and both `outSubWidthC` and `outSubHeightC` are equal to 1. The value of `nnpfc_out_colour_format_idc` should not be equal to 0. When both nnpfc_purpose&0x02 and nnpfc_purpose&0x20 are equal to 0, outSubWidthC and outSubHeightC are presumed to be equal to SubWidthC and SubHeightC, respectively. `nnpfc_pic_width_in_luma_samples` and `nnpfc_pic_height_in_luma_samples` specify the width and height of the luminance sample array of the image generated by applying the NNPF identified by `nnpfc_id` to the cropped, decoded output image, respectively. When `nnpfc_pic_width_in_luma_samples` and `nnpfc_pic_height_in_luma_samples` are not present, they are presumed to be equal to `CroppedWidth` and `CroppedHeight`, respectively. The value of `nnpfc_pic_width_in_luma_samples` should be in the range from `CroppedWidth >> 4` to `CroppedWidth * 16 - 1` (inclusive). The value of `nnpfc_pic_height_in_luma_samples` should be in the range from `CroppedHeight >> 4` to `CroppedHeight * 16 - 1` (inclusive). The nnpfc_num_input_pics_minus1 plus 1 rule is used as the number of decoded output pictures for the input to NNPF. The value of nnpfc_num_input_pics_minus1 should be in the range of 0 to 63 (including the boundary values). When nnpfc_purpose & 0x08 is not equal to 0, the value of nnpfc_num_input_pics_minus1 should be greater than 0. ... nnpfc_out_format_idc being equal to 0 indicates that the sample values output by NNPF are real numbers, where the value range from 0 to 1 (including the boundary values) is linearly mapped to the unsigned integer value range from 0 to (1 << bitDepth) – 1 (including the boundary values) for any desired bit depth bitDepth for subsequent post-processing or display. nnpfc_out_format_idc being equal to 1 indicates that the luminance sample values output by NNPF are unsigned integers in the range from 0 to (1 << (nnpfc_out_tensor_luma_bitdepth_minus8 + 8)) - 1 (including the boundary values), and the chrominance sample values output by NNPF are unsigned integers in the range from 0 to (1 << (nnpfc_out_tensor_chroma_bitdepth_minus8 + 8)) - 1 (including the boundary values). Values of nnpfc_out_format_idc greater than 1 are reserved for future ITU-T|ISO / IEC specifications and should not be present in the bitstream conforming to this version of the document. A decoder conforming to this version of the document should ignore NNPFC SEI messages containing reserved values of nnpfc_out_format_idc. nnpfc_out_tensor_luma_bitdepth_minus8 plus 8 specifies the bit depth of the luminance sample values in the output integer tensor. The value of nnpfc_out_tensor_luma_bitdepth_minus8 should be in the range of 0 to 24 (including the boundary values). nnpfc_out_tensor_chroma_bitdepth_minus8 plus 8 specifies the bit depth of the chrominance sample values in the output integer tensor. The value of nnpfc_out_tensor_chroma_bitdepth_minus8 should be in the range of 0 to 24 (including the boundary values). When nnpfc_purpose & 0x10 is not equal to 0, the value of nnpfc_out_format_idc should be equal to 1, and at least one of the following conditions should be true: -(nnpfc_out_tensor_luma_bitdepth_minus8+8 is greater than BitDepth) Y or (nnpfc_out_tensor_chroma_bitdepth_minus8+8 is greater than BitDepth) C nnpfc_out_order_idc indicates the output order of samples generated by NNPF. In this version of the bitstream conforming to this document, the value of nnpfc_out_order_idc should be in the range of 0 to 3 (inclusive). Values of nnpfc_out_order_idc from 4 to 255 (inclusive) are reserved for future use by ITU-T|ISO / IEC and should not exist in this version of the bitstream conforming to this document. Decoders conforming to this version of this document should ignore NNPFC SEI messages with nnpfc_out_order_idc in the range of 4 to 255 (inclusive). Values of nnpfc_out_order_idc greater than 255 should not exist in this version of the bitstream conforming to this document and are not reserved for future use. ... 6.3 Example 3
[0197] This embodiment refers to items 1, 2, 3, and 4 and all their sub-items as outlined in Section 5 of the previous article. 8.28.1 Characteristics of Neural Network Post-Processing Filters and SEI Message Syntax 8.28.2 Characteristics of Neural Network Post-Processing Filters and Semantics of SEI Messages The Neural Network Post-Processing Filter Characteristics (NNPFC) SEI message specifies the neural networks that can be used as post-processing filters. The use of the specified Neural Network Post-Processing Filter (NNPF) for a specific image is indicated by the Neural Network Post-Processing Filter Activation (NNPFA) SEI message. To use this SEI message, the following variables need to be specified: - Input the image width and height, in units of brightness samples, which are represented as CroppedWidth and CroppedHeight in this article. - The luminance sample array CroppedYPic[idx] and chrominance sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] (if present) of the input image, with index idx in the range of 0 to numInputPics-1 (inclusive), are used as input for NNPF. - BitDepth for the luminance sample array of the input image Y . - Bit depth of the chroma sample array (if any) for the input image. C . - Chroma format indicator, referred to herein as ChromaFormatIdc, as described in sub-entry 7.3. - When nnpfc_auxiliary_inp_idc equals 1, the filter strength control value StrengthControlVal should be a real number in the range of 0 to 1 (inclusive). The input image with index 0 corresponds to the image for which the NNPF defined by the NNPFC SEI message is activated via the NNPFA SEI message. The input images with index i (in the range of 1 to numInputPics-1 (inclusive)) precede the input images with index i-1 in the output order. When an input image with index 0 and nnpfc_purpose&0x08 is not equal to 0 is associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5, all input images are associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5 and the same value as fp_current_frame_is_frame0_flag. The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc, as specified in Table 2. Note 1 – More than one NNPFC SEI message can exist for the same image. When more than one NNPFC SEI message with different values of nnpfc_id exists or is activated for the same image, they can have the same or different values of nnpfc_purpose and nnpfc_mode_idc. nnpfc_purpose indicates the purpose of NNPF, as specified in Table 20. In bitstreams conforming to this version of the document, the value of nnpfc_purpose should be in the range of 0 to 63 (inclusive). Values of nnpfc_purpose from 64 to 65535 (inclusive) are reserved for future use by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages with nnpfc_purpose in the range of 64 to 65535 (inclusive). Table 20 – Definition of nnpfc_purpose Note 2 – When the reserved value of nnpfc_purpose is used by ITU-T|ISO / IEC in the future, the syntax of this SEI message can be extended using the following syntax elements, provided that nnpfc_purpose is equal to that value. When ChromaFormatIdc or nnpfc_purpose&0x02 is not equal to 0, nnpfc_purpose&0x20 should be equal to 0. nnpfc_id contains an identifier that can be used to identify NNPF. The value of nnpfc_id should be between 0 and 2. 32 The range is -2 (inclusive of boundary values). The values of nnpfc_id are 256 to 511 (inclusive of boundary values) and 2. 31 to 2 32 -2 (including boundary values) is reserved for future use by ITU-T|ISO / IEC. Decoders conforming to this document for this version encounter nnpfc_id in the range of 256 to 511 (including boundary values) or 2. 31 to 2 32 When an NNPFC SEI message is received within the range of -2 (including boundary values), the SEI message should be ignored. The following applies when the NNPFC SEI message is the first NNPFC SEI message in decoding order with a specific nnpfc_id value within the current CLVS: - This SEI message specifies the basic NNPF. - This SEI message applies to the currently decoded image and all subsequent decoded images of the current layer (in output order) until the end of the current CLVS. ... When nnpfc_purpose & 0x20 is not equal to 0, nnpfc_out_colour_format_idc specifies the color format output by NNPF, thereby specifying the values of variables outSubWidthC and outSubHeightC. nnpfc_out_colour_format_idc being equal to 1 specifies that the color format output by NNPF is the 4:2:0 format, and both outSubWidthC and outSubHeightC are equal to 2. nnpfc_out_colour_format_idc being equal to 2 specifies that the color format output by NNPF is the 4:2:2 format, and outSubWidthC is equal to 2 and outSubHeightC is equal to 1. nnpfc_out_colour_format_idc being equal to 3 specifies that the color format output by NNPF is the 4:2:4 format, and both outSubWidthC and outSubHeightC are equal to 1. The value of nnpfc_out_colour_format_idc shall not be equal to 0. When both nnpfc_purpose & 0x02 and nnpfc_purpose & 0x20 are equal to 0, outSubWidthC and outSubHeightC are presumed to be equal to SubWidthC and SubHeightC, respectively. nnpfc_num_input_pics_minus1 plus 1 specifies the number of decoded output pictures used as inputs for NNPF. The value of nnpfc_num_input_pics_minus1 shall be in the range of 0 to 63 (including the boundary values). When nnpfc_purpose & 0x08 is not equal to 0, the value of nnpfc_num_input_pics_minus1 shall be greater than 0. ... nnpfc_out_format_idc being equal to 0 indicates that the sample values output by NNPF are real numbers, where the value range from 0 to 1 (including the boundary values) is linearly mapped to the unsigned integer value range from 0 to (1 << bitDepth) – 1 (including the boundary values) for any desired bit depth bitDepth for subsequent post - processing or display. An nnpfc_out_format_idc value of 1 indicates that the luminance sample values output by NNPF are unsigned integers in the range of 0 to (1 << (nnpfc_out_tensor_luma_bitdepth_minus8+8)) - 1 (inclusive of boundary values), and the chrominance sample values output by NNPF are unsigned integers in the range of 0 to (1 << (nnpfc_out_tensor_chroma_bitdepth_minus8+8)) - 1 (inclusive of boundary values). Values of nnpfc_out_format_idc greater than 1 are reserved for future specification by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages containing reserved values of nnpfc_out_format_idc. The value of nnpfc_out_tensor_luma_bitdepth_minus8 plus 8 specifies the bit depth of the luminance sample values in the output integer tensor. The value of nnpfc_out_tensor_luma_bitdepth_minus8 should be in the range of 0 to 24 (inclusive). `nnpfc_out_tensor_chroma_bitdepth_minus8` plus 8 specifies the bit depth of the chroma sample values in the output integer tensor. The value of `nnpfc_out_tensor_chroma_bitdepth_minus8` should be in the range of 0 to 24 (inclusive). When `nnpfc_purpose&0x10` is not equal to 0, the value of `nnpfc_out_format_idc` should be equal to 1, and at least one of the following conditions should be true: -(nnpfc_out_tensor_luma_bitdepth_minus8+8 is greater than BitDepth) Y or (nnpfc_out_tensor_chroma_bitdepth_minus8+8 is greater than BitDepth) C nnpfc_out_order_idc indicates the output order of samples generated by NNPF. In this version of the bitstream conforming to this document, the value of nnpfc_out_order_idc should be in the range of 0 to 3 (inclusive). Values of nnpfc_out_order_idc from 4 to 255 (inclusive) are reserved for future use by ITU-T|ISO / IEC and should not exist in this version of the bitstream conforming to this document. Decoders conforming to this version of this document should ignore NNPFC SEI messages with nnpfc_out_order_idc in the range of 4 to 255 (inclusive). Values of nnpfc_out_order_idc greater than 255 should not exist in this version of the bitstream conforming to this document and are not reserved for future use. When nnpfc_purpose&0x02 is not equal to 0, nnpfc_out_order_idc should not be equal to 3. Table 22 contains informative descriptions of the nnpfc_out_order_idc values. Table 22 – Description of nnpfc_out_order_idc values The process StoreOutputTensors() is used to derive sample values from the output tensor outputTensor, which is a filtered output sample array of FilteredYPic, FilteredCbPic, and FilteredCrPic. The output tensor outputTensor is defined with respect to the given vertical sample coordinates cTop and horizontal sample coordinates cLeft of the top-left sample position of a block of samples included in the input tensor. The process StoreOutputTensors() is defined as follows: `nnpfc_overlap` indicates the horizontal and vertical sample counts of overlap between adjacent input tensors in NNPF. The value of `nnpfc_overlap` should be in the range of 0 to 16,383 (inclusive). `nnpfc_constant_patch_size_flag` equal to 1 indicates that NNPF accepts the precise patch size indicated by `nnpfc_patch_width_minus1` and `nnpfc_patch_height_minus1` as input. The nnpfc_constant_patch_size_flag setting to 0 indicates that NNPF accepts any patch size with a width of inpPatchWidth and a height of inpPatchHeight as input, such that the width of the extended patch (i.e., the patch plus the overlapping area) (which is equal to inpPatchWidth + 2 * nnpfc_overlap) is a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, and the height of the extended patch (which is equal to inpPatchHeight + 2 * nnpfc_overlap) is a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap. Incrementing nnpfc_patch_width_minus1 by 1 indicates the horizontal sample count for the required patch size for the NNPF input, when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_width_minus1 should be in the range of 0 to Min(32766, CroppedWidth-1) (inclusive). Incrementing nnpfc_patch_height_minus1 by 1 indicates the vertical sample count for the required patch size for the NNPF input, when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_height_minus1 should be in the range of 0 to Min(32766, CroppedHeight-1) (inclusive). `nnpfc_extended_patch_width_cd_delta_minus1` plus 1 plus 2 * `nnpfc_overlap` indicates the common divisor of all allowed values for the width of the extended patch required for the NNPF input, when `nnpfc_constant_patch_size_flag` is equal to 0. The value of `nnpfc_extended_patch_width_cd_delta_minus1` should be in the range of 0 to Min(32766, CroppedWidth-1) (inclusive). `nnpfc_extended_patch_height_cd_delta_minus1` plus 1 plus 2 * `nnpfc_overlap` indicates the common divisor of all allowed values for the height of the extended patch required for the NNPF input, when `nnpfc_constant_patch_size_flag` is equal to 0. The value of `nnpfc_extended_patch_height_cd_delta_minus1` should be in the range of 0 to `Min(32766, CroppedHeight-1)` (inclusive of boundary values). This makes the variables inpPatchWidth and inpPatchHeight the width and height of the small block, respectively. If nnpfc_constant_patch_size_flag equals 0, then the following applies: The values of -inpPatchWidth and inpPatchHeight are provided by external means not specified in this document, or set by the post-processor itself. The value of `-inpPatchWidth+2*nnpfc_overlap` should be a positive integer multiple of `nnpfc_extended_patch_width_cd_delta_minus1+1+2*nnpfc_overlap`, and `inpPatchWidth` should be less than or equal to `CroppedWidth`. The value of `inpPatchHeight+2*nnpfc_overlap` should be a positive integer multiple of `nnpfc_extended_patch_height_cd_delta_minus1+1+2*nnpfc_overlap`, and `inpPatchHeight` should be less than or equal to `CroppedHeight`. Otherwise (where `nnpfc_constant_patch_size_flag` equals 1), the value of `inpPatchWidth` is set to equal to `nnpfc_patch_width_minus1+1`, and the value of `inpPatchHeight` is set to equal to `nnpfc_patch_height_minus1+1`. The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, and outPatchCHeight are derived as follows: horCScaling=SubWidthC / outSubWidthC (88) verCScaling=SubHeightC / outSubHeightC (89) outPatchCWidth=outPatchWidth*horCScaling (90) outPatchCHeight=outPatchHeight*verCScaling (91) The requirement for bitstream consistency is that outPatchWidth * CroppedWidth should equal... *inpPatchWidth, and outPatchHeight*CroppedHeight should equal *inpPatchHeight. The nnpfc_padding_type indicates the padding process when referencing sample locations outside the boundaries of the cropped, decoded output image, as described in Table 23. The value of nnpfc_padding_type should be in the range of 0 to 15 (inclusive). Table 23 – Informative description of nnpfc_padding_type values nnpfc_padding_type describe 0 Zero fill 1 Copy fill 2 Reflection fill 3 Surround fill 4 Fixed fill 5..15 reserve nnpfc_luma_padding_val indicates the luminance value to be used for padding when nnpfc_padding_type is equal to 4. nnpfc_cb_padding_val indicates the Cb value to be used for padding when nnpfc_padding_type is equal to 4. nnpfc_cr_padding_val indicates the Cr value to be used for padding when nnpfc_padding_type is equal to 4. The function InpSampleVal(y,x,picHeight,picWidth,croppedPic) takes the vertical sample position y, the horizontal sample position x, the image height picHeight, the image width picWidth, and the sample array croppedPic as input and returns the value of sampleVal as derived below: Note 6 – For the input of the function InpSampleVal(), the vertical position is listed before the horizontal position to be compatible with the input tensor conventions of some inference engines. The following example procedure can be used with NNPF PostProcessingFilter() to generate (multiple) filtered and / or interpolated images in small chunks, containing Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, as indicated by nnpfc_out_order_idc: The order of the images in the stored output tensor is the output order, and the output order generated by applying NNPF in the output order is interpreted as the output order (and does not conflict with the output order of the input images). A value of 1 for nnpfc_complexity_info_present_flag indicates the existence of one or more syntax elements that indicate the complexity of the NNPF associated with nnpfc_id. A value of 0 for nnpfc_complexity_info_present_flag indicates the absence of a syntax element that indicates the complexity of the NNPF associated with nnpfc_id. An `nnpfc_parameter_type_idc` value of 0 indicates that the neural network uses only integer parameters. An `nnpfc_parameter_type_flag` value of 1 indicates that the neural network can use either floating-point or integer parameters. An `nnpfc_parameter_type_idc` value of 2 indicates that the neural network uses only binary parameters. An `nnpfc_parameter_type_idc` value of 3 is reserved for future use by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages with `nnpfc_parameter_type_idc` equal to 3. The values 0, 1, 2, and 3 for nnpfc_log2_parameter_bit_length_minus3 indicate that the neural network does not use parameters with bit lengths greater than 8, 16, 32, and 64, respectively. When nnpfc_parameter_type_idc exists and nnpfc_log2_parameter_bit_length_minus3 does not exist, the neural network does not use parameters with a bit length greater than 1. `nnpfc_num_parameters_idc` indicates the maximum number of neural network parameters in NNPF, expressed as a power of 2048. `nnpfc_num_parameters_idc` equal to 0 indicates that the maximum number of neural network parameters is unknown. The value of `nnpfc_num_parameters_idc` should be in the range of 0 to 52 (inclusive). Values greater than 52 in nnpfc_num_parameters_idc are reserved for future use by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages with nnpfc_num_parameters_idc greater than 52. If the value of nnpfc_num_parameters_idc is greater than zero, then the variable maxNumParameters is deduced as follows: maxNumParameters = (2048 < <nnpfc_num_parameters_idc)–1 (94) The requirement for bitstream consistency is that the number of neural network parameters in NNPF should be less than or equal to maxNumParameters. A value greater than 0 for nnpfc_num_kmac_operations_idc indicates that the maximum number of multiply-accumulate operations per sample in NNPF is less than or equal to nnpfc_num_kmac_operations_idc * 1000. A value of 0 for nnpfc_num_kmac_operations_idc indicates that the maximum number of multiply-accumulate operations in the network is unknown. The value of nnpfc_num_kmac_operations_idc should be between 0 and 2. 32 The range is -2 (including boundary values). `nnpfc_total_kilobyte_size` greater than 0 indicates the total size, in kilobytes, required to store the uncompressed parameters of the neural network. The total size, in bits, is equal to or greater than the sum of the bits used to store each parameter. `nnpfc_total_kilobyte_size` is the total size in bits divided by 8000 and rounded to the nearest integer. `nnpfc_total_kilobyte_size` equal to 0 indicates that the total size required to store the neural network parameters is unknown. The value of `nnpfc_total_kilobyte_size` should be between 0 and 2. 32 The range is -2 (including boundary values). nnpfc_reserved_zero_bit_b should be equal to 0 in the bitstream conforming to this version of the document. The decoder should ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_b is not equal to 0. nnpfc_payload_byte[i] contains the i-th byte of a bitstream conforming to ISO / IEC 15938-17. The sequence of bytes for all current values of i, nnpfc_payload_byte[i], should be a complete bitstream conforming to ISO / IEC 15938-17. 6.4 Example 4
[0198] This embodiment refers to items 1, 2, 3, and 4 and all their sub-items as outlined in Section 5 of the previous article. 8.28.1 Characteristics of Neural Network Post-Processing Filters and SEI Message Syntax 8.28.2 Characteristics of Neural Network Post-Processing Filters and Semantics of SEI Messages The Neural Network Post-Processing Filter Characteristics (NNPFC) SEI message specifies the neural networks that can be used as post-processing filters. The use of the specified Neural Network Post-Processing Filter (NNPF) for a specific image is indicated by the Neural Network Post-Processing Filter Activation (NNPFA) SEI message. To use this SEI message, the following variables need to be specified: - Input the image width and height, in units of brightness samples, which are represented as CroppedWidth and CroppedHeight in this article. - The luminance sample array CroppedYPic[idx] and chrominance sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] (if present) of the input image, with index idx in the range of 0 to numInputPics-1 (inclusive), are used as input for NNPF. - BitDepth for the luminance sample array of the input image Y . - Bit depth of the chroma sample array (if any) for the input image. C . - Chroma format indicator, referred to herein as ChromaFormatIdc, as described in sub-entry 7.3. - When nnpfc_auxiliary_inp_idc equals 1, the filter strength control value StrengthControlVal should be a real number in the range of 0 to 1 (inclusive). The input image with index 0 corresponds to the image for which the NNPF defined by the NNPFC SEI message is activated via the NNPFA SEI message. The input images with index i (in the range of 1 to numInputPics-1 (inclusive)) precede the input images with index i-1 in the output order. When an input image with index 0 and nnpfc_purpose&0x08 is not equal to 0 is associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5, all input images are associated with a frame encapsulation arrangement SEI message with fp_arrangement_type equal to 5 and the same value as fp_current_frame_is_frame0_flag. The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc, as specified in Table 2. Note 1 – More than one NNPFC SEI message can exist for the same image. When more than one NNPFC SEI message with different values of nnpfc_id exists or is activated for the same image, they can have the same or different values of nnpfc_purpose and nnpfc_mode_idc. nnpfc_purpose indicates the purpose of NNPF, as specified in Table 20. In bitstreams conforming to this version of the document, the value of nnpfc_purpose should be in the range of 0 to 63 (inclusive). Values of nnpfc_purpose from 64 to 65535 (inclusive) are reserved for future use by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages with nnpfc_purpose in the range of 64 to 65535 (inclusive). Table 20 – Definition of nnpfc_purpose Note 2 – When the reserved value of nnpfc_purpose is used by ITU-T|ISO / IEC in the future, the syntax of this SEI message can be extended using the following syntax elements, provided that nnpfc_purpose is equal to that value. When ChromaFormatIdc or nnpfc_purpose&0x02 is not equal to 0, nnpfc_purpose&0x20 should be equal to 0. nnpfc_id contains an identifier that can be used to identify NNPF. The value of nnpfc_id should be between 0 and 2. 32 The range is -2 (inclusive of boundary values). The values of nnpfc_id are 256 to 511 (inclusive of boundary values) and 2. 31 to 232 -2 (including boundary values) is reserved for future use by ITU-T|ISO / IEC. Decoders conforming to this document for this version encounter nnpfc_id in the range of 256 to 511 (including boundary values) or 2. 31 to 2 32 When an NNPFC SEI message is received within the range of -2 (including boundary values), the SEI message should be ignored. The following applies when an NNPFC SEI message has a specific nnpfc_id value within the current CLVS and is the first NNPFC SEI message in decoding order: - This SEI message specifies the basic NNPF. - This SEI message applies to the currently decoded image and all subsequent decoded images of the current layer (in output order) until the end of the current CLVS. ... When `nnpfc_purpose&0x20` is not equal to 0, `nnpfc_out_colour_format_idc` specifies the color format of the NNPF output, thus defining the values of the variables `outSubWidthC` and `outSubHeightC`. `nnpfc_out_colour_format_idc` equal to 1 specifies that the NNPF output color format is 4:2:0, and both `outSubWidthC` and `outSubHeightC` are equal to 2. `nnpfc_out_colour_format_idc` equal to 2 specifies that the NNPF output color format is 4:2:2, and both `outSubWidthC` and `outSubHeightC` are equal to 1. `nnpfc_out_colour_format_idc` equal to 3 specifies that the NNPF output color format is 4:2:4, and both `outSubWidthC` and `outSubHeightC` are equal to 1. The value of `nnpfc_out_colour_format_idc` should not be equal to 0. When both nnpfc_purpose&0x02 and nnpfc_purpose&0x20 are equal to 0, outSubWidthC and outSubHeightC are presumed to be equal to SubWidthC and SubHeightC, respectively. The nnpfc_num_input_pics_minus1 plus 1 rule is used as the number of decoded output pictures for the input to the NNPF. The value of nnpfc_num_input_pics_minus1 shall be in the range of 0 to 63 (including the boundary values). When nnpfc_purpose & 0x08 is not equal to 0, the value of nnpfc_num_input_pics_minus1 shall be greater than 0. ... nnpfc_out_format_idc being equal to 0 indicates that the sample values output by the NNPF are real numbers, where the value range from 0 to 1 (including the boundary values) is linearly mapped to the unsigned integer value range from 0 to (1 << bitDepth) – 1 (including the boundary values) for any desired bit depth bitDepth for subsequent post-processing or display. nnpfc_out_format_idc being equal to 1 indicates that the luminance sample values output by the NNPF are unsigned integers in the range of 0 to (1 << (nnpfc_out_tensor_luma_bitdepth_minus8 + 8)) - 1 (including the boundary values), and the chrominance sample values output by the NNPF are unsigned integers in the range of 0 to (1 << (nnpfc_out_tensor_chroma_bitdepth_minus8 + 8)) - 1 (including the boundary values). Values of nnpfc_out_format_idc greater than 1 are reserved for future specifications of ITU-T|ISO / IEC and shall not be present in the bitstream conforming to this version of the document. The decoder conforming to this version of the document shall ignore the NNPFC SEI message containing the reserved value of nnpfc_out_format_idc. nnpfc_out_tensor_luma_bitdepth_minus8 plus 8 specifies the bit depth of the luminance sample values in the output integer tensor. The value of nnpfc_out_tensor_luma_bitdepth_minus8 shall be in the range of 0 to 24 (including the boundary values). nnpfc_out_tensor_chroma_bitdepth_minus8 plus 8 specifies the bit depth of the chrominance sample values in the output integer tensor. The value of nnpfc_out_tensor_chroma_bitdepth_minus8 shall be in the range of 0 to 24 (including the boundary values). When nnpfc_purpose & 0x10 is not equal to 0, the value of nnpfc_out_format_idc shall be equal to 1, and at least one of the following conditions shall be true: -(nnpfc_out_tensor_luma_bitdepth_minus8+8 is greater than BitDepth) Y or (nnpfc_out_tensor_chroma_bitdepth_minus8+8 is greater than BitDepth) C nnpfc_out_order_idc indicates the output order of samples generated by NNPF. In this version of the bitstream conforming to this document, the value of nnpfc_out_order_idc should be in the range of 0 to 3 (inclusive). Values of nnpfc_out_order_idc from 4 to 255 (inclusive) are reserved for future use by ITU-T|ISO / IEC and should not exist in this version of the bitstream conforming to this document. Decoders conforming to this version of this document should ignore NNPFC SEI messages with nnpfc_out_order_idc in the range of 4 to 255 (inclusive). Values of nnpfc_out_order_idc greater than 255 should not exist in this version of the bitstream conforming to this document and are not reserved for future use. When nnpfc_purpose&0x02 is not equal to 0, nnpfc_out_order_idc should not be equal to 3. Table 22 contains informative descriptions of the nnpfc_out_order_idc values. Table 22 – Description of nnpfc_out_order_idc values The process StoreOutputTensors() is used to derive sample values from the output tensor outputTensor, which is a filtered output sample array of FilteredYPic, FilteredCbPic, and FilteredCrPic. The output tensor outputTensor is defined with respect to the given vertical sample coordinates cTop and horizontal sample coordinates cLeft of the top-left sample position of a block of samples included in the input tensor. The process StoreOutputTensors() is defined as follows: nnpfc_overlap indicates the horizontal and vertical sample counts of overlap between adjacent input tensors in NNPF. The value of nnpfc_overlap should be in the range of 0 to 16,383 (inclusive). The nnpfc_constant_patch_size_flag setting being equal to 1 indicates that NNPF accepts the precise patch size as input, as indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. The nnpfc_constant_patch_size_flag setting to 0 indicates that NNPF accepts any patch size with a width of inpPatchWidth and a height of inpPatchHeight as input, such that the width of the extended patch (i.e., the patch plus the overlapping area) (which is equal to inpPatchWidth + 2 * nnpfc_overlap) is a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, and the height of the extended patch (which is equal to inpPatchHeight + 2 * nnpfc_overlap) is a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap. Incrementing nnpfc_patch_width_minus1 by 1 indicates the horizontal sample count for the required patch size for the NNPF input, when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_width_minus1 should be in the range of 0 to Min(32766, CroppedWidth-1) (inclusive). Incrementing nnpfc_patch_height_minus1 by 1 indicates the vertical sample count for the required patch size for the NNPF input, when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_height_minus1 should be in the range of 0 to Min(32766, CroppedHeight-1) (inclusive). `nnpfc_extended_patch_width_cd_delta_minus1` plus 1 plus 2 * `nnpfc_overlap` indicates the common divisor of all allowed values for the width of the extended patch required for the NNPF input, when `nnpfc_constant_patch_size_flag` is equal to 0. The value of `nnpfc_extended_patch_width_cd_delta_minus1` should be in the range of 0 to Min(32766, CroppedWidth-1) (inclusive). `nnpfc_extended_patch_height_cd_delta_minus1` plus 1 plus 2 * `nnpfc_overlap` indicates the common divisor of all allowed values for the height of the extended patch required for the NNPF input, when `nnpfc_constant_patch_size_flag` is equal to 0. The value of `nnpfc_extended_patch_height_cd_delta_minus1` should be in the range of 0 to Min(32766, CroppedHeight-1) (inclusive). Let the variables `inpPatchWidth` and `inpPatchHeight` be the patch size width and patch size height, respectively. If nnpfc_constant_patch_size_flag equals 0, then the following applies: The values of -inpPatchWidth and inpPatchHeight are provided by external means not specified in this document, or set by the post-processor itself. The value of -inpPatchWidth+2*nnpfc_overlap should be a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1+1+2*nnpfc_overlap, and inpPatchWidth should be less than or equal to CroppedWidth. The value of inpPatchHeight+2*nnpfc_overlap should be a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1+1+2*nnpfc_overlap, and inpPatchHeight should be less than or equal to CroppedHeight. Otherwise (nnpfc_constant_patch_size_flag equals 1), the value of inpPatchWidth is set to equal to nnpfc_patch_width_minus1+1, and the value of inpPatchHeight is set to equal to nnpfc_patch_height_minus1+1. The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, and outPatchCHeight are derived as follows: horCScaling=SubWidthC / outSubWidthC (88) verCScaling=SubHeightC / outSubHeightC (89) outPatchCWidth=outPatchWidth*horCScaling (90) outPatchCHeight=outPatchHeight*verCScaling (91) The requirement for bitstream consistency is that outPatchWidth * CroppedWidth should equal... *inpPatchWidth, and outPatchHeight*CroppedHeight should equal *inpPatchHeight. The nnpfc_padding_type indicates the padding process when referencing sample locations outside the boundaries of the cropped, decoded output image, as described in Table 23. The value of nnpfc_padding_type should be in the range of 0 to 15 (inclusive). Table 23 – Informative description of nnpfc_padding_type values nnpfc_padding_type describe 0 Zero fill 1 Copy fill 2 Reflection fill 3 Surround fill 4 Fixed fill 5..15 reserve nnpfc_luma_padding_val indicates the luminance value to be used for padding when nnpfc_padding_type is equal to 4. nnpfc_cb_padding_val indicates the Cb value to be used for padding when nnpfc_padding_type is equal to 4. nnpfc_cr_padding_val indicates the Cr value to be used for padding when nnpfc_padding_type is equal to 4. The function InpSampleVal(y,x,picHeight,picWidth,croppedPic) takes the vertical sample position y, the horizontal sample position x, the image height picHeight, the image width picWidth, and the sample array croppedPic as input and returns the value of sampleVal as derived below: Note 6 – For the input of the function InpSampleVal(), the vertical position is listed before the horizontal position to be compatible with the input tensor conventions of some inference engines. The following example procedure can be used with NNPF PostProcessingFilter() to generate (multiple) filtered and / or interpolated images in small chunks, containing Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, as indicated by nnpfc_out_order_idc: The order of the images in the stored output tensor is the output order, and the output order generated by applying NNPF in the output order is interpreted as the output order (and does not conflict with the output order of the input images). A value of 1 for nnpfc_complexity_info_present_flag indicates the existence of one or more syntax elements that indicate the complexity of the NNPF associated with nnpfc_id. A value of 0 for nnpfc_complexity_info_present_flag indicates the absence of a syntax element that indicates the complexity of the NNPF associated with nnpfc_id. An `nnpfc_parameter_type_idc` value of 0 indicates that the neural network uses only integer parameters. An `nnpfc_parameter_type_flag` value of 1 indicates that the neural network can use either floating-point or integer parameters. An `nnpfc_parameter_type_idc` value of 2 indicates that the neural network uses only binary parameters. An `nnpfc_parameter_type_idc` value of 3 is reserved for future use by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages with `nnpfc_parameter_type_idc` equal to 3. The values 0, 1, 2, and 3 for nnpfc_log2_parameter_bit_length_minus3 indicate that the neural network does not use parameters with bit lengths greater than 8, 16, 32, and 64, respectively. When nnpfc_parameter_type_idc exists and nnpfc_log2_parameter_bit_length_minus3 does not exist, the neural network does not use parameters with a bit length greater than 1. `nnpfc_num_parameters_idc` indicates the maximum number of neural network parameters in NNPF, expressed as a power of 2048. `nnpfc_num_parameters_idc` equal to 0 indicates that the maximum number of neural network parameters is unknown. The value of `nnpfc_num_parameters_idc` should be in the range of 0 to 52 (inclusive). Values greater than 52 in nnpfc_num_parameters_idc are reserved for future use by ITU-T|ISO / IEC and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document should ignore NNPFC SEI messages with nnpfc_num_parameters_idc greater than 52. If the value of nnpfc_num_parameters_idc is greater than zero, then the variable maxNumParameters is deduced as follows: maxNumParameters = (2048 < <nnpfc_num_parameters_idc)–1 (94) The requirement for bitstream consistency is that the number of neural network parameters in NNPF should be less than or equal to maxNumParameters. A value greater than 0 for nnpfc_num_kmac_operations_idc indicates that the maximum number of multiply-accumulate operations per sample in NNPF is less than or equal to nnpfc_num_kmac_operations_idc * 1000. A value of 0 for nnpfc_num_kmac_operations_idc indicates that the maximum number of multiply-accumulate operations in the network is unknown. The value of nnpfc_num_kmac_operations_idc should be between 0 and 2. 32 The range is -2 (including boundary values). `nnpfc_total_kilobyte_size` greater than 0 indicates the total size, in kilobytes, required to store the uncompressed parameters of the neural network. The total size, in bits, is equal to or greater than the sum of the bits used to store each parameter. `nnpfc_total_kilobyte_size` is the total size in bits divided by 8000 and rounded to the nearest integer. `nnpfc_total_kilobyte_size` equal to 0 indicates that the total size required to store the neural network parameters is unknown. The value of `nnpfc_total_kilobyte_size` should be between 0 and 2. 32 The range is -2 (including boundary values). nnpfc_reserved_zero_bit_b should be equal to 0 in the bitstream conforming to this version of the document. The decoder should ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_b is not equal to 0. nnpfc_payload_byte[i] contains the i-th byte of a bitstream conforming to ISO / IEC 15938-17. The sequence of bytes for all current values of i, nnpfc_payload_byte[i], should be a complete bitstream conforming to ISO / IEC 15938-17.
[0199] Further details of embodiments of this disclosure relating to neural network post-processing filters will now be described. As used herein, the terms "neural network post-processing filter" and "neural network post-filter" are used interchangeably. The embodiments of this disclosure should be considered as examples for interpreting general concepts and should not be interpreted in a narrow sense. Furthermore, these embodiments may be applied individually or in any combination.
[0200] Figure 5 A flowchart of a method 500 for video processing according to some embodiments of the present disclosure is shown. Figure 5As shown, at position 502, a conversion between video and video bitstreams is performed. In some embodiments, the conversion may include encoding the video into a bitstream. Alternatively or additionally, the conversion may include decoding the video from the bitstream.
[0201] A neural network post-processing filter (NNPF) is applied to at least one image associated with a video. For example, at least one image can be used as input to the NNPF. In some embodiments, at least one image may include at least one decoded image of the video. Alternatively, at least one image may include at least one cropped decoded image of the video. For example, the decoded image and / or the cropped decoded image may be output by a decoder that decodes the video from a bitstream. In some other embodiments, at least one image may include the output of another NNPF used to filter one or more decoded or cropped decoded images of the video. For example, the NNPF is concatenated with another NNPF. It should be understood that the possible implementations of at least one image associated with a video described herein are merely illustrative and should not be construed as limiting this disclosure in any way.
[0202] The bitstream includes a first indication indicating the purpose of NNPF. For example, the first indication may be included in a Supplemental Enhancement Information (SEI) message or any other suitable video message unit within the bitstream. For example, the first indication may include the syntax element nnpfc_purpose. It should be understood that the names of the indications and / or syntax elements are for illustrative purposes only and not for limitation, and the indications(s) and syntax elements(s) mentioned throughout this disclosure may be represented by any other suitable string other than those mentioned in this disclosure. The scope of this disclosure is not limited in this respect.
[0203] Furthermore, the first candidate among the multiple candidates for this purpose indicates that reducing the width and / or height of at least one image is applicable. In other words, NNPF can support reducing the width and / or height of at least one image. For example, if the first candidate is selected to be included in the purpose of NNPF, then NNPF can be used to reduce the width and / or height of at least one image based on additional information.
[0204] In one example embodiment, the first candidate may be resampling the resolution of at least one image (this may also be simply referred to as resolution resampling). For example, resolution resampling may support resolution downsampling and resolution upsampling. Resolution resampling may indicate that at least one of the following is applicable: reducing the width of at least one image, reducing the height of at least one image, increasing the width of at least one image, or increasing the height of at least one image. In this case, the scaling ratio of the width and / or height of at least one image may be within a predetermined range. For example, and not a limitation, the lower limit of the predetermined range may be equal to 1 / 16. Additionally or alternatively, the upper limit of the predetermined range may be equal to 16. For example, the scaling ratio may be in the range of 1 / 16 to 16, including 1 / 16 and 16. It should be understood that the specific values listed herein are intended to be exemplary and not to limit the scope of this disclosure. For example, the lower limit of the predetermined range may also be equal to 1, 1 / 2, 1 / 4, 1 / 8, etc. Similarly, the upper limit of the predetermined range may also be equal to 1, 2, 4, 8, etc.
[0205] In another example embodiment, the first candidate could be downsampling the resolution of at least one image (this can also be simply referred to as resolution downsampling). Resolution downsampling can indicate that reducing the width and / or height of at least one image is applicable. Alternatively, the purpose of NNPF can be to prohibit both downsampling and upsampling the resolution of at least one image. The downsampling rate of the width and / or height of at least one image can be within a predetermined range. For example, the lower limit of the predetermined range can be equal to 1, 1 / 2, 1 / 4, 1 / 8, or 1 / 16. Additionally or alternatively, the upper limit of the predetermined range can be equal to 1, etc.
[0206] In light of the above, one of the candidate indices for achieving the purpose of NNPF is to reduce the width and / or height of at least one image. Compared to conventional solutions that only support resolution upsampling, the proposed method advantageously enables NNPF to support resolution downsampling. In this way, the functionality of NNPF can be extended and enhanced.
[0207] In some embodiments, a second candidate among a plurality of candidates for this purpose may indicate that at least one of the following is applicable: reducing the luminance bit depth of at least one image, or reducing the chrominance bit depth of at least one image. For example, the bit depth of one or more luminance and / or chrominance sample values of at least one image may be reduced, for example, from 8 bits to 5 bits, etc.
[0208] In one example embodiment, the second candidate could be downsampling the bit depth of at least one image (this can also be simply referred to as bit depth downsampling). Bit depth downsampling indicates that the luminance bit depth and / or chrominance bit depth can be reduced. Alternatively, the purpose of NNPF can be to prohibit both downsampling and upsampling the bit depth of at least one image.
[0209] In another example embodiment, the second candidate could be resampling the bit depth of at least one image (this can also be simply referred to as bit depth resampling). For example, bit depth resampling can support both bit depth downsampling and bit depth upsampling. Bit depth resampling can indicate that at least one of the following is applicable: reducing the luminance bit depth of at least one image, reducing the chroma bit depth of at least one image, increasing the luminance bit depth of at least one image, or increasing the chroma bit depth of at least one image.
[0210] In light of the above, one of the candidates for NNPF is to reduce the luminance bit depth or chrominance bit depth of at least one image. Compared to traditional solutions that only support bit depth upsampling, the proposed method advantageously enables NNPF to support bit depth downsampling. In this way, the functionality of NNPF can be further extended and enhanced.
[0211] In some embodiments, a third candidate among a plurality of candidates for this purpose may indicate at least that downsampling of the chroma format of at least one image is applicable. In one example embodiment, the third candidate may be downsampling of the chroma format of at least one image (this may also be simply referred to as chroma format downsampling or chroma downsampling). Additionally, the purpose of NNPF may prohibit both downsampling and upsampling of the chroma format of at least one image. For example, chroma format downsampling may include changing the chroma format from 4:4:4 to 4:2:2, from 4:4:4 to 4:2:0, or from 4:2:2 to 4:2:0. Additionally or alternatively, chroma format downsampling may include changing the chroma format from 4:4:4 to 4:0:0, from 4:2:2 to 4:0:0, or from 4:2:0 to 4:0:0. It should be understood that the above examples are described for illustrative purposes only. The scope of this disclosure is not limited in this respect.
[0212] In another example embodiment, the third candidate may be resampling the chroma format of at least one image (this may also be simply referred to as chroma format resampling or chroma resampling). For example, chroma format resampling may support both chroma format downsampling and chroma format upsampling. For example, chroma format resampling may include changing the chroma format from 4:4:4 to 4:2:2, from 4:4:4 to 4:2:0, from 4:2:2 to 4:2:0, from 4:2:0 to 4:2:2, from 4:2:0 to 4:2:2, from 4:2:0 to 4:4:4, or from 4:2:2 to 4:4:4.
[0213] Additionally or alternatively, chroma format resampling may include at least one of the following: changing the chroma format from 4:4:4 to 4:2:2, changing the chroma format from 4:4:4 to 4:2:0, changing the chroma format from 4:4:4 to 4:0:0, changing the chroma format from 4:2:2 to 4:2:0, changing the chroma format from 4:2:2 to 4:0:0, and so on. The chroma format can be changed from 4:0:0 to 4:2:0, from 4:0:0 to 4:2:2, from 4:0:0 to 4:4:4, from 4:2:0 to 4:2:2, from 4:2:0 to 4:4:4, or from 4:2:2 to 4:4:4. It should be understood that the above examples are described for illustrative purposes only. The scope of this disclosure is not limited in this respect.
[0214] In light of the above, one of the candidate indices for the purpose of NNPF is to downsample the chroma format of at least one image. Compared to traditional solutions that only support chroma format upsampling, the proposed method advantageously enables NNPF to support chroma format downsampling. In this way, the functionality of NNPF can be further extended and enhanced.
[0215] In some embodiments, a fourth candidate among a plurality of candidates for this purpose may at least indicate that downsampling of the video's picture rate is applicable. As used herein, picture rate may refer to the video's frame rate.
[0216] In one example embodiment, the fourth candidate could be downsampling the video's frame rate (this can also be simply referred to as frame rate downsampling). Alternatively, the purpose of NNPF can be prohibited from including both downsampling and upsampling the video's frame rate. For example, rather than being restrictive, the video's frame rate could be reduced from 60 frames per second (fps) to 30 fps.
[0217] In another example embodiment, the fourth candidate could be resampling the video's image rate (this can also be simply referred to as image rate resampling). For example, image rate resampling could support both image rate downsampling and image rate upsampling. In other words, image rate resampling could indicate that downsampling or upsampling the video's image rate is applicable.
[0218] In light of the above, one of the candidates for NNPF is the indication to downsample the image rate of the video. Compared to traditional solutions that only support image rate upsampling, the proposed method advantageously enables NNPF to support image rate downsampling. In this way, the functionality of NNPF can be further extended and enhanced.
[0219] In some embodiments, the aforementioned target candidates can be combined based on a bitmasking method. For example, each of the plurality of candidates for that target can correspond to a bitmask, and the bitmask can be used to determine the target of the NNPF based on a first indication. Additionally, a bit in the first indication indicates whether the target of the NNPF can include one of the plurality of candidates for that target. For example, and not limitingly, if the value of a first bit (such as bit 0, bit 1, etc.) in the first indication is equal to a first value (such as 1, etc.), then the target of the NNPF can include the first candidate for that target. If the value of the first bit is equal to a second value (such as 0, etc.), then the target of the NNPF may not include the first candidate for that target.
[0220] Furthermore, the purpose of NNPF can be determined based on the result of applying a bitwise operation to the first indication and at least one bit mask. For example, and not limitingly, bit 0 can represent the least significant bit in the syntax element nnpfc_purpose, while bit 1 in the syntax element nnpfc_purpose can correspond to resolution resampling. Bit mask 0x04 can be used to determine whether the purpose of NNPF includes resolution resampling. If (nnpfc_purpose&0x04) equals 0, then the purpose of NNPF does not include resolution resampling. If (nnpfc_purpose&0x04) is greater than 0, then the purpose of NNPF includes resolution resampling. The operator “&” represents a bitwise AND operation. It should be understood that the above description is for illustrative purposes only. The scope of this disclosure is not limited in this respect.
[0221] Given the above, one of the multiple candidates for the purpose of NNPF corresponds to a bit mask used to determine the purpose of NNPF based on a first indication. Compared to conventional solutions, the proposed method advantageously provides a system scheme for transmitting the purpose of NNPF via signaling, thereby supporting possible expansion of the purpose. In this way, potential instabilities and logic problems can be avoided, and encoding / decoding efficiency can be improved.
[0222] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by means of a video processing apparatus. In this method, conversion between video and video bitstreams is performed. A neural network post-processing filter (NNPF) is applied to at least one image associated with the video. The bitstream includes a first indication indicating the purpose of the NNPF, and a first candidate among a plurality of candidates for that purpose indicates at least one of the following is applicable: reducing the width of at least one image, or reducing the height of at least one image.
[0223] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. In this method, conversion between video and video bitstreams is performed. A neural network post-processing filter (NNPF) is applied to at least one image associated with the video. The bitstream includes a first indication indicating the purpose of the NNPF, and a first candidate among a plurality of candidates for that purpose indicates at least one of the following is applicable: reducing the width of at least one image, or reducing the height of at least one image. Furthermore, the bitstream is stored in a non-transitory computer-readable recording medium.
[0224] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.
[0225] Item 1. A method for video processing, comprising: performing a conversion between a video and a bitstream of the video, wherein a neural network post-processing filter (NNPF) is applied to at least one image associated with the video, the bitstream including a first indication of the purpose of the NNPF, and a first candidate among a plurality of candidates for the purpose indicating at least one of the following is applicable: reducing the width of the at least one image, or reducing the height of the at least one image.
[0226] Item 2. The method according to Item 1, wherein the first instruction includes the syntax element nnpfc_purpose.
[0227] Item 3. The method according to any one of items 1-2, wherein the at least one image comprises at least one decoded image of the video or at least one cropped decoded image.
[0228] Item 4. The method according to any one of items 1-3, wherein the first candidate is to resample the resolution of the at least one image.
[0229] Item 5. The method according to Item 4, wherein resampling the resolution of the at least one image indicates that at least one of the following is applicable: reducing the width of the at least one image, reducing the height of the at least one image, increasing the width of the at least one image, or increasing the height of the at least one image.
[0230] Item 6. The method according to any one of items 1-5, wherein the scaling ratio of at least one of the width or height of the at least one image is within a predetermined range.
[0231] Item 7. The method according to Item 6, wherein the lower limit of the predetermined range is equal to 1 / 16.
[0232] Item 8. The method according to any one of Items 6-7, wherein the upper limit of the predetermined range is equal to 16.
[0233] Item 9. The method according to Item 6 or 8, wherein the lower limit of the predetermined range is equal to one of the following: 1, 1 / 2, 1 / 4 or 1 / 8.
[0234] Item 10. The method according to any one of items 6-7 and 9, wherein the upper limit of the predetermined range is equal to one of the following: 1, 2, 4 or 8.
[0235] Item 11. The method according to any one of items 1-3, wherein the first candidate is downsampling the resolution of the at least one image.
[0236] Item 12. The method according to Item 11, wherein the purpose of the NNPF is prohibited from including: downsampling the resolution of the at least one image and upsampling the resolution of the at least one image.
[0237] Item 13. The method according to any one of items 11-12, wherein the downsampling rate of at least one of the width or the height of the at least one image is within a predetermined range.
[0238] Item 14. The method according to Item 13, wherein the lower limit of the predetermined range is equal to one of the following: 1, 1 / 2, 1 / 4, 1 / 8 or 1 / 16.
[0239] Item 15. The method according to any one of Items 13-14, wherein the upper limit of the predetermined range is equal to 1.
[0240] Item 16. The method according to any one of items 1-15, wherein the second candidate among the plurality of candidates for the stated purpose indicates that at least one of the following is applicable: reducing the luminance bit depth of the at least one image, or reducing the chroma bit depth of the at least one image.
[0241] Item 17. The method according to Item 16, wherein the second candidate is a downsampling of the bit depth of the at least one image.
[0242] Item 18. The method according to Item 17, wherein the purpose of the NNPF is prohibited from including: downsampling the bit depth of the at least one image and upsampling the bit depth of the at least one image.
[0243] Item 19. The method according to Item 16, wherein the second candidate is to resample the bit depth of the at least one image.
[0244] Item 20. The method according to Item 19, wherein resampling the bit depth of the at least one image indicates that at least one of the following is applicable: reducing the luminance bit depth of the at least one image, reducing the chroma bit depth of the at least one image, increasing the luminance bit depth of the at least one image, or increasing the chroma bit depth of the at least one image.
[0245] Item 21. The method according to any one of items 1-20, wherein a third candidate among the plurality of candidates for the stated purpose indicates at least that downsampling of the chroma format of the at least one image is applicable.
[0246] Item 22. The method according to Item 21, wherein the third candidate is downsampling the chroma format of the at least one image.
[0247] Item 23. The method according to Item 22, wherein the purpose of the NNPF is prohibited from including: downsampling the chroma format of the at least one image and upsampling the chroma format of the at least one image.
[0248] Item 24. The method according to any one of items 22-23, wherein downsampling the chroma format of the at least one image comprises at least one of: changing the chroma format from a 4:4:4 chroma format to a 4:2:2 chroma format, changing the chroma format from a 4:4:4 chroma format to a 4:2:0 chroma format, or changing the chroma format from a 4:2:2 chroma format to a 4:2:0 chroma format.
[0249] Item 25. The method according to any one of items 22-23, wherein downsampling the chroma format of the at least one image comprises at least one of: changing the chroma format from a 4:4:4 chroma format to a 4:0:0 chroma format, changing the chroma format from a 4:2:2 chroma format to a 4:0:0 chroma format, or changing the chroma format from a 4:2:0 chroma format to a 4:0:0 chroma format.
[0250] Item 26. The method according to Item 21, wherein the third candidate is to resample the chroma format of the at least one image.
[0251] Item 27. The method according to Item 26, wherein resampling the chroma format of the at least one image comprises at least one of the following: changing the chroma format from 4:4:4 to 4:2:2, changing the chroma format from 4:4:4 to 4:2:0, changing the chroma format from 4:2:2 to 4:2:0, changing the chroma format from 4:2:0 to 4:2:2, changing the chroma format from 4:2:0 to 4:2:2, changing the chroma format from 4:2:0 to 4:4:4, or changing the chroma format from 4:2:2 to 4:4:4.
[0252] Item 28. The method according to Item 26, wherein resampling the chroma format of the at least one image comprises at least one of the following: changing the chroma format from 4:4:4 to 4:2:2, changing the chroma format from 4:4:4 to 4:2:0, changing the chroma format from 4:4:4 to 4:0:0, changing the chroma format from 4:2:2 to 4:2:0, and changing the chroma format from 4:2:2 to 4:2:0. The chroma format can be changed from 4:0:0 to 4:2:0, 4:2:2, 4:4:4, 4:2:0 to 4:2:2, 4:4:4, or 4:2:2 to 4:4:4.
[0253] Item 29. The method according to any one of items 1-28, wherein the fourth candidate among the plurality of candidates for the stated purpose indicates at least that downsampling of the image rate of the video is applicable.
[0254] Item 30. The method according to Item 29, wherein the fourth candidate is downsampling the image rate of the video.
[0255] Item 31. The method according to Item 30, wherein the purpose of the NNPF is prohibited from including: downsampling the image rate of the video and upsampling the image rate of the video.
[0256] Item 32. The method according to Item 29, wherein the fourth candidate is to resample the image rate of the video.
[0257] Item 33. The method according to Item 32, wherein the resampling instruction for the image rate of the video: downsampling or upsampling of the image rate of the video is applicable.
[0258] Item 34. The method according to any one of items 1-33, wherein the conversion includes encoding the video into the bitstream.
[0259] Item 35. The method according to any one of items 1-33, wherein the conversion includes decoding the video from the bitstream.
[0260] Item 36. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1-35.
[0261] Item 37. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute instructions of any one of items 1-35.
[0262] Item 38. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: performing a conversion between the video and the video bitstream, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating the purpose of the NNPF, and a first candidate among a plurality of candidates for the purpose indicating at least one of the following is applicable: reducing the width of at least one picture, or reducing the height of at least one picture.
[0263] Item 39. A method for storing a bitstream of video, comprising: performing a conversion between video and video bitstreams, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating the purpose of the NNPF, and a first candidate of a plurality of candidates for the purpose indicating at least one of the following is applicable: reducing the width of at least one picture, or reducing the height of at least one picture; and storing the bitstream in a non-transitory computer-readable recording medium. Example device
[0264] Figure 6 A block diagram of a computing device 600 in which various embodiments of the present disclosure may be implemented is shown. The computing device 600 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0265] It should be understood that, Figure 6 The computing device 600 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0266] like Figure 6As shown, computing device 600 includes general-purpose computing device 600. Computing device 600 may include at least one or more processors or processing units 610, memory 620, storage unit 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660.
[0267] In some embodiments, computing device 600 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server provided by a service provider, a large computing device, etc. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that computing device 600 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).
[0268] Processing unit 610 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 600. Processing unit 610 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0269] Computing device 600 typically includes various computer storage media. Such media can be any media accessible by computing device 600, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 630 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 600.
[0270] The computing device 600 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 6Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0271] Communication unit 640 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in computing device 600 can be implemented by a single computing cluster or by multiple computing machines communicating via communication connections. Therefore, computing device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0272] Input device 650 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 660 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 640, computing device 600 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 600 can also communicate with one or more devices that enable a user to interact with computing device 600, or any device that enables computing device 600 to communicate with one or more other computing devices (e.g., network card, modem, etc.), if needed. Such communication can be performed via an input / output (I / O) interface (not shown).
[0273] In some embodiments, some or all components of computing device 600 may not be integrated into a single device, but may be deployed within a cloud computing architecture. In a cloud computing architecture, components may be provided remotely and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider offers applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed at locations in remote data centers. Cloud computing infrastructure may provide services through shared data centers, although they may appear as a single access point for users. Thus, cloud computing architectures can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, they may be provided from conventional servers or installed directly or otherwise on client devices.
[0274] In embodiments of this disclosure, computing device 600 can be used to implement video encoding / decoding. Memory 620 may include one or more video codec modules 625 having one or more program instructions. These modules can be accessed and executed by processing unit 610 to perform the functions of the various embodiments described herein.
[0275] In an example embodiment of performing video encoding, input device 650 may receive video data as input 670 to be encoded. The video data may be processed, for example, by video codec module 625 to generate an encoded bitstream. The encoded bitstream may be provided as output 680 via output device 660.
[0276] In an example embodiment of performing video decoding, input device 650 may receive an encoded bitstream as input 670. The encoded bitstream may be processed, for example, by video codec module 625 to generate decoded video data. The decoded video data may be provided as output 680 via output device 660.
[0277] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These changes are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A method for video processing, comprising: Perform a conversion between a video and a bitstream of the video, wherein a neural network post-processing filter (NNPF) is applied to at least one image associated with the video, the bitstream including a first indication of the purpose of the NNPF, and the first candidate among a plurality of candidates for the purpose indicates that at least one of the following is applicable: reducing the width of the at least one image, or reducing the height of the at least one image.
2. The method of claim 1, wherein the first instruction includes the syntax element nnpfc_purpose.
3. The method according to any one of claims 1 to 2, wherein the at least one image comprises at least one decoded image of the video or at least one cropped decoded image.
4. The method according to any one of claims 1 to 3, wherein the first candidate is to resample the resolution of the at least one image.
5. The method of claim 4, wherein resampling the resolution of the at least one image indicates that at least one of the following is applicable: Reduce the width of the at least one image. Reduce the height of the at least one image. Increase the width of the at least one image, or Increase the height of the at least one image.
6. The method according to any one of claims 1 to 5, wherein the scaling ratio of at least one of the width or height of the at least one image is within a predetermined range.
7. The method of claim 6, wherein the lower limit of the predetermined range is equal to 1 / 16.
8. The method according to any one of claims 6 to 7, wherein the upper limit of the predetermined range is equal to 16.
9. The method according to claim 6 or 8, wherein the lower limit of the predetermined range is equal to one of the following: 1, 1 / 2, 1 / 4 or 1 / 8.
10. The method according to any one of claims 6 to 7 and 9, wherein the upper limit of the predetermined range is equal to one of the following: 1, 2, 4 or 8.
11. The method according to any one of claims 1 to 3, wherein the first candidate is a downsampling of the resolution of the at least one image.
12. The method of claim 11, wherein the purpose of the NNPF is prohibited by: The resolution of the at least one image is downsampled and the resolution of the at least one image is upsampled.
13. The method according to any one of claims 11 to 12, wherein the downsampling rate of at least one of the width or height of the at least one image is within a predetermined range.
14. The method of claim 13, wherein the lower limit of the predetermined range is equal to one of the following: 1, 1 / 2, 1 / 4, 1 / 8 or 1 / 16.
15. The method according to any one of claims 13 to 14, wherein the upper limit of the predetermined range is equal to 1.
16. The method according to any one of claims 1 to 15, wherein the second candidate among the plurality of candidates for the stated purpose indicates that at least one of the following is applicable: Reduce the luminance bit depth of the at least one image, or Reduce the chroma bit depth of the at least one image.
17. The method of claim 16, wherein the second candidate is a downsampling of the bit depth of the at least one image.
18. The method of claim 17, wherein the purpose of the NNPF is prohibited by: The bit depth of the at least one image is downsampled and the bit depth of the at least one image is upsampled.
19. The method of claim 16, wherein the second candidate is to resample the bit depth of the at least one image.
20. The method of claim 19, wherein resampling the bit depth of the at least one image indicates that at least one of the following is applicable: Reduce the luminance bit depth of the at least one image. Reduce the chroma bit depth of the at least one image. Increase the luminance bit depth of the at least one image, or Increase the chroma bit depth of the at least one image.
21. The method according to any one of claims 1 to 20, wherein a third candidate among the plurality of candidates for the stated purpose indicates at least that downsampling of the chroma format of the at least one image is applicable.
22. The method of claim 21, wherein the third candidate is downsampling the chroma format of the at least one image.
23. The method of claim 22, wherein the purpose of the NNPF is prohibited by: The chroma format of the at least one image is downsampled and the chroma format of the at least one image is upsampled.
24. The method according to any one of claims 22 to 23, wherein downsampling the chroma format of the at least one image comprises at least one of the following: Change the chroma format from 4:4:4 to 4:2:
2. Change the chroma format from 4:4:4 to 4:2:0, or Change the chroma format from 4:2:2 to 4:2:
0.
25. The method according to any one of claims 22 to 23, wherein downsampling the chroma format of the at least one image comprises at least one of the following: Change the chroma format from 4:4:4 to 4:0:
0. Change the chroma format from 4:2:2 to 4:0:0, or Change the chroma format from 4:2:0 to 4:0:
0.
26. The method of claim 21, wherein the third candidate is to resample the chroma format of the at least one image.
27. The method of claim 26, wherein resampling the chroma format of the at least one image comprises at least one of the following: Change the chroma format from 4:4:4 to 4:2:
2. Change the chroma format from 4:4:4 to 4:2:
0. Change the chroma format from 4:2:2 to 4:2:
0. Change the chroma format from 4:2:0 to 4:2:
2. Change the chroma format from 4:2:0 to 4:4:4, or Change the chroma format from 4:2:2 to 4:4:
4.
28. The method of claim 26, wherein resampling the chroma format of the at least one image comprises at least one of the following: Change the chroma format from 4:4:4 to 4:2:
2. Change the chroma format from 4:4:4 to 4:2:
0. Change the chroma format from 4:4:4 to 4:0:
0. Change the chroma format from 4:2:2 to 4:2:
0. Change the chroma format from 4:2:2 to 4:0:
0. Change the chroma format from 4:0:0 to 4:2:
0. Change the chroma format from 4:0:0 to 4:2:
2. Change the chroma format from 4:0:0 to 4:4:
4. Change the chroma format from 4:2:0 to 4:2:
2. Change the chroma format from 4:2:0 to 4:4:4, or Change the chroma format from 4:2:2 to 4:4:
4.
29. The method according to any one of claims 1 to 28, wherein the fourth candidate among the plurality of candidates for the stated purpose indicates at least that downsampling of the image rate of the video is applicable.
30. The method of claim 29, wherein the fourth candidate is to downsample the image rate of the video.
31. The method of claim 30, wherein the purpose of the NNPF is prohibited by: The image rate of the video is downsampled and the image rate of the video is upsampled.
32. The method of claim 29, wherein the fourth candidate is to resample the image rate of the video.
33. The method of claim 32, wherein the resampling instruction for the image rate of the video may involve downsampling or upsampling the image rate of the video.
34. The method according to any one of claims 1 to 33, wherein the conversion comprises encoding the video into the bitstream.
35. The method according to any one of claims 1 to 33, wherein the conversion comprises decoding the video from the bitstream.
36. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 35.
37. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 35.
38. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: Perform a conversion between a video and a bitstream of the video, wherein a neural network post-processing filter (NNPF) is applied to at least one image associated with the video, the bitstream including a first indication of the purpose of the NNPF, and the first candidate among a plurality of candidates for the purpose indicates that at least one of the following is applicable: reducing the width of the at least one image, or reducing the height of the at least one image.
39. A method for storing a bitstream of video, comprising: Perform a conversion between a video and its bitstream, wherein a neural network post-processing filter (NNPF) is applied to at least one image associated with the video, the bitstream including a first indication of the purpose of the NNPF, and the first candidate among a plurality of candidates for the purpose indicates that at least one of the following is applicable: reducing the width of the at least one image, or reducing the height of the at least one image; and The bit stream is stored in a non-transitory computer-readable recording medium.