Method and device for video processing and medium
By improving the syntax elements and signaling of the post-processing filter of neural networks, the problem of insufficient support for the increase in 4:0:0 input and bit depth in the prior art is solved, and more flexible and efficient video encoding and decoding processing is achieved, improving video quality and codec efficiency.
Patent Information
- Application Number
- CN202480006378.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-24
- Filing Date
- 2024-01-03
- Publication Date
- 2025-08-15
AI Technical Summary
The existing neural network post-processing filter characteristics (SEI) design does not support applications with 4:0:0 input in video encoding and decoding, such as the coloring of black and white videos, cannot handle applications with increasing bit depth, the range of syntax elements is not clear, and there are problems in the derivation process of frame rate and output sample point value, resulting in the inability to effectively support multiple video formats and quality improvements.
By re-specifying the syntax element range and signaling of the neural network post-processing filter, it supports the application of 4:0:0 input, allowing bit depth to increase, adjusting frame rate and picture resolution, ensuring the rationality and integrity of syntax elements, improving the derivation process of output sample point values, and supporting the conversion of multiple video formats.
It realizes effective support for 4:0:0 input, improves the quality of video encoding and decoding, supports bit depth increase and frame rate adjustment, ensures flexibility and consistency of video processing, and improves encoding and decoding efficiency.
Smart Images

Figure CN120500844A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly, to a neural network post-processing filter (NNPF). Background Art
[0002] Digital video capabilities are now being used in every aspect of our lives. For video encoding and decoding, various video compression technologies have been proposed, including MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC), and Versatile Video Codec (VVC). However, there is a general desire to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is provided. The method includes performing conversion between a video and a bitstream of the video, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating a purpose of the NNPF, and one of candidates for the purpose is colorization of the at least one picture.
[0005] According to the method of the first aspect of the present disclosure, NNPF is able to support coloring of input to NNPF. Compared with traditional solutions, the proposed method can advantageously support applications that require coloring and applications with input in 4:0:0 format. In this way, the functionality of NNPF becomes diversified and the encoding and decoding quality can be improved.
[0006] In a second aspect, a device for video processing is provided. The device includes a processor and a non-volatile memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.
[0007] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions, the instructions causing a processor to execute the method according to the first aspect of the present disclosure.
[0008] In a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by a video processing apparatus. The method includes performing conversion between the video and the bitstream, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream includes a first indication indicating a purpose of the NNPF, and one of the candidates for the purpose is colorization of the at least one picture.
[0009] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: performing conversion between the video and the bitstream, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating a purpose of the NNPF, and one of candidates for the purpose is colorization of the at least one picture; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings.In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0012] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;
[0013] Figure 2 shows a block diagram illustrating a first example video encoder according to some embodiments of the present disclosure;
[0014] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;
[0015] Figure 4 shows a schematic diagram of a brightness data channel;
[0016] Figure 5 A flowchart showing a method for video processing according to an embodiment of the present disclosure is shown; and
[0017] Figure 6 A block diagram is shown of a computing device in which various embodiments of the present disclosure may be implemented.
[0018] Throughout the drawings, the same or similar reference numbers generally refer to the same or similar elements. DETAILED DESCRIPTION
[0019] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.
[0020] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0021] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment is required to include that particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is intended that such feature, structure, or characteristic, whether or not explicitly described, be applicable to other embodiments and that it is within the knowledge of those skilled in the art to apply such feature, structure, or characteristic.
[0022] It should be understood that although the terms "first" and "second" and the like may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0023] The terms used herein are used only for the purpose of describing specific embodiments and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "comprise," "including," "having," "including," and / or "comprising" when used herein indicate the presence of the features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Sample Environment
[0024] Figure 1is a block diagram illustrating an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0025] The video source 112 may include a source such as a video capture device. Examples of a video capture device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.
[0026] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are coded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be directly transmitted to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0027] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, the destination device 120 being configured to interface with an external display device.
[0028] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.
[0029] Figure 2is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of the video encoder 114 in the system 100 is shown.
[0030] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0031] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.
[0032] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0033] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail in the following sections. Figure 2 are shown separately in the example.
[0034] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0035] The mode selection unit 203 can, for example, select one of a plurality of coding modes (intra-frame coding or inter-frame coding) based on the error result, and provide the resulting intra-frame coded block or inter-frame coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a joint intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).
[0036] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.
[0037] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, "P slices" and "B slices" may refer to portions of a picture consisting of macroblocks that are independent of macroblocks in the same picture.
[0038] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0039] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference pictures in list 0 and list 1 containing multiple reference video blocks, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. The motion estimation unit 204 may output the multiple reference indices and multiple motion vectors for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.
[0040] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0041] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0042] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0043] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0044] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0045] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0046] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0047] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.
[0048] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0049] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0050] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.
[0051] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0052] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 An example of the video decoder 124 in the system 100 is shown.
[0053] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0054] exist Figure 3 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.
[0055] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, which motion information includes motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which includes deriving several most likely candidates based on data from adjacent PBs and reference pictures. The motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially neighboring blocks or temporally neighboring blocks.
[0056] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. Identifiers for the interpolation filters used with sub-pixel precision may be included in the syntax elements.
[0057] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information, and motion compensation unit 302 may use the interpolation filters to produce a prediction block.
[0058] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the encoded video sequence, partition information describing how each macroblock of the picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information used to decode the encoded video sequence. As used herein, in some aspects, "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice can be an entire picture or a region of a picture.
[0059] The intra prediction unit 303 can use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0060] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.
[0061] Some exemplary embodiments of the present disclosure are described in detail below. It should be noted that the section headings used in this document are for ease of understanding and do not limit the embodiments disclosed in a section to that section. In addition, although some embodiments are described with reference to a multifunctional video codec or other specific video codecs, the disclosed technology is also applicable to other video coding and decoding technologies. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps for de-encoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at a different compression bit rate. 1. Brief Overview The present disclosure relates to image / video coding techniques. Specifically, it relates to the definition and signaling of new neural network post-processing filtering objectives, their combination with current objectives, and the scope and constraints of various syntax elements. The ideas can be applied alone or in various combinations to video bitstreams encoded and decoded by any codec (e.g., the Versatile Video Codec (VVC) standard and / or the Versatile SEI Message (VSEI) standard for encoding and decoding video bitstreams). 2. Abbreviation APS Adaptive Parameter Set AU Access Unit CLVS codec layer video sequence CLVSS codec layer video sequence starts CRC Cyclic Redundancy Check CVS codec video sequence FIR Finite Impulse Response IRAP Intra-frame Random Access Point NAL Network Abstraction Layer PPS Picture Parameter Set PU picture unit RASL Random Access Skipped Preamble SEI Supplemental Enhancement Information STSA Stepwise Temporal Sublayer Access VCL video codec layer VSEI Versatile Supplementary Enhancement Information (Recommendation ITU-T H.274 | ISO / IEC 23002-7) VUI Video Availability Information VVC Versatile Video Codec (Recommendation ITU-T H.266 | ISO / IEC 23090-3) 3. Introduction 3.1. Video Codec Standards Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. The two organizations jointly developed the H.262 / MPEG-2 Video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction and transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). When the Versatile Video Codec (VVC) project officially launched, JVET was later renamed the Joint Video Experts Team (JVET). VVC is a new codec standard that aims to reduce bit rate by 50% compared to HEVC. The standard was finalized by JVET at its 19th meeting, which ended on July 1, 2020. The Versatile Video Codec (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) and the associated Versatile Supplementary Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) for codec video bitstreams have been designed for the widest range of applications, including traditional uses such as television broadcasting, video conferencing or playback from storage media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, composition and merging of content from multiple codec video bitstreams, multi-view video, scalable layer encoding and decoding, and viewport-adaptive 360° immersive media. The Essential Video Codec (EVC) standard (ISO / IEC 23094-1) is another video codec standard recently developed by MPEG. 3.2. General SEI messages and SEI messages in VVC and VSEI SEI messages assist with processes related to decoding, display, or other purposes. However, no SEI messages are required for constructing luma or chroma samples through the decoding process. Standard-compliant decoders do not need to process this information to achieve output order consistency. Some SEI messages are required for checking bitstream consistency and output timing decoder consistency. Other SEI messages are not required for checking bitstream consistency. Annex D of VVC specifies the syntax and semantics of SEI message payloads for some SEI messages, and specifies the use of SEI messages and VUI parameters for which syntax and semantics are specified in ITU-T H.274 | ISO / IEC 23002-7. 3.3. Signaling of Neural Network Post-Processing Filters An excerpt from the specifications of two SEI messages used for signaling of neural network post-processing filters is shown below. 4. Question The current design of the Neural Network Post-Processing Filter Characteristics (NNPFC) SEI message has the following problems. 1) It does not support applications with 4:0:0 input, e.g. for colorization of black and white video. 2) It does not support applications that may require an increase in bit depth, for example, from SDR to HDR. 3) The range of the syntax element nnpfc_num_input_pics_minus2 is unspecified. 4) The scope of the syntax element nnpfc_interpolated_pics[i] is unspecified. 5) All instances of nnpfc_interpolated_pics[i] can be equal to 0, which results in no frame rate change from the input video to the output. It contradicts the purpose of frame rate upsampling. 6) numOutputPics does not take the input picture into account, which makes it impossible to change the input picture in some applications. 7) The procedure StoreOutputTensors() used to derive the sample values in the filtered output sample array from the output tensor outputTensor uses the wrong scope in the procedure. 8) When the purpose indicates coloring, some syntax elements are missing, some syntax elements exist but are useless, some constraints are missing, and some conversion processes are missing. 9) Whether the syntax elements for luma or chroma bit depth are signaled shall depend on the input or output method and order. 10) Some syntax and semantics may need to be adjusted to enable non-4:0:0 input for shading. 11) When the input is in 4:0:0 format for rendering, the chroma sample positions should be specified. 5. Detailed solution In order to solve the above problems, the following methods are disclosed. The solutions should be considered as examples to explain the general concept and should not be interpreted in a narrow sense. In addition, these solutions can be applied alone or combined in any way. 1) To address issue 1, new purposes or sub-purposes can be specified to describe the behavior of NNPF for 4:0:0 input. a. In one example, alternatively, further, a combination of chroma format change and / or frame rate and / or picture resolution and / or bit depth change for 4:0:0 input is also allowed, for example, indicated in the SEI message. b. nnpfc_out_order_idc should not be equal to 0 when the input is in 4:0:0 format and there is no color format change. c. In one example, when the NNPF input for which the value of the variable ChromaFormatIdc is equal to 0 is in 4:0:0 format, the NNPF purpose should not be chroma upsampling. i. Alternatively, when the NNPF purpose is chroma upsampling, the NNPF input should not be in 4:0:0 format. d. In one example, when the NNPF input is in 4:0:0 format, the syntax element(s) related to the chroma input tensor are not signaled. i. In one example, when the NNPF input is in 4:0:0 format, no syntax element is signaled to indicate the bit depth of the chroma sample values in the input integer tensor. ii. In one example, when the NNPF input is in 4:0:0 format, the syntax element(s) indicating the chroma type and / or chroma fill value are not signaled. e. In one example, the aforementioned condition “when NNPF input is in 4:0:0 format” can be replaced with “when NNPF input is in 4:4:4 format, separate plane codec is applied, and the current input is the luma component”. 2) To solve problem 2, a new purpose may be specified for bit depth increase. a. In one example, a syntax element may be signaled to indicate the difference between the input bit depth and the output bit depth. i. In one example, the syntax element nnpfc_delta_bitdepth_minus1 is signaled, and nnpfc_delta_bitdepth_minus1 plus 1 specifies the output bit depth minus the input bit depth. b. In one example, when the goal is bit depth increase, assuming both input and output are in integer value format, the network output must have a higher bit depth than the network input. c. In one example, when the goal is bit depth increase, assuming the network output is in integer value format, the network output must have a higher bit depth than the bit depth of the decoded video output by the video decoder. d. In one example, when the purpose is related to bit depth increase, it is required that both the network input and output are in integer value format, and the network output must have a higher bit depth than the network input. i. In one example, the network output having a higher bit depth than the network input means that there is no color component whose output bit depth is less than the input bit depth, and that for at least one color component, the output bit depth is greater than the input bit depth. ii. In one example, the network output having a higher bit depth than the network input means that for all color components, the bit depth values of the output are greater than the bit depth values of the input. e. In one example, when the purpose is related to bit depth increase, it is required that both the network input and output are in integer value format, and the network output must have a higher bit depth than the decoded video output. i. In one example, the network output having a higher bit depth than the decoded video output means that there are no color components for which the network output has a bit depth less than the decoded video output, and that for at least one color component, the network output has a bit depth greater than the decoded video output. ii. In one example, the network output having a higher bit depth than the decoded video output means that for all color components, the bit depth values of the network output are greater than the bit depth values of the decoded video output. f. In one example, when the purpose is related to bit depth increase, the network output is required to be in integer value format and the network output must have a higher bit depth than the decoded video output. i. In one example, the network output having a higher bit depth than the decoded video output means that there are no color components for which the network output has a bit depth less than the decoded video output, and that for at least one color component, the network output has a bit depth greater than the decoded video output. ii. In one example, the network output having a higher bit depth than the decoded video output means that for all color components, the network output has a bit depth value greater than the decoded video output bit depth value. 3) To address issue 3, the range of the syntax element nnpfc_num_input_pics_minus2 is specified as 0 to N (inclusive). a. In one example, N is 2. b. In one example, N is 6. c. In one example, N is 14. d. In one example, N is 30. e. In one example, N is 62. f. In one example, N is 126. g. In one example, N is 254. h. Alternatively, N is 2^K–2, where K is a positive integer. 4) To solve problem 4, the range of the syntax element nnpfc_interpolated_pics[i] is specified to be 0 to N (including the boundary values). a. In one example, N is 1. b. In one example, N is 2. c. In one example, N is 4. d. In one example, N is 8. e. In one example, N is 16. f. In one example, N is 32. g. In one example, N is 64. h. Alternatively, N is 2^K, where K is a positive integer. 5) To solve problem 5, the constraint is that for i in the range of 0 to nnpfc_num_input_pics_minus2 (including the boundary values), at least one of nnpfc_interpolated_pics[i] must be greater than 0. a. Alternatively, nnpfc_interpolated_pics[nnpfc_num_input_pics_minus2] is derived to be non-zero when nnpfc_interpolated_pics[i] is equal to 0 for i in the range of 0 to nnpfc_num_input_pics_minus2–1. i. In one example, nnpfc_interpolated_pics[nnpfc_num_input_pics_minus2] is derived as 1. 6) To solve problem 6, it is proposed to consider the number of input pictures during the calculation of numOutputPics. a. In one example, numOutputPics is the number of input pictures plus the total number of inserted pictures. 7) In order to solve problem 7, it is proposed to use numOutputPics in the process StoreOutputTensors() to derive the sample values in the filtered output sample array from the output tensor outputTensor. 8) To address issue 2, the NNPF objective may indicate bit depth increase (also called bit depth upsampling), possibly together with other types of upsampling. Additionally, one or more of the following aspects apply. a. In one example, regardless of whether the NNPF purpose indicates bit depth increase, it is required that when the bit depth of a color component output by the network is greater than the bit depth of the corresponding color component input to the network, the bit depth of each of the other color components output by the network must be greater than or equal to the bit depth of the corresponding color component input to the network. b. In one example, when the destination indicates increased bit depth, both the network input and output are required to be in integer value format, and the network output must have a higher bit depth than the network input. i. In one example, the network output having a higher bit depth than the network input means that for each color component, the bit depth of the network output is greater than the bit depth of the corresponding color component of the network input. ii. In one example, the network output has a higher bit depth than the network input means that for at least one color component, the bit depth of the network output is greater than the bit depth of the corresponding color component of the network input. c. In one example, when the purpose indicates increased bit depth, both the network input and output are required to be in integer-valued format, and the network output must have a higher bit depth than the cropped output picture output by the video decoder. i. In one example, the network output has a higher bit depth than the cropped output picture means that for each color component, the bit depth of the network output is greater than the bit depth of the corresponding color component of the cropped output picture. ii. In one example, the network output has a higher bit depth than the cropped output picture means that for at least one color component, the bit depth of the network output is greater than the bit depth of the corresponding color component of the cropped output picture. d. In one example, when the destination indicates increased bit depth, the network output is required to be in integer-valued format and the network output must have a higher bit depth than the cropped output picture output by the video decoder. i. In one example, the network output has a higher bit depth than the cropped output picture means that for each color component, the bit depth of the network output is greater than the bit depth of the corresponding color component of the cropped output picture. ii. In one example, the network output has a higher bit depth than the cropped output picture means that for at least one color component, the bit depth of the network output is greater than the bit depth of the corresponding color component of the cropped output picture. 9) To address issue 8, one or more of the following syntactic or semantic changes are specified: a. In one example, the value of the bit of nnpf_purpose (eg, corresponding to 0x020 (ie, bit 5)) indicates whether the purpose is coloring. i. In one example, when nnpfc_purpose & 0x20 is not equal to 0, it indicates that the purpose is coloring. ii. In one example, when nnpfc_purpose & 0x20 is equal to 0, it indicates that the purpose is no coloring or non-coloring. b. In one example, when the purpose indicates coloring, the input to the network is required to be in 4:0:0 format, where the value of the variable ChromaFormatIdc is equal to 0. c. In one example, when the intent indicates shading, there should be chroma output by the NNPF. i. In one example, when the purpose indicates coloring, nnpfc_out_order_idc should not be equal to 0. d. In one example, when the destination indicates shading, the chroma component sample values may be inferred to be equal to 0, 0.5, or 1 when the tensor input is real. e. In one example, when the destination indicates shading, when the tensor input is an integer, the chroma component sample value can be inferred to be equal to 0 or 1 << (inpTensorBitDepth C -1) or 1<<(inpTensorBitDepth C )-1, where inpTensorBitDepth C is the bit depth of the chroma sample values in the input integer tensor. f. In one example, when the intent indicates shading and ChromaFormatIdc is equal to 0, there may only be luma input. i. In one example, when the destination indicates coloring and ChromaFormatIdc is equal to 0, nnpfc_inp_order_idc shall be equal to 0. g. In one example, when the intent indicates chroma and ChromaFormatIdc is equal to 0, chroma-only input is not allowed. i. In one example, when the purpose indicates coloring and ChromaFormatIdc is equal to 0, nnpfc_inp_order_idc should not be equal to 1. h. In one example, when the destination indicates shading, it may be required that the destination does not also indicate chroma upsampling. i. In one example, when nnpfc_purpose&0x20 is not equal to 0, nnpfc_purpose&0x02 is required to be equal to 0. i. In one example, when the destination indicates chroma upsampling, it may be required that the destination does not also indicate shading. i. In one example, when nnpfc_purpose&0x02 is not equal to 0, nnpfc_purpose&0x20 is required to be equal to 0. j. In one example, the purpose of indicating shading and the purpose of indicating chroma upsampling shall be mutually exclusive. i. In one example, ((nnpfc_purpose & 0x02) != 0) && ((nnpfc_purpose & 0x20) != 0) must be equal to 0. k. In one example, when the destination indicates colorization, an indication is signaled in the NNPFC SEI message (eg, a two-bit syntax element named nnpfc_out_colour_format_idc) to specify whether the color format of the output of the NNPF is 4:2:0, 4:2:2, or 4:4:4 format. 1. In one example, when the purpose indicates coloring, an indication is signaled in the NNPFC SEI message (eg, a two-bit syntax element named nnpfc_out_colour_format_idc) to specify the values of the variables outSubWidthC and outSubHeightC. i. In one example, further, the following are specified: nnpfc_out_colour_fomrat_idc equal to 1 specifies that both outSubWidthC and outSubHeightC are equal to 2. nnpfc_out_colour_fomrat_idc equal to 2 specifies that both outSubWidthC and outSubHeightC are equal to 1. nnpfc_out_colour_fomrat_idc equal to 3 specifies that both outSubWidthC and outSubHeightC are equal to 1. The value of nnpfc_out_colour_fomrat_idc shall not be equal to 0. ii. In one example, further specifying that when nnpfc_purpose&0x02 and nnpfc_purpose&0x20 are both equal to 0, outSubWidthC is inferred to be equal to SubWidthC, and outSubHeightC is inferred to be equal to SubHeightC. 10) To address issue 9, one or more of the following syntactic or semantic changes are specified: a. When the method of sample array ordering of the cropped decoded output picture that is one of the input pictures to the post-processing filter indicates that there is no luma matrix, the syntax element indicating the luma input bit depth shall not be signaled. i. In one example, when nnpfc_inp_order_idc is equal to 1, nnpfc_inp_tensor_luma_bitdepth_minus8 should not be signaled. b. When the method of sample array ordering of the cropped decoded output picture that is one of the input pictures to the post-processing filter indicates that no chroma matrix is present, the syntax element indicating the chroma input bit depth shall not be signaled. i. In one example, when nnpfc_inp_order_idc is equal to 0, nnpfc_inp_tensor_chroma_bitdepth_minus8 should not be signaled. c. In one example, syntax element(s) indicating the method of sample array ordering of a cropped decoded output picture that is one of the input pictures to a post-processing filter should be signaled before syntax element(s) indicating luma and / or chroma input bit depth. i. In one example, nnpfc_inp_order_idc should be signaled before nnpfc_inp_tensor_luma_bitdepth_minus8 and / or nnpfc_inp_tensor_chroma_bitdepth_minus8 in the syntax table. d. When the output order of the samples produced from the post-processing filter indicates that no luma matrix is present, the syntax element indicating the luma input bit depth shall not be signaled. i. In one example, when nnpfc_out_order_idc is equal to 1, nnpfc_out_tensor_luma_bitdepth_minus8 should not be signaled. e. When the output order of the samples produced from the post-processing filter indicates that no chroma matrix is present, the syntax element indicating the chroma input bit depth shall not be signaled. i. In one example, when nnpfc_out_order_idc is equal to 0, nnpfc_out_tensor_chroma_bitdepth_minus8 should not be signaled. f. In one example, the syntax element(s) indicating the output order of samples produced from the post-processing filters should be signaled before the syntax element(s) indicating the luma and / or chroma output bit depth. i. In one example, nnpfc_out_order_idc should be signaled before nnpfc_out_tensor_luma_bitdepth_minus8 and / or nnpfc_out_tensor_chroma_bitdepth_minus8 in the syntax table. 11) To address issue 10, one or more of the following syntactic or semantic changes are specified: a. When the NNPF destination indicates chroma, the output chroma format may be indicated by the syntax element only if the input is in 4:0:0 format. b. Alternatively, when the NNPF destination indicates shading, the output chroma format may be indicated by the syntax element regardless of the input chroma format. 12) To address issue 11, one or more of the following syntactic or semantic changes are specified: a. A flag is signaled to indicate whether the chroma sample positions are signaled. b. When the destination indicates chroma and the input is in 4:0:0 format, the flag described in the above sub-item shall be true (indicating that the chroma sample positions are being signaled). c. When the chroma sample positions are not signaled in the NNPF SEI message, they are assumed to be the same as in the VUI. 6. Examples The following are some example embodiments of the solution aspects outlined in Section 5 above. Most of the relevant parts have been added or modified Underline Shown, and some of the deleted parts are shown with strikethrough. There may be some other changes that are editorial in nature and therefore not highlighted. Example 1 This embodiment addresses solution item 1 and all its subitems outlined in subsection 5 above. 6.2. Example 2 This embodiment addresses Solution Item 2 and all its subitems outlined in Section 5 above. 6.3. Example 3 This embodiment addresses solution items 1 and 2 outlined in subsection 5 above, and all their subitems. 6.4. Example 4 This embodiment addresses solution item 3 and all its subitems outlined in subsection 5 above. 6.5. Example 5 This embodiment addresses solution item 4 and all its subitems outlined in subsection 5 above. 6.6. Example 6 This embodiment addresses Solution Item 5 and all its subitems outlined in Section 5 above. 6.7. Example 7 This embodiment addresses solution item 6 and all its subitems outlined in subsection 5 above. 6.8. Example 8 This embodiment addresses solution item 7 and all its subitems outlined in subsection 5 above. 6.9. Example 9 This embodiment addresses Solution Item 2 and all its subitems outlined in Section 5 above. 6.10. Example 10 This embodiment is directed to solution item 11 and all its subitems outlined in subsection 5 above. Example 11 This embodiment is directed to solution item 11 and all its subitems outlined in subsection 5 above. 6.12. Example 12 This embodiment is directed to solution item 12 and all its subitems outlined in subsection 5 above.
[0062] More details of embodiments of the present disclosure related to neural network post-processing filters will be described below. As used herein, the terms "neural network post-processing filter" and "neural network post filter" can be used interchangeably. The embodiments of the present disclosure should be considered as examples to explain general concepts and should not be interpreted in a narrow sense. In addition, these embodiments can be applied alone or combined in any way.
[0063] Figure 5 FIG. 5 is a flow chart showing a method 500 for video processing according to some embodiments of the present disclosure. Figure 5 As shown, at 502, conversion between a video and a bitstream of the video is performed. In some embodiments, the conversion may include encoding the video into a bitstream. Alternatively or additionally, the conversion may include decoding the video from the bitstream.
[0064] A neural network post-processing filter (NNPF) is applied to at least one picture associated with a video. For example, the at least one picture may be used as input to the NNPF. In some embodiments, the at least one picture may include at least one decoded picture of the video. Alternatively, the at least one picture may include at least one cropped decoded picture of the video. For example, the decoded picture and / or the cropped decoded picture may be output by a decoder that decodes the video from a bitstream. In some further embodiments, the at least one picture may include the output of another NNPF for filtering one or more decoded pictures or cropped decoded pictures of the video. For example, the NNPF is spliced with another NNPF. It should be understood that the possible implementations of the at least one picture associated with the video described herein are illustrative only and should not be construed in any way as limiting the present disclosure.
[0065] The bitstream includes a first indication indicating the purpose of the NNPF. For example, the first indication may include a syntax element nnpfc_purpose. For example, the first indication may be included in a supplemental enhancement information (SEI) message or any other suitable video message unit in the bitstream. It should be understood that the names of the indications and / or syntax elements are for illustration only and not limitation, and the (multiple) indications and (multiple) syntax elements mentioned throughout this disclosure may be represented by any other suitable string other than the strings mentioned in this disclosure. The scope of the present disclosure is not limited in this respect.
[0066] Furthermore, one of the candidates for the purpose is colorization of at least one image. For example, if the purpose of the NNPF includes colorization, a 4:0:0 format image can be processed to output a 4:2:0 format image, etc. In another example, a black and white image can be converted to a color image using the NNPF.
[0067] Given the above, NNPF is able to support colorization of NNPF input. Compared to traditional solutions, the proposed method can advantageously support applications that require colorization and applications with input in 4:0:0 format. Thus, in this way, the functionality of NNPF is diversified and the encoding and decoding quality can be improved.
[0068] In one example embodiment, the plurality of candidates for the purpose of NNPF may include a change in the color format of at least one picture. For example, the change in color format may include a change from a color format to another color format having a chroma downsampling rate that is less than the chroma downsampling rate of the color format, such as a change from a 4:0:0 format to a 4:2:0 format, a change from a 4:2:0 format to a 4:2:2 format, etc. As used herein, the terms "chroma format" and "color format" may be used interchangeably.
[0069] In another example embodiment, the plurality of candidates for the purpose of the NNPF may include a change in the picture resolution of at least one picture. For example, the change in picture resolution may include an increase in picture resolution. Additionally or alternatively, the plurality of candidates for the purpose of the NNPF may include a change in the picture rate of at least one picture. For example, the change in picture rate may include an increase in picture rate.
[0070] In some additional or alternative embodiments, the plurality of candidates for the purpose of NNPF may include a change in the bit depth of sample values in at least one picture. For example, the change in bit depth may include an increase in bit depth. In this case, the candidate for the purpose may also be referred to as a bit depth increase or bit depth upsampling.
[0071] As described above, each of the colorization of the input of the NNPF, the change of the color format, the change of the image resolution, the change of the image rate, and the change of the bit depth can be a candidate for the purpose of the NNPF. It should be understood that the multiple candidates for the purpose of the NNPF can also include any other suitable process, such as improving visual quality, etc. The scope of the present disclosure is not limited in this respect.
[0072] In some embodiments, the purpose of the NNPF may include a combination of multiple candidates for the purpose. For example, the purpose of the NNPF is allowed to include a combination of at least two of the following: a change in the chroma format of at least one picture, a change in the picture rate of at least one picture, a change in the resolution of at least one picture, a change in the bit depth of the sample values in at least one picture, or coloring of at least one picture.
[0073] In some embodiments, the value of the bit in the first indication may indicate whether the purpose of the NNPF includes the coloring of at least one picture. In this case, if the result of applying a bitwise AND operation to the first indication and the first bit mask is not equal to the first value (such as 0, etc.), the purpose of the NNPF may include the coloring of at least one picture. If the result of applying a bitwise AND operation to the first indication and the first bit mask is equal to the first value, the purpose of the NNPF does not include the coloring of at least one picture. The proposed method advantageously provides a system solution for transmitting the purpose of the NNPF by signaling, thereby supporting possible expansion of the purpose. In this way, potential instabilities and logical problems can be avoided, and encoding and decoding efficiency can be improved.
[0074] For the purpose of illustration, it is assumed that the value of bit 5 in the syntax element nnpfc_purpose indicates whether the purpose of the NNPF includes shading of at least one picture, while bit 0 may represent the least significant bit in the syntax element nnpfc_purpose. If (nnpfc_purpose & 0x20) is not equal to 0, it is determined that the purpose of the NNPF includes shading. If (nnpfc_purpose & 0x20) is equal to 0, it is determined that the purpose of the NNPF does not include shading. The operator "&" represents a bitwise AND operation. It should be understood that the specific values described herein are intended to be examples and not to limit the scope of the present disclosure.
[0075] In some embodiments, if the purpose of the first indication for indicating the NNPF includes coloring of at least one picture, the chroma format of the at least one picture may be a 4:0:0 format, which may also be referred to as "monochrome." When the chroma format of the at least one picture is a 4:0:0 format, the value of the variable ChromaFormatIdc may be equal to 0.
[0076] In some embodiments, if the first indication is used to indicate that the purpose of the NNPF includes chroma upsampling of at least one picture, then the first indication is used to indicate that the purpose of the NNPF does not include coloring of at least one picture. For example, the purpose of the NNPF is not allowed to include chroma upsampling and coloring at the same time. If the result of applying a bitwise AND operation to the first indication and the second bit mask (e.g., 0x02, etc.) is not equal to the first value (e.g., 0, etc.), then the result of applying a bitwise AND operation to the first indication and the first bit mask (e.g., 0x20, etc.) is equal to the first value. For example, when nnpfc_purpose&0x02 is not equal to 0, it is required that nnpfc_purpose&0x20 should be equal to 0.
[0077] In some embodiments, if the first indication is used to indicate that the purpose of the NNPF includes coloring of at least one picture, the bitstream may also include a second indication indicating the color format of the output of the NNPF. In addition, the second indication may further be used to indicate the values of the variables outSubWidthC and outSubHeightC. For example, the second indication may include the syntax element nnpfc_out_colour_format_idc. The second indication is included in the Neural Network Post-Processing Filter Characteristics (NNPFC) Supplemental Enhancement Information (SEI) message in the bitstream.
[0078] For example, a second indication equal to a second value (such as 1) may be used to indicate that variables outSubWidthC and outSubHeightC are both equal to 2. A second indication equal to a third value (such as 2) may be used to indicate that variable outSubWidthC is equal to 2 and variable outSubHeightC is equal to 1. A second indication equal to a fourth value (such as 3) may be used to indicate that variables outSubWidthC and outSubHeightC are both equal to 1. In addition, the second indication is not allowed to be equal to a fifth value (such as 0).
[0079] In some embodiments, if the first indication is used to indicate that the purpose of the NNPF does not include shading of at least one picture and chroma upsampling of at least one picture, the color format of the output of the NNPF is the same as the color format of the at least one picture. For example, if the result of applying a bitwise AND operation to the first indication and a first bit mask (such as 0x20, etc.) is equal to a first value (such as 0, etc.) and the result of applying a bitwise AND operation to the first indication and a second bit mask (such as 0x02) is equal to the first value, then the variable outSubWidthC is equal to the variable SubWidthC, and the variable outSubHeightC is equal to the variable SubHeightC. For example, if (nnpf_purpose & 0x02) and (nnpfc_purpose & 0x20) are both equal to 0, then outSubWidthC is inferred to be equal to SubWidthC, and outSubHeightC is inferred to be equal to SubHeightC.
[0080] In some embodiments, the bitstream may further include a third indication for indicating the number of at least one picture. The value of the third indication is within a first predetermined range. For example, the first predetermined range is a range from 0 to N (including boundary values), where N is an integer such as 2, 6, 14, 30, 62, 126, 254, or (2^K-2), and K is a positive integer. By way of example and not limitation, the third indication may include syntax elements named nnpfc_num_input_pics_minus2, nnpfc_num_input_pics_minus1, etc.
[0081] In some embodiments, the bitstream may further include a fourth indication for indicating the number of interpolated pictures generated by the NNPF between the i-th picture and the (i+1)-th picture in the at least one picture, where i is an integer. The value of the fourth indication is within a second predetermined range. For example, the second predetermined range is a range from 0 to M (including boundary values), where M is an integer such as 1, 2, 4, 8, 16, 32, 64, or 2^K, and K is a positive integer. By way of example and not limitation, the fourth indication may include a syntax element nnpfc_interpolated_pics[i].
[0082] Additionally or alternatively, at least one interpolated picture is generated by the NNPF between two pictures in the at least one picture. For example, for at least one value of i in the range of 0 to an upper limit (inclusive), at least one of the fourth indications is greater than 0, and the upper limit is equal to the number of the at least one picture minus 2. For example, if all syntax elements nnpfc_interpolated_pics[i] for i in the range of 0 to K1 are equal to 0, then nnpfc_interpolated_pics[K2] is a sixth value greater than 0, such as 1, etc., wherein K1 is equal to the number of the at least one picture minus 3, and K2 is equal to the number of the at least one picture minus 2.
[0083] In some embodiments, the number of pictures in the output of the NNPF is determined based on the number of at least one picture. For example, the number of pictures in the output of the NNPF is determined based on the number of at least one picture and the total number of inserted pictures generated by the NNPF. For purposes of illustration and not limitation, an example of this is shown in Example 7 above.
[0084] In some embodiments, the process for determining the sample values in the filtered output sample array from the output tensor of the NNPF is performed using the number of pictures in the output tensor of the NNPF. For illustration and not limitation, an example of this is shown in the above embodiment 8.
[0085] In some embodiments, if the input tensor of the NNPF is in real format, and the first indication is used to indicate that the purpose of the NNPF includes coloring of at least one picture, the sample value in the chroma component of the input tensor can be equal to a predetermined real number, such as 0, 0.5, 0.9, 1, etc.
[0086] Additionally or alternatively, if the input tensor of the NNPF is in the format of integers, and the first indication is used to indicate that the purpose of the NNPF includes the shading of at least one picture, the sample value in the chroma component of the input tensor may be equal to a predetermined integer, such as 0. Alternatively, if the input tensor of the NNPF is in the format of integers, and the first indication is used to indicate that the purpose of the NNPF includes the shading of at least one picture, the sample value in the chroma component may be determined based on the bit depth of the chroma sample value in the input tensor, such as 1<<(inpTensorBitDepth C -1) or 1<<(inpTensorBitDepth C )-1, where inpTensorBitDepth C Represents the bit depth of the chroma sample values in the input tensor.
[0087] In some embodiments, if the color format of at least one picture is monochrome and the first indication is used to indicate that the purpose of the NNPF includes shading of at least one picture, then only the luma matrix is present in the input tensor of the NNPF. For example, if the variable ChromaFormatIdc is equal to 0 and the first indication is used to indicate that the purpose of the NNPF includes shading of at least one picture, then the syntax element nnpfc_inp_order_idc shall be equal to 0.
[0088] In some embodiments, if the color format of at least one picture is monochrome and the first indication is used to indicate that the purpose of the NNPF includes coloring of at least one picture, at least one matrix other than the chroma matrix is present in the input tensor of the NNPF. That is, in this case, only chroma input is allowed. For example, if the variable ChromaFormatIdc is equal to 0 and the first indication is used to indicate that the purpose of the NNPF includes coloring of at least one picture, the syntax element nnpfc_inp_order_idc should not be equal to 1.
[0089] In some embodiments, if the first indication is used to indicate that the purpose of the NNPF includes colorization of at least one picture, then the purpose of the NNPF does not include chroma upsampling. For example, when (nnpfc_purpose & 0x20) is not equal to 0, it is required that (nnpfc_purpose & 0x02) should be equal to 0.
[0090] In some further embodiments, the purpose of the NNPF including colorization of at least one picture and the purpose of the NNPF including chroma upsampling are mutually exclusive. For example, if the purpose of the NNPF includes chroma upsampling, the purpose of the NNPF does not include colorization of at least one picture. If the purpose of the NNPF includes colorization of at least one picture, the purpose of the NNPF does not include chroma upsampling.
[0091] In view of the above situation, the solutions according to some embodiments of the present disclosure can advantageously avoid potential instability and logic problems, thereby improving encoding and decoding efficiency.
[0092] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. In the method, conversion between the video and the bitstream is performed. A neural network post-processing filter (NNPF) is applied to at least one picture associated with the video. The bitstream includes a first indication indicating a purpose of the NNPF, and one of the candidates for the purpose is a coloring of the at least one picture.
[0093] According to some further embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In the method, conversion between a video and a bitstream is performed. A neural network post-processing filter (NNPF) is applied to at least one picture associated with the video. The bitstream includes a first indication indicating a purpose of the NNPF, and one of the candidates for the purpose is a coloring of at least one picture. In addition, the bitstream is stored in a non-transitory computer-readable recording medium.
[0094] The embodiments of the present disclosure may be described according to the following items, the features of which may be combined in any reasonable way.
[0095] Item 1. A method for video processing, comprising: performing a conversion between a video and a bitstream of the video, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating a purpose of the NNPF, and one of the candidates for the purpose is a colorization of the at least one picture.
[0096] Clause 2. The method of clause 1, wherein the first indication comprises a syntax element nnpfc_purpose.
[0097] Item 3. A method according to any one of Items 1 to 2, wherein the purpose of the NNPF is allowed to include a combination of at least two of the following: a change in the chroma format of the at least one picture, a change in the picture rate of the at least one picture, a change in the resolution of the at least one picture, a change in the bit depth of the sample values in the at least one picture, or the coloring of the at least one picture.
[0098] Clause 4. A method according to any one of clauses 1 to 3, wherein the value of a bit in the first indication indicates whether the purpose of the NNPF includes the coloring of the at least one picture.
[0099] Item 5. A method according to Item 4, wherein if the result of applying the bitwise AND operation to the first indication and the first bit mask is not equal to a first value, the purpose of the NNPF includes the coloring of the at least one picture, or if the result of applying the bitwise AND operation to the first indication and the first bit mask is equal to the first value, the purpose of the NNPF does not include the coloring of the at least one picture.
[0100] Clause 6. The method of clause 5, wherein the first bit mask is 0x20, or the first value is 0.
[0101] Item 7. A method according to any one of Items 1 to 6, wherein if the purpose of the first indication for indicating the NNPF includes the coloring of the at least one picture, the chroma format of the at least one picture is a 4:0:0 format.
[0102] Item 8. A method according to any one of Items 1 to 7, wherein if the first indication is used to indicate that the purpose of the NNPF includes chroma upsampling of the at least one picture, then the first indication is used to indicate that the purpose of the NNPF does not include the coloring of the at least one picture.
[0103] Item 9. The method of Item 8, wherein if a result of applying a bitwise AND operation to the first indication and the second bit mask is not equal to a first value, then a result of applying a bitwise AND operation to the first indication and the first bit mask is equal to the first value.
[0104] Item 10. The method of Item 9, wherein the first bit mask is 0x20, the second bit mask is 0x02, or the first value is 0.
[0105] Item 11. A method according to any one of Items 1 to 10, wherein if the first indication is used to indicate that the purpose of the NNPF includes the coloring of the at least one picture, the bitstream also includes a second indication for indicating the color format of the output of the NNPF.
[0106] Clause 12. The method of clause 11, wherein the second indication is further used to indicate values of variables outSubWidthC and outSubHeightC.
[0107] Item 13. A method according to Item 12, wherein the second indication equal to the second value is used to indicate that the variables outSubWidthC and outSubHeightC are both equal to 2, or the second indication equal to the third value is used to indicate that the variable outSubWidthC is equal to 2 and the variable outSubHeightC is equal to 1, or the second indication equal to the fourth value is used to indicate that the variables outSubWidthC and outSubHeightC are both equal to 1, or the second indication is not allowed to be equal to the fifth value.
[0108] Item 14. The method of Item 13, wherein the second value is 1, the third value is 2, the fourth value is 3, or the fifth value is 0.
[0109] Item 15. A method according to any one of Items 11 to 14, wherein the second indication is included in a Neural Network Post-Processing Filter Characteristics (NNPFC) Supplemental Enhancement Information (SEI) message in the bitstream.
[0110] Clause 16. A method according to any of clauses 11 to 15, wherein the second indication comprises a syntax element nnpfc_out_colour_format_idc.
[0111] Item 17. A method according to any one of Items 1 to 16, wherein if the first indication is used to indicate that the purpose of the NNPF does not include the shading of the at least one picture and the chroma upsampling of the at least one picture, the color format of the output of the NNPF is the same as the color format of the at least one picture.
[0112] Item 18. A method according to Item 17, wherein if the result of applying a bitwise AND operation to the first indication and the first bit mask is equal to a first value and the result of applying a bitwise AND operation to the first indication and the second bit mask is equal to the first value, then the variable outSubWidthC is equal to the variable SubWidthC, and the variable outSubHeightC is equal to the variable SubHeightC.
[0113] Item 19. The method of Item 18, wherein the first bit mask is 0x20, the second bit mask is 0x02, or the first value is 0.
[0114] Item 20. A method according to any one of Items 1 to 19, wherein the bitstream further comprises a third indication for indicating the number of the at least one picture, and a value of the third indication is within a first predetermined range.
[0115] Item 21. The method of Item 20, wherein the first predetermined range is a range from 0 to N, inclusive, and N is an integer.
[0116] Item 22. The method of Item 21, wherein N is one of the following: 2, 6, 14, 30, 62, 126, 254, or (2^K-2), and K is a positive integer.
[0117] Item 23. A method according to any one of Items 1 to 22, wherein the bitstream further includes a fourth indication for indicating the number of inserted pictures generated by the NNPF between the i-th picture and the (i+1)-th picture in the at least one picture, i is an integer, and the value of the fourth indication is within a second predetermined range.
[0118] Clause 24. The method of clause 23, wherein the fourth indication comprises a syntax element nnpfc_interpolated_pics[i].
[0119] Item 25. A method according to any one of Items 23 to 24, wherein the second predetermined range is a range from 0 to M, inclusive, and M is an integer.
[0120] Item 26. The method of any one of Items 23 to 25, wherein at least one inserted picture is generated by the NNPF between two pictures of the at least one picture.
[0121] Item 27. A method according to any one of Items 23 to 26, wherein for at least one value of i in the range of 0 to an upper limit, inclusive, at least one of the fourth indications is greater than 0, and the upper limit is equal to the number of the at least one picture minus 2.
[0122] Clause 28. A method according to any of clauses 1 to 27, wherein the number of pictures in the output of the NNPF is determined based on the number of the at least one picture.
[0123] Item 29. The method of Item 28, wherein the number of pictures in the output of the NNPF is determined based on the number of the at least one picture and the total number of inserted pictures generated by the NNPF.
[0124] Item 30. A method according to any one of Items 1 to 29, wherein the process for determining the sample values in the filtered output sample array from the output tensor of the NNPF is performed by using the number of pictures in the output tensor of the NNPF.
[0125] Item 31. A method according to any one of Items 1 to 2, wherein the chroma format of the at least one picture is a 4:0:0 format.
[0126] Item 32. A method according to item 24, wherein if the syntax elements nnpfc_interpolated_pics[i] for i in the range of 0 to K1 are all equal to 0, then nnpfc_interpolated_pics[K2] is a sixth value greater than 0, K1 is equal to the number of the at least one picture minus 3, and K2 is equal to the number of the at least one picture minus 2.
[0127] Item 33. The method of Item 32, wherein the sixth value is 1.
[0128] Item 34. The method of Item 25, wherein M is one of the following: 1, 2, 4, 8, 16, 32, 64, or 2^K, where K is a positive integer.
[0129] Item 35. A method according to any one of Items 1 to 34, wherein if an input tensor of the NNPF is in real format and the first indication is used to indicate that the purpose of the NNPF includes the shading of the at least one picture, then the sample value in the chroma component of the input tensor is equal to a predetermined real number.
[0130] Item 36. A method according to any one of Items 1 to 35, wherein if an input tensor of the NNPF is in integer format and the first indication is used to indicate that the purpose of the NNPF includes the shading of the at least one picture, then the sample value in the chroma component of the input tensor is equal to a predetermined integer or is determined based on the bit depth of the chroma sample value in the input tensor.
[0131] Item 37. A method according to any one of Items 1 to 36, wherein if the color format of the at least one picture is monochrome and the first indication is used to indicate that the purpose of the NNPF includes the coloring of the at least one picture, then only the luminance matrix is present in the input tensor of the NNPF.
[0132] Item 38. The method of Item 37, wherein if a variable ChromaFormatIdc is equal to 0 and the first indication is for indicating that the purpose of the NNPF includes the coloring of the at least one picture, then the syntax element nnpfc_inp_order_idc is equal to 0.
[0133] Item 39. A method according to any one of Items 1 to 38, wherein if the color format of the at least one picture is monochrome and the first indication is used to indicate that the purpose of the NNPF includes the coloring of the at least one picture, then at least one matrix other than the chroma matrix is present in the input tensor of the NNPF.
[0134] Item 40. The method of Item 39, wherein if a variable ChromaFormatIdc is equal to 0 and the first indication is for indicating that the purpose of the NNPF includes the coloring of the at least one picture, then the syntax element nnpfc_inp_order_idc is not equal to 1.
[0135] Item 41. A method according to any one of Items 1 to 40, wherein if the first indication is used to indicate that the purpose of the NNPF includes the coloring of the at least one picture, then the purpose of the NNPF does not include chroma upsampling.
[0136] Item 42. A method according to any one of Items 1 to 41, wherein the purpose of the NNPF comprising the colorization of the at least one picture and the purpose of the NNPF comprising chroma upsampling are mutually exclusive.
[0137] Item 43. The method of Item 41, wherein if the purpose of the NNPF includes the chroma upsampling, the purpose of the NNPF does not include the coloring of the at least one picture.
[0138] Item 44. A method according to any one of Items 1 to 43, wherein the at least one picture comprises at least one decoded picture of the video or at least one cropped decoded picture.
[0139] Item 45. The method of any one of Items 1 to 44, wherein the converting comprises encoding the video into the bitstream.
[0140] Item 46. A method according to any one of Items 1 to 44, wherein the converting comprises decoding the video from the bitstream.
[0141] Item 47. An apparatus for video processing, comprising a processor and non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1 to 46.
[0142] Item 48. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 46.
[0143] Item 49. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: performing a conversion between the video and the bitstream, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating a purpose of the NNPF, and one of the candidates for the purpose is a colorization of the at least one picture.
[0144] Item 50. A method for storing a bitstream of a video, comprising: performing a conversion between the video and the bitstream, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating a purpose of the NNPF, and one of the candidates for the purpose is a colorization of the at least one picture; and storing the bitstream in a non-transitory computer-readable recording medium. Example device
[0145] Figure 6 A block diagram of a computing device 600 in which various embodiments of the present disclosure may be implemented is shown. The computing device 600 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0146] It should be understood that Figure 6 The computing device 600 shown in FIG. 6 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the disclosed embodiments.
[0147] like Figure 6 As shown, computing device 600 comprises a general computing device 600. Computing device 600 may include at least one or more processors or processing units 610, memory 620, storage unit 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660.
[0148] In some embodiments, the computing device 600 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, a large computing device, etc. provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 600 can support any type of interface to the user (such as a "wearable" circuit device, etc.).
[0149] The processing unit 610 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capability of the computing device 600. The processing unit 610 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0150] The computing device 600 typically includes various computer storage media. Such media can be any media accessible by the computing device 600, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory), or any combination thereof. The storage unit 630 can be any removable or non-removable medium and can include machine-readable media, such as memory, flash drive, disk, or other media that can be used to store information and / or data and can be accessed in the computing device 600.
[0151] The computing device 600 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 6 Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.
[0152] The communication unit 640 communicates with another computing device via a communication medium. In addition, the functionality of the components in the computing device 600 can be implemented by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0153] Input device 650 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and the like. Output device 660 may be one or more of various output devices, such as a display, speaker, printer, and the like. With the aid of communication unit 640, computing device 600 may also communicate with one or more external devices (not shown), such as storage devices and display devices, one or more devices that enable a user to interact with computing device 600, or, if desired, any device that enables computing device 600 to communicate with one or more other computing devices (e.g., a network card, a modem, and the like). Such communication may be performed via an input / output (I / O) interface (not shown).
[0154] In some embodiments, some or all components of the computing device 600 may also be arranged in a cloud computing architecture rather than being integrated into a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides an application via a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on servers in a remote location. Computing resources in a cloud computing environment can be consolidated or distributed across remote data centers. Cloud computing infrastructure can provide services through shared data centers, although to users, they appear as a single access point. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider in a remote location. Alternatively, the components and functionality described herein can be provided by a conventional server or installed directly or otherwise on a client device.
[0155] In an embodiment of the present disclosure, the computing device 600 may be used to implement video encoding / decoding. The memory 620 may include one or more video encoding / decoding modules 625 having one or more program instructions. These modules are accessible and executable by the processing unit 610 to perform the functions of the various embodiments described herein.
[0156] In an example embodiment performing video encoding, an input device 650 may receive video data as input to be encoded 670. The video data may be processed, for example, by a video codec module 625 to generate an encoded bitstream. The encoded bitstream may be provided as output 680 via an output device 660.
[0157] In an example embodiment performing video decoding, an input device 650 may receive an encoded bitstream as input 670. The encoded bitstream may be processed, for example, by a video codec module 625 to generate decoded video data. The decoded video data may be provided as output 680 via an output device 660.
[0158] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such variations are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for video processing, comprising: Performing conversion between a video and a bitstream of the video, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating a purpose of the NNPF, and one of the candidates for the purpose is a colorization of the at least one picture. The method according to claim 1 , wherein the first indication comprises a syntax element nnpfc_purpose.
3. The method according to any one of claims 1 to 2, wherein the purpose of the NNPF is allowed to include a combination of at least two of the following: a change in the chroma format of the at least one picture, a change in the picture rate of the at least one picture, a change in the resolution of the at least one picture, a change in the bit depth of the sample values in the at least one picture, or said coloring of said at least one picture. 4 . The method according to claim 1 , wherein a value of a bit in the first indication indicates whether the purpose of the NNPF includes the coloring of the at least one picture.
5. The method of claim 4 , wherein if a result of applying a bitwise AND operation to the first indication and a first bit mask is not equal to a first value, the purpose of the NNPF comprises the coloring of the at least one picture, or If the result of applying the bitwise AND operation on the first indication and the first bit mask is equal to the first value, the purpose of the NNPF does not include the coloring of the at least one picture. The method of claim 5 , wherein the first bit mask is 0x20, or the first value is 0.
7. The method according to any one of claims 1 to 6, wherein if the purpose of the first indication for indicating the NNPF includes the coloring of the at least one picture, the chroma format of the at least one picture is a 4:0:0 format.
8. The method according to any one of claims 1 to 7, wherein if the first indication is used to indicate that the purpose of the NNPF includes chroma upsampling of the at least one picture, then the first indication is used to indicate that the purpose of the NNPF does not include the coloring of the at least one picture. 9 . The method of claim 8 , wherein if a result of applying a bitwise AND operation to the first indication and the second bit mask is not equal to a first value, a result of applying a bitwise AND operation to the first indication and the first bit mask is equal to the first value. 10 . The method of claim 9 , wherein the first bit mask is 0x20, the second bit mask is 0x02, or the first value is 0.
11. A method according to any one of claims 1 to 10, wherein if the first indication is used to indicate that the purpose of the NNPF includes the coloring of the at least one picture, the bitstream also includes a second indication for indicating the color format of the output of the NNPF. 12 . The method according to claim 11 , wherein the second indication is further used to indicate values of variables outSubWidthC and outSubHeightC.
13. The method according to claim 12, wherein the second indication equal to the second value is used to indicate that the variables outSubWidthC and outSubHeightC are both equal to 2, or The second indication equal to the third value is used to indicate that the variable outSubWidthC is equal to 2 and the variable outSubHeightC is equal to 1, or The second indication equal to the fourth value is used to indicate that the variables outSubWidthC and outSubHeightC are both equal to 1, or The second indication is not allowed to be equal to a fifth value. The method of claim 13 , wherein the second value is 1, the third value is 2, the fourth value is 3, or the fifth value is 0.
15. The method of any one of claims 11 to 14, wherein the second indication is included in a Neural Network Post-Processing Filter Characteristics (NNPFC) Supplemental Enhancement Information (SEI) message in the bitstream.
16. The method according to any one of claims 11 to 15, wherein the second indication comprises a syntax element nnpfc_out_colour_format_idc.
17. A method according to any one of claims 1 to 16, wherein if the first indication is used to indicate that the purpose of the NNPF does not include the coloring of the at least one picture and the chroma upsampling of the at least one picture, then the color format of the output of the NNPF is the same as the color format of the at least one picture.
18. The method of claim 17 , wherein if a result of applying a bitwise AND operation to the first indication and a first bit mask is equal to a first value and a result of applying a bitwise AND operation to the first indication and a second bit mask is equal to the first value, then a variable outSubWidthC is equal to a variable SubWidthC, and a variable outSubHeightC is equal to a variable SubHeightC.
19. The method of claim 18, wherein the first bit mask is 0x20, the second bit mask is 0x02, or the first value is 0.
20. The method according to any one of claims 1 to 19, wherein the bitstream further comprises a third indication for indicating the number of the at least one picture, and a value of the third indication is within a first predetermined range.
21. The method of claim 20, wherein the first predetermined range is a range from 0 to N, inclusive, and N is an integer.
22. The method of claim 21, wherein N is one of the following: 2, 6, 14, 30, 62, 126, 254, or (2^K-2), and K is a positive integer.
23. The method according to any one of claims 1 to 22, wherein the bitstream further includes a fourth indication, the fourth indication being used to indicate the number of inserted pictures generated by the NNPF between the i-th picture and the (i+1)-th picture in the at least one picture, i being an integer, and the value of the fourth indication being within a second predetermined range.
24. The method of claim 23, wherein the fourth indication comprises a syntax element nnpfc_interpolated_pics[i].
25. The method according to any one of claims 23 to 24, wherein the second predetermined range is a range from 0 to M, inclusive, and M is an integer.
26. The method according to any one of claims 23 to 25, wherein at least one inserted picture is generated by the NNPF between two pictures of the at least one picture.
27. A method according to any one of claims 23 to 26, wherein for at least one value of i in the range of 0 to an upper limit, inclusive, at least one of the fourth indications is greater than 0, and the upper limit is equal to the number of the at least one picture minus 2.
28. The method according to any one of claims 1 to 27, wherein the number of pictures in the output of the NNPF is determined based on the number of the at least one picture.
29. The method of claim 28, wherein the number of pictures in the output of the NNPF is determined based on the number of the at least one picture and the total number of inserted pictures generated by the NNPF.
30. The method of any one of claims 1 to 29, wherein the process for determining sample values in a filtered output sample array from an output tensor of the NNPF is performed by using the number of pictures in the output tensor of the NNPF.
31. The method according to any one of claims 1 to 2, wherein the chroma format of the at least one picture is a 4:0:0 format.
32. The method of claim 24, wherein if the syntax elements nnpfc_interpolated_pics[i] for i in the range of 0 to K1 are all equal to 0, then nnpfc_interpolated_pics[K2] is a sixth value greater than 0, K1 is equal to the number of the at least one picture minus 3, and K2 is equal to the number of the at least one picture minus 2. The method of claim 32 , wherein the sixth value is 1.
34. The method of claim 25, wherein M is one of the following: 1, 2, 4, 8, 16, 32, 64, or 2^K, and K is a positive integer.
35. A method according to any one of claims 1 to 34, wherein if the input tensor of the NNPF is in real format, and the first indication is used to indicate that the purpose of the NNPF includes the coloring of the at least one picture, then the sample value in the chroma component of the input tensor is equal to a predetermined real number.
36. A method according to any one of claims 1 to 35, wherein if the input tensor of the NNPF is in integer format, and the first indication is used to indicate that the purpose of the NNPF includes the shading of the at least one picture, then the sample value in the chroma component of the input tensor is equal to a predetermined integer or is determined based on the bit depth of the chroma sample value in the input tensor.
37. A method according to any one of claims 1 to 36, wherein if the color format of the at least one picture is monochrome, and the first indication is used to indicate that the purpose of the NNPF includes the coloring of the at least one picture, then only the brightness matrix is present in the input tensor of the NNPF.
38. The method of claim 37, wherein if a variable ChromaFormatIdc is equal to 0 and the first indication is used to indicate that the purpose of the NNPF includes the coloring of the at least one picture, then a syntax element nnpfc_inp_order_idc is equal to 0.
39. A method according to any one of claims 1 to 38, wherein if the color format of the at least one picture is monochrome, and the first indication is used to indicate that the purpose of the NNPF includes the coloring of the at least one picture, then at least one matrix different from the chroma matrix is present in the input tensor of the NNPF.
40. The method of claim 39, wherein if a variable ChromaFormatIdc is equal to 0 and the first indication is used to indicate that the purpose of the NNPF includes the coloring of the at least one picture, then a syntax element nnpfc_inp_order_idc is not equal to 1.
41. The method of any one of claims 1 to 40, wherein if the first indication is used to indicate that the purpose of the NNPF includes the coloring of the at least one picture, then the purpose of the NNPF does not include chroma upsampling.
42. The method according to any one of claims 1 to 41, wherein the purpose of the NNPF comprising the coloring of the at least one picture and the purpose of the NNPF comprising chroma upsampling are mutually exclusive.
43. The method of claim 41, wherein if the purpose of the NNPF includes the chroma upsampling, the purpose of the NNPF does not include the coloring of the at least one picture.
44. The method of any one of claims 1 to 43, wherein the at least one picture comprises at least one decoded picture or at least one cropped decoded picture of the video.
45. The method of any one of claims 1 to 44, wherein the converting comprises encoding the video into the bitstream.
46. The method of any one of claims 1 to 44, wherein the converting comprises decoding the video from the bitstream.
47. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1 to 46.
48. A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to execute the method according to any one of claims 1 to 46.
49. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: Performing a conversion between the video and the bitstream, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating a purpose of the NNPF, and one of the candidates for the purpose is a colorization of the at least one picture.
50. A method for storing a bitstream of a video, comprising: performing conversion between the video and the bitstream, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream comprising a first indication indicating a purpose of the NNPF, and one of the candidates for the purpose is a colorization of the at least one picture; as well as The bitstream is stored in a non-transitory computer-readable recording medium.