Method and device for video processing and medium
By introducing a neural network post-processing filter (NNPF) in the video encoding and decoding process, the problem of insufficient encoding and decoding efficiency in the prior art is solved, and a more efficient video processing effect is achieved.
Patent Information
- Application Number
- CN202480006872.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-07
- Filing Date
- 2024-01-05
- Publication Date
- 2025-08-12
AI Technical Summary
The existing video encoding and codec technology has room for improvement in encoding and codec efficiency and performance, especially in the multi-function video encoding and codec (VVC) standard, which requires further improvement of video processing efficiency.
The bitstream of the video unit is converted using a neural network post-processing filter (NNPF) and applied to a set of pictures in sequence. The conversion is performed through the output of the NNPF to improve the encoding and decoding efficiency and performance.
Through the application of NNPF, the encoding and decoding efficiency and performance of video processing are improved, and the requirements of multi-functional video encoding and decoding standards are met.
Smart Images

Figure CN120476584A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly, to neural network post-processing filters with multiple picture inputs. Background Art
[0002] Digital video capabilities are now being used in every aspect of our lives. For video encoding and decoding, various video compression technologies have been proposed, including MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC), and Versatile Video Codec (VVC). However, there is a general desire to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is provided. The method includes: for conversion between a video unit and a bitstream of the video unit, determining that a neural network post-processing filter (NNPF) is activated for a set of pictures associated with the video unit; applying the NNPF to one or more pictures in the set of pictures according to a sequence; and performing conversion based on an output of the NNPF. In this manner, how the NNPF is applied is specified, and codec efficiency and performance are improved.
[0005] In a second aspect, a device for video processing is provided. The device includes a processor and a non-volatile memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions, which enable a processor to execute the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video, the bitstream of the video being generated by a method performed by an apparatus for video processing. The method includes: determining that a neural network post-processing filter (NNPF) is activated for a set of pictures associated with a video unit of the video; applying the NNPF to one or more pictures in the set of pictures according to a sequence; and generating a bitstream based on an output of the NNPF.
[0008] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: determining that a neural network post-processing filter (NNPF) is activated for a set of pictures related to a video unit of the video; applying the NNPF to one or more pictures in the set of pictures according to a sequence; generating a bitstream based on an output of the NNPF; and storing the bitstream in a non-transitory computer-readable medium.
[0009] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings.In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0011] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;
[0012] Figure 2 shows a block diagram illustrating a first example video encoder according to some embodiments of the present disclosure;
[0013] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;
[0014] Figure 4 Schematic diagram of deriving four luminance channels (right) from the luminance component (left) when nnpfc_inp_order_idc is equal to 3;
[0015] Figure 5 A flowchart showing a method for video processing according to an embodiment of the present disclosure is shown; and
[0016] Figure 6 A block diagram is shown of a computing device in which various embodiments of the present disclosure may be implemented.
[0017] Throughout the drawings, the same or similar reference numbers generally refer to the same or similar elements. DETAILED DESCRIPTION
[0018] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.
[0019] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0020] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment will include that particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is intended that such feature, structure, or characteristic, whether or not explicitly described, be applicable to other embodiments and that it is within the knowledge of those skilled in the art to apply that feature, structure, or characteristic.
[0021] It should be understood that although the terms "first" and "second" and the like may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0022] The terms used herein are used only for the purpose of describing specific embodiments and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "comprise," "including," "having," "including," and / or "comprising" when used herein indicate the presence of the features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Sample Environment
[0023] Figure 1is a block diagram illustrating an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 can include a source device 110 and a destination device 120. The source device 110 can also be referred to as a video encoding device, and the destination device 120 can also be referred to as a video decoding device. In operation, the source device 110 can be configured to generate encoded video data, and the destination device 120 can be configured to decode the encoded video data generated by the source device 110. The source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0024] The video source 112 may include a source such as a video capture device. Examples of a video capture device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.
[0025] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include a coded picture and associated data. The coded picture is a coded representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be directly sent to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0026] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, the destination device 120 being configured to interface with an external display device.
[0027] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.
[0028] Figure 2is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of the video encoder 114 in the system 100 is shown.
[0029] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0030] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.
[0031] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0032] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail in the following sections. Figure 2 are shown separately in the example.
[0033] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0034] The mode selection unit 203 can, for example, select one of a plurality of codec modes (intra-frame codec or inter-frame codec) based on the error result, and provide the resulting intra-frame coded block or inter-frame coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).
[0035] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the buffer 213 other than the picture associated with the current video block.
[0036] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, "P slices" and "B slices" may refer to portions of a picture consisting of macroblocks that are independent of macroblocks in the same picture.
[0037] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0038] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference pictures in list 0 and list 1 containing multiple reference video blocks, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. The motion estimation unit 204 may output the multiple reference indices and multiple motion vectors for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.
[0039] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0040] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0041] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0042] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0043] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0044] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0045] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0046] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.
[0047] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0048] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0049] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.
[0050] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0051] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 An example of the video decoder 124 in the system 100 is shown.
[0052] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0053] exist Figure 3 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.
[0054] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information, which includes motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, including deriving several most likely candidates based on data from adjacent PBs and reference pictures. The motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially neighboring blocks or temporally neighboring blocks.
[0055] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. Identifiers for the interpolation filters used with sub-pixel precision may be included in the syntax elements.
[0056] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information, and motion compensation unit 302 may use the interpolation filters to produce a prediction block.
[0057] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the encoded video sequence, partition information describing how each macroblock of the picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information used to decode the encoded video sequence. As used herein, in some aspects, "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice can be an entire picture or a region of a picture.
[0058] The intra prediction unit 303 can use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0059] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.
[0060] Some exemplary embodiments of the present disclosure are described in detail below. It should be noted that the section headings used in this document are for ease of understanding and do not limit the embodiments disclosed in a section to that section. In addition, although some embodiments are described with reference to a multifunctional video codec or other specific video codec, the disclosed technology is also applicable to other video codec technologies. In addition, although some embodiments describe the video coding and decoding steps in detail, it should be understood that the corresponding decoding steps of the decoding will be implemented by the decoder. In addition, the term video processing includes video coding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at a different compression bit rate. 1. Brief Overview The present disclosure relates to image / video codec technology. In particular, the present disclosure relates to the use of neural network post-processing filters with input of multiple pictures. In this document, the use and control can be applied in video units (e.g., pictures / slices / CTUs). For video bitstreams encoded and decoded by any codec (e.g., the Versatile Video Codec (VVC) standard and / or the Versatile SEI Message (VSEI) standard for encoded video bitstreams), the concepts can be applied individually or in various combinations. 2. Abbreviation APS Adaptive Parameter Set AU Access Unit CLVS Codec layer video sequence CLVSS Codec layer video sequence start CRC Cyclic Redundancy Check CVS encoded and decoded video sequence FIR Finite Impulse Response IRAP Intra-frame random access point NAL Network Abstraction Layer PPS picture parameter set PU picture unit RASL Random Access Skip Preamble SEI Supplemental Enhancement Information STSA Stepped Temporal Sublayer Access VCL video codec layer VSEI Versatile Supplementary Enhancement Information (Recommendation ITU-T H.274 | ISO / IEC 23002-7) VUI Video Availability Information VVC Versatile Video Codec (Recommendation ITU-T H.266 | ISO / IEC 23090-3) 3. Introduction 3.1 Video Codec Standards Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new approaches and incorporated them into reference software called the Joint Exploration Model (JEM). When the Versatile Video Codec (VVC) project officially began, JVET was later renamed the Joint Video Experts Team (JVET). VVC is a new codec standard that aims to reduce bit rate by 50% compared to HEVC. The standard was finalized by JVET at the 19th JVET meeting that ended on July 1, 2020. The Versatile Video Codec (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) and the associated Versatile Supplementary Enhancement Information (VSEI) standard for coded video bitstreams (ITU-T H.274 | ISO / IEC 23002-7) have been designed for the widest range of applications, including conventional uses such as television broadcasting, video conferencing, or playback from storage media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, composition and merging of content from multiple coded video bitstreams, multi-view video, scalable layered codecs, and viewport-adaptive 360° immersive media. The Essential Video Codec (EVC) standard (ISO / IEC 23094-1) is another video codec standard that has been recently developed by MPEG. 3.2. General SEI messages and SEI messages in VVC and VSEI SEI messages assist processes related to decoding, display, or other purposes. However, no SEI messages are required to construct luma or chroma samples from the decoding process. A compliant decoder does not need to process this information for output order conformance. Some SEI messages are required to check bitstream conformance and output timing decoder conformance. Other SEI messages are not required to check bitstream conformance. Annex D of VVC specifies the syntax and semantics of SEI message payloads for some SEI messages, and specifies the use of SEI messages and VUI parameters for which syntax and semantics are specified in ITU-T H.274 | ISO / IEC 23002-7. 3.3. Signaling of Neural Network Post-Processing Filters WG 05 output document N0158 and JVET-AB2006 include provisions for two SEI messages for signaling of neural network post-processing filters, as shown below. 8.28 Neural Network Post-Processing Filter Characteristics SEI Message 8.28.1 Neural Network Post-Processing Filter Characteristics SEI Message Syntax 8.28.2 Neural Network Post-Processing Filter Characteristics SEI Message Semantics The Neural Network Post-Processing Filter Characteristic (NNPFC) SEI message specifies a neural network that can be used as a post-processing filter. The use of a specified post-processing filter for a particular picture is indicated using the Neural Network Post-Processing Filter Activation SEI message. Use of this SEI message requires the definition of the following variables: - The width and height of the cropped decoded output picture, in units of luma samples, denoted by CroppedWidth and CroppedHeight respectively. - The luma sample array CroppedYPic[idx] and the chroma sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] (if present) of the cropped decoded output picture, where idx is in the range of 0 to numInputPics–1 (inclusive), which are used as input to the post-processing filters. - BitDepth of the luma sample array used for the cropped decoded output picture Y . - Bit depth of the chroma sample array (if any) used for the cropped decoded output picture C . - Chroma format indicator, denoted herein by ChromaFormatIdc, as described in clause 7.3. - When nnpfc_auxiliary_inp_idc is equal to 1, the filter strength control value StrengthControlVal. The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc as specified in Table 2. NOTE 1 - More than one NNPFC SEI message may exist for the same picture. When more than one NNPFC SEI message with different values of nnpfc_id is present or activated for the same picture, they may have the same or different values of nnpfc_purpose and nnpfc_mode_idc. nnpfc_id contains an identification number that can be used to identify the post-processing filter. The value of nnpfc_id must be between 0 and 2. 32 -2 (including the boundary value). The value of nnpfc_id is from 256 to 511 (including the boundary value) and from 2 31 to 2 32 The value -2 (inclusive) is reserved for future use by ITU-T|ISO / IEC. A decoder conforming to this version of this document encountering a nnpfc_id with a value in the range 256 to 511 (inclusive) or 31 to 2 32 When an NNPFC SEI message with an nnpfc_id in the range of –2 (inclusive) is received, the SEI message shall be ignored. When the NNPFC SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, the following applies: - This SEI message specifies the base post-processing filters. - This SEI message is related to the current decoded picture and all subsequent decoded pictures of the current layer, in output order until the end of the current CLVS. When an NNPFC SEI message is a repetition of a previous NNPFC SEI message in decoding order in the current CLVS, subsequent semantics apply as if the SEI message was the only NNPFC SEI message with the same content within the current CLVS. When the NNPFC SEI message is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, the following applies: - This SEI message defines updates relative to the preceding base post-processing filter in decoding order with the same nnpfc_id value. - This SEI message is related to the current decoded picture and all subsequent decoded pictures of the current layer, in output order until the end of the current CLVS, or the next NNPFC SEI message (in output order) with this specific nnpfc_id value within the current CLVS. nnpfc_mode_idc equal to 0 indicates that the SEI message contains an ISO / IEC 15938-17 bitstream that specifies a base post-processing filter or an update relative to a base post-processing filter with the same nnpfc_id value. When the NNPFC SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, nnpfc_mode_idc equal to 1 specifies that the base post-processing filter associated with the nnpfc_id value is a neural network identified by a URI indicated by nnpfc_uri, whose format is identified by the tag URI nnpfc_tag_uri. When the NNPFC SEI message is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, nnpfc_mode_idc equal to 1 specifies that updates relative to the base post-processing filter with the same nnpfc_id value are defined by a URI indicated by nnpfc_uri whose format is identified by the tag URI nnpfc_tag_uri. In bitstreams conforming to this version of this document, the value of nnpfc_mode_idc shall be in the range of 0 to 1, inclusive. Values of nnpfc_mode_idc from 2 to 255, inclusive, are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_mode_idc in the range of 2 to 255, inclusive. Values of nnpfc_mode_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use. When this SEI message is the first NNPFC SEI message in decoding order with a specific nnpfc_id value within the current CLVS, the post-processing filter PostProcessingFilter() is assigned to be the same as the base post-processing filter. When this SEI message is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, the post-processing filter PostProcessingFilter() is obtained by applying updates to the base post-processing filter, the updates being defined by this SEI message. Updates are not cumulative, but each update is applied on the base post-processing filter, which is the post-processing filter specified by the first NNPFC SEI message in decoding order with a specific nnpfc_id value within the current CLVS. nnpfc_reserved_zero_bit_a shall be equal to 0 in bitstreams conforming to this version of this document. Decoders shall ignore NNPFC SEI messages with nnpfc_reserved_zero_bit_a not equal to 0. nnpfc_tag_uri contains a tag URI with syntax and semantics as specified by IETF RFC 4151 that identifies the format and associated information about the neural network that is used as a base post-processing filter or as an update to a base post-processing filter that has the same nnpfc_id value as specified by nnpfc_uri. NOTE 2 – nnpfc_tag_uri enables unique identification of the format of the neural network data specified by nnrpf_uri. No central registration authority is required. nnpfc_tag_uri equal to "tag:iso.org,2023:15938-17" indicates that the neural network data identified by nnpfc_uri complies with ISO / IEC 15938-17. nnpfc_uri contains a URI with syntax and semantics as specified in IETF Internet Standard 66 that identifies a neural network that is used as a base post-processing filter or an update relative to a base post-processing filter with the same nnpfc_id value. nnpfc_formatting_and_purpose_flag equal to 1 specifies that syntax elements related to filter purpose, input format, output format and complexity are present. nnpfc_formatting_and_purpose_flag equal to 0 specifies that syntax elements related to filter purpose, input format, output format and complexity are not present. When this SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, nnpfc_formatting_and_purpose_flag shall be equal to 1. When this SEI message is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, nnpfc_formatting_and_purpose_flag must be equal to 0. nnpfc_purpose indicates the purpose of the post-processing filter, as specified in Table 2-1. In bitstreams conforming to this version of this document, the value of nnpfc_purpose shall be in the range of 0 to 5, inclusive. Values of nnpfc_purpose from 6 to 1023, inclusive, are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_purpose in the range of 6 to 1203, inclusive. Values of nnpfc_purpose greater than 1023 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use. Table 2-1- Definition of nnpfc_purpose NOTE 3 – When the reserved value of nnpfc_purpose is used by ITU-T | ISO / IEC in the future, the syntax of this SEI message may be extended with syntax elements whose presence is conditional on nnpfc_purpose being equal to this value. When SubWidthC is equal to 1 and SubHeightC is equal to 1, nnpfc_purpose does not need to be equal to 2 or 4. nnpfc_out_sub_c_flag equal to 1 specifies that outSubWidthC is equal to 1 and outSubHeightC is equal to 1. nnpfc_out_sub_c_flag equal to 0 specifies that outSubWidthC is equal to 2 and outSubHeightC is equal to 1. When nnpfc_out_sub_c_flag is not present, outSubWidthC is inferred to be equal to SubWidthC, and outSubHeightC is inferred to be equal to SubHeightC. When ChromaFormatIdc is equal to 2 and nnpfc_out_sub_c_flag is present, the value of nnpfc_out_sub_c_flag shall be equal to 1. nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples specify the width and height, respectively, of the luma sample array of the picture, resulting from applying the post-processing filter identified by nnpfc_id to the cropped decoded output picture. When nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples are not present, they are inferred to be equal to CroppedWidth and CroppedHeight, respectively. The value of nnpfc_pic_width_in_luma_samples shall be in the range of CroppedWidth to CroppedWidth*16–1, inclusive. The value of nnpfc_pic_height_in_luma_samples shall be in the range of CroppedHeight to CroppedHeight*16–1, inclusive. nnpfc_num_input_pics_minus2 plus 2 specifies the number of decoded output pictures used as input to the post-processing filters. nnpfc_interpolated_pics[i] specifies the number of interpolated pictures generated by the post-processing filter between the i-th picture and the (i+1)-th picture used as input to the post-processing filter. The variable numInputPics, which specifies the number of pictures used as input to the post-processing filter, and the variable numOutputPics, which specifies the total number of images produced by the post-processing filter, are derived as follows: nnpfc_component_last_flag is equal to 1 to indicate that the last dimension in the input tensor inputTensor of the post-processing filter and the output tensor outputTensor generated by the post-processing filter is used for the current channel. nnpfc_component_last_flag is equal to 0 to indicate that the third dimension in the input tensor inputTensor of the post-processing filter and the output tensor outputTensor generated by the post-processing filter is used for the current channel. NOTE 4 – The first dimension in the input and output tensors is used for batch indexing, which is a practice in some neural network frameworks. Although the formulas in the semantics of this SEI message use a batch size corresponding to a batch index equal to 0, it is up to the post-processing implementation to determine the batch size used as input to the neural network inference. Note 5 – For example, when nnpfc_inp_order_idc is equal to 3 and nnpfc_auxiliary_inp_idc is equal to 1, there are 7 channels in the input tensor, including four luma matrices, two chroma matrices, and one auxiliary input matrix. In this case, the procedure DeriveInputTensors() will derive each of these 7 channels of the input tensor one by one, and when a particular channel of these channels is processed, that channel is called the current channel during the process. nnpfc_inp_format_idc indicates a method for converting the sample values of the cropped decoded output picture into the input values of the post-processing filter. When nnpfc_inp_format_idc is equal to 0, the input values of the post-processing filter are real numbers, and the functions InpY() and InpC() are defined as follows: InpY( x ) = x ÷ ( ( 1 << BitDepthY ) – 1) (77) InpC( x )= x ÷ ( ( 1 << BitDepthC ) - 1 ) (78) When nnpfc_inp_format_idc is equal to 1, the input values of the post-processing filter are unsigned integers, and the functions InpY() and InpC() are defined as follows: The variable inpTensorBitDepth is derived from the syntax element nnpfc_inp_tensor_bitdepth_minus8 specified as follows. Values of nnpfc_inp_format_idc greater than 1 are reserved for future specification by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages containing reserved values of nnpfc_inp_format_idc. nnpfc_inp_tensor_bitdepth_minus8 specifies the bit depth of the luma sample values in the input integer tensor plus 8. The value of inpTensorBitDepth is derived as follows: inpTensorBitDepth = nnpfc_inp_tensor_bitdepth_minus8 + 8 (81) A bitstream conformance requirement is that the value of nnpfc_inp_tensor_bitdepth_minus8 must be in the range 0 to 24 (inclusive). nnpfc_inp_order_idc indicates a method of ordering the sample array of the cropped decoded output picture into one of the input pictures of the post-processing filter. In bitstreams conforming to this version of this document, the value of nnpfc_inp_order_idc shall be in the range of 0 to 3, inclusive. Values of nnpfc_inp_order_idc from 4 to 255, inclusive, are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 4 to 255, inclusive. Values of nnpfc_inp_order_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use. When ChromaFormatIdc is not equal to 1, nnpfc_inp_order_idc does not need to be equal to 3. Table 2-2 contains informative descriptions of the values of nnpfc_inp_order_idc. Table 2-2-Description of nnpfc_inp_order_idc values A tile is a rectangular array of samples of a component (eg, luma or chroma components) from a picture. nnpfc_auxiliary_inp_idc greater than 0 indicates that auxiliary input data is present in the input tensor of the neural network post-processing filter. nnpfc_auxiliary_inp_idc equal to 0 indicates that auxiliary input data is not present in the input tensor. nnpfc_auxiliary_inp_idc equal to 1 specifies that the auxiliary input data is derived as specified in Equation 82. In bitstreams conforming to this version of this document, the value of nnpfc_auxiliary_inp_idc shall be in the range of 0 to 1, inclusive. Values of nnpfc_inp_order_idc from 2 to 255, inclusive, are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 2 to 255, inclusive. Values of nnpfc_inp_order_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use. The procedure DeriveInputTensors() is used to derive the input tensor inputTensor for a given vertical sample coordinate cTop and a horizontal sample coordinate cLeft, which specifies the top left sample position of a small block of samples included in the input tensor. The procedure DeriveInputTensors() is defined as follows: nnpfc_separate_colour_description_present_flag equal to 1 indicates that a different combination of color primaries, transfer characteristics, and matrix coefficients for the picture produced by the post-processing filter is specified in the SEI message syntax structure. nnfpc_separate_colour_description_present_flag equal to 0 indicates that the combination of color primaries, transfer characteristics, and matrix coefficients for the picture produced by the post-processing filter is the same as indicated in the VUI parameters for CLVS. nnpfc_colour_primaries has the same semantics as specified in subclause 7.3 for the vui_colour_primaries syntax element, except as follows: -nnpfc_colour_primaries specifies the color primaries of the picture produced by applying the neural network post-processing filters specified in the SEI message, instead of the color primaries used for CLVS. - When nnpfc_colour_primaries is not present in the NNPFC SEI message, the value of nnpfc_colour_primaries is inferred to be equal to vui_colour_primaries. nnpfc_transfer_characteristics has the same semantics as specified in subclause 7.3 for the vui_transfer_characteristics syntax element, except as follows: -nnpfc_transfer_characteristics specifies the transfer characteristics of the picture produced by applying the neural network post-processing filters specified in the SEI message, rather than the transfer characteristics used for CLVS. - When nnpfc_transfer_characteristics is not present in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is inferred to be equal to vui_transfer_characteristics. nnpfc_matrix_coeffs has the same semantics as specified for the vui_matrix_coeffs syntax element in subclause 7.3, except as follows: -nnpfc_matrix_coeffs specifies the matrix coefficients of the picture produced by applying the neural network post-processing filters specified in the SEI message, rather than the matrix coefficients used for CLVS. - When nnpfc_matrix_coeffs is not present in the NNPFC SEI message, the value of nnpfc_matrix_coeffs is inferred to be equal to vui_matrix_coeffs. - The allowed values of nnpfc_matrix_coeffs are not constrained by the chroma format of the decoded video picture indicated by the value of ChromaFormatIdc for the semantics of the VUI parameter. - When nnpfc_matrix_coeffs is equal to 0, nnpfc_out_order_idc need not be equal to 1 or 3. nnpfc_out_format_idc being equal to 0 indicates that the sample values output by the post - processing filter are real numbers, where values in the range from 0 to 1 (including the boundary values) are linearly mapped to unsigned integer values in the range from 0 to (1 << bitDepth) – 1 (including the boundary values) for any desired bit depth bitDepth used for subsequent post - processing or display. nnpfc_out_format_flag being equal to 1 indicates that the sample values output by the post - processing filter are unsigned integers in the range from 0 to (1 << (nnpfc_out_tensor_bitdepth_minus8 + 8)) - 1 (including the boundary values). Values of nnpfc_out_format_idc greater than 1 are reserved for future ITU - T|ISO / IEC specifications and need not be present in the bitstream conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages containing reserved values of nnpfc_out_format_idc. nnpfc_out_tensor_bitdepth_minus8 plus 8 specifies the bit depth of the sample values in the output integer tensor. The value of nnpfc_out_tensor_bitdepth_minus8 shall be in the range from 0 to 24 (including the boundary values). nnpfc_out_order_idc indicates the output order of the samples produced by the post - processing filter. In the bitstream conforming to this version of this document, the value of nnpfc_out_order_idc shall be in the range from 0 to 3 (including the boundary values). Values of nnpfc_out_order_idc from 4 to 255 (including the boundary values) are reserved for future ITU - T| ISO / IEC use and need not be present in the bitstream conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages having nnpfc_out_order_idc in the range from 4 to 255 (including the boundary values). nnpfc_out_order_idc values greater than 255 need not be present in the bitstream conforming to this version of this document and are not reserved for future use. When nnpfc_purpose is equal to 2 or 4, nnpfc_out_order_idc need not be equal to 3. Table 2-3 contains informative descriptions of the nnpfc_out_order_idc values. Table 2-3 - Description of nnpfc_out_order_idc values The procedure StoreOutputTensors() is used to derive the sample values in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor for given vertical sample coordinates cTop and horizontal sample coordinates cLeft that specify the top-left sample position of the block of samples included in the input tensor. The procedure StoreOutputTensors() is defined as follows: nnpfc_constant_patch_size_flag equal to 1 indicates that the post-processing filter accepts the exact patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1 as input. nnpfc_constant_patch_size_flag equal to 0 indicates that the post-processing filter accepts an arbitrary patch size as input that is a positive integer multiple of the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. nnpfc_patch_width_minus1+1, indicates the horizontal sample count of the required patch size as input to the post-processing filter when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_width_minus1 shall be in the range of 0 to Min(32766, CroppedWidth-1), inclusive. nnpfc_patch_height_minus1+1, indicates the vertical sample count of the patch size required as input to the post-processing filter when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_height_minus1 shall be in the range of 0 to Min(32766, Croppedheight-1), inclusive. Let the variables inpPatchWidth and inpPatchHeight be the patch size width and patch size height respectively. If nnpfc_constant_patch_size_flag is equal to 0, the following applies: – The values of inpPatchWidth and inpPatchHeight are provided by external means not specified in this document or set by the post-process itself. –inpPatchWidth must be a positive integer multiple of nnpfc_patch_width_minus1+1 and must be less than or equal to CroppedWidth. inpPatchHeight must be a positive integer multiple of nnpfc_patch_height_minus1+1 and must be less than or equal to CroppedHeight. Otherwise (nnpfc_constant_patch_size_flag is equal to 1), the value of inpPatchWidth is set equal to nnpfc_patch_width_minus1+1, and the value of inpPatchHeight is set equal to nnpfc_patch_height_minus1+1. nnpfc_overlap indicates the horizontal and vertical sample counts of overlap of adjacent input tensors to the post-processing filter. The value of nnpfc_overlap must be in the range of 0 to 16383 (inclusive). The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows: outPatchWidth=(nnpfc_pic_width_in_luma_samples*inpPatchWidth) / CroppedWidth(84) outPatchHeight=(nnpfc_pic_height_in_luma_samples*inpPatchHeight) / CroppedHeight(85) horCScaling = SubWidthC / outSubWidthC (86) verCScaling = SubHeightC / outSubHeightC (87) outPatchCWidth = outPatchWidth * horCScaling (88) outPatchCHeight = outPatchHeight * verCScaling (89) overlapSize = nnpfc_overlap (90) The bitstream conformance requirement is that outPatchWidth*CroppedWidth shall be equal to nnpfc_pic_width_in_luma_samples*inpPatchWidth, and outPatchHeight*CroppedHeight shall be equal to nnpfc_pic_height_in_luma_samples*inpPatchHeight. nnpfc_padding_type indicates the padding process when referencing sample positions outside the boundaries of the cropped decoded output picture, as described in Tables 2-4. The value of nnpfc_padding_type shall be in the range of 0 to 15 (inclusive). Table 2-4 - Informative description of nnpfc_padding_type values nnpfc_luma_padding_val indicates the luma value to be used for padding when nnpfc_padding_type is equal to 4. nnpfc_cb_padding_val indicates the Cb value to be used for padding when nnpfc_padding_type is equal to 4. nnpfc_cr_padding_val indicates the Cr value to be used for padding when nnpfc_padding_type is equal to 4. The function InpSampleVal(y,x,picHeight,picWidth,croppedPic), whose input is the vertical sample point position y, the horizontal sample point position x, the picture height picHeight, the picture width picWidth and the sample array croppedPic, returns the value of sampleVal derived as follows: Note 6 – For the input to the function InpSampleVal(), the vertical position is listed before the horizontal position for compatibility with the input tensor convention of some inference engines. The following example process may be used to filter the cropped decoded output picture tile-wise using a post-processing filter PostProcessingFilter() to generate a filtered picture containing Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated by nnpfc_out_order_idc. nnpfc_complexity_info_present_flag equal to 1 specifies that one or more syntax elements are present that indicate the complexity of the post-processing filter associated with nnpfc_id. nnpfc_complexity_info_present_flag equal to 0 specifies that no syntax elements are present that indicate the complexity of the post-processing filter associated with nnpfc_id. nnpfc_parameter_type_idc equal to 0 indicates that the neural network uses only integer parameters. nnpfc_parameter_type_flag equal to 1 indicates that the neural network can use floating-point parameters or integer parameters. nnpfc_parameter_type_idc equal to 2 indicates that the neural network uses only binary parameters. nnpfc_parameter_type_idc equal to 3 is reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_parameter_type_idc equal to 3. nnpfc_log2_parameter_bit_length_minus3 equal to 0, 1, 2, and 3 respectively indicates that the neural network does not use parameters with bit lengths greater than 8, 16, 32, and 64. When nnpfc_parameter_type_idc is present and nnpfc_log2_parameter_bit_length_minus3 is not present, the neural network does not use parameters with bit lengths greater than 1. nnpfc_num_parameters_idc indicates the maximum number of neural network parameters for the post-processing filters, in powers of 2048. nnpfc_num_parameters_idc equal to 0 indicates that the maximum number of neural network parameters is unknown. The value of nnpfc_num_parameters_idc shall be in the range of 0 to 52, inclusive. Values of nnpfc_num_parameters_idc greater than 52 are reserved for future use by ITU-T | ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_num_parameters_idc greater than 52. If the value of nnpfc_num_parameters_idc is greater than 0, the variable maxNumParameters is derived as follows: maxNumParameters = (2048 << nnpfc_num_parameters_idc) - 1 (93) The bitstream conformance requirement is that the number of neural network parameters of the post-processing filters must be less than or equal to maxNumParameters. nnpfc_num_kmac_operations_idc greater than 0 indicates that the maximum number of multiply-accumulate operations per sample of the post-processing filter is less than or equal to nnpfc_num_kmac_operations_idc * 1000. nnpfc_num_kmac_operations_idc equal to 0 indicates that the maximum number of multiply-accumulate operations of the network is unknown. The value of nnpfc_num_kmac_operations_idc shall be between 0 and 2. 32 The range is -1 (including the boundary value). nnpfc_total_kilobyte_size is greater than 0 to indicate the total size in kilobytes required to store the uncompressed parameters for the neural network. The total size in bits is the number of bits equal to or greater than the sum of the bits used to store each parameter. nnpfc_total_kilobyte_size is the total size in bits divided by 8000 and rounded up. nnpfc_total_kilobyte_size is equal to 0 to indicate that the total size required to store the parameters for the neural network is unknown. The value of nnpfc_total_kilobyte_size must be between 0 and 2. 32 The range is -1 (including the boundary value). nnpfc_reserved_zero_bit_b shall be equal to 0 in bitstreams conforming to this version of this document. The decoder shall ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_b is not equal to 0. nnpfc_payload_byte[i] contains the i-th byte of a bitstream conforming to ISO / IEC 15938-17. The byte sequence nnpfc_payload_byte[i] for all present values of i shall be a complete bitstream conforming to ISO / IEC 15938-17. 8.29 Neural Network Post-Processing Filter Activation SEI Message 8.29.1 Neural Network Post-Processing Filter Activation SEI Message Syntax 8.29.2 Neural Network Post-Processing Filter Activation SEI Message Semantics The Neural Network Post-Processing Filter Activation (NNPFA) SEI message activates or deactivates the possibility of using the target neural network post-processing filter identified by nnpfa_target_id for post-processing filtering of a set of pictures. NOTE 1—Multiple NNPFA SEI messages may exist for the same picture, for example, when the post-processing filters are intended for different purposes or filter different color components. nnpfa_target_id indicates the target neural network post-processing filter, which is specified by one or more neural network post-processing filter characteristics SEI messages related to the current picture and has nnpfc_id equal to nnpfa_target_id. The value of nnpfa_target_id must be between 0 and 2 32 -2 (including the boundary value). The value of nnpfa_target_id is between 256 and 511 (including the boundary value), and 2 31 to 2 32 -2 (inclusive) are reserved for future use by ITU-T|ISO / IEC. A decoder conforming to this version of this document will not be used when encountering a decoder with nnpfa_target_id in the range of 256 to 511 (inclusive) or 2 31 to 2 32 When an NNPFA SEI message is received within the range of -2 (including the boundary value), the SEI message shall be ignored. An NNPFA SEI message with a specific value of nnpfa_target_id need not be present in the current PU unless one or both of the following conditions are true: – in the current CLVS, there is an NNPFC SEI message with nnpfc_id equal to the specific value of nnpfa_target_id that appears in the PU that precedes the current PU in decoding order; – There is an NNPFC SEI message whose nnpfc_id is equal to the specific value of nnpfa_target_id in the current PU. When a PU includes both an NNPFC SEI message with a specific value of nnpfc_id and an NNPFA SEI message with nnpfa_target_id equal to the specific value of nnpfc_id, the NNPFC SEI message shall precede the NNPFA SEI message in decoding order. nnpfa_cancel_flag equal to 1 indicates that the persistence of the target neural network post-processing filter established by any previous NNPFA SEI message (the message has the same nnpfa_target_id as the current SEI message) is cancelled, that is, the target neural network post-processing filter is not used again unless it is activated again by another NNPFA SEI message (with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 0). nnpfa_cancel_flag equal to 0 indicates that nnpfa_persistence_flag is set afterwards. nnpfa_persistence_flag specifies the persistence of the target neural network post-processing filter for the current layer. nnpfa_persistence_flag equal to 0 specifies that the target neural network post-processing filter can only be used for post-processing filtering for the current picture. nnpfa_persistence_flag equal to 1 specifies that the target neural network post-processing filter can be used for post-processing filtering for the current picture in the current layer and all subsequent pictures in output order until one or more of the following conditions are true: – A new CLVS starts for the current layer; – End of bitstream; A picture in the current layer is associated with an NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1, and is output after the current picture in output order. NOTE 2—The target neural network post-processing filter is not used for the subsequent picture in the current layer that has the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1. The NNPFA SEI message is associated with the 4. Question The current design of the Neural Network Post-Processing Filter Characteristics (NNPFC) SEI message and the Neural Network Post-Processing Filter Activation (NNPFA) SEI message has the following issues: 1) NNPF can be activated for a group of pictures. However, it is unclear how to apply NNPF to all pictures in the group for which NNPF is activated, especially when multiple input pictures are specified. For example, should NNPF be applied one picture at a time? For another example, in what order should NNPF be applied to pictures? 2) For the NNPF purpose of picture rate upsampling, the number of interpolated pictures between each pair of adjacent input pictures is specified. However, typically, even when more than two input pictures are used for interpolation, some pictures are interpolated only between a pair of input pictures. In addition, performing more interpolation than just interpolating some pictures between a pair of input pictures may result in pictures between a particular pair of input pictures being interpolated multiple times, making the process even more ambiguous and confusing. 3) When a neural network post-processing filter (NNPF) takes one picture as input, it is not clearly specified which picture is used as the input picture when the NNPF is applied to a specific current picture for which the NNPF is activated. 4) When a neural network post-processing filter (NNPF) takes multiple pictures as input, when the NNPF is applied to a specific picture for which the NNPF is activated, it is not clearly specified which pictures are used as input pictures. 5) When a neural network post-processing filter (NNPF) takes multiple pictures as input, it is possible that the total number of available pictures is less than the number of input pictures when the NNPF is applied to a specific current picture for which the NNPF is activated. This can happen when the total number of pictures for which the NNPF is activated is less than the number of input pictures, or when the current picture is the first or last picture among the pictures for which the NNPF is activated. However, it is unclear how to apply the NNPF in such a scenario. 5. Detailed solution To solve the above problems, the following methods are disclosed. The solutions should be considered as examples to explain the general concept and should not be interpreted in a narrow sense. In addition, these solutions can be applied alone or combined in any way. 1) To address issue 1, one or more of the following aspects are specified: a. In one example, it is specified that NNPF is activated for a set of pictures (denoted as TargetPictures), NNPF is applied to each picture in TargetPictures. NNPF is applied to one picture at a time in the decoding order of the pictures in TargetPictures. b. In one example, alternatively, it is specified that for NNPF, a set of pictures (denoted as TargetPictures) is activated, NNPF is applied to each group consisting of a fixed number of consecutive pictures in TargetPictures. NNPF is applied to one group at a time in the decoding order of pictures in TargetPictures. c. In one example, alternatively, it is specified that for NNPF, a group of pictures (denoted as TargetPictures) is activated, NNPF is applied to each picture in TargetPictures, and NNPF is applied to one picture at a time in the output order of the pictures in TargetPictures. d. In one example, alternatively, it is specified that for NNPF, a set of pictures (denoted as TargetPictures) is activated, NNPF is applied to each group consisting of a fixed number of consecutive pictures in TargetPictures. NNPF is applied one group at a time in the output order of the pictures in TargetPictures. 2) To solve Problem 2, it is specified that for the NNPF purpose of picture rate upsampling, even when more than two input pictures are utilized for interpolation, some pictures are interpolated only between a pair of input pictures. a. In one example, based on the current syntax of the NNPFC SEI message, nnpfc_interpolated_pics[i] is required to be in the range of 0 to nnpfc_num_input_pics_minus2 Only one i (including the boundary values) is greater than 0; that is, nnpfc_interpolated_pics[i] must be equal to 0 for all other values of i. b. In one example, it is provided that even when more than two input pictures are utilized for interpolation, some pictures are interpolated only between a pair of middle input pictures. i. In one example, based on the current syntax of the NNPFC SEI message, it is required that nnpfc_interpolated_pics[i] must be greater than 0 only when i is equal to 0 to nnpfc_num_input_pics_minus2 / 2; That is, nnpfc_interpolated_pics[i] must be equal to 0 for all other values of i. ii. In one example, the syntax of the NNPFC SEI message is changed so that only one instance of nnpfc_interpolated_pics[i] is signaled, for example, using the syntax element name nnpfc_interpolated_pics_minus1, and nnpfc_interpolated_pics_minus1 plus 1 specifies the number of interpolated pictures between a pair of middle input pictures. iii. In one example, nnpfc_num_input_pics_minus2 is required to be an even number, and a pair of middle input pictures is the nnpfc_num_input_pics_minus2 / 2th input picture and the (nnpfc_num_input_pics_minus2 / 2+1) input images. 3) To solve problem 3, it is stipulated that when NNPF takes a picture as input, when NNPF is applied to a specific current picture for which NNPF is activated, the input picture is the current picture itself. 4) To solve Problem 4, the picture used as the input picture can be determined, specified, or transmitted via a signal. a. In one example, the identity of the picture used as input is transmitted via a signal. i. In one example, each picture order count (POC) value of an input picture is signaled. ii. In one example, the difference in POC values in the input pictures may be signaled. 1. In one example, the POC difference between each input picture and the picture for which NNPF is activated is signaled. 2. In one example, the minimum POC value in the input picture is signaled, and the difference of the remaining POC values from the minimum POC value is signaled. 3. In one example, the maximum POC value in the input picture is signaled, and the difference of the remaining POC values from the maximum POC value is signaled. 4. In one example, the middle POC value in the input picture is signaled, and the difference of the remaining POC values from the middle POC value is signaled. b. In one example, input pictures are specified and listed in output order. i. In one example, the pictures before the current picture are listed in the output order. The pictures after the current picture are listed in the output order. ii. In one example, pictures before the current picture are listed in decoding order, and pictures after the current picture are listed in output order. iii. In one example, pictures before the current picture are listed in output order, and pictures after the current picture are listed in decoding order. iv. In one example, pictures before the current picture are listed in decoding order. Pictures after the current picture are listed in decoding order. 5) To solve problem 2, the interpolated image generated by the post-processing filter can be specified or transmitted via a signal. a. In one example, one or more syntax elements are signaled to specify interpolated pictures. i. In one example, the minimum and maximum POC values of the interpolated picture are signaled. ii. In one example, the minimum POC value of the interpolated pictures is signaled, and the maximum POC value of the interpolated pictures can be derived from the number of interpolated pictures, which is signaled in the NNPFC SEI message. b. In one example, the interpolated image can be determined by the input image. i. In one example, the interpolated image is located between a pair of intermediate input images. 1. In one example, with and The input picture indexed by consists of a pair of intermediate input pictures, where N is an integer representing the number of input pictures, and Represents the largest integer less than or equal to x. ii. In one example, the interpolated picture is located between the first pair of input pictures. iii. In one example, the interpolated image is located between the last pair of input images. iv. In one example, a picture is interpolated between any pair of consecutive input pictures, and the identity of the pair of pictures is signaled. 1. In one example, the POC value of the input picture in the pair of pictures is signaled. 2. In one example, the index of the first input picture in the pair of pictures is signaled. 6) To solve problem 5, one or more of the following methods can be used to generate unusable input images. a. In one example, a fill method may be used. i. In one example, unavailable pictures may be replaced with the closest available picture in order of decoding order. ii. In one example, unavailable pictures may be replaced with available pictures with the lowest QP in order of decoding order. iii. In one example, unavailable images may be filled with default values. 1. In one example, the default value is normalized. a. In one example, the default value is 0. b. In one example, the default value is 0.5. c. In one example, the default value is 1. 2. In one example, the default value is not normalized and ranges from 0 to N-1 (inclusive), where N is an integer. a. In one example, furthermore, N is defined as 1<<(nnpfc_inp_tensor_bitdepth_minus8+8). b. In one example, unavailable images can be interpolated using existing available images. i. In one example, a bilinear filter can be used for interpolation. ii. In one example, a bicubic filter may be used for interpolation. iii. In one example, a Lanczos filter may be used for interpolation. iv. In one example, a neural network based interpolation filter may be used for interpolation. 6. Examples The following are some example embodiments for aspects of the solution outlined in Section 5 above. Most relevant sections have been added or modified to Underline Highlighted, and some of the deleted parts are marked with There may be some other changes that are editorial in nature and therefore not highlighted. Example 1 This example addresses solution items 4, 5, and 6 outlined in Section 5 above, and all of their sub-items. 8.29.1 Neural Network Post-Processing Filter Activation SEI Message Syntax 8.29.2 Neural Network Post-Processing Filter Activation SEI Message Semantics … nnpfa_num_input_pics_minus1 plus 1 specifies the number of decoded output pictures. The number of pics to be used as input to the post-processing filter. The value of nnpfa_num_input_pics_minus1 must be in the range of 0 to 15. range (including the boundary value). nnpfa_input_pics_poc[i] specifies the POC value of the input image. The value must be between 0 and 2 32 The range is -1 (including the boundary value). nnpfa_least_output_pics_poc specifies the minimum POC value of the interpolated picture. The value of output_pics_poc must be between 0 and 2. 32 The range is -1 (including the boundary value). nnpfa_largest_output_pics_poc specifies the maximum POC value of the interpolated picture. The value of nnpfa_largest_output_pics_poc must be between 0 and 2 32 The range is -1 (including the boundary value).
[0061] The term “video unit” or “codec unit” or “block” may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, a TB.
[0062] Figure 5 FIG. 5 is a flow chart of a method 500 for video processing according to an embodiment of the present disclosure. The method 500 is implemented during conversion between a video unit of a video and a bitstream of the video.
[0063] At block 510 , for conversion between a video unit of a video and a bitstream of the video unit, a neural network post-processing filter (NNPF) is determined to be activated for a set of pictures associated with the video unit.
[0064] At block 520 , the NNPF is applied to one or more pictures in a set of pictures according to a sequence.
[0065] At block 530, conversion is performed based on the output of the NNPF. In some embodiments, the conversion may include encoding the video unit into a bitstream. Alternatively or additionally, the conversion may include decoding the video unit from the bitstream. In this way, it specifies how to apply the NNPF and improves codec efficiency and performance.
[0066] In some embodiments, the NNPF is applied to each picture in a set of pictures in the output order of the pictures in the set of pictures, and the NNPF is applied to one picture at a time. For example, BitstreamToFilter is decoded and the list CroppedDecodedPictures is set to a list of cropped decoded pictures in output order, the cropped decoded pictures being produced by decoding BitstreamToFilter. In addition, for each cropped decoded picture in CroppedDecodedPictures and for each cropped decoded picture one or more NNPFs are activated, the filtering process for one picture is repeatedly called in output order.
[0067] In some embodiments, the NNPF is applied to each picture in a group of pictures in the decoding order of the pictures in the group of pictures, and the NNPF is applied one picture at a time. In some other embodiments, the NNPF is applied to each group of a fixed number of consecutive pictures in the group of pictures in the decoding order of the pictures in the group of pictures, and the NNPF is applied one group at a time. Alternatively, the NNPF is applied to each group of a fixed number of consecutive pictures in the group of pictures in the output order of the pictures in the group of pictures, and the NNPF is applied one group at a time.
[0068] In some embodiments, if the NNPF takes a picture as input, and if the NNPF is applied to the current picture for which the NNPF is activated, the input picture of the NNPF is the current picture. For example, the filtered picture and / or the interpolated picture is generated by the NNPF by applying the NNPF process specified in the semantics of the NNPFC SEI message to the current picture in small blocks.
[0069] In some embodiments, the picture used as the input picture of the NNPF is determined. Alternatively, the picture used as the input picture of the NNPF is specified. In some other embodiments, the picture used as the input picture of the NNPF is indicated.
[0070] In some embodiments, the input pictures are specified and listed in output order. For example, the order of the pictures in ListNnpfOutputPics is the output order.
[0071] In some embodiments, the one or more pictures before the current picture are listed in output order, and the one or more pictures after the current picture are listed in output order. In some other embodiments, the one or more pictures before the current picture are listed in decoding order, and the one or more pictures after the current picture are listed in output order. Alternatively, the one or more pictures before the current picture are listed in output order, and the one or more pictures after the current picture are listed in decoding order. In some embodiments, the one or more pictures before the current picture are listed in decoding order, and the one or more pictures after the current picture are listed in decoding order.
[0072] In some embodiments, one or more identifiers related to the pictures used as input are indicated. For example, each picture order count (POC) value of the input pictures is indicated.
[0073] In some embodiments, differences related to POC values in the input pictures are indicated. For example, the POC difference between each input picture and the picture for which NNPF is activated is indicated. As another example, the minimum POC value in the input picture is indicated, and the difference between the remaining POC value and the minimum POC value is indicated. As another example, the maximum POC value in the input picture is indicated, and the difference between the remaining POC value and the maximum POC value is indicated. As another example, the medium POC value in the input picture is indicated, and the difference between the remaining POC value and the medium POC value is indicated.
[0074] In some embodiments, for NNPF purposes of picture rate upsampling, one or more pictures are interpolated between a pair of input pictures, even if more than two input pictures are used for interpolation. For example, the syntax of the post-processing filter characteristics (NNPFC) supplemental enhancement information (SEI) message requires that each picture in the i-th NNPFC interpolated picture (denoted as nnpfc_interpolated_pics[i]) is greater than 0 for only one i in the range of 0 to the number of NNPFC input pictures minus 2 (denoted as nnpfc_num_input_pics_minus2), inclusive, where i is an integer. In other words, the i-th NNPFC interpolated picture (i.e., nnpfc_interpolated_pics[i]) is equal to 0 for all other values of i.
[0075] In some other embodiments, for NNPF purposes of picture rate upsampling, one or more pictures are interpolated between a pair of intermediate input pictures, even if more than two input pictures are used for interpolation. For example, the syntax of the NNPFC SEI message requires that each picture in the i-th NNPFC interpolated picture (denoted as nnpfc_interpolated_pics[i]) be greater than 0 for i if and only if i is equal to 0 to the number of NNPFC input pictures minus 2 divided by 2 (denoted as nnpfc_num_input_pics_minus2 / 2), where i is an integer. In other words, the i-th NNPFC interpolated picture (i.e., nnpfc_interpolated_pics[i]) is equal to 0 for all other values of i.
[0076] In some embodiments, only one instance of the i-th NNPFC interpolated picture is indicated in the NNPFC SEI message. For example, a syntax element denoted nnpfc_interpolated_pics_minus1 may be indicated in the NNPFC SEI message. In this case, nnpfc_interpolated_pics_minus1 plus 1 may specify the number of interpolated pictures between a pair of intermediate input pictures. In some other embodiments, nnpfc_num_input_pics_minus2 is an even number, and the pair of intermediate input pictures is the nnpfc_num_input_pics_minus2 / 2th input picture and the (nnpfc_num_input_pics_minus2 / 2+1)th input picture.
[0077] In some embodiments, an interpolated picture generated by a post-processing filter is specified. Alternatively, an interpolated picture generated by an NNPF is indicated.
[0078] In some embodiments, one or more syntax elements are indicated to specify the interpolated pictures. For example, a minimum POC value and a maximum POC value for the interpolated pictures are indicated. As another example, a minimum POC value for the interpolated pictures is indicated, and a maximum POC value for the interpolated pictures is derived based on the number of interpolated pictures indicated in the NNPFC SEI message.
[0079] In some embodiments, the interpolated picture is determined based on the input pictures. For example, the interpolated picture is between a pair of intermediate input pictures. In an example embodiment, and The input picture indexed by includes a pair of intermediate input pictures, where N is an integer and represents the number of input pictures, and Represents the largest integer less than or equal to x.
[0080] In some embodiments, the interpolated picture is between the first pair of input pictures. In some other embodiments, the interpolated picture is between the last pair of input pictures.
[0081] In some embodiments, a picture is interpolated between a pair of consecutive input pictures, and the identities of the pair of consecutive input pictures are indicated. For example, the POC values of the input pictures in the pair of consecutive input pictures are indicated. As another example, the index of the first input picture in the pair of consecutive input pictures is indicated.
[0082] In some embodiments, padding is used to generate unavailable input pictures. For example, unavailable pictures are replaced with the closest available picture in decoding order. As another example, unavailable pictures are replaced with the available picture with the lowest quantization parameter (QP) in decoding order.
[0083] In some embodiments, unavailable images are filled with default values. For example, the default values are normalized. In one example, the default value is 0. Alternatively, the default value is 0.5. As another example, the default value is 1.
[0084] In some embodiments, the default value is not normalized and is in the range from 0 to N-1 (inclusive), where N is an integer. In some embodiments, N is specified as 1<<(nnpfc_inp_tensor_bitdepth_minus8+8).
[0085] In some embodiments, the unavailable pictures are interpolated using existing available pictures. For example, a bilinear filter is used to interpolate the unavailable pictures. As another example, a bicubic filter is used to interpolate the unavailable pictures. As another example, a Lanczos filter is used to interpolate the unavailable pictures. For example, a neural network-based interpolation filter is used to interpolate the unavailable pictures.
[0086] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video, and the bitstream of the video is generated by a method performed by a device for video processing. The method includes: determining that a neural network post-processing filter (NNPF) is activated for a group of pictures related to a video unit of the video; applying the NNPF to one or more pictures in the group of pictures according to a sequence; and generating a bitstream based on the output of the NNPF.
[0087] According to yet other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. The method includes: determining that a neural network post-processing filter (NNPF) is activated for a set of pictures related to a video unit of the video; applying the NNPF to one or more pictures in the set of pictures according to a sequence; generating a bitstream based on an output of the NNPF; and storing the bitstream in a non-transitory computer-readable medium.
[0088] The embodiments of the present disclosure may be described according to the following items, features of which may be combined in any reasonable way.
[0089] Item 1. A method for video processing, comprising: for conversion between a video unit of a video and a bitstream of the video unit, determining that a neural network post-processing filter (NNPF) is activated for a set of pictures related to the video unit; applying the NNPF to one or more pictures in the set of pictures according to a sequence; and performing conversion based on an output of the NNPF.
[0090] Item 2. The method of Item 1, wherein the NNPF is applied to each picture in the set of pictures in an output order of the pictures in the set of pictures, and the NNPF is applied one picture at a time.
[0091] Item 3. The method of Item 1, wherein the NNPF is applied to each picture in a group of pictures in the decoding order of the pictures in the group of pictures, and the NNPF is applied one picture at a time.
[0092] Item 4. The method of Item 1, wherein the NNPF is applied to each group comprising a fixed number of consecutive pictures in a group of pictures in decoding order of the pictures in the group of pictures, and the NNPF is applied to one group at a time.
[0093] Item 5. The method of Item 1, wherein the NNPF is applied to each group comprising a fixed number of consecutive pictures in the group of pictures in an output order of the pictures in the group of pictures, and the NNPF is applied to one group at a time.
[0094] Item 6. The method of any one of Items 1 to 5, wherein if the NNPF takes a picture as input, and if the NNPF is applied to a current picture for which the NNPF is activated, the input picture of the NNPF is the current picture.
[0095] Item 7. A method according to any one of items 1 to 6, wherein the picture used as input picture to the NNPF is determined, or wherein the picture used as input picture to the NNPF is specified, or wherein the picture used as input picture to the NNPF is indicated.
[0096] Item 8. The method of Item 7, wherein input pictures are specified and listed in output order.
[0097] Item 9. The method of Item 8, wherein one or more pictures before the current picture are listed in output order, and one or more pictures after the current picture are listed in output order.
[0098] Item 10. The method of Item 8, wherein one or more pictures before the current picture are listed in decoding order, and one or more pictures after the current picture are listed in output order.
[0099] Item 11. The method of Item 8, wherein one or more pictures before the current picture are listed in output order, and one or more pictures after the current picture are listed in decoding order.
[0100] Item 12. The method of Item 8, wherein one or more pictures before the current picture are listed in decoding order, and one or more pictures after the current picture are listed in decoding order.
[0101] Clause 13. The method of clause 7, wherein one or more identifications related to the picture used as input are indicated.
[0102] Clause 14. The method of clause 13, wherein each picture order count (POC) value of the input pictures is indicated.
[0103] Item 15. The method of Item 13, wherein differences with respect to POC values in the input pictures are indicated.
[0104] Item 16. The method of Item 15, wherein a POC difference between each input picture and the picture for which the NNPF is activated is indicated.
[0105] Item 17. The method of Item 15, wherein a minimum POC value in the input picture is indicated, and a difference between the remaining POC values and the minimum POC value is indicated.
[0106] Item 18. The method of Item 15, wherein a maximum POC value in the input picture is indicated, and a difference between the remaining POC values and the maximum POC value is indicated.
[0107] Item 19. The method of Item 15, wherein a median POC value in the input picture is indicated, and a difference between the remaining POC value and the median POC value is indicated.
[0108] Item 20. The method of any one of items 1 to 19, wherein for NNPF purposes of picture rate upsampling, one or more pictures are interpolated between a pair of input pictures, even if more than two input pictures are utilized for the interpolation.
[0109] Item 21. The method of Item 20, wherein the syntax of the post-processing filter characteristics (NNPFC) supplemental enhancement information (SEI) message requires that each NNPFC interpolated picture in the i-th NNPFC interpolated picture is greater than 0 for only one i in the range of 0 to the number of NNPFCs of the input picture minus 2 (including the boundary values), the i-th NNPFC interpolated picture is represented as nnpfc_interpolated_pics[i], the number of NNPFCs of the input picture minus 2 is represented as nnpfc_num_input_pics_minus2, where i is an integer.
[0110] Item 22. The method of Item 21, wherein the i-th NNPFC interpolation picture is equal to 0 for all other values of i.
[0111] Item 23. The method of any one of Items 1 to 19, wherein for NNPF purposes of picture rate upsampling, one or more pictures are interpolated between a pair of intermediate input pictures, even if more than two input pictures are utilized for the interpolation.
[0112] Item 24. The method of item 23, wherein the syntax of the NNPFC SEI message requires that each NNPFC interpolated picture in the i-th NNPFC interpolated picture is greater than 0 if and only if i is equal to 0 to the value of the number of NNPFCs of the input picture minus 2 divided by 2, the i-th NNPFC interpolated picture is denoted as nnpfc_interpolated_pics[i], the value of the number of NNPFCs of the input picture minus 2 divided by 2 is denoted as nnpfc_num_input_pics_minus2 / 2, where i is an integer.
[0113] Item 25. The method of Item 23, wherein only one instance of the i-th NNPFC interpolated picture is indicated in the NNPFC SEI message, the only one instance of the i-th NNPFC interpolated picture being denoted as nnpfc_interpolated_pics[i].
[0114] Clause 26. The method of clause 25, wherein a syntax element denoted nnpfc_interpolated_pics_minus1 is indicated in the NNPFC SEI message, and wherein nnpfc_interpolated_pics_minus1 plus 1 specifies the number of interpolated pictures between a pair of middle input pictures.
[0115] Item 27. The method of Item 23, wherein nnpfc_num_input_pics_minus2 is an even number, and a pair of middle input pictures is the nnpfc_num_input_pics_minus2 / 2th input picture and the (nnpfc_num_input_pics_minus2 / 2+1)th input picture.
[0116] Item 28. The method of any one of Items 1 to 27, wherein an interpolated picture generated by a post-processing filter is specified, or wherein an interpolated picture generated by an NNPF is indicated.
[0117] Clause 29. The method of clause 28, wherein one or more syntax elements are indicated to specify an interpolated picture.
[0118] Item 30. The method of Item 29, wherein a minimum POC value of the interpolated picture and a maximum POC value of the interpolated picture are indicated.
[0119] Clause 31. The method of clause 29, wherein a minimum POC value of the interpolated pictures is indicated, and a maximum POC value of the interpolated pictures is derived based on a number of interpolated pictures indicated in the NNPFC SEI message.
[0120] Item 32. The method of Item 28, wherein the interpolated picture is determined based on the input picture.
[0121] Item 33. The method of Item 32, wherein the interpolated image is between a pair of intermediate input images.
[0122] Item 34. The method according to Item 33, wherein the index is and The input picture of comprises a pair of intermediate input pictures, where N is an integer and represents the number of input pictures, and Represents the largest integer less than or equal to x.
[0123] Item 35. The method of Item 32, wherein the interpolated picture is between the first pair of input pictures.
[0124] Item 36. The method of Item 32, wherein the interpolated image is between the last pair of input images.
[0125] Item 37. The method of Item 32, wherein the picture is interpolated between a pair of consecutive input pictures, and the identities of the pair of consecutive input pictures are indicated.
[0126] Item 38. The method of Item 37, wherein the POC value of the input picture in a pair of consecutive input pictures is indicated.
[0127] Item 39. The method of Item 37, wherein an index of a first input picture of a pair of consecutive input pictures is indicated.
[0128] Item 40. A method according to any one of Items 1 to 39, wherein padding is used to generate an unusable input image.
[0129] Item 41. The method of Item 40, wherein the unavailable picture is replaced with the closest available picture in order of decoding order.
[0130] Item 42. The method of Item 40, wherein the unavailable picture is replaced with an available picture having the lowest quantization parameter (QP) in order in decoding order.
[0131] Item 43. The method of Item 40, wherein unavailable images are filled with default values.
[0132] Item 44. The method of Item 43, wherein the default value is normalized.
[0133] Item 45. The method of Item 44, wherein the default value is 0, or wherein the default value is 0.5, or wherein the default value is 1.
[0134] Item 46. The method of Item 43, wherein the default value is not normalized and is in the range from 0 to N-1 (inclusive), where N is an integer.
[0135] Item 47. The method of Item 46, wherein N is defined as 1<<(nnpfc_inp_tensor_bitdepth_minus8+8).
[0136] Item 48. The method of any one of Items 1 to 39, wherein the unavailable pictures are interpolated using existing available pictures.
[0137] Item 49. The method of Item 48, wherein a bilinear filter is used for interpolation of unavailable images.
[0138] Item 50. The method of Item 48, wherein a bicubic filter is used for interpolation of unavailable images.
[0139] Item 51. The method of Item 48, wherein a Lanczos filter is used for interpolation of unavailable images.
[0140] Item 52. The method of Item 48, wherein a neural network based interpolation filter is used for interpolation of unavailable images.
[0141] Item 53. The method of any one of Items 1 to 52, wherein converting comprises encoding the video unit into a bitstream.
[0142] Item 54. The method of any one of Items 1 to 52, wherein converting comprises decoding the video unit from a bitstream.
[0143] Item 55. An apparatus for video processing, comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1 to 54.
[0144] Item 56. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 54.
[0145] Item 57. A non-transitory computer-readable recording medium storing a bitstream of a video, the bitstream of the video being generated by a method performed by an apparatus for video processing, wherein the method comprises: determining that a neural network post-processing filter (NNPF) is activated for a set of pictures associated with a video unit of the video; applying the NNPF to one or more pictures in the set of pictures according to a sequence; and generating a bitstream based on an output of the NNPF.
[0146] Item 58. A method for storing a bitstream of a video, comprising: determining that a neural network post-processing filter (NNPF) is activated for a set of pictures associated with a video unit of the video; applying the NNPF to one or more pictures in the set of pictures according to a sequence; generating a bitstream based on an output of the NNPF; and storing the bitstream in a non-transitory computer-readable medium. Example device
[0147] Figure 6 A block diagram of a computing device 600 in which various embodiments of the present disclosure may be implemented is shown. The computing device 600 may be implemented as, or included in, the source device 60 (or video encoder 64 or 200) or the destination device 120 (or video decoder 124 or 300).
[0148] It should be understood that Figure 6 The computing device 600 shown in FIG. 6 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the disclosed embodiments.
[0149] like Figure 6As shown, computing device 600 comprises a general computing device 600. Computing device 600 may include at least one or more processors or processing units 610, memory 620, storage unit 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660.
[0150] In some embodiments, the computing device 600 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, a large computing device, etc. provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 600 can support any type of interface to the user (such as a "wearable" circuit device, etc.).
[0151] Processing unit 610 may be a physical processor or a virtual processor and may implement various processes based on programs stored in memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of computing device 600. Processing unit 610 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0152] The computing device 600 typically includes various computer storage media. Such media can be any media accessible by the computing device 600, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory), or any combination thereof. The storage unit 630 can be any removable or non-removable medium and can include machine-readable media, such as memory, flash drive, disk, or other media that can be used to store information and / or data and can be accessed in the computing device 600.
[0153] The computing device 600 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 6Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.
[0154] The communication unit 640 communicates with another computing device via a communication medium. In addition, the functionality of the components in the computing device 600 can be implemented by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0155] Input device 650 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and the like. Output device 660 may be one or more of various output devices, such as a display, speaker, printer, and the like. With the aid of communication unit 640, computing device 600 may also communicate with one or more external devices (not shown), such as storage devices and display devices, one or more devices that enable a user to interact with computing device 600, or, if desired, any device that enables computing device 600 to communicate with one or more other computing devices (e.g., a network card, a modem, and the like). Such communication may be performed via an input / output (I / O) interface (not shown).
[0156] In some embodiments, some or all components of the computing device 600 may also be arranged in a cloud computing architecture rather than being integrated into a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides an application via a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on servers in a remote location. Computing resources in a cloud computing environment can be consolidated or distributed across remote data centers. Cloud computing infrastructure can provide services through shared data centers, although to users, they appear as a single access point. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider in a remote location. Alternatively, the components and functionality described herein can be provided by a conventional server or installed directly or otherwise on a client device.
[0157] In an embodiment of the present disclosure, the computing device 600 may be used to implement video encoding / decoding. The memory 620 may include one or more video encoding / decoding modules 625 having one or more program instructions. These modules are accessible and executable by the processing unit 610 to perform the functions of the various embodiments described herein.
[0158] In an example embodiment performing video encoding, an input device 650 may receive video data as input to be encoded 670. The video data may be processed, for example, by a video codec module 625 to generate an encoded bitstream. The encoded bitstream may be provided as output 680 via an output device 660.
[0159] In an example embodiment performing video decoding, an input device 650 may receive an encoded bitstream as input 670. The encoded bitstream may be processed, for example, by a video codec module 625 to generate decoded video data. The decoded video data may be provided as output 680 via an output device 660.
[0160] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such variations are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for video processing, comprising: For conversion between a video unit of a video and a bitstream of the video unit, determining a neural network post-processing filter (NNPF) to be activated for a set of pictures associated with the video unit; Applying the NNPF to one or more pictures in the set of pictures according to a sequence; as well as The conversion is performed based on the output of the NNPF. 2 . The method of claim 1 , wherein the NNPF is applied to each picture in the set of pictures in an output order of the pictures in the set of pictures, and the NNPF is applied to one picture at a time.
3. The method of claim 1, wherein the NNPF is applied to each picture in the set of pictures in a decoding order of the pictures in the set of pictures, and the NNPF is applied one picture at a time.
4. The method of claim 1, wherein the NNPF is applied to each group including a fixed number of consecutive pictures in the group of pictures in the decoding order of the pictures in the group of pictures, and the NNPF is applied to one group at a time. 5 . The method of claim 1 , wherein the NNPF is applied to each group including a fixed number of consecutive pictures in the set of pictures in an output order of the pictures in the set of pictures, and the NNPF is applied to one group at a time.
6. The method according to any one of claims 1 to 5, wherein if the NNPF takes a picture as input, and if the NNPF is applied to a current picture for which the NNPF is activated, the input picture of the NNPF is the current picture.
7. The method according to any one of claims 1 to 6, wherein a picture used as an input picture of the NNPF is determined, or where a picture to be used as an input picture of the NNPF is specified, or Among them, the pictures used as input pictures of the NNPF are indicated. The method of claim 7 , wherein the input pictures are specified and listed in output order. 9 . The method of claim 8 , wherein one or more pictures before the current picture are listed in output order, and one or more pictures after the current picture are listed in output order.
10. The method of claim 8, wherein one or more pictures before the current picture are listed in decoding order, and one or more pictures after the current picture are listed in output order. 11 . The method of claim 8 , wherein one or more pictures before the current picture are listed in output order, and one or more pictures after the current picture are listed in decoding order.
12. The method of claim 8, wherein one or more pictures before the current picture are listed in decoding order, and one or more pictures after the current picture are listed in decoding order.
13. The method of claim 7, wherein one or more identifiers associated with the picture used as input are indicated. The method of claim 13 , wherein each picture order count (POC) value of the input picture is indicated. The method of claim 13 , wherein differences related to POC values in the input pictures are indicated. The method of claim 15 , wherein a POC difference between each input picture and the picture for which the NNPF is activated is indicated. 17 . The method of claim 15 , wherein a minimum POC value in the input picture is indicated, and a difference between remaining POC values and the minimum POC value is indicated.
18. The method of claim 15, wherein a maximum POC value in the input picture is indicated, and a difference between remaining POC values and the maximum POC value is indicated.
19. The method of claim 15, wherein a median POC value in the input picture is indicated, and a difference between a residual POC value and the median POC value is indicated.
20. The method according to any one of claims 19 to 19, wherein for NNPF purposes of picture rate upsampling, one or more pictures are interpolated between a pair of input pictures, even if more than two input pictures are utilized for the interpolation.
21. The method of claim 20, wherein the syntax of a post-processing filter characteristic (NNPFC) supplemental enhancement information (SEI) message requires that each of the i-th NNPFC interpolated pictures is greater than 0 for only one i in the range of 0 to the number of NNPFCs of the input picture minus 2 (including boundary values), the i-th NNPFC interpolated picture being denoted as nnpfc_interpolated_pics[i], the number of NNPFCs of the input picture minus 2 being denoted as nnpfc_num_input_pics_minus2, where i is an integer.
22. The method of claim 21, wherein the i-th NNPFC interpolated picture is equal to 0 for all other values of i.
23. The method according to any one of claims 1 to 19, wherein for NNPF purposes of picture rate upsampling, one or more pictures are interpolated between a pair of intermediate input pictures, even if more than two input pictures are utilized for the interpolation.
24. The method of claim 23, wherein the syntax of the NNPFC SEI message requires that each NNPFC interpolated picture in the i-th NNPFC interpolated picture is greater than 0 for a value from 0 to the number of NNPFCs of the input picture minus 2 divided by 2, the i-th NNPFC interpolated picture being denoted as nnpfc_interpolated_pics[i], the value of the number of NNPFCs of the input picture minus 2 divided by 2 being denoted as nnpfc_num_input_pics_minus2 / 2, where i is an integer. 25 . The method of claim 23 , wherein only one instance of the i-th NNPFC interpolated picture is indicated in the NNPFC SEI message, the only one instance of the i-th NNPFC interpolated picture being denoted as nnpfc_interpolated_pics[i].
26. The method of claim 25, wherein a syntax element denoted nnpfc_interpolated_pics_minus1 is indicated in the NNPFC SEI message, and where nnpfc_interpolated_pics_minus1 plus 1 specifies the number of interpolated pictures between the pair of middle input pictures. 27 . The method of claim 23 , wherein nnpfc_num_input_pics_minus2 is an even number, and the pair of middle input pictures is a nnpfc_num_input_pics_minus2 / 2th input picture and a (nnpfc_num_input_pics_minus2 / 2+1)th input picture.
28. A method according to any one of claims 1 to 27, wherein the interpolated picture generated by the post-processing filter is specified, or wherein the interpolated picture generated by the NNPF is indicated.
29. The method of claim 28, wherein one or more syntax elements are indicated to specify the interpolated picture.
30. The method of claim 29, wherein a minimum POC value of the interpolated picture and a maximum POC value of the interpolated picture are indicated. 31 . The method of claim 29 , wherein a minimum POC value of the interpolated pictures is indicated, and a maximum POC value of the interpolated pictures is derived based on the number of the interpolated pictures indicated in an NNPFC SEI message.
32. The method of claim 28, wherein the interpolated picture is determined based on an input picture.
33. The method of claim 32, wherein the interpolated picture is between the pair of intermediate input pictures.
34. The method according to claim 33, wherein the index is and The input picture includes the pair of middle input pictures, where N is an integer and represents the number of the input pictures, and Represents the largest integer less than or equal to x.
35. The method of claim 32, wherein the interpolated picture is between a first pair of input pictures.
36. The method of claim 32, wherein the interpolated picture is between a last pair of input pictures.
37. The method of claim 32, wherein a picture is interpolated between a pair of consecutive input pictures, and the identities of the pair of consecutive input pictures are indicated.
38. The method of claim 37, wherein a POC value of an input picture in the pair of consecutive input pictures is indicated.
39. The method of claim 37, wherein an index of a first input picture of the pair of consecutive input pictures is indicated.
40. The method according to any one of claims 1 to 39, wherein padding is used to generate an unusable input image.
41. The method of claim 40, wherein the unavailable picture is replaced with the closest available picture in sequence in decoding order.
42. The method of claim 40, wherein the unavailable picture is replaced with an available picture having a lowest quantization parameter (QP) in order in decoding order. The method of claim 40 , wherein the unavailable pictures are filled with default values.
44. The method of claim 43, wherein the default value is normalized.
45. The method according to claim 44, wherein the default value is 0, or The default value is 0.5, or The default value is 1.
46. The method of claim 43, wherein the default value is not normalized and is in the range from 0 to N-1 (inclusive), where N is an integer. The method of claim 46 , wherein the N is specified as 1<<(nnpfc_inp_tensor_bitdepth_minus8+8).
48. The method of any one of claims 1 to 39, wherein unavailable pictures are interpolated using existing available pictures.
49. The method of claim 48, wherein a bilinear filter is used for interpolation of the unavailable picture.
50. The method of claim 48, wherein a bicubic filter is used for interpolation of the unavailable picture.
51. The method of claim 48, wherein a Lanczos filter is used for interpolation of the unavailable picture.
52. The method of claim 48, wherein a neural network based interpolation filter is used for interpolation of the unavailable picture.
53. The method of any one of claims 1 to 52, wherein the converting comprises encoding the video unit into the bitstream.
54. The method of any one of claims 1 to 52, wherein the converting comprises decoding the video unit from the bitstream.
55. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1 to 54.
56. A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to execute the method according to any one of claims 1 to 54.
57. A non-transitory computer-readable recording medium storing a bit stream of a video, wherein the bit stream of the video is generated by a method performed by an apparatus for video processing, wherein the method comprises: determining a neural network post-processing filter (NNPF) to be activated for a set of pictures associated with a video unit of the video; Applying the NNPF to one or more pictures in the set of pictures according to a sequence; as well as The bitstream is generated based on the output of the NNPF.
58. A method for storing a bitstream of a video, comprising: determining a neural network post-processing filter (NNPF) to be activated for a set of pictures associated with a video unit of the video; Applying the NNPF to one or more pictures in the set of pictures according to a sequence; generating the bitstream based on an output of the NNPF; as well as The bitstream is stored in a non-transitory computer-readable medium.