Method and device for video processing and medium
By redesigning the syntax elements and signaling of the neural network post-processing filter, the problems of increasing bit depth and changing frame rate in video encoding and decoding in the prior art are solved, and effective conversion from standard dynamic range to high dynamic range is realized, and the encoding and decoding quality is improved.
Patent Information
- Application Number
- CN202480006377.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-24
- Filing Date
- 2024-01-03
- Publication Date
- 2025-08-22
AI Technical Summary
The existing neural network post-processing filter characteristic (SEI) designs do not support black and white video coloring, bit depth increase, unspecified syntax element range, and frame rate change in video encoding and decoding, resulting in the inability to effectively apply to the conversion from standard dynamic range to high dynamic range.
By respecifying the syntax element range and signaling of the neural network post-processing filter, the application of 4:0:0 input is supported, allowing the bit depth to increase, and clarifying the changes in frame rate and bit depth in signal transmission, ensuring the rationality and integrity of the syntax elements, solving the problems of missing and inconsistency of syntax elements.
It realizes the effective application of neural network post-processing filters in video encoding and decoding, supports the conversion from standard dynamic range to high dynamic range, and improves the quality and flexibility of encoding and decoding.
Smart Images

Figure CN120530643A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly, to a neural network post-processing filter (NNPF). Background Art
[0002] Digital video capabilities are now being used in every aspect of our lives. For video encoding and decoding, various video compression technologies have been proposed, including MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC), and Versatile Video Codec (VVC). However, there is a general desire to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is provided. The method includes performing conversion between a video and a bitstream of the video, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating a purpose of the NNPF, and one of candidates for the purpose is increasing a bit depth of sample values in the at least one picture.
[0005] According to the method of the first aspect of the present disclosure, NNPF is able to support an increase in the bit depth of the input to NNPF. Compared with traditional solutions, the proposed method can advantageously support applications that require an increase in bit depth, such as conversion from standard dynamic range (SDR) to high dynamic range (HDR). In this way, the functionality of NNPF becomes diversified and the encoding and decoding quality can be improved.
[0006] In a second aspect, a device for video processing is provided. The device includes a processor and a non-volatile memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.
[0007] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions, the instructions causing a processor to execute the method according to the first aspect of the present disclosure.
[0008] In a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by a video processing apparatus. The method includes performing conversion between the video and the bitstream, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating a purpose of the NNPF, and one of candidates for the purpose is to increase the bit depth of sample values in the at least one picture.
[0009] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: performing conversion between the video and the bitstream, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating a purpose of the NNPF, and one of candidates for the purpose is to increase the bit depth of sample values in the at least one picture; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings.In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0012] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;
[0013] Figure 2 shows a block diagram illustrating a first example video encoder according to some embodiments of the present disclosure;
[0014] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;
[0015] Figure 4 shows a schematic diagram of a brightness data channel;
[0016] Figure 5 A flowchart showing a method for video processing according to an embodiment of the present disclosure is shown; and
[0017] Figure 6 A block diagram is shown of a computing device in which various embodiments of the present disclosure may be implemented.
[0018] Throughout the drawings, the same or similar reference numbers generally refer to the same or similar elements. DETAILED DESCRIPTION
[0019] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.
[0020] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0021] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment is required to include that particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is intended that such feature, structure, or characteristic, whether or not explicitly described, be applicable to other embodiments and that it is within the knowledge of those skilled in the art to apply such feature, structure, or characteristic.
[0022] It should be understood that although the terms "first" and "second" and the like may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0023] The terms used herein are used only for the purpose of describing specific embodiments and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "comprise," "including," "having," "including," and / or "comprising" when used herein indicate the presence of the features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Sample Environment
[0024] Figure 1is a block diagram illustrating an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0025] The video source 112 may include a source such as a video capture device. Examples of a video capture device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.
[0026] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are coded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be directly transmitted to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0027] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, the destination device 120 being configured to interface with an external display device.
[0028] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.
[0029] Figure 2is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of the video encoder 114 in the system 100 is shown.
[0030] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0031] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.
[0032] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0033] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail in the following sections. Figure 2 are shown separately in the example.
[0034] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0035] The mode selection unit 203 can, for example, select one of a plurality of coding modes (intra-frame coding or inter-frame coding) based on the error result, and provide the resulting intra-frame coded block or inter-frame coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a joint intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).
[0036] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.
[0037] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, "P slices" and "B slices" may refer to portions of a picture consisting of macroblocks that are independent of macroblocks in the same picture.
[0038] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0039] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference pictures in list 0 and list 1 containing multiple reference video blocks, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. The motion estimation unit 204 may output the multiple reference indices and multiple motion vectors for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.
[0040] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0041] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0042] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0043] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0044] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0045] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0046] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0047] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.
[0048] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0049] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0050] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.
[0051] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0052] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 An example of the video decoder 124 in the system 100 is shown.
[0053] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0054] exist Figure 3 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.
[0055] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, which motion information includes motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which includes deriving several most likely candidates based on data from adjacent PBs and reference pictures. The motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially neighboring blocks or temporally neighboring blocks.
[0056] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. Identifiers for the interpolation filters used with sub-pixel precision may be included in the syntax elements.
[0057] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information, and motion compensation unit 302 may use the interpolation filters to produce a prediction block.
[0058] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the encoded video sequence, partition information describing how each macroblock of the picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information used to decode the encoded video sequence. As used herein, in some aspects, "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice can be an entire picture or a region of a picture.
[0059] The intra prediction unit 303 can use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0060] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.
[0061] Some exemplary embodiments of the present disclosure are described in detail below. It should be noted that the section headings used in this document are for ease of understanding and do not limit the embodiments disclosed in a section to that section. In addition, although some embodiments are described with reference to a multifunctional video codec or other specific video codecs, the disclosed technology is also applicable to other video coding and decoding technologies. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps for de-encoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at a different compression bit rate. 1. Brief Overview The present disclosure relates to image / video coding techniques. Specifically, it relates to the definition and signaling of new neural network post-processing filtering objectives, their combination with current objectives, and the scope and constraints of various syntax elements. The ideas can be applied alone or in various combinations to video bitstreams encoded and decoded by any codec (e.g., the Versatile Video Codec (VVC) standard and / or the Versatile SEI Message (VSEI) standard for encoding and decoding video bitstreams). 2. Abbreviation APS Adaptive Parameter Set AU Access Unit CLVS codec layer video sequence CLVSS codec layer video sequence starts CRC Cyclic Redundancy Check CVS codec video sequence FIR Finite Impulse Response IRAP Intra-frame Random Access Point NAL Network Abstraction Layer PPS picture parameter set PU picture unit RASL Random Access Skipped Preamble SEI Supplemental Enhancement Information STSA Stepwise Temporal Sublayer Access VCL video codec layer VSEI Versatile Supplementary Enhancement Information (Recommendation ITU-T H.274 | ISO / IEC 23002-7) VUI Video Availability Information VVC Versatile Video Codec (Recommendation ITU-T H.266 | ISO / IEC 23090-3) 3. Introduction 3.1. Video Codec Standards Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. The two organizations jointly developed the H.262 / MPEG-2 Video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction and transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). When the Versatile Video Codec (VVC) project officially launched, JVET was later renamed the Joint Video Experts Team (JVET). VVC is a new codec standard that aims to reduce bit rate by 50% compared to HEVC. The standard was finalized by JVET at its 19th meeting, which ended on July 1, 2020. The Versatile Video Codec (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) and the associated Versatile Supplementary Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) for codec video bitstreams have been designed for the widest range of applications, including traditional uses such as television broadcasting, video conferencing or playback from storage media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, composition and merging of content from multiple codec video bitstreams, multi-view video, scalable layer encoding and decoding, and viewport-adaptive 360° immersive media. The Essential Video Codec (EVC) standard (ISO / IEC 23094-1) is another video codec standard recently developed by MPEG. 3.2. General SEI messages and SEI messages in VVC and VSEI SEI messages assist with processes related to decoding, display, or other purposes. However, no SEI messages are required for constructing luma or chroma samples through the decoding process. Standard-compliant decoders do not need to process this information to achieve output order consistency. Some SEI messages are required for checking bitstream consistency and output timing decoder consistency. Other SEI messages are not required for checking bitstream consistency. Annex D of VVC specifies the syntax and semantics of SEI message payloads for some SEI messages, and specifies the use of SEI messages and VUI parameters for which syntax and semantics are specified in ITU-T H.274 | ISO / IEC 23002-7. 3.3. Signaling of Neural Network Post-Processing Filters An excerpt from the specifications of two SEI messages used for signaling of neural network post-processing filters is shown below. 4. Question The current design of the Neural Network Post-Processing Filter Characteristics (NNPFC) SEI message has the following problems. 1) It does not support applications with 4:0:0 input, e.g., for colorization of black and white video. 2) It does not support applications that may require an increase in bit depth, for example, from SDR to HDR. 3) The range of the syntax element nnpfc_num_input_pics_minus2 is unspecified. 4) The scope of the syntax element nnpfc_interpolated_pics[i] is unspecified. 5) All instances of nnpfc_interpolated_pics[i] can be equal to 0, which results in no frame rate change from the input video to the output. It contradicts the purpose of frame rate upsampling. 6) numOutputPics does not take the input picture into account, which makes it impossible to change the input picture in some applications. 7) The procedure StoreOutputTensors() used to derive the sample values in the filtered output sample array from the output tensor outputTensor uses the wrong scope in the procedure. 8) When the purpose indicates coloring, some syntax elements are missing, some syntax elements exist but are useless, some constraints are missing, and some conversion processes are missing. 9) Whether the syntax elements for luma or chroma bit depth are signaled shall depend on the input or output method and order. 10) Some syntax and semantics may need to be adjusted to enable non-4:0:0 input for shading. 11) When the input is in 4:0:0 format for rendering, the chroma sample positions should be specified. 5. Detailed solution In order to solve the above problems, the following methods are disclosed. The solutions should be considered as examples to explain the general concept and should not be interpreted in a narrow sense. In addition, these solutions can be applied alone or combined in any way. 1) To address issue 1, new purposes or sub-purposes can be specified to describe the behavior of NNPF for 4:0:0 input. a. In one example, alternatively, further, a combination of chroma format change and / or frame rate and / or picture resolution and / or bit depth change for 4:0:0 input is also allowed, for example, indicated in the SEI message. b. nnpfc_out_order_idc should not be equal to 0 when the input is in 4:0:0 format and there is no color format change. c. In one example, when the NNPF input for which the value of the variable ChromaFormatIdc is equal to 0 is in 4:0:0 format, the NNPF purpose should not be chroma upsampling. i. Alternatively, when the NNPF purpose is chroma upsampling, the NNPF input should not be in 4:0:0 format. d. In one example, when the NNPF input is in 4:0:0 format, the syntax element(s) related to the chroma input tensor are not signaled. i. In one example, when the NNPF input is in 4:0:0 format, no syntax element is signaled to indicate the bit depth of the chroma sample values in the input integer tensor. ii. In one example, when the NNPF input is in 4:0:0 format, the syntax element(s) indicating the chroma type and / or chroma fill value are not signaled. e. In one example, the aforementioned condition “when NNPF input is in 4:0:0 format” can be replaced with “when NNPF input is in 4:4:4 format, separate plane codec is applied, and the current input is the luma component”. 2) To solve problem 2, a new purpose may be specified for bit depth increase. a. In one example, a syntax element may be signaled to indicate the difference between the input bit depth and the output bit depth. i. In one example, the syntax element nnpfc_delta_bitdepth_minus1 is signaled, and nnpfc_delta_bitdepth_minus1 plus 1 specifies the output bit depth minus the input bit depth. b. In one example, when the goal is bit depth increase, assuming both input and output are in integer-valued format, the network output must have a higher bit depth than the network input. c. In one example, when the goal is bit depth increase, assuming the network output is in integer value format, the network output must have a higher bit depth than the bit depth of the decoded video output by the video decoder. d. In one example, when the purpose is related to bit depth increase, it is required that both the network input and output are in integer value format, and the network output must have a higher bit depth than the network input. i. In one example, the network output having a higher bit depth than the network input means that there is no color component whose output bit depth is less than the input bit depth, and that for at least one color component, the output bit depth is greater than the input bit depth. ii. In one example, the network output having a higher bit depth than the network input means that for all color components, the bit depth values of the output are greater than the bit depth values of the input. e. In one example, when the purpose is related to bit depth increase, it is required that both the network input and output are in integer value format, and the network output must have a higher bit depth than the decoded video output. i. In one example, the network output having a higher bit depth than the decoded video output means that there are no color components for which the network output has a bit depth less than the decoded video output, and that for at least one color component, the network output has a bit depth greater than the decoded video output. ii. In one example, the network output having a higher bit depth than the decoded video output means that for all color components, the bit depth values of the network output are greater than the bit depth values of the decoded video output. f. In one example, when the purpose is related to bit depth increase, the network output is required to be in integer value format and the network output must have a higher bit depth than the decoded video output. i. In one example, the network output having a higher bit depth than the decoded video output means that there are no color components for which the network output has a bit depth less than the decoded video output, and that for at least one color component, the network output has a bit depth greater than the decoded video output. ii. In one example, the network output having a higher bit depth than the decoded video output means that for all color components, the bit depth values of the network output are greater than the bit depth values of the decoded video output. 3) To solve problem 3, the range of the syntax element nnpfc_num_input_pics_minus2 is specified to be 0 to N (including the boundary values). a. In one example, N is 2. b. In one example, N is 6. c. In one example, N is 14. d. In one example, N is 30. e. In one example, N is 62. f. In one example, N is 126. g. In one example, N is 254. h. Alternatively, N is 2^K–2, where K is a positive integer. 4) To solve problem 4, the range of the syntax element nnpfc_interpolated_pics[i] is specified to be 0 to N (including the boundary values). a. In one example, N is 1. b. In one example, N is 2. c. In one example, N is 4. d. In one example, N is 8. e. In one example, N is 16. f. In one example, N is 32. g. In one example, N is 64. h. Alternatively, N is 2^K, where K is a positive integer. 5) To solve problem 5, the constraint is that for i in the range of 0 to nnpfc_num_input_pics_minus2 (including the boundary values), at least one of nnpfc_interpolated_pics[i] must be greater than 0. a. Alternatively, nnpfc_interpolated_pics[nnpfc_num_input_pics_minus2] is derived to be non-zero when nnpfc_interpolated_pics[i] is equal to 0 for i in the range of 0 to nnpfc_num_input_pics_minus2–1. i. In one example, nnpfc_interpolated_pics[nnpfc_num_input_pics_minus2] is derived as 1. 6) To solve problem 6, it is proposed to consider the number of input pictures during the calculation of numOutputPics. a. In one example, numOutputPics is the number of input pictures plus the total number of inserted pictures. 7) In order to solve problem 7, it is proposed to use numOutputPics in the process StoreOutputTensors() to derive the sample values in the filtered output sample array from the output tensor outputTensor. 8) To address issue 2, the NNPF objective may indicate bit depth increase (also called bit depth upsampling), possibly together with other types of upsampling. Additionally, one or more of the following aspects apply. a. In one example, regardless of whether the NNPF purpose indicates bit depth increase, it is required that when the bit depth of a color component output by the network is greater than the bit depth of the corresponding color component input to the network, the bit depth of each of the other color components output by the network must be greater than or equal to the bit depth of the corresponding color component input to the network. b. In one example, when the destination indicates increased bit depth, both the network input and output are required to be in integer value format, and the network output must have a higher bit depth than the network input. i. In one example, the network output having a higher bit depth than the network input means that for each color component, the bit depth of the network output is greater than the bit depth of the corresponding color component of the network input. ii. In one example, the network output has a higher bit depth than the network input means that for at least one color component, the bit depth of the network output is greater than the bit depth of the corresponding color component of the network input. c. In one example, when the purpose indicates increased bit depth, both the network input and output are required to be in integer-valued format, and the network output shall have a higher bit depth than the cropped output picture output by the video decoder. i. In one example, the network output has a higher bit depth than the cropped output picture means that for each color component, the bit depth of the network output is greater than the bit depth of the corresponding color component of the cropped output picture. ii. In one example, the network output has a higher bit depth than the cropped output picture means that for at least one color component, the bit depth of the network output is greater than the bit depth of the corresponding color component of the cropped output picture. d. In one example, when the destination indicates increased bit depth, the network output is required to be in integer-valued format and the network output must have a higher bit depth than the cropped output picture output by the video decoder. i. In one example, the network output has a higher bit depth than the cropped output picture means that for each color component, the bit depth of the network output is greater than the bit depth of the corresponding color component of the cropped output picture. ii. In one example, the network output has a higher bit depth than the cropped output picture means that for at least one color component, the bit depth of the network output is greater than the bit depth of the corresponding color component of the cropped output picture. 9) To address issue 8, one or more of the following syntactic or semantic changes are specified: a. In one example, the value of the bit of nnpf_purpose (eg, corresponding to 0x020 (ie, bit 5)) indicates whether the purpose is coloring. i. In one example, when nnpfc_purpose & 0x20 is not equal to 0, it indicates that the purpose is coloring. ii. In one example, when nnpfc_purpose & 0x20 is equal to 0, it indicates that the purpose is no coloring or non-coloring. b. In one example, when the purpose indicates coloring, the input to the network is required to be in 4:0:0 format, where the value of the variable ChromaFormatIdc is equal to 0. c. In one example, when the intent indicates shading, there should be chroma output by the NNPF. i. In one example, when the purpose indicates coloring, nnpfc_out_order_idc should not be equal to 0. d. In one example, when the destination indicates shading, the chroma component sample values may be inferred to be equal to 0, 0.5, or 1 when the tensor input is real. e. In one example, when the destination indicates shading, when the tensor input is an integer, the chroma component sample value can be inferred to be equal to 0 or 1 << (inpTensorBitDepth C -1) or 1<<(inpTensorBitDepth C )-1, where inpTensorBitDepth C is the bit depth of the chroma sample values in the input integer tensor. f. In one example, when the intent indicates shading and ChromaFormatIdc is equal to 0, there may only be luma input. i. In one example, when the destination indicates coloring and ChromaFormatIdc is equal to 0, nnpfc_inp_order_idc shall be equal to 0. g. In one example, when the intent indicates chroma and ChromaFormatIdc is equal to 0, chroma-only input is not allowed. i. In one example, when the purpose indicates coloring and ChromaFormatIdc is equal to 0, nnpfc_inp_order_idc should not be equal to 1. h. In one example, when the destination indicates shading, it may be required that the destination does not also indicate chroma upsampling. i. In one example, when nnpfc_purpose&0x20 is not equal to 0, nnpfc_purpose&0x02 is required to be equal to 0. i. In one example, when the destination indicates chroma upsampling, it may be required that the destination does not also indicate shading. i. In one example, when nnpfc_purpose&0x02 is not equal to 0, nnpfc_purpose&0x20 is required to be equal to 0. j. In one example, the purpose of indicating shading and the purpose of indicating chroma upsampling shall be mutually exclusive. i. In one example, ((nnpfc_purpose & 0x02) != 0) && ((nnpfc_purpose & 0x20) != 0) must be equal to 0. k. In one example, when the destination indicates colorization, an indication is signaled in the NNPFC SEI message (eg, a two-bit syntax element named nnpfc_out_colour_format_idc) to specify whether the color format of the output of the NNPF is 4:2:0, 4:2:2, or 4:4:4 format. 1. In one example, when the purpose indicates coloring, an indication is signaled in the NNPFC SEI message (eg, a two-bit syntax element named nnpfc_out_colour_format_idc) to specify the values of the variables outSubWidthC and outSubHeightC. i. In one example, further, the following are specified: nnpfc_out_colour_fomrat_idc equal to 1 specifies that both outSubWidthC and outSubHeightC are equal to 2. nnpfc_out_colour_fomrat_idc equal to 2 specifies that both outSubWidthC and outSubHeightC are equal to 1. nnpfc_out_colour_fomrat_idc equal to 3 specifies that both outSubWidthC and outSubHeightC are equal to 1. The value of nnpfc_out_colour_fomrat_idc shall not be equal to 0. ii. In one example, further specifying that when nnpfc_purpose&0x02 and nnpfc_purpose&0x20 are both equal to 0, outSubWidthC is inferred to be equal to SubWidthC, and outSubHeightC is inferred to be equal to SubHeightC. 10) To address issue 9, one or more of the following syntactic or semantic changes are specified: a. When the method of sample array ordering of the cropped decoded output picture that is one of the input pictures to the post-processing filter indicates that there is no luma matrix, the syntax element indicating the luma input bit depth shall not be signaled. i. In one example, when nnpfc_inp_order_idc is equal to 1, nnpfc_inp_tensor_luma_bitdepth_minus8 should not be signaled. b. When the method of sample array ordering of the cropped decoded output picture that is one of the input pictures to the post-processing filter indicates that no chroma matrix is present, the syntax element indicating the chroma input bit depth shall not be signaled. i. In one example, when nnpfc_inp_order_idc is equal to 0, nnpfc_inp_tensor_chroma_bitdepth_minus8 should not be signaled. c. In one example, syntax element(s) indicating the method of sample array ordering of a cropped decoded output picture that is one of the input pictures to a post-processing filter should be signaled before syntax element(s) indicating luma and / or chroma input bit depth. i. In one example, nnpfc_inp_order_idc should be signaled before nnpfc_inp_tensor_luma_bitdepth_minus8 and / or nnpfc_inp_tensor_chroma_bitdepth_minus8 in the syntax table. d. When the output order of the samples produced from the post-processing filter indicates that no luma matrix is present, the syntax element indicating the luma input bit depth shall not be signaled. i. In one example, when nnpfc_out_order_idc is equal to 1, nnpfc_out_tensor_luma_bitdepth_minus8 should not be signaled. e. When the output order of the samples produced from the post-processing filter indicates that no chroma matrix is present, the syntax element indicating the chroma input bit depth shall not be signaled. i. In one example, when nnpfc_out_order_idc is equal to 0, nnpfc_out_tensor_chroma_bitdepth_minus8 should not be signaled. f. In one example, the syntax element(s) indicating the output order of samples produced from the post-processing filters should be signaled before the syntax element(s) indicating the luma and / or chroma output bit depth. i. In one example, nnpfc_out_order_idc should be signaled before nnpfc_out_tensor_luma_bitdepth_minus8 and / or nnpfc_out_tensor_chroma_bitdepth_minus8 in the syntax table. 11) To address issue 10, one or more of the following syntactic or semantic changes are specified: a. When the NNPF destination indicates chroma, the output chroma format may be indicated by the syntax element only if the input is in 4:0:0 format. b. Alternatively, when the NNPF destination indicates shading, the output chroma format may be indicated by the syntax element regardless of the input chroma format. 12) To address issue 11, one or more of the following syntactic or semantic changes are specified: a. A flag is signaled to indicate whether the chroma sample positions are signaled. b. When the destination indicates chroma and the input is in 4:0:0 format, the flag described in the above sub-item shall be true (indicating that the chroma sample positions are being signaled). c. When the chroma sample positions are not signaled in the NNPF SEI message, they are assumed to be the same as in the VUI. 6. Examples The following are some example embodiments of the solution aspects outlined in Section 5 above. Most of the relevant parts have been added or modified Underline Shown, and some of the deleted parts are shown with strikethrough. There may be some other changes that are editorial in nature and therefore not highlighted. Example 1 This embodiment addresses solution item 1 and all its subitems outlined in subsection 5 above. 6.2. Example 2 This embodiment addresses Solution Item 2 and all its subitems outlined in Section 5 above. 6.3. Example 3 This embodiment addresses solution items 1 and 2 outlined in subsection 5 above, and all their subitems. 6.4. Example 4 This embodiment addresses solution item 3 and all its subitems outlined in subsection 5 above. 6.5. Example 5 This embodiment addresses solution item 4 and all its subitems outlined in subsection 5 above. 6.6. Example 6 This embodiment addresses Solution Item 5 and all its subitems outlined in Section 5 above. 6.7. Example 7 This embodiment addresses solution item 6 and all its subitems outlined in subsection 5 above. 6.8. Example 8 This embodiment addresses solution item 7 and all its subitems outlined in subsection 5 above. 6.9. Example 9 This embodiment addresses Solution Item 2 and all its subitems outlined in Section 5 above. 6.10. Example 10 This embodiment is directed to solution item 11 and all its subitems outlined in subsection 5 above. Example 11 This embodiment is directed to solution item 11 and all its subitems outlined in subsection 5 above. 6.12. Example 12 This embodiment is directed to solution item 12 and all its subitems outlined in subsection 5 above.
[0062] More details of embodiments of the present disclosure related to neural network post-processing filters will be described below. As used herein, the terms "neural network post-processing filter" and "neural network post filter" can be used interchangeably. The embodiments of the present disclosure should be considered as examples to explain general concepts and should not be interpreted in a narrow sense. In addition, these embodiments can be applied alone or combined in any way.
[0063] Figure 5 FIG. 5 is a flow chart of a method 500 for video processing according to some embodiments of the present disclosure. Figure 5 As shown, at 502, conversion between a video and a bitstream of the video is performed. In some embodiments, the conversion may include encoding the video into a bitstream. Alternatively or additionally, the conversion may include decoding the video from the bitstream.
[0064] A neural network post-processing filter (NNPF) is applied to at least one picture associated with a video. For example, the at least one picture may be used as input to the NNPF. In some embodiments, the at least one picture may include at least one decoded picture of the video. Alternatively, the at least one picture may include at least one cropped decoded picture of the video. For example, the decoded picture and / or the cropped decoded picture may be output by a decoder that decodes the video from a bitstream. In some further embodiments, the at least one picture may include the output of another NNPF for filtering one or more decoded pictures or cropped decoded pictures of the video. For example, the NNPF is spliced with another NNPF. It should be understood that the possible implementations of the at least one picture associated with the video described herein are illustrative only and should not be construed in any way as limiting the present disclosure.
[0065] The bitstream includes a first indication indicating the purpose of the NNPF. For example, the first indication may be included in a supplemental enhancement information (SEI) message or any other suitable video message unit in the bitstream. For example, the first indication may include a syntax element nnpfc_purpose. It should be understood that the names of the indications and / or syntax elements are for illustration only and not limitation, and that the (multiple) indications and (multiple) syntax elements mentioned throughout this disclosure may be represented by any other suitable string other than the strings mentioned in this disclosure. The scope of the present disclosure is not limited in this respect.
[0066] In addition, one of the candidates for the purpose is to increase the bit depth of the sample values in at least one picture. For example, if the purpose of the NNPF includes increasing the bit depth, the bit depth of the sample values in the output of the NNPF can be greater than that of the at least one picture. As used herein, this candidate for the purpose may also be referred to as bit depth increase or bit depth upsampling.
[0067] In view of the above, NNPF is able to support an increase in the bit depth of the input to NNPF. Compared with traditional solutions, the proposed method can advantageously support applications that require increased bit depth, such as conversion from standard dynamic range (SDR) to high dynamic range (HDR). In this way, the functionality of NNPF becomes diversified and the encoding and decoding quality can be improved.
[0068] In some embodiments, if the sample values in the output of the NNPF are in the format of integer values, and the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the bit depth of the sample values in the output is higher than the bit depth of the sample values in the at least one picture. For example, in a case where the at least one picture includes a decoded picture output by a video decoder, when the purpose is to increase the bit depth, assuming that the network output is in the format of integer values, the network output should have a higher bit depth than the bit depth of the decoded picture output by the video decoder.
[0069] In some embodiments, if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, and the bit depth of the sample values in the output is higher than the bit depth of the sample values in at least one picture.
[0070] In an example embodiment, if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, and the bit depth of the sample values in at least one color component of the output is higher than the bit depth of the sample values in at least one corresponding color component of at least one picture. For example, in a case where at least one picture includes a cropped output picture output by a video decoder, when the purpose indicates an increase in the bit depth, the network output is required to be in the format of integer values, and for at least one color component, the bit depth of the network output is greater than the bit depth of the corresponding color component of the cropped output picture.
[0071] In an example embodiment, if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, and the bit depth of the sample values in each color component of the output is higher than the bit depth of the sample values in each corresponding color component of at least one picture.
[0072] In some embodiments, if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the following requirements should be met: 1) the sample values in the output of the NNPF are in the format of integer values; 2) the sample values in at least one picture are in the format of integer values; 3) the bit depth of the sample values in the output is higher than the bit depth of the sample values in at least one picture. For example, the at least one picture may include at least one decoded picture of a video. Alternatively, the at least one picture may include the output of another NNPF. The scope of the present disclosure is not limited in this respect.
[0073] In an example embodiment, if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, and the sample values in the at least one picture are in the format of integer values. In addition, the bit depth of the sample values in each color component of the output is higher than the bit depth of the sample values in each corresponding color component of the at least one picture.
[0074] In another example embodiment, if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, and the sample values in the at least one picture are in the format of integer values. In addition, the bit depth of the sample values in the at least one color component of the output is higher than the bit depth of the sample values in the at least one corresponding color component of the at least one picture.
[0075] In some embodiments, if each of the sample values in at least one picture and the sample values in the output of the NNPF is in the format of an integer value, and the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the bit depth of the sample values in the output is higher than the bit depth of the sample values in the at least one picture.
[0076] In some embodiments, if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the following requirements should be met: 1) the sample values in the output of the NNPF are in the format of integer values; 2) the sample values in at least one picture are in the format of integer values; 3) the bit depth of the sample values in each color component of the output is higher than or equal to the bit depth of the sample values in each corresponding color component of the at least one picture; 4) the bit depth of the sample values in at least one color component of the output is higher than the bit depth of the sample values in at least one corresponding color component of the at least one picture. For example, the at least one picture may include at least one decoded picture of the video. Alternatively, the at least one picture may include the output of another NNPF. The scope of the present disclosure is not limited in this respect.
[0077] In some further embodiments, if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the following requirements should be met: 1) the sample values in the output of the NNPF are in the format of integer values; 2) the bit depth of the sample values in each color component of the output is higher than or equal to the bit depth of the sample values in each corresponding color component of at least one picture; 3) the bit depth of the sample values in at least one color component of the output is higher than the bit depth of the sample values in at least one corresponding color component of at least one picture. For example, the at least one picture may include at least one decoded picture of a video. Alternatively, the at least one picture may include the output of another NNPF. The scope of the present disclosure is not limited in this respect.
[0078] In some embodiments, the bitstream may further include a second indication indicating a difference between the bit depth of the sample values in the output of the NNPF and the bit depth of the sample values in the at least one picture. By way of example and not limitation, the second indication may include a syntax element nnpfc_delta_bitdepth_minus1. Additionally, the value of the syntax element nnpfc_delta_bitdepth_minus1 plus one equals the difference between the bit depth of the sample values in the output and the bit depth of the sample values in the at least one picture.
[0079] In view of the foregoing, the solutions according to some embodiments of the present disclosure can advantageously avoid potential instabilities and logic problems, thereby improving encoding and decoding efficiency.
[0080] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a video, and a method executed by a video processing device generates a bitstream. In the method, conversion between the video and the bitstream is performed. A neural network post-processing filter (NNPF) is applied to at least one picture associated with the video. The bitstream includes a first indication indicating the purpose of the NNPF, and one of the candidates for the purpose is to increase the bit depth of the sample values in at least one picture.
[0081] According to some further embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In the method, conversion between a video and a bitstream is performed. A neural network post-processing filter (NNPF) is applied to at least one picture associated with the video. The bitstream includes a first indication indicating a purpose of the NNPF, and one of the candidates for the purpose is to increase the bit depth of sample values in at least one picture. In addition, the bitstream is stored in a non-transitory computer-readable recording medium.
[0082] The embodiments of the present disclosure may be described according to the following items, features of which may be combined in any reasonable way.
[0083] Item 1. A method for video processing, comprising: performing conversion between a video and a bitstream of the video, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating a purpose of the NNPF, and one of the candidates for the purpose is to increase the bit depth of sample values in the at least one picture.
[0084] Clause 2. The method of clause 1, wherein the first indication comprises a syntax element nnpfc_purpose.
[0085] Item 3. A method according to any of Items 1 to 2, wherein the at least one picture comprises at least one decoded picture or at least one cropped decoded picture of the video.
[0086] Item 4. A method according to any one of Items 1 to 3, wherein if the sample values in the output of the NNPF are in the format of integer values, and the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, then the bit depth of the sample values in the output is higher than the bit depth of the sample values in the at least one picture.
[0087] Item 5. A method according to any one of Items 1 to 3, wherein if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, and the bit depth of the sample values in the output is higher than the bit depth of the sample values in the at least one picture.
[0088] Item 6. A method according to any one of Items 1 to 3, wherein if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, and the bit depth of the sample values in at least one color component of the output is higher than the bit depth of the sample values in at least one corresponding color component of the at least one picture.
[0089] Item 7. A method according to any one of Items 1 to 3, wherein if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, and the bit depth of the sample values in each color component of the output is higher than the bit depth of the sample values in each corresponding color component of the at least one picture.
[0090] Item 8. A method according to any one of Items 1 to 3, wherein if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, the sample values in the at least one picture are in the format of integer values, and the bit depth of the sample values in the output is higher than the bit depth of the sample values in the at least one picture.
[0091] Item 9. A method according to any one of Items 1 to 3, wherein if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, the sample values in the at least one picture are in the format of integer values, and the bit depth of the sample values in each color component of the output is higher than the bit depth of the sample values in each corresponding color component of the at least one picture.
[0092] Item 10. A method according to any one of Items 1 to 3, wherein if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, the sample values in the at least one picture are in the format of integer values, and the bit depth of the sample values in at least one color component of the output is higher than the bit depth of the sample values in at least one corresponding color component of the at least one picture.
[0093] Item 11. A method according to any one of Items 1 to 3, wherein if each of the sample values in the at least one picture and the sample values in the output of the NNPF is in the format of an integer value, and the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, then the bit depth of the sample value in the output is higher than the bit depth of the sample value in the at least one picture.
[0094] Item 12. A method according to any one of Items 1 to 3, wherein if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, the sample values in the at least one picture are in the format of integer values, the bit depth of the sample values in each color component of the output is higher than or equal to the bit depth of the sample values in each corresponding color component of the at least one picture, and the bit depth of the sample values in at least one color component of the output is higher than the bit depth of the sample values in at least one corresponding color component of the at least one picture.
[0095] Item 13. A method according to any one of Items 1 to 3, wherein if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, the bit depth of the sample values in each color component of the output is higher than or equal to the bit depth of the sample values in each corresponding color component of the at least one picture, and the bit depth of the sample values in at least one color component of the output is higher than the bit depth of the sample values in at least one corresponding color component of the at least one picture.
[0096] Item 14. A method according to any one of Items 1 to 13, wherein the bitstream further comprises a second indication for indicating a difference between a bit depth of sample values in the output of the NNPF and the bit depth of the sample values in the at least one picture.
[0097] Item 15. The method of Item 14, wherein the second indication comprises a syntax element nnpfc_delta_bitdepth_minus1, and a value of the syntax element nnpfc_delta_bitdepth_minus1 plus one is equal to the difference between the bit depth of the sample values in the output and the bit depth of the sample values in the at least one picture.
[0098] Item 16. The method of any one of Items 1 to 15, wherein the converting comprises encoding the video into the bitstream.
[0099] Item 17. A method according to any one of Items 1 to 15, wherein the converting comprises decoding the video from the bitstream.
[0100] Item 18. An apparatus for video processing, comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1 to 17.
[0101] Item 19. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 17.
[0102] Item 20. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: performing a conversion between the video and the bitstream, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating a purpose of the NNPF, and one of the candidates for the purpose is to increase the bit depth of sample values in the at least one picture.
[0103] Item 21. A method for storing a bitstream of a video, comprising: performing a conversion between the video and the bitstream, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating a purpose of the NNPF, and one of candidates for the purpose is to increase the bit depth of sample values in the at least one picture; and storing the bitstream in a non-transitory computer-readable recording medium. Example device
[0104] Figure 6 A block diagram of a computing device 600 in which various embodiments of the present disclosure may be implemented is shown. The computing device 600 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0105] It should be understood that Figure 6 The computing device 600 shown in FIG. 6 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the disclosed embodiments.
[0106] like Figure 6 As shown, computing device 600 comprises a general computing device 600. Computing device 600 may include at least one or more processors or processing units 610, memory 620, storage unit 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660.
[0107] In some embodiments, the computing device 600 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, a large computing device, etc. provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 600 can support any type of interface to the user (such as a "wearable" circuit device, etc.).
[0108] The processing unit 610 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capability of the computing device 600. The processing unit 610 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0109] The computing device 600 typically includes various computer storage media. Such media can be any media accessible by the computing device 600, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory), or any combination thereof. The storage unit 630 can be any removable or non-removable medium and can include machine-readable media, such as memory, flash drive, disk, or other media that can be used to store information and / or data and can be accessed in the computing device 600.
[0110] The computing device 600 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 6 Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.
[0111] The communication unit 640 communicates with another computing device via a communication medium. In addition, the functionality of the components in the computing device 600 can be implemented by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 600 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0112] Input device 650 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and the like. Output device 660 may be one or more of various output devices, such as a display, speaker, printer, and the like. With the aid of communication unit 640, computing device 600 may also communicate with one or more external devices (not shown), such as storage devices and display devices, one or more devices that enable a user to interact with computing device 600, or, if desired, any device that enables computing device 600 to communicate with one or more other computing devices (e.g., a network card, a modem, and the like). Such communication may be performed via an input / output (I / O) interface (not shown).
[0113] In some embodiments, some or all components of the computing device 600 may also be arranged in a cloud computing architecture rather than being integrated into a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides an application via a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on servers in a remote location. Computing resources in a cloud computing environment can be consolidated or distributed across remote data centers. Cloud computing infrastructure can provide services through shared data centers, although to users, they appear as a single access point. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider in a remote location. Alternatively, the components and functionality described herein can be provided by a conventional server or installed directly or otherwise on a client device.
[0114] In an embodiment of the present disclosure, the computing device 600 may be used to implement video encoding / decoding. The memory 620 may include one or more video encoding / decoding modules 625 having one or more program instructions. These modules are accessible and executable by the processing unit 610 to perform the functions of the various embodiments described herein.
[0115] In an example embodiment performing video encoding, an input device 650 may receive video data as input to be encoded 670. The video data may be processed, for example, by a video codec module 625 to generate an encoded bitstream. The encoded bitstream may be provided as output 680 via an output device 660.
[0116] In an example embodiment performing video decoding, an input device 650 may receive an encoded bitstream as input 670. The encoded bitstream may be processed, for example, by a video codec module 625 to generate decoded video data. The decoded video data may be provided as output 680 via an output device 660.
[0117] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such variations are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for video processing, comprising: Performing conversion between a video and a bitstream of the video, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating a purpose of the NNPF, and one of candidates for the purpose is increasing the bit depth of sample values in the at least one picture. The method according to claim 1 , wherein the first indication comprises a syntax element nnpfc_purpose. 3 . The method according to claim 1 , wherein the at least one picture comprises at least one decoded picture of the video or at least one cropped decoded picture.
4. The method according to any one of claims 1 to 3, wherein if the sample values in the output of the NNPF are in the format of integer values, and the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, then the bit depth of the sample values in the output is higher than the bit depth of the sample values in the at least one picture.
5. The method according to any one of claims 1 to 3, wherein if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, and the bit depth of the sample values in the output is higher than the bit depth of the sample values in the at least one picture.
6. A method according to any one of claims 1 to 3, wherein if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, and the bit depth of the sample values in at least one color component of the output is higher than the bit depth of the sample values in at least one corresponding color component of the at least one picture.
7. A method according to any one of claims 1 to 3, wherein if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, and the bit depth of the sample values in each color component of the output is higher than the bit depth of the sample values in each corresponding color component of the at least one picture.
8. The method according to any one of claims 1 to 3, wherein if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, the sample values in the at least one picture are in the format of integer values, and the bit depth of the sample values in the output is higher than the bit depth of the sample values in the at least one picture.
9. A method according to any one of claims 1 to 3, wherein if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, the sample values in the at least one picture are in the format of integer values, and the bit depth of the sample values in each color component of the output is higher than the bit depth of the sample values in each corresponding color component of the at least one picture.
10. The method according to any one of claims 1 to 3, wherein if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, the sample values in the at least one picture are in the format of integer values, and the bit depth of the sample values in at least one color component of the output is higher than the bit depth of the sample values in at least one corresponding color component of the at least one picture.
11. A method according to any one of claims 1 to 3, wherein if each of the sample values in the at least one picture and the sample values in the output of the NNPF is in the format of an integer value, and the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, then the bit depth of the sample value in the output is higher than the bit depth of the sample value in the at least one picture.
12. A method according to any one of claims 1 to 3, wherein if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, the sample values in the at least one picture are in the format of integer values, the bit depth of the sample values in each color component of the output is higher than or equal to the bit depth of the sample values in each corresponding color component of the at least one picture, and the bit depth of the sample values in at least one color component of the output is higher than the bit depth of the sample values in at least one corresponding color component of the at least one picture.
13. A method according to any one of claims 1 to 3, wherein if the first indication is used to indicate that the purpose of the NNPF includes increasing the bit depth, the sample values in the output of the NNPF are in the format of integer values, the bit depth of the sample values in each color component of the output is higher than or equal to the bit depth of the sample values in each corresponding color component of the at least one picture, and the bit depth of the sample values in at least one color component of the output is higher than the bit depth of the sample values in at least one corresponding color component of the at least one picture.
14. The method according to any one of claims 1 to 13, wherein the bitstream further comprises a second indication for indicating a difference between a bit depth of a sample value in the output of the NNPF and the bit depth of the sample value in the at least one picture. 15 . The method of claim 14 , wherein the second indication comprises a syntax element nnpfc_delta_bitdepth_minus1 , and a value of the syntax element nnpfc_delta_bitdepth_minus1 plus one is equal to the difference between the bit depth of the sample values in the output and the bit depth of the sample values in the at least one picture.
16. The method of any one of claims 1 to 15, wherein the converting comprises encoding the video into the bitstream.
17. The method of any one of claims 1 to 15, wherein the converting comprises decoding the video from the bitstream.
18. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 17.
19. A non-transitory computer-readable storage medium storing instructions, wherein the instructions cause a processor to execute the method according to any one of claims 1 to 17.
20. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: Performing conversion between the video and the bitstream, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream including a first indication indicating a purpose of the NNPF, and one of the candidates for the purpose is to increase the bit depth of sample values in the at least one picture.
21. A method for storing a bitstream of a video, comprising: performing conversion between the video and the bitstream, wherein a neural network post-processing filter (NNPF) is applied to at least one picture associated with the video, the bitstream comprising a first indication indicating a purpose of the NNPF, and wherein one of candidates for the purpose is increasing a bit depth of sample values in the at least one picture; as well as The bitstream is stored in a non-transitory computer-readable recording medium.