Neural network post-processing filter purpose with picture rate up-sampling
By introducing neural network post-processing filter characteristic signaling, the video data and bitstream conversion is optimized, and the problem of inefficient bandwidth usage in video encoding and decoding is solved, and more efficient video transmission is achieved.
Patent Information
- Application Number
- CN202480006858.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-14
- Filing Date
- 2024-01-04
- Publication Date
- 2025-08-12
AI Technical Summary
When processing video data, existing video encoding and decoding technologies are difficult to effectively manage the signaling of neural network post-processing filters, resulting in inefficient bandwidth usage, especially in high-resolution and high-frame rate video transmission, which increases the network burden.
By introducing neural network postprocessing filter characteristics (NNPFC) signaling, supplementary enhancement information (SEI) messages are used to determine the purpose of neural network postprocessing filter (NNPF), and the conversion process between video data and bitstream is optimized, including image rate upsampling and interpolated image number management.
It improves bandwidth usage efficiency during video encoding and decoding, reduces network burden, and improves the quality and efficiency of video transmission.
Smart Images

Figure CN120476586A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefits of International Patent Application No. PCT / CN2023 / 072219, filed on January 14, 2023, which claims priority to International Patent Application No. PCT / CN2023 / 070334, filed on January 4, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure relates to the generation, storage, and consumption of digital audio-visual media information in file formats. Background Art
[0004] Digital video accounts for the largest share of bandwidth usage on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, bandwidth demands for digital video usage are likely to continue to grow. Summary of the Invention
[0005] The first aspect relates to a method for processing video data, comprising: determining a neural network post-processing filter (NNPF) purpose based on a neural network post-processing filter characteristic (NNPFC) supplemental enhancement information (SEI) message; and performing conversion between visual media data and a bitstream based on the NNPF purpose.
[0006] A second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform any one of the aforementioned aspects.
[0007] A third aspect relates to a non-transitory computer-readable medium, comprising a computer program product for use with a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, so that when executed by a processor, the video codec device performs the method according to any one of the preceding aspects.
[0008] A fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method includes: determining a neural network post-processing filter (NNPF) purpose based on a neural network post-processing filter characteristic (NNPFC) supplemental enhancement information (SEI) message; and generating a bitstream based on the determination.
[0009] A fifth aspect relates to a method for storing a bitstream of a video, comprising: determining a neural network post-processing filter (NNPF) purpose based on a neural network post-processing filter characteristic (NNPFC) supplemental enhancement information (SEI) message; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] The sixth aspect relates to the method, device or system described in the present disclosure.
[0011] For the sake of clarity, any of the foregoing embodiments may be combined with any one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.
[0012] These and other features will become more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] For a more complete understanding of this disclosure, reference is now made to the following brief description taken in connection with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.
[0014] Figure 1 An example of deriving a luma channel from a luma component is shown.
[0015] Figure 2 is a block diagram showing an example video processing system.
[0016] Figure 3 is a block diagram of an example video processing device.
[0017] Figure 4 is a flow chart of an example method of video processing.
[0018] Figure 5 is a block diagram illustrating an example video encoding and decoding system.
[0019] Figure 6 is a block diagram illustrating an example encoder.
[0020] Figure 7 is a block diagram illustrating an example decoder.
[0021] Figure 8 is a schematic diagram of an example encoder. DETAILED DESCRIPTION
[0022] It should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or yet to be developed. The present disclosure should not be limited in any way to the illustrative implementations, drawings, and examples shown below, including the exemplary designs and implementations shown and described herein, but may be modified within the scope of the appended claims and their full scope of equivalents.
[0023] The section headings used in this disclosure are for ease of understanding and are not intended to limit the applicability of the techniques and embodiments disclosed in each section to only that section. In addition, the use of H.266 terminology in some descriptions is for ease of understanding only and is not intended to limit the scope of the disclosed embodiments. Therefore, the embodiments described herein are also applicable to other video codec protocols and designs. In this disclosure, editorial changes to the Versatile Video Codec (VVC) specification and / or the International Organization for Standardization (ISO) Base Media File Format (ISOBMFF) standard are shown in bold italics to indicate deleted text and in bold to indicate added text.
[0024] 1. Preliminary Discussion
[0025] The present disclosure relates to image / video codec technology. In particular, the present disclosure relates to the definition and signaling of neural network post-processing filter (NNPF) objectives, which have picture rate upsampling and other types of upsampling, more efficient signaling of the number of interpolated pictures, the order of multiple types of upsampling, and the number of input pictures for any NNPF purpose. These concepts can be applied alone or in various combinations to video bitstreams encoded and decoded by any codec, such as the Versatile Video Codec (VVC) standard and / or the Versatile Supplementary Enhancement Information (SEI) messages for encoding and decoding video bitstreams (VSEI) standard.
[0026] 2. Abbreviation
[0027] The following abbreviations may be used in the present disclosure: adaptation parameter set (APS), access unit (AU), codec layer video sequence (CLVS), codec layer video sequence start (CLVSS), cyclic redundancy check (CRC), codec video sequence (CVS), finite impulse response (FIR), intra random access point (IRAP), network abstraction layer (NAL), picture parameter set (PPS), picture unit (PU), random access skip leading (RASL) picture, supplemental enhancement information (SEI), step-by-step temporal sublayer access (STSA), video codec layer (VCL), generic supplemental enhancement information (VSEI) as described in Rec. ITU-T H.274 | ISO / IEC 23002-7, generic video codec (VVC) as described in Rec. ITU-T H.266 | ISO / IEC 23090-3.
[0028] 3. Further Discussion
[0029] 3.1 Video Codec Standards
[0030] Video codec standards have evolved primarily through the development of standards within the International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) and the International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC). ITU-T produced H.261 and H.263, ISO / IEC produced Moving Picture Experts Group (MPEG)-1 and MPEG-4 Visual, and the two organizations jointly produced H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / High Efficiency Video Codec (HEVC) [1]. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore video codec technologies beyond High Efficiency Video Codec (HEVC), the Video Codec Experts Group (VCEG) and the Moving Picture Experts Group (MPEG) established the Joint Video Exploration Team (JVET). Furthermore, JVET has adopted a variety of methodologies and incorporated them into a reference software called the Joint Exploration Model (JEM) [2]. JVET was later renamed the Joint Video Experts Team (JVET) when the Versatile Video Codec (VVC) project was officially launched. VVC[3] is a codec standard that aims to reduce the bit rate by 50% compared to HEVC.
[0031] The Versatile Video Codec (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) [3] and the related Versatile Supplementary Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) [4] for codec video bitstreams are intended for the widest range of applications, including simple uses such as television broadcasting, video conferencing or stored media playback, as well as more advanced use cases such as adaptive bitrate streaming, video region extraction, composition and merging of content from multiple codec video bitstreams, multi-view video, scalable layered codecs and viewport-adaptive 360° immersive media.
[0032] The Essential Video Codec (EVC) standard (ISO / IEC 23094-1) is another video codec standard being developed by MPEG.
[0033] 3.2 SEI Message Overview and SEI Messages in VVC and VSEI
[0034] SEI messages facilitate processes related to decoding, display, or other purposes. However, no SEI messages are required to construct luma or chroma samples from the decoding process. A compliant decoder does not need to process this information to conform to the output order. Some SEI messages are required to check bitstream conformance and output timing decoder conformance. Other SEI messages are not required to check bitstream conformance.
[0035] Appendix D of VVC specifies the syntax and semantics of the SEI message payload of some SEI messages and specifies the use of SEI messages and VUI parameters. Its syntax and semantics are specified in ITU-T H.274 | ISO / IEC 23002-7.
[0036] 3.3 Signaling of Neural Network Post-Processing Filters
[0037] The WG 05 output documents N0158[5] and JVET-AB2006[6] include specifications for two SEI messages for neural network post-processing filter signaling, as shown below.
[0038] 8.28 Neural Network Post-Processing Filter Characteristics SEI Message
[0039] 8.28.1 Neural Network Post-Processing Filter Characteristics SEI Message Syntax
[0040]
[0041]
[0042]
[0043] 8.28.2 Neural Network Post-Processing Filter Characteristics SEI Message Semantics
[0044] The Neural Network Post-Processing Filter Characteristics (NNPFC) SEI message specifies a neural network that can be used as a post-processing filter. The use of a specified post-processing filter for a particular picture is indicated by the Neural Network Post-Processing Filter Activation SEI message.
[0045] Using this SEI message requires defining the following variables:
[0046] – The width and height of the cropped decoded output picture in units of luma samples, denoted by CroppedWidth and CroppedHeight respectively.
[0047] – Crops the luma sample array CroppedYPic[idx] and the chroma sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] (if present) of the decoded output picture, where idx is in the range 0 to numInputPics-1 (inclusive), to be used as input to the post-processing filters.
[0048] – Clip the bit depth BitDepthY of the luminance sample array of the decoded output picture.
[0049] – Clip the bit depth BitDepthC of the chroma sample array of the decoded output picture (if any).
[0050] – The chroma format indicator, denoted by ChromaFormatIdc, as described in subclause 7.3.
[0051] – When nnpfc_auxiliary_inp_idc is equal to 1, the filter strength control value StrengthControlVal shall be a real number in the range 0 to 1 (inclusive).
[0052] The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc as specified in Table 2. NOTE 1 – More than one NNPFC SEI message may be present for the same picture. When more than one NNPFC SEI message with different nnpfc_id values is present or activated for the same picture, they may have the same or different nnpfc_purpose and nnpfc_mode_idc values.
[0053] nnpfc_id contains an identification number that can be used to identify the post-processing filter. The value of nnpfc_id shall be in the range of 0 to 232-2 (inclusive). The values of nnpfc_id from 256 to 511 (inclusive) and from 231 to 232-2 (inclusive) are reserved for future use by ITU-T|ISO / IEC. Decoders conforming to this version of this document that encounter an NNPFC SEI message with a nnpfc_id in the range of 256 to 511 (inclusive) or 231 to 232-2 (inclusive) shall ignore the SEI message.
[0054] When the NNPFC SEI message is the first NNPFC SEI message with a specific nnpfc_id value in decoding order in the current CLVS, the following rules apply:
[0055] – This SEI message specifies the base post-processing filters.
[0056] – This SEI message is related to the current decoded picture and all subsequent decoded pictures of the current layer, in output order, until the end of the current CLVS.
[0057] When an NNPFC SEI message repeats a previous NNPFC SEI message in decoding order within the current CLVS, the subsequent semantics apply as if that SEI message was the only NNPFC SEI message with the same content within the current CLVS.
[0058] When the NNPFC SEI message is not the first NNPFC SEI message with a specific nnpfc_id value in decoding order in the current CLVS, the following rules apply:
[0059] – This SEI message defines the updates relative to the previous base post-processing filter with the same nnpfc_id value in decoding order.
[0060] – This SEI message pertains to the current decoded picture and all subsequent decoded pictures of the current layer, in output order, until the end of the current CLVS, or the end of the next NNPFC SEI message with a specific nnpfc_id value, in output order, within the current CLVS.
[0061] nnpfc_mode_idc equal to 0 indicates that this SEI message contains an ISO / IEC 15938-17 bitstream that specifies a base post-processing filter, or an update relative to a base post-processing filter with the same nnpfc_id value.
[0062] When the NNPFC SEI message is the first NNPFC SEI message with a particular nnpfc_id value in decoding order in the current CLVS, nnpfc_mode_idc equal to 1 specifies that the base post-processing filter associated with the nnpfc_id value is a neural network identified by the URI indicated by nnpfc_uri, whose format is identified by the tag URI nnpfc_tag_uri.
[0063] When the NNPFC SEI message is not the first NNPFC SEI message with a particular nnpfc_id value in decoding order in the current CLVS, nnpfc_mode_idc equal to 1 specifies that updates relative to the base post-processing filter with the same nnpfc_id value are defined by the URI indicated by nnpfc_uri, whose format is identified by the tag URI nnpfc_tag_uri.
[0064] In bitstreams conforming to this version of this document, the value of nnpfc_mode_idc shall be in the range of 0 to 1, inclusive. Values of nnpfc_mode_idc from 2 to 255, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not appear in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_mode_idc in the range of 2 to 255, inclusive. Values of nnpfc_mode_idc greater than 255 shall not appear in bitstreams conforming to this version of this document and are not reserved for future use.
[0065] When this SEI message is the first NNPFC SEI message with a specific nnpfc_id value in decoding order in the current CLVS, the post-processing filter PostProcessingFilter() is assigned the same value as the base post-processing filter.
[0066] When this SEI message is not the first NNPFC SEI message with a particular nnpfc_id value in decoding order in the current CLVS, the post-processing filter PostProcessingFilter() is obtained by applying the updates defined by this SEI message to the base post-processing filter.
[0067] Updates are not cumulative, but each update is applied to the base post-processing filter specified by the first NNPFC SEI message with a specific nnpfc_id value in decoding order in the current CLVS.
[0068] In bitstreams conforming to this version of this document, nnpfc_reserved_zero_bit_a shall be equal to 0. A decoder shall ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_a is not equal to 0.
[0069] nnpfc_tag_uri contains a tag URI whose syntax and semantics are as specified in IETF RFC 4151, identifying the format and related information about the neural network used as a base post-processing filter, or an update relative to a base post-processing filter with the same nnpfc_id value specified by nnpfc_uri.
[0070] NOTE 2 – nnpfc_tag_uri is able to uniquely identify the format of neural network data specified by nnrpf_uri without the need for a central registration authority.
[0071] nnpfc_tag_uri equal to "tag:iso.org,2023:15938-17" indicates that the neural network data identified by nnpfc_uri complies with ISO / IEC 15938-17.
[0072] nnpfc_uri contains a URI whose syntax and semantics are as described in IETF Internet Standard 66, identifying a neural network used as a base post-processing filter, or an update relative to a base post-processing filter with the same nnpfc_id value.
[0073] nnpfc_formatting_and_purpose_flag equal to 1 specifies that syntax elements related to filter purpose, input formatting, output formatting, and complexity are present. nnpfc_formatting_and_purpose_flag equal to 0 specifies that syntax elements related to filter purpose, input formatting, output formatting, and complexity are not present.
[0074] When this SEI message is the first NNPFC SEI message with a particular nnpfc_id value in decoding order in the current CLVS, nnpfc_formatting_and_purpose_flag shall be equal to 1. When this SEI message is not the first NNPFC SEI message with a particular nnpfc_id value in decoding order in the current CLVS, nnpfc_formatting_and_purpose_flag shall be equal to 0.
[0075] nnpfc_purpose indicates the purpose of the post-processing filter specified in Table 20.
[0076] In bitstreams conforming to this version of this document, the value of nnpfc_purpose shall be in the range of 0 to 5, inclusive. Values of nnpfc_purpose from 6 to 1023, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not appear in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_purpose in the range of 6 to 1203, inclusive. Values of nnpfc_purpose greater than 1023 shall not appear in bitstreams conforming to this version of this document and are not reserved for future use.
[0077] Table 20 - Definition of nnpfc_purpose
[0078]
[0079]
[0080] NOTE 3 – When ITU-T | ISO / IEC uses the reserved value of nnpfc_purpose in the future, the syntax of this SEI message may be extended with a syntax element with nnpfc_purpose equal to that value as a condition for the presence of that syntax element.
[0081] When SubWidthC is equal to 1 and SubHeightC is equal to 1, nnpfc_purpose should not be equal to 2 or 4.
[0082] nnpfc_out_sub_c_flag equal to 1 specifies that outSubWidthC is equal to 1 and outSubHeightC is equal to 1. nnpfc_out_sub_c_flag equal to 0 specifies that outSubWidthC is equal to 2 and outSubHeightC is equal to 1. When nnpfc_out_sub_c_flag is not present, outSubWidthC is inferred to be equal to SubWidth, and outSubHeightC is inferred to be equal to SubHeightC. When ChromaFormatIdc is equal to 2 and nnpfc_out_sub_c_flag is present, the value of nnpfc_out_sub_c_flag shall be equal to 1.
[0083] nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples specify the width and height, respectively, of the array of luma samples of the picture resulting from applying the post-processing filter identified by nnpfc_id to the cropped decoded output picture. When nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples are not present, they are inferred to be equal to CroppedWidth and CroppedHeight, respectively. The value of nnpfc_pic_width_in_luma_samples shall be in the range of CroppedWidth to CroppedWidth*16-1, inclusive. The value of nnpfc_pic_height_in_luma_samples shall be in the range of CroppedHeight to CroppedHeight*16-1, inclusive.
[0084] nnpfc_num_input_pics_minus2 plus 2 specifies the number of decoded output pictures used as input to the post-processing filters.
[0085] nnpfc_interpolated_pics[i] specifies the number of interpolated pictures generated by the post-processing filter between the i-th and (i+1)-th pictures used as input to the post-processing filter.
[0086] The variable numInputPics specifies the number of pictures used as input to the post-processing filter, and the variable numOutputPics specifies the total number of pictures produced by the post-processing filter. The variables numInputPics and numOutputPics are derived as follows:
[0087]
[0088] nnpfc_component_last_flag is equal to 1 to indicate that the last dimension in the input tensor inputTensor of the post-processing filter and the output tensor outputTensor produced by the post-processing filter is used for the current channel. nnpfc_component_last_flag is equal to 0 to indicate that the third dimension in the input tensor inputTensor of the post-processing filter and the output tensor outputTensor produced by the post-processing filter is used for the current channel.
[0089] NOTE 4 – The first dimension in the input and output tensors is used for the batch index, which is a practice in some neural network frameworks. Although the formulas in the semantics of this SEI message use the batch size corresponding to a batch index equal to 0, it is up to the post-processing implementation to determine the batch size used as input to the neural network inference.
[0090] Note 5 – For example, when nnpfc_inp_order_idc is equal to 3 and nnpfc_auxiliary_inp_idc is equal to 1, there are 7 channels in the input tensor, including four luma matrices, two chroma matrices, and one auxiliary input matrix. In this case, the procedure DeriveInputTensors() will derive each of these 7 channels of the input tensor one by one, and when processing a specific channel among these channels, that channel is called the current channel in the procedure.
[0091] nnpfc_inp_format_idc indicates the method of converting the sample values of the cropped decoded output picture into the input values of the post-processing filter. When nnpfc_inp_format_idc is equal to 0, the input values of the post-processing filter are real numbers, and the functions InpY() and InpC() are defined as follows:
[0092] InpY( x ) = x ÷ ( ( 1 << BitDepthY ) - 1 ) (77)
[0093] InpC( x )= x ÷ ( ( 1 << BitDepthC ) - 1 ) (78)
[0094] When nnpfc_inp_format_idc is equal to 1, the input values to the post-processing filters are unsigned integers, and the functions InpY() and InpC() are defined as follows:
[0095]
[0096] The variable inpTensorBitDepth is derived from the syntax element nnpfc_inp_tensor_bitdepth_minus8 as described below.
[0097] Values of nnpfc_inp_format_idc greater than 1 are reserved by ITU-T|ISO / IEC for future specification and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages containing reserved values of nnpfc_inp_format_idc.
[0098] nnpfc_inp_tensor_bitdepth_minus8 specifies the bit depth of the luma sample values in the input integer tensor plus 8. The value of inpTensorBitDepth is derived as follows:
[0099] inpTensorBitDepth = nnpfc_inp_tensor_bitdepth_minus8 + 8 (81)
[0100] A requirement for bitstream conformance is that the value of nnpfc_inp_tensor_bitdepth_minus8 should be in the range 0 to 24, inclusive.
[0101] nnpfc_inp_order_idc indicates the method for ordering the sample array of the cropped decoded output picture as one of the input pictures to the post-processing filter.
[0102] In bitstreams conforming to this version of this document, the value of nnpfc_inp_order_idc shall be in the range of 0 to 3, inclusive. Corresponding values of nnpfc_inp_order_idc from 4 to 255, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not appear in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 4 to 255, inclusive. Values of nnpfc_inp_order_idc greater than 255 shall not appear in bitstreams conforming to this version of this document and are not reserved for future use.
[0103] When ChromaFormatIdc is not equal to 1, nnpfc_inp_order_idc shall not be equal to 3.
[0104] Table 21 contains informative descriptions of the nnpfc_inp_order_idc values.
[0105] Table 21 - Description of nnpfc_inp_order_idc values
[0106]
[0107] Figure 1 An example of deriving the luma channel from the luma component is shown, for example when nnpfc_inp_order_idc is equal to 3.
[0108] A tile is a rectangular array of samples from a component of a picture (eg, luma or chroma).
[0109] nnpfc_auxiliary_inp_idc greater than 0 indicates the presence of auxiliary input data in the input tensor of the neural network post-processing filter. nnpfc_auxiliary_inp_idc equal to 0 indicates the absence of auxiliary input data in the input tensor. nnpfc_auxiliary_inp_idc equal to 1 specifies that the auxiliary input data is derived according to Equation 82.
[0110] In bitstreams conforming to this version of this document, the value of nnpfc_auxiliary_inp_idc shall be in the range of 0 to 1, inclusive. Values of nnpfc_inp_order_idc from 2 to 255, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not appear in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 2 to 255, inclusive. Values of nnpfc_inp_order_idc greater than 255 shall not appear in bitstreams conforming to this version of this document and are not reserved for future use.
[0111] The DeriveInputTensors() procedure derives the input tensor inputTensor given the vertical sample coordinate cTop and horizontal sample coordinate cLeft that specify the top left sample position of the sample patch contained in the input tensor, as follows:
[0112]
[0113]
[0114]
[0115]
[0116] nnpfc_separate_colour_description_present_flag equal to 1 indicates that a different combination of color primaries, transfer characteristics, and matrix coefficients for the picture produced by the post-processing filters is specified in the SEI message syntax structure. nnfpc_separate_colour_description_present_flag equal to 0 indicates that the combination of color primaries, transfer characteristics, and matrix coefficients for the picture produced by the post-processing filters is the same as indicated in the VUI parameters of the CLVS.
[0117] nnpfc_colour_primaries has the same semantics as specified in subclause 7.3 for the vui_colour_primaries syntax element, with the following exceptions:
[0118] –nnpfc_colour_primaries specifies the color primaries of the picture obtained by applying the neural network post-processing filters specified in the SEI message, instead of the primaries used for CLVS.
[0119] – When nnpfc_colour_primaries is not present in the NNPFC SEI message, the value of nnpfc_colour_primaries is inferred to be equal to vui_colour_primaries.
[0120] nnpfc_transfer_characteristics has the same semantics as specified in subclause 7.3 for the vui_transfer_characteristics syntax element, with the following exceptions:
[0121] –nnpfc_transfer_characteristics specifies the transfer characteristics of the picture obtained by applying the neural network post-processing filters specified in the SEI message, instead of the transfer characteristics used for CLVS.
[0122] – When nnpfc_transfer_characteristics is not present in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is inferred to be equal to vui_transfer_characteristics.
[0123] nnpfc_matrix_coeffs has the same semantics as specified for the vui_matrix_coeffs syntax element in subclause 7.3, with the following exceptions:
[0124] –nnpfc_matrix_coeffs specifies the matrix coefficients of the picture obtained by applying the neural network post-processing filters specified in the SEI message, instead of the matrix coefficients used for CLVS.
[0125] – When nnpfc_matrix_coeffs is not present in the NNPFC SEI message, the value of nnpfc_matrix_coeffs is inferred to be equal to vui_matrix_coeffs.
[0126] – The allowed values of nnpfc_matrix_coeffs are not constrained by the chroma format of the decoded video picture, which is indicated by the ChromaFormatIdc value for the semantics of the VUI parameters.
[0127] – When nnpfc_matrix_coeffs is equal to 0, nnpfc_out_order_idc shall not be equal to 1 or 3.
[0128] nnpfc_out_format_idc being equal to 0 indicates that the sample values output by the post - processing filter are real numbers, where the value range from 0 to 1 (including the end values) is linearly mapped to the unsigned integer value range from 0 to (1<<bitDepth)–1 (including the end values) for any required bit depth bitDepth for subsequent post - processing or display.
[0129] nnpfc_out_format_flag being equal to 1 indicates that the sample values output by the post - processing filter are unsigned integers within the range from 0 to (1<<(nnpfc_out_tensor_bitdepth_minus8 + 8)) - 1 (including the end values).
[0130] Values of nnpfc_out_format_idc greater than 1 are reserved by ITU - T|ISO / IEC for future specifications and shall not appear in a bitstream compliant with this version of this document. Decoders compliant with this version of this document shall ignore NNPFC SEI messages containing reserved values of nnpfc_out_format_idc.
[0131] nnpfc_out_tensor_bitdepth_minus8 plus 8 specifies the bit depth of the sample values in the output integer tensor. The value of nnpfc_out_tensor_bitdepth_minus8 shall be in the range from 0 to 24 (including the end values).
[0132] nnpfc_out_order_idc indicates the output order of the samples produced by the post - processing filter.
[0133] In bitstreams conforming to this version of this document, the value of nnpfc_out_order_idc shall be in the range of 0 to 3, inclusive. Values of nnpfc_out_order_idc from 4 to 255, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not appear in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_out_order_idc in the range of 4 to 255, inclusive. Values of nnpfc_out_order_idc greater than 255 shall not appear in bitstreams conforming to this version of this document and are not reserved for future use.
[0134] When nnpfc_purpose is equal to 2 or 4, nnpfc_out_order_idc should not be equal to 3.
[0135] Table 22 contains informative descriptions of the nnpfc_out_order_idc values.
[0136] Table 22 - Description of nnpfc_out_order_idc values
[0137]
[0138] The StoreOutputTensors() procedure, for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft that specify the top-left sample position of the sample patch included in the input tensor, is used to derive the sample values in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor, as follows:
[0139]
[0140]
[0141]
[0142] nnpfc_constant_patch_size_flag equal to 1 indicates that the post-processing filter accepts as input exactly the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. nnpfc_constant_patch_size_flag equal to 0 indicates that the post-processing filter accepts as input any patch size that is a positive integer multiple of the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1.
[0143] When nnpfc_constant_patch_size_flag is equal to 1, nnpfc_patch_width_minus1+1 indicates the horizontal sample count of the patch size required for the post-processing filter input. The value of nnpfc_patch_width_minus1 should be in the range of 0 to Min(32766, CroppedWidth-1), inclusive.
[0144] nnpfc_patch_height_minus1+1 indicates the vertical sample count of the patch size required for the post-processing filter input when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_height_minus1 should be in the range of 0 to Min(32766, CroppedHeight-1), inclusive.
[0145] Let the variables inpPatchWidth and inpPatchHeight be the width and height of the patch size respectively.
[0146] If nnpfc_constant_patch_size_flag is equal to 0, the following applies:
[0147] The values for –inpPatchWidth and inpPatchHeight are either provided by external means not specified in this document, or set by the post-processor itself.
[0148] –inpPatchWidth should be a positive integer multiple of nnpfc_patch_width_minus1+1 and should be less than or equal to CroppedWidth. inpPatchHeight should be a positive integer multiple of nnpfc_patch_height_minus1+1 and should be less than or equal to CroppedHeight.
[0149] Otherwise (nnpfc_constant_patch_size_flag is equal to 1), the value of inpPatchWidth is set equal to nnpfc_patch_width_minus1+1, and the value of inpPatchHeight is set equal to nnpfc_patch_height_minus1+1.
[0150] nnpfc_overlap indicates the horizontal and vertical sample counts of overlap of adjacent input tensors to the post-processing filter. The value of nnpfc_overlap should be in the range of 0 to 16383 (inclusive).
[0151] The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows:
[0152] outPatchWidth = ( nnpfc_pic_width_in_luma_samples * inpPatchWidth ) / CroppedWidth (84)
[0153] outPatchHeight = ( nnpfc_pic_height_in_luma_samples * inpPatchHeight ) / CroppedHeight (85)
[0154] horCScaling = SubWidthC / outSubWidthC (86)
[0155] verCScaling = SubHeightC / outSubHeightC (87)
[0156] outPatchCWidth = outPatchWidth * horCScaling (88)
[0157] outPatchCHeight = outPatchHeight * verCScaling (89)
[0158] overlapSize = nnpfc_overlap (90)
[0159] The bitstream conformance requirement is that outPatchWidth*CroppedWidth shall be equal to nnpfc_pic_width_in_luma_samples*inpPatchWidth, and outPatchHeight*CroppedHeight shall be equal to nnpfc_pic_height_in_luma_samples*inpPatchHeight.
[0160] nnpfc_padding_type indicates the padding process when referencing sample positions outside the boundaries of the cropped decoded output picture, as described in Table 23. The value of nnpfc_padding_type shall be in the range of 0 to 15 (inclusive).
[0161] Table 23 - Informative description of nnpfc_padding_type values
[0162]
[0163]
[0164] nnpfc_luma_padding_val indicates the luma value used for padding when nnpfc_padding_type is equal to 4.
[0165] nnpfc_cb_padding_val indicates the Cb value used for padding when nnpfc_padding_type is equal to 4.
[0166] nnpfc_cr_padding_val indicates the Cr value used for padding when nnpfc_padding_type is equal to 4.
[0167] The function InpSampleVal(y,x,picHeight,picWidth,croppedPic) takes as input the vertical sample position y, the horizontal sample position x, the picture height picHeight, the picture width picWidth, and the sample array croppedPick, and returns the sampleVal value derived as follows:
[0168] NOTE 6 – For the input to the function InpSampleVal(), the vertical position is listed before the horizontal position to be compatible with the input tensor convention of some inference engines.
[0169]
[0170] The following example process can be used to filter the cropped decoded output picture on a tile-by-tile basis using the post-processing filter PostProcessingFilter() to generate a filtered picture containing the Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated by nnpfc_out_order_idc.
[0171]
[0172] nnpfc_complexity_info_present_flag equal to 1 specifies that one or more syntax elements indicating the complexity of the post-processing filter associated with nnpfc_id are present. nnpfc_complexity_info_present_flag equal to 0 specifies that no syntax elements indicating the complexity of the post-processing filter associated with nnpfc_id are present.
[0173] nnpfc_parameter_type_idc equal to 0 indicates that the neural network uses only integer parameters. nnpfc_parameter_type_flag equal to 1 indicates that the neural network can use floating-point or integer parameters. nnpfc_parameter_type_idc equal to 2 indicates that the neural network uses only binary parameters. nnpfc_parameter_type_idc equal to 3 is reserved for future use by ITU-T|ISO / IEC and shall not appear in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_parameter_type_idc equal to 3.
[0174] nnpfc_log2_parameter_bit_length_minus3 equal to 0, 1, 2, and 3 indicates that the neural network does not use parameters with bit lengths greater than 8, 16, 32, and 64, respectively. When nnpfc_parameter_type_idc is present and nnpfc_log2_parameter_bit_length_minus3 is not present, the neural network does not use parameters with bit lengths greater than 1.
[0175] nnpfc_num_parameters_idc indicates the maximum number of neural network parameters for the post-processing filter, in units of powers of 2048. nnpfc_num_parameters_idc equal to 0 indicates that the maximum number of neural network parameters is unknown. The value nnpfc_num_parameters_idc shall be in the range of 0 to 52, inclusive. Values of nnpfc_num_parameters_idc greater than 52 are reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_num_parameters_idc greater than 52.
[0176] If the value of nnpfc_num_parameters_idc is greater than zero, the variable maxNumParameters is derived as follows:
[0177] maxNumParameters = (2048 << nnpfc_num_parameters_idc) - 1 (93)
[0178] The number of neural network parameters of the post-processing filters should be less than or equal to maxNumParameters, which is a requirement for bitstream conformance.
[0179] nnpfc_num_kmac_operations_idc greater than 0 indicates that the maximum number of multiply-accumulate operations per sample of the post-processing filter is less than or equal to nnpfc_num_kmac_operations_idc * 1000. nnpfc_num_kmac_operations_idc equal to 0 indicates that the maximum number of multiply-accumulate operations of the network is unknown. The value of nnpfc_num_kmac_operations_idc should be between 0 and 2. 32 The range is -1 (inclusive).
[0180] nnpfc_total_kilobyte_size greater than 0 indicates the total size in kilobytes required to store the uncompressed parameters of the neural network. The total size in bits is a number equal to or greater than the sum of the bits used to store each parameter. nnpfc_total_kilobyte_size is the total size in bits divided by 8000, rounded up. nnpfc_total_kilobyte_size equal to 0 indicates that the total size required to store the parameters of the neural network is unknown. The value of nnpfc_total_kilobyte_size should be between 0 and 2. 32 The range is -1 (inclusive).
[0181] In bitstreams conforming to this version of this document, nnpfc_reserved_zero_bit_b shall be equal to 0. A decoder shall ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_b is not equal to 0.
[0182] nnpfc_payload_byte[i] contains the i-th byte of a bitstream conforming to ISO / IEC 15938-17. For all values of i present, the byte sequence nnpfc_payload_byte[i] shall be a complete bitstream conforming to ISO / IEC 15938-17.
[0183] 8.29 Neural Network Post-Processing Filter Activation SEI Message
[0184] 8.29.1 Neural Network Post-Processing Filter Activation SEI Message Syntax
[0185]
[0186]
[0187] 8.29.2 Neural Network Post-Processing Filter Activation SEI Message Semantics
[0188] The Neural Network Post-Processing Filter Activation (NNPFA) SEI message activates or deactivates the possible use of the target neural network post-processing filter identified by nnpfa_target_id for post-processing filtering of a set of pictures.
[0189] NOTE 1 – Several NNPFA SEI messages may exist for the same picture, for example when post-processing filters are used for different purposes or filter different color components.
[0190] nnpfa_target_id indicates the target neural network post-processing filter, specified by one or more neural network post-processing filter characteristics SEI messages associated with the current picture and with nnpfc_id equal to nnfpa_target_id.
[0191] The value of nnpfa_target_id should be between 0 and 2 32 -2 (inclusive). From 256 to 511 (inclusive) and from 2 31 to 2 32 nnpfa_target_id values of -2 (inclusive) are reserved for future use by ITU-T|ISO / IEC. A decoder conforming to this version of this document shall not encounter an nnpfa_target_id between 256 and 511 (inclusive) or 2 31 to 2 32 When an NNPFA SEI message is received with a value in the range -2 (inclusive), the SEI message shall be ignored.
[0192] An NNPFA SEI message with a specific value of nnpfa_target_id shall not be present in the current PU unless one or both of the following conditions are met:
[0193] – In the current CLVS, an NNPFC SEI message with nnpfc_id equal to a specific value of nnpfa_target_id exists in a PU that precedes the current PU in decoding order.
[0194] – There is an NNPFC SEI message with nnpfc_id equal to a specific value of nnpfa_target_id in the current PU.
[0195] When a PU includes both an NNPFC SEI message with a specific value of nnpfc_id and an NNPFA SEI message with nnpfa_target_id equal to the specific value of nnpfc_id, the NNPFC SEI message shall precede the NNPFA SEI message in decoding order.
[0196] nnpfa_cancel_flag equal to 1 indicates that the persistence of the target neural network post-processing filter established by any previous NNPFA SEI message with the same nnpfa_target_id as the current SEI message is canceled, i.e., the target neural network post-processing filter is not used again unless it is activated by another NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 0. nnpfa_cancel_flag equal to 0 indicates that it is followed by nnpfa_persistence_flag.
[0197] nnpfa_persistence_flag specifies the persistence of the target neural network post-processing filters of the current layer.
[0198] nnpfa_persistence_flag equal to 0 specifies that the target neural network post-processing filter can only be used for post-processing filtering of the current image.
[0199] nnpfa_persistence_flag equal to 1 specifies that the target neural network post-processing filters can be used to post-process the current picture and all subsequent pictures of the current layer in output order until one or more of the following conditions are met:
[0200] – A new CLVS starts for the current layer.
[0201] – End of bitstream.
[0202] – Output the picture in the current layer that is associated with the NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1 immediately following the current picture in output order.
[0203] NOTE 2 – The target neural network post-processing filter is not applicable to this subsequent picture in the current layer that is associated with an NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1.
[0204] 4. Technical Problems Solved by the Disclosed Embodiments
[0205] The example design of the Neural Network Post-Processing Filter Characteristic (NNPFC) SEI message has the following issues:
[0206] First, neural network post-processing filter (NNPF) objectives are defined for chroma upsampling only, resolution upsampling only, both chroma upsampling and resolution upsampling, and picture rate upsampling only. However, picture rate upsampling combined with other types of upsampling can also be used for video applications. Therefore, there is a need for an NNPF objective that can utilize picture rate upsampling combined with other types of upsampling via signal transmission.
[0207] Secondly, when picture rate upsampling is used, a value is signaled to indicate the number of interpolated pictures between each pair of consecutive input images. However, in most cases, the number of interpolated pictures between each pair of consecutive input pictures is the same. Therefore, it is desirable to make signaling more efficient in common cases by signaling only a single value to indicate the number of interpolated pictures between each pair of consecutive input pictures.
[0208] Third, when performing multiple types of upsampling (such as chroma upsampling, resolution upsampling, and picture rate upsampling), the order may need to be explicitly specified or signaled.
[0209] Fourth, when the neural network post-processing filter (NNPF) purpose is specified as only visual quality improvement, only chroma upsampling, only resolution upsampling, or chroma upsampling and resolution upsampling, only one input picture is used. However, for these purposes, multiple input pictures may be beneficial.
[0210] 5. List of solutions and implementation examples
[0211] To address the above problems, the following methods are disclosed. These aspects should be considered as examples to explain the general concept and should not be interpreted narrowly. In addition, these examples can be applied alone or in combination in any way.
[0212] 1) To address issue 1, one or more of the following new objectives are defined:
[0213] a. In one example, new objectives are defined for picture rate upsampling and chroma upsampling, but not for resolution upsampling.
[0214] b. In one example, new objectives are defined for picture rate upsampling and resolution upsampling, but not for chroma upsampling.
[0215] c. In one example, new objectives are defined for picture rate upsampling, resolution upsampling, and chroma upsampling.
[0216] d. In one example, new objectives are defined for picture rate upsampling and visual quality improvement.
[0217] 2) To solve problem 2, an indication is transmitted by signaling to indicate whether the number of interpolated pictures between each pair of consecutive input pictures is the same.
[0218] a. Furthermore, in one example, when the number of interpolated pictures between each pair of consecutive input pictures is indicated to be the same, an indication of the number of interpolated pictures between each pair of consecutive input pictures is signaled.
[0219] i. Furthermore, in one example, the number of interpolated pictures between each pair of consecutive input pictures minus one is signaled.
[0220] 1. Furthermore, in one example, the value of the number minus one is constrained to be in the range of 0 to N (inclusive), where N is an integer.
[0221] a. In addition, in one example, N is specified to be 1, 3, 7, 15, 31, or 63.
[0222] 3) To solve problem 2, an indication is transmitted by signaling to indicate whether the fixed output frame rate is the same, where the same means that the number of interpolated pictures between each pair of consecutive input pictures is the same.
[0223] a. Furthermore, in one example, when indicating that the number of interpolated pictures between each pair of consecutive input pictures is the same, an indication of the output frame rate or the ratio between the output frame rate and the input frame rate is also signaled.
[0224] i. Additionally, in one example, the output frame rate minus the input frame rate is signaled.
[0225] 1. Additionally, alternatively, the output frame rate is subtracted from the input frame rate minus 1 via signal transmission.
[0226] 2. Additionally, in one example, the output frame rate minus the input frame rate may be signaled via the ue(v) or u(N) codec syntax elements.
[0227] ii. Additionally, in one example, a ratio between the output frame rate and the input frame rate is signaled.
[0228] 1. Furthermore, alternatively, the ratio between the output frame rate and the input frame rate is reduced by 1 by signal transmission.
[0229] 2. Furthermore, in one example, multiple syntax elements may be used to signal the ratio between the output frame rate and the input frame rate minus K (eg, K=0 or 1).
[0230] a. Alternatively, the ratio between the output frame rate and the input frame rate can be signaled using two ue(v) codec syntax elements, for example, using frr_a_minusK and frr_b_minusL, and the ratio is set equal to (frr_a_minusK+K)÷
[0231] (frr_a_minusK+K-(frr_b_minusL+L)), where K and L are integer values, for example, both are 1.
[0232] 4) To address issue 3, when multiple types of upsampling (such as chroma upsampling, resolution upsampling, and picture rate upsampling) are performed, the order of the multiple types of upsampling may be determined, specified, or transmitted through a signal.
[0233] a. In one example, the order may be predefined or fixed.
[0234] i. In one example, it is specified that when both picture rate upsampling and resolution upsampling are performed, picture rate upsampling should be performed after / before resolution upsampling.
[0235] ii. In one example, it is specified that when both chroma upsampling and resolution upsampling are performed, chroma upsampling should be performed before / after resolution upsampling.
[0236] iii. In one example, it is specified that when both chroma upsampling and picture rate upsampling are performed, picture rate upsampling should be performed after / before chroma upsampling.
[0237] iv. In one example, the above inventions can be combined in any manner.
[0238] b. In one example, the order may be transmitted by signaling. For example, at least one syntax element may be transmitted by signaling to indicate the order of different upsampling methods.
[0239] c. In one example, the order can be adaptive.
[0240] i. In one example, the order may depend on the codec mode / statistics of the video unit (eg, prediction mode, qp, temporal layer, slice type, etc.).
[0241] d. In one example, when executing multiple types of NNPF, the order may depend on the priority of the NNPF purpose.
[0242] i. In one example, priority can be transmitted through signaling.
[0243] 1. In one example, priority is signaled via syntax elements of the ue(v) or u(N) codec.
[0244] 2. In one example, the value of the priority is constrained to be in the range of 0 to N (inclusive), where N is an integer.
[0245] 3. In one example, a higher priority NNPF is executed before a lower priority NNPF.
[0246] 5) To solve problem 4, the number of input images can be specified for any neural network post-processing filter (NNPF) purpose.
[0247] a. In one example, for the purpose of visual quality improvement, the number of input pictures can be specified.
[0248] b. In one example, for the purpose of chroma upsampling, the number of input pictures can be specified.
[0249] c. In one example, for the purpose of resolution upsampling, the number of input pictures can be specified.
[0250] d. In one example, for the purpose of chroma upsampling and resolution upsampling, the number of input pictures can be specified.
[0251] e. In one example, by signaling the number of input pictures minus 1 (e.g., denoted as nnpfc_num_input_pics_minus1), the number of input pictures (e.g., denoted as numInputPics) is calculated as:
[0252] numInputPics=nnpfc_num_input_pics_minus1+1.
[0253] i. In one example, when picture rate upsampling is used, the value of nnpfc_num_input_pics_minus1 should be greater than 0.
[0254] f. In one example, the number of input pictures may be different when processing different frames. The number of input pictures may be signaled for each frame.
[0255] g. In one example, specific or adaptive rules may be used to determine which images are used as input for NNPF purposes.
[0256] i. In one example, it may depend on the codec mode / statistics of the video unit (eg, prediction mode, qp, temporal layer, slice type, etc.).
[0257] ii. In one example, it may be predefined or fixed.
[0258] 1. In one example, all reconstructed pictures in decoding order can be used as input.
[0259] 2. In one example, all reconstructed images in the order they were displayed can be used as input.
[0260] iii. In one example, it can be transmitted via a signal.
[0261] 1. In one example, a list of picture indices is signaled for each frame, which indicates the pictures to be used as input.
[0262] 6. Examples
[0263] The following are some example embodiments of the various aspects summarized in Section 5. Most of the relevant parts that have been added or modified are shown in bold, and some deleted parts are shown in italic bold. There may be some other changes of an editorial nature, so they are not highlighted.
[0264] 6.1 First embodiment
[0265] This embodiment applies to items 1 and 2 summarized in Section 5 and all their sub-items.
[0266] 8.28.1 Neural Network Post-Processing Filter Characteristics SEI Message Syntax
[0267]
[0268]
[0269] 8.28.2 Neural Network Post-Processing Filter Characteristics SEI Message Semantics ...
[0271] nnpfc_purpose indicates the purpose of the post-processing filter specified in Table 20.
[0272] In bitstreams conforming to this version of this document, the value of nnpfc_purpose shall be between 0 and The value of nnpfc_purpose is within the range of (including the end value). to 1023 (inclusive) are reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore nnpfc_purpose in arrive NNPFC SEI messages within the range (inclusive) of nnpfc_purpose. Values of nnpfc_purpose greater than 1023 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use.
[0273] Table 20 - Definition of nnpfc_purpose
[0274]
[0275] NOTE 3 – When ITU-T | ISO / IEC uses the reserved value of nnpfc_purpose in the future, the syntax of this SEI message may be extended with a syntax element with nnpfc_purpose equal to that value as a condition for the presence of that syntax element.
[0276] When SubWidthC equals 1 and SubHeightC equals 1, nnpfc_purpose should not equal 2 .
[0277] nnpfc_out_sub_c_flag equal to 1 specifies that outSubWidthC is equal to 1 and outSubHeightC1 is equal to 1. nnpfc_out_sub_c_flag equal to 0 specifies that outSubWidthC is equal to 2 and outSubHeightC is equal to 1. When nnpfc_out_sub_c_flag is not present, outSubWidthC is inferred to be equal to SubWidth, and outSubHeightC is inferred to be equal to SubHeightC. When ChromaFormatIdc is equal to 2 and nnpfc_out_sub_c_flag is present, the value of nnpfc_out_sub_c_flag shall be equal to 1.
[0278] nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples specify the width and height, respectively, of the array of luma samples of the picture resulting from applying the post-processing filter identified by nnpfc_id to the cropped decoded output picture. When nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples are not present, they are inferred to be equal to CroppedWidth and CroppedHeight, respectively. The value of nnpfc_pic_width_in_luma_samples shall be in the range of CroppedWidth to CroppedWidth*16-1, inclusive. The value of nnpfc_pic_height_in_luma_samples shall be in the range of CroppedHeight to CroppedHeight*16-1, inclusive.
[0279] nnpfc_num_input_pics_minus2 plus 2 specifies the number of decoded output pictures used as input to the post-processing filters.
[0280]
[0281] nnpfc_interpolated_pics[i] specifies The number of interpolated pictures generated by the post-processing filter between the i-th and (i+1)-th pictures used as input to the post-processing filter.
[0282] The variable numInputPics specifies the number of pictures used as input to the post-processing filter, and the variable numOutputPics specifies the total number of pictures produced by the post-processing filter. The variables numInputPics and numOutputPics are derived as follows:
[0283]
[0284] ...
[0286] nnpfc_out_order_idc indicates the output order of samples produced by the post-processing filter.
[0287] In bitstreams conforming to this version of this document, the value of nnpfc_out_order_idc shall be in the range of 0 to 3, inclusive. Values of nnpfc_out_order_idc from 4 to 255, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not appear in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_out_order_idc in the range of 4 to 255, inclusive. Values of nnpfc_out_order_idc greater than 255 shall not appear in bitstreams conforming to this version of this document and are not reserved for future use.
[0288] When nnpfc_purpose equals 2 nnpfc_out_order_idc should not be equal to 3.
[0289] Table 22 contains informative descriptions of the nnpfc_out_order_idc values.
[0290] Table 22 - Description of nnpfc_out_order_idc values
[0291]
[0292] The StoreOutputTensors() procedure, for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft that specify the top-left sample position of the sample patch included in the input tensor, is used to derive the sample values in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor, as follows:
[0293]
[0294]
[0295]
[0296] ...
[0298] 6.2 Example 2
[0299] This example applies to items 1 and 2 summarized in Section 5 and all their subitems, excluding item 1.d.
[0300] 8.28.1 Neural Network Post-Processing Filter Characteristics SEI Message Syntax
[0301]
[0302] 8.28.2 Neural Network Post-Processing Filter Characteristics SEI Message Semantics ...
[0304] nnpfc_purpose indicates the purpose of the post-processing filter specified in Table 20.
[0305] In bitstreams conforming to this version of this document, the value of nnpfc_purpose shall be in the range 0 to 8.5, inclusive. Values of nnpfc_purpose 9.6 to 1023, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not appear in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_purpose in the range 9.6 to 1023.1203, inclusive. Values of nnpfc_purpose greater than 1023 shall not appear in bitstreams conforming to this version of this document and are not reserved for future use.
[0306] Table 20 - Definition of nnpfc_purpose
[0307]
[0308] NOTE 3 – When ITU-T | ISO / IEC uses the reserved value of nnpfc_purpose in the future, the syntax of this SEI message may be extended with a syntax element with nnpfc_purpose equal to that value as a condition for the presence of that syntax element.
[0309] When SubWidthC equals 1 and SubHeightC equals 1, nnpfc_purpose should not equal 2 .
[0310] nnpfc_out_sub_c_flag equal to 1 specifies outSubWidthC equal to 1 and outSubHeightC equal to 1. nnpfc_out_sub_c_flag equal to 0 specifies outSubWidthC equal to 2 and outSubHeightC equal to 1. When nnpfc_out_sub_c_flag is not present, outSubWidthC is inferred to be equal to SubWidth, and outSubHeightC is inferred to be equal to SubHeightC. When ChromaFormatIdc is equal to 2 and nnpfc_out_sub_c_flag is present, the value of nnpfc_out_sub_c_flag shall be equal to 1.
[0311] nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples specify the width and height, respectively, of the array of luma samples of the picture resulting from applying the post-processing filter identified by nnpfc_id to the cropped decoded output picture. When nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples are not present, they are inferred to be equal to CroppedWidth and CroppedHeight, respectively. The value of nnpfc_pic_width_in_luma_samples shall be in the range of CroppedWidth to CroppedWidth*16-1, inclusive. The value of nnpfc_pic_height_in_luma_samples shall be in the range of CroppedHeight to CroppedHeight*16-1, inclusive.
[0312] nnpfc_num_input_pics_minus2 plus 2 specifies the number of decoded output pictures used as input to the post-processing filters.
[0313]
[0314] nnpfc_interpolated_pics[i] specifies The number of interpolated pictures generated by the post-processing filter between the i-th and (i+1)-th pictures used as input to the post-processing filter.
[0315] The variable numInputPics specifies the number of pictures used as input to the post-processing filter, and the variable numOutputPics specifies the total number of pictures produced by the post-processing filter. The variables numInputPics and numOutputPics are derived as follows:
[0316] ...
[0318] nnpfc_out_order_idc indicates the output order of samples produced by the post-processing filter.
[0319] In bitstreams conforming to this version of this document, the value of nnpfc_out_order_idc shall be in the range of 0 to 3, inclusive. Values of nnpfc_out_order_idc from 4 to 255, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not appear in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_out_order_idc in the range of 4 to 255, inclusive. Values of nnpfc_out_order_idc greater than 255 shall not appear in bitstreams conforming to this version of this document and are not reserved for future use.
[0320] When nnpfc_purpose equals 2 nnpfc_out_order_idc should not be equal to 3.
[0321] Table 22 contains informative descriptions of the nnpfc_out_order_idc values.
[0322] Table 22 - Description of nnpfc_out_order_idc values
[0323]
[0324] The StoreOutputTensors() procedure, for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft that specify the top-left sample position of the sample patch included in the input tensor, is used to derive the sample values in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor, as follows:
[0325]
[0326]
[0327] ...
[0329] 6.3 Example 3
[0330] This embodiment applies to items 1, 2, and 5 summarized in Section 5 and all their sub-items.
[0331] 8.28.1 Neural Network Post-Processing Filter Characteristics SEI Message Syntax
[0332]
[0333] 8.28.2 Neural Network Post-Processing Filter Characteristics SEI Message Semantics ...
[0335] nnpfc_purpose indicates the purpose of the post-processing filter specified in Table 20.
[0336] In bitstreams conforming to this version of this document, the value of nnpfc_purpose shall be in the range of 0 to 9.5, inclusive. Values of nnpfc_purpose from 10.6 to 1023, inclusive, are reserved for future use by ITU-T | ISO / IEC and shall not appear in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_purpose in the range of 10.6 to 1023.1203, inclusive. Values of nnpfc_purpose greater than 1023 shall not appear in bitstreams conforming to this version of this document and are not reserved for future use.
[0337] Table 20 - Definition of nnpfc_purpose
[0338]
[0339] NOTE 3 – When ITU-T | ISO / IEC uses the reserved value of nnpfc_purpose in the future, the syntax of this SEI message may be extended with a syntax element with nnpfc_purpose equal to that value as a condition for the presence of that syntax element.
[0340] When SubWidthC equals 1 and SubHeightC equals 1, nnpfc_purpose should not equal 2 .
[0341] nnpfc_out_sub_c_flag equal to 1 specifies that outSubWidthC is equal to 1 and outSubHeightC1 is equal to 1. nnpfc_out_sub_c_flag equal to 0 specifies that outSubWidthC is equal to 2 and outSubHeightC is equal to 1. When nnpfc_out_sub_c_flag is not present, outSubWidthC is inferred to be equal to SubWidth, and outSubHeightC is inferred to be equal to SubHeightC. When ChromaFormatIdc is equal to 2 and nnpfc_out_sub_c_flag is present, the value of nnpfc_out_sub_c_flag shall be equal to 1.
[0342] nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples specify the width and height, respectively, of the array of luma samples of the picture resulting from applying the post-processing filter identified by nnpfc_id to the cropped decoded output picture. When nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples are not present, they are inferred to be equal to CroppedWidth and CroppedHeight, respectively. The value of nnpfc_pic_width_in_luma_samples shall be in the range of CroppedWidth to CroppedWidth*16-1, inclusive. The value of nnpfc_pic_height_in_luma_samples shall be in the range of CroppedHeight to CroppedHeight*16-1, inclusive.
[0343] add Specifies the number of decoded output pictures used as input to the post-processing filters.
[0344] nnpfc_interpolated_pics[i] specifies The number of interpolated pictures generated by the post-processing filter between the i-th and (i+1)-th pictures used as input to the post-processing filter.
[0345] The variable numInputPics specifies the number of pictures used as input to the post-processing filter, and the variable numOutputPics specifies the total number of pictures produced by the post-processing filter. The variables numInputPics and numOutputPics are derived as follows:
[0346] ...
[0348] nnpfc_out_order_idc indicates the output order of samples produced by the post-processing filter.
[0349] In bitstreams conforming to this version of this document, the value of nnpfc_out_order_idc shall be in the range of 0 to 3, inclusive. Values of nnpfc_out_order_idc from 4 to 255, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not appear in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_out_order_idc in the range of 4 to 255, inclusive. Values of nnpfc_out_order_idc greater than 255 shall not appear in bitstreams conforming to this version of this document and are not reserved for future use.
[0350] When nnpfc_purpose equals 2 nnpfc_out_order_idc should not be equal to 3.
[0351] Table 22 contains informative descriptions of the nnpfc_out_order_idc values.
[0352] Table 22 - Description of nnpfc_out_order_idc values
[0353]
[0354] The StoreOutputTensors() procedure, for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft that specify the top-left sample position of the sample patch included in the input tensor, is used to derive the sample values in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor, as follows:
[0355]
[0356]
[0357] ...
[0359] 7. References
[0360] [1] ITU-T and ISO / IEC, “High-efficiency video coding”, Rec. ITU-T H.265 | ISO / IEC 23008-2 (effective version).
[0361] [2] J. Chen, E. Alshina, G.J. Sullivan, J.-R. Ohm, and J. Boyce, “Algorithmic Description of the Joint Exploration Test Model 7 (JEM7)”, JVET-G1001, August 2017.
[0362] [3] Rec. ITU-T H.266 | ISO / IEC 23090-3, “Versatile Video Codec”, 2022.
[0363] [4] Rec. ITU-T Rec. H.274 | ISO / IEC 23002-7, “Multifunctional Supplementary Enhancement Information Message for Coded Video Bitstreams”, 2022.
[0364] [5] ISO / IEC JTC 1 / SC 29 / WG 05 Output Document N0158, “ISO / IEC 23002-7:202x Text (Edition 2) DAM 1 Information technology — MPEG video techniques — Part 7: Versatile supplementary enhancement information messages for coded and decoded video bitstreams, Amendment 1: Additional SEI messages”, October 2022.
[0365] [6] S. McCarthy, T. Chujoh, M. Hannuksela, G. Sullivan, and Y.-K. Wang (eds.), “Additional SEI Messages for VSEI (Draft 3)”, JVET Output Paper JVET-AB 2006, publicly available online: https: / / www.jvet-experts.org / doc_end_user / current_document.php?id=12215.
[0366] Figure 2is a block diagram illustrating an example video processing system 4000 in which various embodiments disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or may be received in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a passive optical network (PON), and wireless interfaces such as Wi-Fi or a cellular interface.
[0367] System 4000 may include a codec component 4004 that can implement the various codecs or encoding methods described in this disclosure. Codec component 4004 can reduce the average bit rate of the video from input 4002 to the output of codec component 4004 to generate a codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. The output of codec component 4004 can be stored or sent via connected communication, as represented by component 4006. The bitstream (or codec) representation of the stored or transmitted video received at input 4002 can be used by component 4008 to generate pixel values or transmit to display interface 4010. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that the codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the codec results will be performed by the decoder.
[0368] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE) interface, etc. The embodiments described in the present disclosure may be embodied in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0369] Figure 3is a block diagram of an example video processing device 4100. Device 4100 can be used to implement one or more methods described herein. Device 4100 can be embodied in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. Device 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. Processor(s) 4102 can be configured to implement one or more methods described in the present disclosure. Memory(s) 4104 can be used to store data and code for implementing the methods and embodiments described herein. Video processing circuitry 4106 can be used to implement some embodiments described in the present disclosure in hardware circuitry. In some embodiments, video processing circuitry 4106 can be at least partially included in processor 4102, such as a graphics coprocessor.
[0370] Figure 4 4 is a flow chart of an example method 4200 for video processing. At step 4202, the method 4200 determines a neural network post-processing filter (NNPF) objective based on a neural network post-processing filter characteristic (NNPFC) supplemental enhancement information (SEI) message. At step 4204, conversion between visual media data and a bitstream is performed based on the NNPF objective.
[0371] It should be noted that method 4200 can be implemented in an apparatus for processing video data, the apparatus comprising a processor and non-transitory memory having instructions thereon, such as video encoder 4400, video decoder 4500, and / or encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform method 4200. Furthermore, method 4200 can be performed by a non-transitory computer-readable medium comprising a computer program product for use with a video codec apparatus. The computer program product comprises computer-executable instructions stored on a non-transitory computer-readable medium, such that when executed by the processor, the video codec apparatus performs method 4200.
[0372] Figure 5 4 is a block diagram illustrating an example video codec system 4300 in which embodiments of the present disclosure may be utilized. The video codec system 4300 may include a source device 4310 and a destination device 4320. The source device 4310 generates encoded video data and may be referred to as a video encoding device. The destination device 4320 may decode the encoded video data generated by the source device 4310 and may be referred to as a video decoding device.
[0373] Source device 4310 may include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. Video source 4312 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and associated data. A codec picture is a codec representation of a picture. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to target device 4320 via network 4330 via I / O interface 4316. The encoded video data may also be stored on storage medium / server 4340 for access by target device 4320.
[0374] Target device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. I / O interface 4326 may include a receiver and / or a modem. I / O interface 4326 may obtain encoded video data from source device 4310 or storage medium / server 4340. Video decoder 4324 may decode the encoded video data. Display device 4322 may display the decoded video data to a user. Display device 4322 may be integrated with target device 4320, or may be external to target device 4320 and configured to interface with an external display device.
[0375] The video encoder 4314 and the video decoder 4324 may operate according to a video compression standard, such as the HEVC standard, the VVC standard, and other current and / or additional standards.
[0376] Figure 6 is a block diagram illustrating an example of a video encoder 4400, which may be Figure 5 4. Video encoder 4314 in system 4300 shown in FIG. Video encoder 4400 can be configured to perform any or all embodiments of the present disclosure. Video encoder 4400 includes multiple functional components. The embodiments described in this disclosure can be shared between various components of video encoder 4400. In some examples, a processor can be configured to perform any or all embodiments described in this disclosure.
[0377] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402, a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a cache 4413 and an entropy coding unit 4414, and the prediction unit may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405 and an intra-frame prediction unit 4406.
[0378] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is a picture in which the current video block is located.
[0379] Furthermore, some components, such as the motion estimation unit 4404 and the motion compensation unit 4405 , may be highly integrated but are represented separately in the example of the video encoder 4400 for purposes of explanation.
[0380] The segmentation unit 4401 may segment a picture into one or more video blocks. The video encoder 4400 and the video decoder 4500 may support various video block sizes.
[0381] The mode selection unit 4403 can, for example, select one of the coding modes (intra or inter) based on the error result, and provide the resulting intra or inter coded block to the residual generation unit 4407 to generate residual block data and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 4403 can select a combination of intra and inter prediction (CIIP) modes, where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 4403 can also select a resolution of motion vectors for the block (e.g., sub-pixel or integer pixel precision).
[0382] To perform inter-frame prediction on the current video block, the motion estimation unit 4404 may generate motion information of the current video block by comparing the current video block with one or more reference frames from the buffer 4413. The motion compensation unit 4405 may determine a predicted video block for the current video block based on motion information and decoded samples of pictures from the buffer 4413 other than the picture associated with the current video block.
[0383] The motion estimation unit 4404 and the motion compensation unit 4405 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0384] In some examples, motion estimation unit 4404 may perform unidirectional prediction for the current video block, and motion estimation unit 4404 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 4404 may then generate a reference index indicating the reference picture in list 0 or list 1 containing the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0385] In other examples, the motion estimation unit 4404 may perform bidirectional prediction for the current video block. The motion estimation unit 4404 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in the reference pictures in list 1. The motion estimation unit 4404 may then generate reference indexes indicating the reference pictures containing the reference video blocks in list 0 and list 1, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 4404 may output the reference index and motion vector for the current video block as motion information for the current video block. The motion compensation unit 4405 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0386] In some examples, motion estimation unit 4404 may output a complete set of motion information for use in the decoding process of a decoder. In some examples, motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, motion estimation unit 4404 may reference motion information of another video block to signal the motion information of the current video block. For example, motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0387] In one example, the motion estimation unit 4404 may indicate, in a syntax structure associated with the current video block, a value that indicates to the video decoder 4500 that the current video block has the same motion information as another video block.
[0388] In another example, the motion estimation unit 4404 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0389] As discussed above, the video encoder 4400 can predictively signal motion vectors.Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and merge mode signaling.
[0390] Intra-frame prediction unit 4406 can perform intra-frame prediction on the current video block. When intra-frame prediction unit 4406 performs intra-frame prediction on the current video block, intra-frame prediction unit 4406 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0391] The residual generation unit 4407 may generate residual data for the current video block by subtracting the predicted video blocks of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0392] In other examples, there may be no residual data for the current video block, such as in skip mode, and the residual generation unit 4407 may not perform a subtraction operation.
[0393] Transform processing unit 4408 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0394] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0395] The inverse quantization unit 4410 and the inverse transform unit 4411 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 4412 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block for storage in the buffer 4413.
[0396] After the reconstruction unit 4412 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.
[0397] The entropy coding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives data, the entropy coding unit 4414 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0398] Figure 7 is a block diagram illustrating an example of a video decoder 4500, which may be Figure 5 4. The video decoder 4324 in the system 4300 shown in FIG. Video decoder 4500 can be configured to perform any or all embodiments of the present disclosure. In the example shown, video decoder 4500 includes multiple functional components. The embodiments described in this disclosure can be shared between various components of video decoder 4500. In some examples, the processor can be configured to perform any or all embodiments described in this disclosure.
[0399] In the example shown, video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, video decoder 4500 may perform a decoding process that is substantially the opposite of the encoding process described with respect to video encoder 4400.
[0400] The entropy decoding unit 4501 can retrieve a coded bitstream. The coded bitstream can include entropy-encoded video data (e.g., coded video data blocks). The entropy decoding unit 4501 can decode the entropy-encoded video data, and based on the entropy-decoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. For example, the motion compensation unit 4502 can determine such information by performing AMVP and merge mode.
[0401] The motion compensation unit 4502 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of an interpolation filter for use with sub-pixel precision may be included in a syntax element.
[0402] The motion compensation unit 4502 may calculate interpolated values of sub-integer pixels of a reference block using interpolation filters used by the video encoder 4400 during encoding of the video block. The motion compensation unit 4502 may determine the interpolation filters used by the video encoder 4400 based on received syntax information and use the interpolation filters to generate a prediction block.
[0403] The motion compensation unit 4502 can use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) slices of the coded video sequence, partitioning information describing how each macroblock of the pictures of the coded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame codec block, and other information used to decode the coded video sequence.
[0404] The intra prediction unit 4503 can form a prediction block from spatially neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 4504 inversely quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 4501. The inverse transform unit 4505 applies an inverse transform.
[0405] The reconstruction unit 4506 can add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 4507, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces decoded video for presentation on a display device.
[0406] Figure 8 is a schematic diagram of an example encoder 4600. The encoder 4600 is suitable for implementing techniques for VVC. The encoder 4600 includes three loop filters, namely a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike the DF 4602, which uses a predefined filter, the SAO 4604 and the ALF 4606 use the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter, respectively, taking advantage of the codec side information of the signaled offset and filter coefficients. The ALF 4606 is located at the last processing stage for each picture and can be seen as a tool that attempts to catch and repair artifacts created by previous stages.
[0407] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using reference pictures obtained from a reference picture cache 4612. The residual block from the inter-frame prediction or intra-frame prediction is fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed into an entropy codec component 4618. The entropy codec component 4618 performs entropy coding and decoding on the prediction results and quantized transform coefficients and sends them to a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed into an inverse quantization component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 can output images to the DF 4602 , SAO 4604 , and ALF 4606 for filtering before the images are stored in the reference picture cache 4612 .
[0408] The following provides a list of preferred solutions for some examples.
[0409] The following solutions show examples of the embodiments discussed herein.
[0410] 1. A method for processing media data, comprising: determining a neural network post-processing filter (NNPF) purpose based on a neural network post-processing filter characteristic (NNPFC) supplemental enhancement information (SEI) message; and performing conversion between visual media data and a bitstream based on the NNPF purpose.
[0411] 2. The method according to solution 1, wherein the NNPF purpose is both picture rate upsampling and chroma upsampling, rather than resolution upsampling.
[0412] 3. The method according to any one of solutions 1-2, wherein the NNPF purpose is both picture rate upsampling and resolution upsampling, but not chroma upsampling.
[0413] 4. A method according to any one of solutions 1-3, wherein the NNPF purpose is all picture rate upsampling, resolution upsampling and chroma upsampling.
[0414] 5. A method according to any one of solutions 1-4, wherein the bitstream includes an indication of whether the number of interpolated pictures between each pair of consecutive input pictures is the same.
[0415] 6. A method according to any one of solutions 1-5, wherein, when the indication indicates that the number of interpolated pictures between each pair of consecutive input pictures is the same, the bitstream also includes an indication of the number of interpolated pictures between each pair of consecutive input pictures.
[0416] 7. A method according to any one of solutions 1-6, wherein the bitstream includes the number of interpolated pictures between each pair of consecutive input pictures minus one.
[0417] 8. A method according to any of solutions 1-7, wherein the number of interpolated pictures between each pair of consecutive input pictures minus one is limited to the range of 0 to N, where N is an integer.
[0418] 9. A method according to any one of solutions 1-8, wherein N is specified to be 1, 3, 7, 15, 31 or 63.
[0419] 10. The method of any of solutions 1-9, wherein the bitstream includes an indication of whether the fixed output frame rate is constant.
[0420] 11. A method according to any of solutions 1-10, wherein when the number of interpolated pictures between each pair of consecutive input pictures is the same, an indication of the output frame rate is included in the bitstream.
[0421] 12. The method of any of solutions 1-11, wherein the output frame rate is signaled as the output frame rate minus the input frame rate.
[0422] 13. The method of any of solutions 1-12, wherein the output frame rate is signaled as a ratio between the output frame rate and the input frame rate.
[0423] 14. The method according to any of the solutions 1-13, wherein the conversion comprises multiple types of upsampling, and the order of the types of upsampling is determined.
[0424] 15. The method according to any of solutions 1-14, wherein the order of the types of upsampling is indicated in the bitstream.
[0425] 16. The method according to any of solutions 1-15, wherein when both picture rate upsampling and resolution upsampling are performed, picture rate upsampling is performed before resolution upsampling.
[0426] 17. The method according to any of solutions 1-16, wherein when both chroma upsampling and resolution upsampling are performed, chroma upsampling is performed before resolution upsampling.
[0427] 18. The method according to any of solutions 1-17, wherein when both chroma upsampling and picture rate upsampling are performed, picture rate upsampling is performed after chroma upsampling.
[0428] 19. The method according to any of solutions 1-18, wherein the order of the types of upsampling is adaptive based on the codec mode or statistics of the corresponding video unit.
[0429] 20. A method according to any of solutions 1-19, wherein when performing multiple types of NNPF, the order of upsampling the types depends on the priority of the NNPF purpose.
[0430] 21. The method according to any of the solutions 1-20, wherein the number of input pictures is specified for the purpose of visual quality improvement.
[0431] 22. The method according to any of the solutions 1-21, wherein for the purpose of chroma upsampling, the number of input pictures is specified.
[0432] 23. The method according to any of the solutions 1-22, wherein the number of input pictures is specified for the purpose of resolution upsampling.
[0433] 24. The method according to any of the solutions 1-23, wherein the number of input pictures is specified for the purpose of chroma upsampling and resolution upsampling.
[0434] 25. The method according to any of the solutions 1-24, wherein the number of input pictures minus 1 (nnpfc_num_input_pics_minus1) is signaled, and the number of input pictures (numInputPics) is calculated as: numInputPics=nnpfc_num_input_pics_minus1+1.
[0435] 26. The method according to any of the solutions 1-25, wherein the number of input pictures is signaled for each frame.
[0436] 27. A method according to any one of solutions 1-26, wherein the picture used as input for NNPF purposes is applied according to a rule, and the rule indicates that the picture used as input for NNPF purposes depends on the codec mode of the video unit, is predefined, or is transmitted by signal.
[0437] 28. A device for processing video data, comprising: a processor; and a non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of Solutions 1 to 27.
[0438] 29. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, so that when executed by a processor, the video codec device performs a method according to any one of Solutions 1 to 27.
[0439] 30. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method comprises: determining a neural network post-processing filter (NNPF) purpose based on a neural network post-processing filter characteristic (NNPFC) supplemental enhancement information (SEI) message; and generating a bitstream based on the determination.
[0440] 31. A method for storing a bitstream of a video, comprising: determining a neural network post-processing filter (NNPF) purpose based on a neural network post-processing filter characteristic (NNPFC) supplemental enhancement information (SEI) message; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0441] 32. A method, apparatus or system as described in the present disclosure.
[0442] In the solution described herein, an encoder can conform to the format rules by generating a codec representation according to the format rules. In the solution described herein, a decoder can use the format rules to parse syntax elements in the codec representation to produce decoded video, knowing whether the syntax elements are present according to the format rules.
[0443] In the present disclosure, the term "video processing" may refer to video encoding, video decoding, video compression or video decompression. For example, a video compression algorithm may be applied during conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of the current video block may, for example, correspond to bits that are co-located or distributed at different positions within the bitstream, as defined by the syntax. For example, a macroblock may be encoded based on error residual values from a transform and a codec, and bits in a header and other fields in the bitstream may also be used. Furthermore, during conversion, the decoder may parse the bitstream based on this determination, knowing whether some fields are present or not, as described in the above solution. Similarly, the encoder may determine whether to include certain syntax fields, and generate the codec representation accordingly by including or excluding the syntax fields in the codec representation.
[0444] The disclosed solutions, examples, embodiments, modules, and functional operations, as well as other solutions, examples, embodiments, modules, and functional operations described in this disclosure, can be implemented in digital electronic circuitry or computer software, firmware, or hardware, including the structures disclosed in this disclosure and their structural equivalents, or a combination of one or more thereof. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by a data processing apparatus or to control the operation of the data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that effects a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.
[0445] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file containing other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located in one location or distributed across multiple locations and interconnected by a communications network.
[0446] The processes and logic flows described in this disclosure can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC).
[0447] By way of example, processors suitable for executing computer programs include both general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from a read-only memory or a random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer may also include, or be operatively coupled to, one or more mass storage devices (e.g., magnetic, magneto-optical, or optical disks) for storing data, receiving data from them or transferring data to them, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and compact disk read-only memory (CD ROM) and digital versatile disk read-only memory (DVD-ROM) disks. The processor and memory may be supplemented by, or incorporated in, special-purpose logic circuitry.
[0448] While this disclosure contains many details, these should not be construed as limitations on the scope of any subject matter or what may be claimed, but rather as descriptions of features that may be unique to particular embodiments of the disclosure. Certain features described in this disclosure in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually in multiple implementations or in any suitable subcombination. Furthermore, while features may be described above as functioning in certain combinations and even initially claimed as such, in some cases one or more features in a claimed combination may be excised from the combination, and a claimed combination may be directed to a subcombination or variations of a subcombination.
[0449] Similarly, while operations are depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this disclosure should not be understood as requiring such separation in all embodiments.
[0450] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this disclosure.
[0451] A first component is directly coupled to a second component when there are no intervening components between the first and second components other than a wire, trace, or another medium. A first component is indirectly coupled to a second component when there are intervening components between the first and second components other than a wire, trace, or another medium. The term "coupled" and its variations encompass both direct and indirect couplings. Unless otherwise specified, the use of the term "about" is intended to include a range of ±10% of the following figure.
[0452] Although several embodiments are provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples should be considered illustrative rather than restrictive, and the present invention is not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.
[0453] In addition, without departing from the scope of the present disclosure, the techniques, systems, subsystems, and methods described and illustrated as discrete or separate in various embodiments may be combined or integrated with other systems, modules, techniques, or methods. Other items shown or discussed as coupled may be directly connected or indirectly coupled or communicated through some interface, device, or intermediate component, whether electrically, mechanically, or otherwise. Other examples of changes, substitutions, and modifications may be determined by those skilled in the art and may be made without departing from the spirit and scope of the present disclosure.
Claims
1. A method for processing media data, comprising: Determine the purpose of the neural network post-processing filter (NNPF) based on the neural network post-processing filter characteristic (NNPFC) supplemental enhancement information (SEI) message; as well as Performing conversion between visual media data and bitstream based on the NNPF object, Among them, the number of input pictures is specified for the NNPF purpose.
2. The method according to claim 1, wherein The number of input pictures is specified for purposes including visual quality improvement.
3. The method according to any one of claims 1 to 2, wherein The number of input pictures is specified for purposes including chroma upsampling.
4. The method according to any one of claims 1 to 3, wherein The number of input pictures is specified for purposes including resolution upsampling.
5. The method according to any one of claims 1 to 4, wherein The number of input pictures is specified for purposes including chroma upsampling and resolution upsampling.
6. The method according to any one of claims 1 to 5, wherein The number of input pictures minus 1 (nnpfc_num_input_pics_minus1) is signaled in the bitstream, and the number of input pictures (numInputPics) is calculated as follows: numInputPics=nnpfc_num_input_pics_minus1+1.
7. The method according to any one of claims 1 to 6, wherein When picture rate upsampling is used, the value of npfc_num_input_pics_minus1 is constrained to be greater than 0.
8. The method according to any one of claims 1 to 7, wherein The number of input pictures is signaled for each frame.
9. The method according to any one of claims 1 to 8, wherein The pictures used as input for NNPF purposes are applied according to a rule, and the rule indicates that the pictures used as input for NNPF purposes depend on the codec mode of the video unit, are predefined or transmitted through a signal.
10. The method according to any one of claims 1 to 9, wherein The pictures used as input for NNPF purposes are all reconstructed pictures in decoding order or all reconstructed pictures in display order.
11. The method according to any one of claims 1 to 10, wherein An index list of pictures is signaled for each frame, and the signaled index list of pictures indicates that the pictures are used as input for NNPF purposes.
12. The method according to any one of claims 1 to 11, wherein The conversion includes multiple types of upsampling, and the order of the upsampling types is determined.
13. The method according to any one of claims 1 to 12, wherein The order of upsampling types is indicated in the bitstream.
14. The method according to any one of claims 1 to 13, wherein When both picture rate upsampling and resolution upsampling are performed, the picture rate upsampling is performed before the resolution upsampling.
15. The method according to any one of claims 1 to 13, wherein When both picture rate upsampling and resolution upsampling are performed, the picture rate upsampling is performed after the resolution upsampling.
16. The method according to any one of claims 1 to 15, wherein When both chroma upsampling and resolution upsampling are performed, the chroma upsampling is performed before the resolution upsampling.
17. The method according to any one of claims 1 to 15, wherein When both chroma upsampling and resolution upsampling are performed, the chroma upsampling is performed after the resolution upsampling.
18. The method according to any one of claims 1 to 17, wherein When both chroma upsampling and picture rate upsampling are performed, the picture rate upsampling is performed after the chroma upsampling.
19. The method according to any one of claims 1 to 17, wherein: When both chroma upsampling and picture rate upsampling are performed, the picture rate upsampling is performed before the chroma upsampling.
20. The method according to any one of claims 1 to 19, wherein The order of upsampling types is adaptive based on the codec mode or statistics of the corresponding video unit.
21. The method according to any one of claims 1 to 20, wherein When performing multiple types of NNPF, the order of upsampling types depends on the priority of the NNPF purpose.
22. The method according to any one of claims 1 to 21, wherein The priorities of multiple NNPF destinations are signaled in the bitstream.
23. The method according to any one of claims 1 to 22, wherein The priority of the NNPF destination is signaled as an unsigned integer syntax element or an unsigned integer exponent Golomb codec syntax element.
24. The method according to any one of claims 1 to 23, wherein The value of the priority of the NNPF purpose is constrained to be in the range of 0 to N, inclusive, where N is an integer.
25. The method according to any one of claims 1 to 24, wherein NNPF objectives with higher priority are executed before NNDF objectives with lower priority.
26. A device for processing video data, comprising: processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 25.
27. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, the video codec device performs the method according to any one of claims 1 to 25.
28. A non-transitory computer-readable recording medium storing thereon a bit stream of a video generated by executing the method by a video processing apparatus, wherein: The method comprises: determining a neural network post-processing filter (NNPF) purpose based on a neural network post-processing filter characteristic (NNPFC) supplemental enhancement information (SEI) message; and generating a bitstream based on the determination, Among them, the number of input pictures is specified for the NNPF purpose.
29. A method for storing a video bitstream, comprising: Determine the purpose of the neural network post-processing filter (NNPF) based on the neural network post-processing filter characteristic (NNPFC) supplemental enhancement information (SEI) message; generating the bitstream based on the determination; as well as storing the bitstream in a non-transitory computer-readable recording medium, Among them, the number of input pictures is specified for the NNPF purpose.