System and method for signaling neural network post-filter overlap and sublayer frame rate upsampling information in video coding
Patent Information
- Application Number
- JP2023035692
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-11
- Filing Date
- 2023-03-08
- Publication Date
- 2026-02-19
AI Technical Summary
Existing video encoding standards lack efficient methods for signaling neural network post-filter parameter information, particularly for frame rate upsampling, which is crucial for enhancing video quality and adaptability.
The proposed techniques involve signaling neural network post-filter characteristic messages, including syntax elements that specify the number of interpolated pictures for each sublayer, enabling better frame rate upsampling and temporal scalability in video encoding.
Enhances video quality by allowing for more effective frame rate upsampling and temporal scalability, improving the adaptability and flexibility of video encoding systems.
Smart Images

Figure 00000066_0000 
Figure 00000066_0001 
Figure 00000067_0000
Abstract
Description
[Technical field]
[0001] (Related Applications) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 436,521, filed December 31, 2022, which is incorporated by reference in its entirety.
[0002] FIELD This disclosure relates to video encoding, and more particularly, to techniques for signaling neural network postfilter parameter information for encoded video. [Background technology]
[0003] Digital video capabilities may be incorporated into a wide range of devices, including digital televisions, laptop or desktop computers, tablet computers, digital recording devices, digital media players, video gaming devices, cellular phones, including so-called smart phones, medical imaging devices, and the like. Digital video may be encoded according to a video encoding standard. A video encoding standard defines a format for a compliant bitstream that encapsulates the encoded video data. A compliant bitstream is a data structure that can be received and decoded by a video decoding device to generate recovered video data. A video encoding standard may incorporate video compression techniques. Examples of video encoding standards include ISO / IEC MPEG-4 Visual and ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) and High-Efficiency Video Coding (HEVC). HEVC is described in High Efficiency Video Coding (HEVC), Rec. ITU-T H.265 (December 2016), which is incorporated herein by reference and is referred to herein as ITU-T H.265. Extensions and improvements to ITU-T H.265 are being considered for the development of next generation video coding standards. For example, the ITU-T Video Coding Experts Group (VCEG) and ISO / IEC (Moving Picture Experts Group (MPEG), collectively known as the Joint Video Exploration Team (JVET)) have standardized video coding techniques with compression capabilities that significantly exceed those of the current HEVC standard.The Joint Exploration Model 7 (JEM 7), Algorithm Description of Joint Exploration Test Model 7 (JEM 7), ISO / IEC JTC1 / SC29 / WG11 Document: JVET-G1001, July 2017, Torino, IT describes coding features that were under collaborative test model study by JVET as having the potential to improve video coding technology beyond the capabilities of ITU-T H.265, which is incorporated herein by reference. It should be noted that the coding features of JEM 7 are implemented in the JEM reference software. As used herein, the term JEM may collectively refer to the algorithms included in JEM 7 and the implementation of the JEM reference software. Additionally, in response to the "Joint Call for Proposals on Video Compression with Capabilities beyond HEVC" jointly issued by VCEG and MPEG, multiple descriptions of video coding tools have been published by various groups in the past 10 years. th The initial draft text of the video coding specification, derived from multiple descriptions of video coding tools, was proposed at the Meeting of ISO / IEC JTC1 / SC29 / WG11 16-20 April 2018, San Diego, CA. thThis development of a video coding standard by VCEG and MPEG is called the Versatile Video Coding (VVC) project. "Versatile Video Coding (Draft 10)", 20th Meeting of ISO / IEC JTC1 / SC29 / WG11 7-16 October 2020, Teleconference, document JVET-T2001-v2 represents the current iteration of the draft text of the video coding specification corresponding to the VVC project, which is incorporated herein by reference and referred to as JVET-T2001.
[0004] Video compression techniques allow for reducing data requirements for storing and transmitting video data. Video compression techniques may reduce data requirements by exploiting inherent redundancy in a video sequence. Video compression techniques may subdivide a video sequence into successively smaller portions (i.e., groups of pictures within a video sequence, pictures within groups of pictures, regions within pictures, sub-regions within regions, etc.). Intra-prediction coding techniques (e.g., spatial prediction techniques within a picture) and inter-prediction techniques (i.e., inter-picture techniques (temporal)) may be used to generate difference values between a unit of video data being coded and a reference unit of video data. The difference values may be referred to as residual data. The residual data may be coded as quantized transform coefficients. Syntax elements may be associated with the residual data and the reference coding units (e.g., intra-prediction mode index, and motion information). The residual data and syntax elements may be entropy coded. The entropy coded residual data and syntax elements may be included in a data structure that forms a compliant bitstream. Summary of the Invention
[0005] In general, this disclosure describes various techniques for encoding video data. In particular, this disclosure describes techniques for signaling neural network postfilter parameter information for encoded video data. It should be noted that although the techniques of this disclosure are described with respect to ITU-T H.264, ITU-T H.265, JEM, and JVET-T2001, the techniques of this disclosure are generally applicable to video encoding. For example, the encoding techniques described herein may be incorporated into video encoding systems (including video encoding systems based on future video encoding standards) that include video block structures, intra-prediction techniques, inter-prediction techniques, transform techniques, filtering techniques, and / or entropy encoding techniques other than those included in ITU-T H.265, JEM, and JVET-T2001. Thus, references to ITU-T H.264, ITU-T H.265, JEM, and / or JVET-T2001 are for illustrative purposes and should not be construed to limit the scope of the techniques described herein. Furthermore, it should be noted that the incorporation by reference of documents herein is for explanatory purposes and should not be construed as limiting or creating ambiguity with respect to the terms used herein. For example, if an incorporated reference provides a definition of a term that differs from that of another incorporated reference and / or as that term is used herein, that term should be construed to broadly include each corresponding definition and / or to include each specific definition instead.
[0006] In one embodiment, a method for encoding video data includes signaling a neural network post-filter characteristic message and signaling, in the neural network post-filter characteristic message, a syntax element that specifies a number of interpolated pictures generated by a post-processing filter for each of a number of sub-layers associated with the neural network post-filter characteristic message.
[0007] In one embodiment, a device comprises one or more processors configured to signal a neural network post-filter characteristic message and signal, in the neural network post-filter characteristic message, a syntax element that specifies a number of interpolated pictures generated by a post-processing filter for each of a number of sub-layers associated with the neural network post-filter characteristic message.
[0008] In one embodiment, a non-transitory computer-readable storage medium includes instructions stored thereon that, when executed, cause one or more processors of the device to signal a neural network post-filter characteristic message and signal in the neural network post-filter characteristic message a syntax element that specifies a number of interpolated pictures generated by a post-processing filter for each of a number of sub-layers associated with the neural network post-filter characteristic message.
[0009] In one embodiment, an apparatus comprises means for signaling a neural network post-filter characteristic message; and means for signaling, in the neural network post-filter characteristic message, a syntax element that specifies a number of interpolated pictures generated by a post-processing filter for each of a number of sub-layers associated with the neural network post-filter characteristic message.
[0010] In one embodiment, a method for decoding video data includes receiving a neural network post-filter characteristic message and parsing from the neural network post-filter characteristic message a syntax element that specifies a number of interpolated pictures generated by a post-processing filter for each of a number of sub-layers associated with the neural network post-filter characteristic message.
[0011] In one embodiment, a device comprises one or more processors configured to receive a neural network post-filter characteristic message and parse from the neural network post-filter characteristic message a syntax element that specifies a number of interpolated pictures generated by a post-processing filter for each of a number of sub-layers associated with the neural network post-filter characteristic message.
[0012] In one embodiment, a non-transitory computer-readable storage medium includes instructions stored thereon that, when executed, cause one or more processors of the device to receive a neural network post-filter characteristic message and parse from the neural network post-filter characteristic message a syntax element that specifies a number of interpolated pictures generated by a post-processing filter for each of a number of sub-layers associated with the neural network post-filter characteristic message.
[0013] In one embodiment, an apparatus comprises means for receiving a neural network post-filter characteristic message and means for parsing from the neural network post-filter characteristic message a syntax element that specifies a number of interpolated pictures generated by a post-processing filter for each of a number of sub-layers associated with the neural network post-filter characteristic message.
[0014] The details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims. [Brief description of the drawings]
[0015] [Figure 1] FIG. 1 is a block diagram illustrating an example of a system that may be configured to encode and decode video data in accordance with one or more techniques of this disclosure. [Diagram 2] 1 is a conceptual diagram illustrating encoded video data and corresponding data structures in accordance with one or more techniques of this disclosure. [Diagram 3] 1 is a conceptual diagram illustrating a data structure that encapsulates encoded video data and corresponding metadata in accordance with one or more techniques of this disclosure. [Figure 4] A conceptual diagram illustrating an example of components that may be included in an implementation of a system that may be configured to encode and decode video data in accordance with one or more techniques of this disclosure. [Diagram 5] FIG. 1 is a block diagram illustrating an example of a video encoding device that may be configured to encode video data in accordance with one or more techniques of this disclosure. [Figure 6] 1 is a block diagram illustrating an example of a video decoding device that may be configured to decode video data in accordance with one or more techniques of this disclosure. [Figure 7] FIG. 2 is a conceptual diagram illustrating an example of a packed data channel for a luma component, in accordance with one or more techniques of this disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] Video content includes a video sequence consisting of a series of frames (or pictures). A series of frames may also be referred to as a group of pictures (GOP). Each video frame or picture may be divided into one or more regions. A region may be defined according to a base unit (e.g., a video block) and a set of rules that define the region. For example, the rule that defines a region may be that the region must be an integer number of video blocks arranged in a rectangle. Furthermore, the video blocks in a region may be ordered according to a scan pattern (e.g., a raster scan). As used herein, the term video block may generally refer to an area of a picture, or more specifically, may refer to a maximal array of sample values that may be predictively coded, its subdivisions, and / or corresponding structures. Furthermore, the term current video block may refer to a portion of a picture being coded or decoded. A video block may be defined as an array of sample values. It should be noted that in some cases, a pixel value may be described as including sample values for each component of the video data, which may also be referred to as color components (e.g., luma component (Y) and chroma components (Cb and Cr), or red, green, and blue color components). It should be noted that in some cases, the terms pixel value and sample value are used interchangeably. Furthermore, in some cases, a pixel or sample may be referred to as a pel. A video sampling format, which may be referred to as a chroma format, may be defined as the number of chroma samples included in a video block relative to the number of luma samples included in the video block. For example, for a 4:2:0 format, the sampling rate for the luma component is twice the sampling rate of the chroma components in both the horizontal and vertical directions.
[0017] A video coding device may perform predictive coding on video blocks and their subdivisions. The video blocks and their subdivisions may be referred to as nodes. ITU-T H.264 specifies macroblocks containing 16x16 luma samples. That is, in ITU-T H.264, a picture is divided into macroblocks. ITU-T H.265 specifies a similar coding tree unit (CTU) structure, sometimes called a largest coding unit (LCU). In ITU-T H.265, a picture is divided into CTUs. In ITU-T H.265, for a picture, the CTU size may be set to contain 16x16, 32x32, or 64x64 luma samples. In ITU-T H.265, a CTU is composed of a respective coding tree block (CTB) for each component of video data (e.g., luma (Y) and chroma (Cb and Cr). It should be noted that a video having one luma component and two corresponding chroma components may be described as having two channels, i.e., a luma channel and a chroma channel. Furthermore, in ITU-T H.265, a CTU may be partitioned according to a quad-tree (QT) partitioning structure, such that the CTB of a CTU is partitioned into coding blocks (CBs). That is, in ITU-T H.265, a CTU may be partitioned into quad-tree leaf nodes. According to ITU-T H.265, one luma CB, together with two corresponding chroma CBs and related syntax elements, is called a coding unit (CU). In ITU-T H.265, a minimum allowed size of a CB may be signaled. In H.265, the smallest allowed size of luma CB is 8x8 luma samples. In ITU-T H.265, the decision to code a picture portion using intra or inter prediction is made at the CU level.
[0018] In ITU-T H.265, a CU is associated with a prediction unit structure whose root is the CU. In ITU-T H.265, the prediction unit structure allows for splitting the luma CB and chroma CB for the purpose of generating corresponding reference samples. That is, in ITU-T H.265, the luma CB and chroma CB may be split into respective luma and chroma prediction blocks (PBs), where a PB includes a block of sample values to which the same prediction is applied. In ITU-T H.265, the CB may be split into 1, 2, or 4 PBs. ITU-T H.265 supports PB sizes from 64x64 samples to 4x4 samples. In ITU-T H.265, square PBs are supported for intra prediction, where the CB may form a PB or the CB may be split into 4 square PBs. In addition to square PB, ITU-T H.265 supports rectangular PB for inter prediction, where CB may be bisected vertically or horizontally to form PB. In addition, ITU-T H.265 supports four asymmetric PB division for inter prediction, where CB is divided into two PBs by one-quarter of CB height (at top or bottom) or width (at left or right). Using intra prediction data (e.g., intra prediction mode syntax element) or inter prediction data (e.g., motion data syntax element) corresponding to PB, reference sample value and / or predicted sample value for PB are generated.
[0019] JEM specifies a CTU with a maximum size of 256x256 luma samples. JEM specifies a quad tree plus binary tree (QTBT) block structure. In JEM, the QTBT structure allows a quad tree leaf node to be further split by a binary tree structure (BT). That is, in JEM, the binary tree structure allows a quad tree leaf node to be split recursively vertically or horizontally. In JVET-T2001, the CTU is split according to a quad tree plus multitype tree (QTMT or QT+MTT) structure. QTMT in JVET-T2001 is similar to QTBT in JEM. However, in JVET-T2001, in addition to exhibiting a binary split, the multitype tree can exhibit a so-called ternary (or triple tree, TT) split. The ternary split splits a block into three blocks vertically or horizontally. In the case of vertical TT division, the block is divided at a quarter of its width from the left edge and at a quarter of its width from the right edge, and in the case of horizontal TT division, the block is divided at a quarter of its height from the top edge and at a quarter of its height from the bottom edge.
[0020] As mentioned above, each video frame or picture may be divided into one or more regions. For example, according to ITU-T H.265, each video frame or picture may be divided to include one or more slices, and further divided to include one or more tiles, where each slice includes a sequence of CTUs (e.g., in raster scan order), and a tile is a sequence of CTUs corresponding to a rectangular portion of a picture. It should be noted that a slice, in ITU-T H.265, is a sequence of one or more slice segments starting with an independent slice segment and including all subsequent dependent slice segments (if any) preceding the next independent slice segment (if any). A slice segment, like a slice, is a sequence of CTUs. Thus, in some cases, the terms slice and slice segment may be used interchangeably to indicate a sequence of CTUs arranged in raster scan order. It should be further noted that in ITU-T H.265, a tile may consist of CTUs included in two or more slices, and a slice may consist of CTUs included in two or more tiles. However, ITU-T H.265 specifies that one or both of the following conditions must be met: (1) all CTUs in a slice belong to the same tile, and (2) all CTUs in a tile belong to the same slice.
[0021] For JVET-T2001, a slice is instead only required to consist of an integer number of CTUs, but is instead required to consist of an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile. Note that in JVET-T2001, the slice design does not include slice segments (i.e., there are no independent / dependent slice segments). Thus, in JVET-T2001, a picture may contain a single tile, where the single tile is contained within a single slice, or a picture may contain multiple tiles, where multiple tiles (or their CTU rows) may be contained within one or more slices. In JVET-T2001, the division of a picture into tiles is specified by specifying the height of each of the tile rows and the width of each of the tile columns. Thus, in JVET-T2001, a tile is a rectangular region of CTUs within a particular tile row and a particular tile column location. Further, it should be noted that JVET-T2001 specifies the case where a picture may be divided into sub-pictures, where a sub-picture is a rectangular region of a CTU within a picture. The top-left CTU of a sub-picture may be located at any CTU position within a picture, with the sub-picture being constrained to contain one or more slices. Thus, unlike tiles, sub-pictures are not necessarily limited to a specific row and column position. It should be noted that a sub-picture may be useful for encapsulating an area of interest within a picture, and the sub-bitstream extraction process may be used only to decode and display the specific area of interest. That is, as described in more detail below, a bitstream of coded video data includes a sequence of network abstraction layer (NAL) units, where the NAL units encapsulate the coded video data (i.e., video data corresponding to a slice of a picture) or the NAL units encapsulate metadata (e.g., parameter sets) used to decode the video data, and the sub-bitstream extraction process forms a new bitstream by removing one or more NAL units from the bitstream.
[0022] FIG. 2 is a conceptual diagram illustrating an example of a picture in a picture group divided according to tiles, slices, and subpictures. It should be noted that the techniques described herein may be applicable to tiles, slices, subpictures, their subdivisions, and / or their equivalent structures. That is, the techniques described herein may be generally applicable regardless of how a picture is divided into regions. For example, in some cases, the techniques described herein may be applicable when tiles may be divided into so-called bricks, where a brick is a rectangular region of a CTU row in a particular tile. Further, for example, in some cases, the techniques described herein may be applicable when one or more tiles may be included in a so-called tile group, where a tile group includes an integer number of adjacent tiles. In the example shown in FIG. 2, Pic 3 is a set of 16 tiles (i.e., Tile 0 From Tile 15 ) and three slices (i.e. Slice 0 From Slice 2 In the example shown in Figure 2, Slice 0 The four tiles (i.e., Tile 0 From Tile 3 ), Slice 1 There are eight tiles (i.e., Tile 4 From Tile 11 ), Slice 2 The four tiles (i.e., Tile 12 From Tile 15 ) Furthermore, as shown in the example of Figure 2, 3 The image contains two subpictures (i.e., 0 and Subpicture 1 ), where Subpicture 0 Slice 0 and Slice 1 Includes Subpicture 1 Slice 2As mentioned above, subpictures can be useful for encapsulating regions of interest within a picture, and the sub-bitstream extraction process can be used to selectively decode (and display) the regions of interest. For example, referring to FIG. 2, a Subpicture 0 may correspond to the action part of a sporting event presentation (e.g., a view of the field), and Subpicture 1 Slice may correspond to a scrolling banner displayed during a sporting event presentation. Organizing pictures into subpictures in this manner may allow the viewer to disable the display of the scrolling banner. That is, through the sub-bitstream extraction process, 2 NAL units may be removed from the bitstream (and thus may not be decoded and / or displayed), 0 NAL Units and Slices 1 The NAL units may be decoded and displayed. The encapsulation of slices of a picture into respective NAL unit data structures and sub-bitstream extraction are described in further detail below.
[0023] For intra-prediction coding, the intra-prediction mode may specify the location of the reference sample in the picture. In ITU-T H.265, the defined possible intra-prediction modes include planar (i.e., surface fitting) prediction mode, DC (i.e., monotonic global averaging) prediction mode, and 33 angular prediction modes (predMode: 2-34). In JEM, the defined possible intra-prediction modes include planar, DC, and 65 angular prediction modes. It should be noted that the planar and DC prediction modes may be referred to as non-directional prediction modes, and the angular prediction modes may be referred to as directional prediction modes. It should be noted that the techniques described herein may be generally applicable regardless of the number of defined possible prediction modes.
[0024] For inter-predictive coding, a reference picture is determined and a motion vector (MV) identifies samples in the reference picture used to generate a prediction for the current video block. For example, the current video block may be predicted using reference sample values located in one or more previously coded pictures, and the motion vector is used to indicate the position of the reference block relative to the current video block. The motion vector may be, for example, a horizontal displacement component of the motion vector (i.e., the MV x ), the vertical displacement component of the motion vector (i.e., MV y), and the resolution of the motion vectors (e.g., ¼ pixel precision, ½ pixel precision, 1 pixel precision, 2 pixel precision, 4 pixel precision). Previously decoded pictures, which may include pictures output before or after the current picture, may be organized into one or more reference picture lists and identified using reference picture index values. Furthermore, in inter-predictive coding, uni-prediction refers to generating a prediction using sample values from a single reference picture, and bi-prediction refers to generating a prediction using respective sample values from two reference pictures. That is, in uni-prediction, a single reference picture and corresponding motion vector are used to generate a prediction for the current video block, and in bi-prediction, a first reference picture and corresponding first motion vector, and a second reference picture and corresponding second motion vector are used to generate a prediction for the current video block. In bi-prediction, the respective sample values are combined (e.g., added, rounded, clipped, or averaged according to a weight) to generate a prediction. A picture and its regions may be classified based on what kind of prediction modes may be used to code its video blocks. That is, for a region having a B type (e.g., B slice), bi-prediction mode, uni-prediction mode, and intra-prediction mode may be used, for a region having a P type (e.g., P slice), uni-prediction mode and intra-prediction mode may be used, and for a region having an I type (e.g., I slice), only intra-prediction mode may be used. As mentioned above, the reference picture is identified by a reference index. For example, in a P slice, there may be a single reference picture list RefPicList0, and in a B slice, in addition to RefPicList0, there may be a second independent reference picture list RefPicList1. It should be noted that uni-prediction in a B slice may use one of RefPicList0 or RefPicList1 to generate a prediction. It should also be noted that during the decoding process, at the start of decoding a picture, a reference picture list is generated from a previously decoded picture stored in a decoded picture buffer (DPB).
[0025] Furthermore, the coding standard may support various modes of motion vector prediction. Motion vector prediction allows a value of a motion vector for a current video block to be derived based on another motion vector. For example, a set of candidate blocks with associated motion information may be derived from spatial and temporal neighboring blocks to the current video block. Furthermore, generated (or default) motion information may be used for motion vector prediction. Examples of motion vector prediction include advanced motion vector prediction (AMVP), temporal motion vector prediction (TMVP), the so-called "combined" mode, as well as "skip" and "direct" motion estimation. Furthermore, other examples of motion vector prediction include advanced temporal motion vector prediction (ATMVP) and spatial-temporal motion vector prediction (STMVP). In motion vector prediction, both the video encoding device and the video decoding device perform the same process to derive a set of candidates. Thus, for the current video block, the same set of candidates is generated during encoding and decoding.
[0026] As mentioned above, in inter-prediction coding, reference samples in a previously coded picture are used to code a video block in a current picture. A previously coded picture that is available for use as a reference when coding a current picture is called a reference picture. It should be noted that the decoding order does not necessarily correspond to the picture output order, i.e., the temporal order of pictures in a video sequence. In ITU-T H.265, when a picture is decoded, it is stored in a decoded picture buffer (DPB) (which may be called a frame buffer, a reference buffer, a reference picture buffer, etc.). In ITU-T H.265, pictures stored in the DPB are removed from the DPB when they are output and are no longer needed for coding a subsequent picture. In ITU-T H.265, the decision of whether a picture should be removed from the DPB is performed once per picture after decoding the slice header, i.e., at the beginning of the decoding of the picture. For example, referring to FIG. 2, the picture 2 is Pic 1 Similarly, Pic 3 is Pic 0 , and assuming that the picture numbers correspond to the decoding order, the DPB is populated as follows: 0 After decrypting, the DPB is 0}, Pic 1 At the beginning of decoding, the DPB is 0}, Pic 1 After decrypting, the DPB is 0 ,Pic 1}, Pic 2 At the beginning of decoding, the DPB is 0 ,Pic 1}. Next, Pic 2 is Pic 1 It is decrypted by referring to Pic 2 After decrypting, the DPB is 0 ,Pic 1 ,Pic 2}. Pic 3At the beginning of decoding, the picture Pic 0 and Pic 1 is Pic 3 (or any subsequent pictures not shown) and are therefore marked for removal from the DPB. 1 and Pic 2 Assuming that the output is, the DPB is {Pic 0}. Then, Pic 3 is Pic 0 The process of marking pictures for removal from the DPB is sometimes called Reference Picture Set (RPS) management.
[0027] As mentioned above, the intra prediction data or the inter prediction data is used to generate reference sample values for a block of sample values. The difference between the sample values included in the current PB or another type of picture substructure and the associated reference sample (e.g., the reference sample generated using the prediction) may be referred to as residual data. The residual data may include a respective array of difference values corresponding to each component of the video data. The residual data may be in the pixel domain. A transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), an integer transform, a wavelet transform, or a conceptually similar transform, may be applied to the array of difference values to generate transform coefficients. It is noted that in ITU-T H.265 and JVET-T2001, a CU is associated with a transform tree structure with its root at the CU level. The transform tree is divided into one or more transform units (TUs). That is, the array of difference values may be partitioned (e.g., four 8×8 transforms may be applied to a 16×16 array of residual values) for purposes of generating transform coefficients. For each component of video data, such subdivision of the difference values may be referred to as a Transform Block (TB). Note that in some cases, a core transform and subsequent secondary transforms may be applied at the video encoder to generate the transform coefficients. For the video decoder, the order of the transforms is reversed.
[0028] The quantization process may be performed directly on the transform coefficients or on the residual sample values (e.g., in the case of palette coding quantization). Quantization approximates the transform coefficients with amplitudes limited to a particular set of values. Quantization essentially scales the transform coefficients to change the amount of data required to represent a group of transform coefficients. Quantization may include division of the transform coefficients (or values resulting from adding an offset value to the transform coefficients) by a quantization scale factor and any associated rounding function (e.g., rounding to the nearest integer). Quantized transform coefficients are sometimes referred to as coefficient level values. Inverse quantization (or "dequantization") may include multiplication of coefficient level values by a quantization scale factor and any mutual rounding or offset addition operations. It should be noted that as used herein, the term quantization process may refer in some cases to division by a quantization scale factor to generate level values and in some cases to multiplication by a quantization scale factor to recover transform coefficients. That is, the quantization process may refer in some cases to quantization and in some cases to inverse quantization. Furthermore, while in some of the following examples the quantization process is described with respect to arithmetic operations associated with decimal number representation, it should be noted that such description is for illustrative purposes and should not be construed as limiting. For example, the techniques described herein may be implemented in devices using binary arithmetic, etc. For example, the multiplication and division operations described herein may be implemented using bit shifting operations, etc.
[0029] The quantized transform coefficients and syntax elements (e.g., syntax elements indicating a coding structure of a video block) may be entropy coded according to an entropy coding technique. The entropy coding process includes coding the values of the syntax elements using a lossless data compression algorithm. Examples of entropy coding techniques include content adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), probability interval partitioning entropy coding (PIPE), etc. The entropy coded quantized transform coefficients and the corresponding entropy coded syntax elements may form a compliant bitstream that can be used to reproduce the video data at a video decoding device. The entropy coding process, e.g., CABAC, may include performing binarization on the syntax elements. Binarization refers to the process of converting the value of a syntax element into a series of one or more bits. These bits are sometimes called "bins." The binarization may include one or a combination of the following encoding techniques: fixed-length encoding, unary encoding, shortened unary encoding, shortened Rice encoding, Golomb encoding, k-th exponential Golomb encoding, and Golomb-Rice encoding. For example, the binarization may include representing an integer value of 5 of a syntax element as 00000101 using an 8-bit fixed-length binarization technique, or representing an integer value of 5 as 11110 using a unary encoding binarization technique. As used herein, each of the terms fixed-length encoding, unary encoding, shortened unary encoding, shortened Rice encoding, Golomb encoding, k-th exponential Golomb encoding, and Golomb-Rice encoding may refer to a general implementation of these techniques and / or a more specific implementation of these encoding techniques. For example, an implementation of Golomb-Rice encoding may be specifically defined according to a video encoding standard.In a CABAC example, for a particular bin, the context provides a most probable state (MPS) value for the bin (i.e., the MPS for the bin is one of 0 or 1) and a probability value that the bin is the MPS or least probably state (LPS). For example, the context may indicate that the MPS of the bin is 0 and the probability that the bin is 1 is 0.3. Note that the context may be determined based on values of previously coded bins, including bins in the current syntax element and previously coded syntax elements. For example, values of syntax elements associated with neighboring video blocks may be used to determine the context of the current bin.
[0030] As mentioned above, the sample values of the reconstructed block may differ from the sample values of the encoded current video block. Furthermore, it should be noted that in some cases, encoding video data block by block may result in artifacts (e.g., so-called blocking artifacts, banding artifacts, etc.). For example, blocking artifacts may cause encoded block boundaries of the reconstructed video data to be visually perceptible to a user. In this manner, the reconstructed sample values may be modified to minimize the difference between the encoded sample values of the current video block and the reconstructed block and / or to minimize artifacts introduced by the video encoding process. Such modification may be generally referred to as filtering. It should be noted that filtering may occur as part of an in-loop filtering process or a post-loop (or post-filtering) filtering process. In an in-loop filtering process, the sample values resulting from the filtering process may be used for the predicted video block (e.g., stored in a reference frame buffer for subsequent encoding in a video encoding device and subsequent decoding in a video decoding device). In a post-loop filtering process, the sample values resulting from the filtering process are simply output as part of the decoding process (e.g., not used for subsequent encoding). For example, in a video decoding device, in an in-loop filtering process, sample values resulting from filtering the reconstructed block are used for subsequent decoding (e.g., stored in a reference buffer) and output (e.g., to a display), whereas in a post-loop filtering process, the reconstructed block is used for subsequent decoding, and sample values resulting from filtering the reconstructed block are output and not used for subsequent decoding.
[0031] Deblocking (or de-blocking), deblock filtering, or applying a deblocking filter refers to a process of smoothing (i.e., making the boundary less perceptible by a viewer) the boundary of adjacent reconstructed video blocks. Smoothing the boundary of adjacent reconstructed video blocks may include modifying sample values contained in a row or column adjacent to the boundary. JVET-T2001 specifies when a deblocking filter is applied to reconstructed sample values as part of an in-loop filtering process. In addition to applying a deblocking filter as part of an in-loop filtering process, JVET-T2001 specifies when sample adaptive offset (SAO) filtering may be applied in the in-loop filtering process. In general, SAO is a process of modifying deblocked sample values in a region by conditionally adding an offset value. Another type of filtering process includes the so-called adaptive loop filter (ALF). ALF with block-based adaptation is specified in JEM. In JEM, ALF is applied after the SAO filter. It should be noted that the ALF may be applied to the reconstructed samples independently of other filtering techniques. The process for applying the ALF specified in the JEM in a video coding device may be summarized as follows: (1) each 2×2 block of the luma component of the reconstructed picture is classified according to a classification index, (2) a set of filter coefficients is derived for each classification index, (3) a filtering decision is determined for the luma component, (4) a filtering decision is determined for the chroma components, and (5) the filter parameters (e.g., coefficients and decisions) are signaled. JVET-T2001 specifies deblocking, SAO, and ALF filters that may be described as generally based on the deblocking, SAO, and ALF filters specified in ITU-T H.265 and JEM.
[0032] It is noted that JVET-T2001 is referred to as a pre-release version of ITU-T H.266 and is therefore a near-finished draft of the video coding standard resulting from the VVC project, and is therefore sometimes referred to as the first version of the VVC standard (alternatively, VVC or VVC version 1 or ITU-H.266). It is noted that during the VVC project, convolutional neural network (CNN)-based techniques that showed potential for artifact removal and objective quality improvement were investigated, but it was decided not to include such techniques in the VVC standard. However, CNN-based techniques are currently being considered for extending and / or improving VVC. Some CNN-based techniques are related to post-filtering. For example, "AHG11: Content-adaptive neural network post-filter", 26th Meeting of ISO / IEC JTC1 / SC29 / WG11 20-29 April 2022, Teleconference, document JVET-Z0082-v2 (referred to herein as JVET-Z0082) describes a content-adaptive neural network based post-filter. Note that in JVET-Z0082, content adaptation is achieved by overfitting a NN post-filter to a test video. Further note that the result of the overfitting process in JVET-Z0082 is a weight update.JVET-Z0082 describes when the weight updates are coded in ISO / IEC FDIS 15938-17. Information technology - Multimedia content description interface - Part 17: Compression of neural networks for multimedia content description and analysis and Test Model of Incremental Compression of Neural Networks for Multimedia Content Description and Analysis (INCTM), N0179. February 2022, which are sometimes collectively referred to as the MPEG NNR (Neural Network Representation) or Neural Network Coding (NNC) standards. JVET-Z0082 further describes when the coded weight updates are signaled in the video bitstream as NNR postfilter SEI messages. "AHG9: NNR post-filter SEI message", 26th Meeting of ISO / IEC JTC1 / SC29 / WG11 20-29 April 2022, Teleconference, document JVET-Z0052-v1 (referred to herein as JVET-Z0052) describes the NNR post-filter SEI message utilized by JVET-Z0082. The elements of the NN post-filter described in JVET-Z0082 and the NNR post-filter SEI message described in JVET-Z0052 were adopted in "Additional SEI messages for VSEI (Draft 2)", 27th Meeting of ISO / IEC JTC1 / SC29 / WG11 13-22 July 2022, Teleconference, document JVET-AA2006-v2 (referred to herein as JVET-AA2006). JVET-AA2006 specifies a versatile supplemental enhancement information message for coded video bitstreams (VSEI).JVET-AA2006 specifies syntax and semantics for the neural network postfilter characteristics SEI message and for the neural network postfilter activation SEI message. The neural network postfilter characteristics SEI message specifies the neural networks that can be used as post-processing filters. The use of the specified post-processing filters for a particular picture is indicated in the neural network postfilter activation SEI message. Furthermore, "Information technology - MPEG video technologies - Part 7: Versatile supplemental enhancement information messages for coded video bitstreams, AMENDMENT 1: Additional SEI messages" (28th Meeting of ISO / IEC JTC1 / SC29 / WG5 9, November 2022, Mainz, DE) documents JVET-AB2006, m61498 (referred to herein as JVET-AB2006). JVET-AB2006 is described in more detail below. The techniques described herein provide techniques for signaling neural network postfilter messages.
[0033] For formulas used herein, the following arithmetic operators may be used:
[0034] [Table 1]
[0035] Additionally, the following mathematical functions may be used: Log2(x), the base 2 logarithm of x;
[0036]
number
[0037] With respect to the example syntax used herein, the following definitions of logical operators may apply: x&&y The Boolean logic "product" of x and y x||y The Boolean "union" of x and y ! Boolean logic "no" x?y:zIf x is true or not equal to 0, then evaluate the value of y, else evaluate the value of z.
[0038] In addition, the following relational operators may be applied:
[0039] [Table 2]
[0040] Furthermore, in the syntax descriptors used herein, it should be noted that the following descriptors may apply: -b(8): A byte with an arbitrary pattern of bits (8 bits). The parsing process of this descriptor is specified by the return value of the function read_bits(8). -f(n): fixed pattern bit sequence using n bits written (left to right), left bit first. The parsing process of this descriptor is specified by the return value of the function read_bits(n). -se(v): A signed integer zeroth-order Exp-Golomb encoded syntax element, left bit first. -tb(v): A shortened binary using up to maxVal bits, where maxVal is defined in the semantics of the syntax element. -tu(v): A shorthand unary that uses up to maxVal bits, where maxVal is defined in the semantics of the syntax element. -u(n): An unsigned integer using n bits. When n is "v" in the syntax table, the number of bits varies depending on the values of other syntax elements. The parsing process of this descriptor is specified by the return value of the function read_bits(n), which is interpreted as a binary representation of an unsigned integer written most significant bit first. -ue(v): An unsigned integer zeroth-order Exp-Golomb encoded syntax element, left bit first.
[0041] As mentioned above, video content includes a video sequence consisting of a series of pictures, and each picture may be divided into one or more regions. In JVET-T2001, the coded representation of a picture includes the VCL NAL units of a particular layer within an AU, including all CTUs of the picture. For example, referring again to FIG. 2, Pic 3 The coded representation of is composed of three coded slice NAL units (i.e., Slice 0 NAL unit, Slice 1 NAL Units and Slices 2It should be noted that the term video coding layer (VCL) NAL unit is used as a generic term for coded slice NAL units, i.e., VCL NAL is a generic term that includes all types of slice NAL units. As mentioned above and explained in more detail below, NAL units may encapsulate metadata used to decode video data. NAL units that encapsulate metadata used to decode a video sequence are generally referred to as non-VCL NAL units. Thus, in JVET-T2001, a NAL unit can be a VCL NAL unit or a non-VCL NAL unit. It should be noted that a VCL NAL unit includes slice header data that provides information used to decode a particular slice. Thus, in JVET-T2001, information used to decode video data, which in some cases may be referred to as metadata, is not limited to being included in a non-VCL NAL unit. JVET-T2001 specifies the case where a picture unit (PU) is a set of NAL units associated with each other according to a specified classification rule, consecutive in decoding order, and containing exactly one coded picture, and an access unit (AU) is a set of PUs containing coded pictures that belong to different layers and are associated at the same time for output from the DPB. JVET-T2001 further specifies the case where a layer is a set of VCL NAL units and associated non-VCL NAL units that all have a specific value of the layer identifier. Furthermore, in JVET-T2001, a PU consists of zero or one PH NAL unit, one coded picture containing one or more VCL NAL units, and zero or more other non-VCL NAL units.Furthermore, in JVET-T2001, a coded video sequence (CVS) is a sequence of AUs consisting of, in decoding order, a CVSS AU followed by zero or more AUs that are not CVSS AUs, including all subsequent AUs up to but not including any subsequent AUs that are CVSS AUs, where a coded video sequence start (CVSS) AU is an AU with a PU for each layer in the CVS, and a coded picture in each present picture unit is a coded layer video sequence start (CLVSS) picture. In JVET-T2001, a coded layer video sequence (CLVS) is a sequence of PUs in the same layer consisting, in decoding order, of a CLVSS PU followed by zero or more PUs that are not CLVSS PUs, including all subsequent PUs up to but not including any subsequent PUs that are CLVSS PUs. That is, in JVET-T2001, a bitstream may be described as including a sequence of AUs that form one or more CVSs.
[0042] Multi-layer video coding allows a video presentation to be decoded / displayed as a presentation corresponding to a base layer of video data and as one or more additional presentations corresponding to enhancement layers of the video data. For example, the base layer may allow a video presentation to be presented having a basic level of quality (e.g., high resolution rendering and / or 30 Hz frame rate), and the enhancement layer may allow a video presentation to be presented having an increased level of quality (e.g., ultra-high resolution rendering and / or 60 Hz frame rate). The enhancement layer may be coded by referencing the base layer. That is, for example, a picture in the enhancement layer may be coded (e.g., using inter-layer prediction techniques) by referencing one or more pictures in the base layer (including scaled versions thereof). It should be noted that layers may also be coded independently of each other. In this case, there may be no inter-layer prediction between two layers. Each NAL unit may include an identifier indicating the layer of video data with which the NAL unit is associated. As mentioned above, a sub-bitstream extraction process may be used to decode and display only a particular region of interest of a picture. Additionally, the sub-bitstream extraction process may be used to decode and display only a particular layer of a video. Sub-bitstream extraction may refer to a process in which a device receiving a compliant or conforming bitstream forms a new compliant or conforming bitstream by discarding and / or modifying data in the received bitstream. For example, sub-bitstream extraction may be used to form a new compliant or conforming bitstream that corresponds to a particular representation (e.g., a higher quality representation) of the video.
[0043] In JVET-T2001, each of a video sequence, a GOP, a picture, a slice, and a CTU may be associated with metadata that describes video coding properties, and some types of metadata may be encapsulated in non-VCL NAL units. JVET-T2001 defines parameter sets that may be used to describe video data and / or video coding properties. In particular, JVET-T2001 includes four types of parameter sets: video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), and adaptation parameter set (APS), where an SPS applies to zero or more entire CVSs, a PPS applies to zero or more entire coded pictures, an APS applies to zero or more slices, and a VPS may be optionally referenced by an SPS. A PPS applies to the individual coded pictures that reference it. In JVET-T2001, parameter sets may be encapsulated as non-VCL NAL units and / or signaled as messages. JVET-T2001 also includes a picture header (PH) encapsulated as a non-VCL NAL unit. In JVET-T2001, the picture header applies to all slices of a coded picture. JVET-T2001 further allows decoding capability information (DCI) and supplemental enhancement information (SEI) messages to be signaled. In JVET-T2001, the DCI and SEI messages assist processes related to decoding, display, or other purposes, but the DCI and SEI messages may not be required to create luma or chroma samples following the decoding process. In JVET-T2001, the DCI and SEI messages may be signaled in the bitstream using non-VCL NAL units. Furthermore, the DCI and SEI messages may be conveyed by some mechanism other than by being present in the bitstream (i.e., signaled out-of-band).
[0044] FIG. 3 shows an example of a bitstream including multiple CVSs, where the CVSs include AUs, and the AUs include picture units. The example shown in FIG. 3 corresponds to an example of encapsulating slice NAL units shown in the example of FIG. 2 in a bitstream. In the example shown in FIG. 3, Pic 3 The corresponding picture units for are three VCL NAL coded slice NAL units, namely, Slice 0 NAL unit, Slice 1 NAL Units and Slices 2 3 includes a PPS NAL unit and two non-VCL NAL units, namely, a PPS NAL unit and a PH NAL unit. Note that in FIG. 3, the header is a NAL unit header (i.e., should not be confused with a slice header). Note further that in FIG. 3, other non-VCL NAL units not shown may be included in the CVS, such as an SPS NAL unit, a VPS NAL unit, an SEI message NAL unit, etc. 3 The PPS NAL units used to decode the PPS NAL units may be found elsewhere in the bitstream, e.g., 0 Note that the PH syntax structure may be included in the picture unit corresponding to the current PU, or may be provided by an external mechanism. As described in more detail below, in JVET-T2001, the PH syntax structure may be present in the slice header of a VCL NAL unit or in the PH NAL unit of the current PU.
[0045] JVET-T2001 defines NAL unit header semantics that specify the type of raw byte sequence payload (RBSP) data structure contained in a NAL unit. Table 1 shows the syntax of the NAL unit header defined in JVET-T2001.
[0046] [Table 3]
[0047] JVET-T2001 specifies the following definitions for each syntax element shown in Table 1. forbidden_zero_bit shall be equal to 0. nuh_reserved_zero_bit shall be equal to 0. A value of 1 for nuh_reserved_zero_bit may be specified in the future by ITU-T|ISO / IEC. Although the value of nuh_reserved_zero_bit is required to be equal to 0 in this version of this specification, decoders conforming to this version of this specification shall allow a value of nuh_reserved_zero_bit equal to 1 to appear in the syntax and shall ignore (i.e., remove from the bitstream and discard) NAL units with nuh_reserved_zero_bit equal to 1. nuh_layer_id specifies the identifier of the layer to which a VCL NAL unit belongs, or the identifier of the layer to which a non-VCL NAL unit applies. Values of nuh_layer_id shall be in the range of 0 to 55, inclusive. Other values of nuh_layer_id are reserved for future use by ITU-T | ISO / IEC. Although values of nuh_layer_id are required to be in the range of 0 to 55, inclusive, in this version of this specification, decoders conforming to this version of this specification shall allow values of nuh_layer_id greater than 55 to appear in the syntax and shall ignore (i.e., remove from the bitstream and discard) NAL units with nuh_layer_id greater than 55. The value of nuh_layer_id shall be the same for all VCL NAL units of a coded picture. The value of nuh_layer_id of a coded picture or PU is the value of nuh_layer_id of the VCL NAL units of the coded picture or PU. When nal_unit_type is equal to PH_NUT or FD_NUT, nuh_layer_id shall be equal to the nuh_layer_id of the associated VCL NAL unit. When nal_unit_type is equal to EOS_NUT, nuh_layer_id shall be equal to one of the nuh_layer_id values of a layer present in CVS. NOTE – The values of nuh_layer_id in the DCI, OPI, VPS, AUD, and EOB NAL units are not constrained. nuh_temporal_id_plus1 minus 1 specifies the temporal identifier for the NAL unit. The value of nuh_temporal_id_plus1 shall not be equal to 0. The variable TemporalId is derived as follows: TemporalId=nuh_temporal_id_plus1-1 When nal_unit_type is in the range from IDR_W_RADL to RSV_IRAP_11, inclusive, TemporalId shall be equal to 0. When nal_unit_type is equal to STSA_NUT and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, TemporalId shall be greater than 0. The value of TemporalId shall be the same for all VCL NAL units of an AU. The value of TemporalId of a coded picture, PU, or AU is the value of TemporalId of the VCL NAL units of the coded picture, PU, or AU. The value of TemporalId of a sublayer representation is the largest value of TemporalId of all VCL NAL units in the sublayer representation. The values of TemporalId for non-VCL NAL units are constrained as follows: - if nal_unit_type is equal to DCI_NUT, OPI_NUT, VPS_NUT, or SPS_NUT, TemporalId shall be equal to 0 and the TemporalId of the AU containing the NAL unit shall be equal to 0. Otherwise, if nal_unit_type is equal to PH_NUT, then TemporalId shall be equal to the TemporalId of the PU that contains the NAL unit. Otherwise, if nal_unit_type is equal to EOS_NUT or EOB_NUT, TemporalId shall be equal to 0. Otherwise, if nal_unit_type is equal to AUD_NUT, FD_NUT, PREFIX_SEI_NUT, or SUFFIX_SEI_NUT, then TemporalId shall be equal to the TemporalId of the AU that contains the NAL unit. Otherwise, when nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, TemporalId shall be greater than or equal to the TemporalId of the PU that contains the NAL unit. NOTE - When the NAL unit is a non-VCL NAL unit, the value of TemporalId is equal to the minimum of the TemporalId values of all AUs to which the non-VCL NAL unit applies. When nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, TemporalId may be greater than or equal to the TemporalId of the containing AU, since all PPSs and APSs may be included at the beginning of the bitstream (e.g., when they are transported out-of-band and the receiver places them at the beginning of the bitstream), and the first coded picture has TemporalId equal to 0. nal_unit_type specifies the NAL unit type, i.e., the type of RBSP data structure contained in the NAL unit as specified in Table 2. NAL units with nal_unit_type in the range UNSPEC28 to UNSPEC31, inclusive, whose semantics are unspecified, SHALL NOT affect the decoding process specified in this specification. NOTE - NAL unit types in the range of UNSPEC_28 to UNSPEC_31 may be used as determined by the application. The decoding process for these values of nal_unit_type is not specified in this specification. Since different applications may use these NAL unit types for different purposes, it is expected that particular care will be taken in the design of encoders that generate NAL units with these nal_unit_type values, and in the design of decoders that interpret the content of NAL units with these nal_unit_type values. This specification does not define the management of these values. These nal_unit_type values may only be suitable for use in contexts where "collisions" of usage (i.e., different definitions of the meaning of the NAL unit content for the same nal_unit_type value) are not significant or possible, or are managed, e.g., defined or managed in a controlling application or transport specification, or by controlling the environment in which the bitstream is delivered. For purposes other than determining the amount of data in a DU of the bitstream, a decoder SHALL ignore (remove from the bitstream and discard) the content of all NAL units that use reserved values of nal_unit_type. NOTE - This requirement allows for the future definition of compatible extensions to this specification.
[0048] [Table 4] NOTE - A Clean Random Access (CRA) picture may have an associated RASL or RADL picture present in the bitstream. NOTE - An Instantaneous Decoding Refresh (IDR) picture with nal_unit_type equal to IDR_N_LP does not have an associated leading picture present in the bitstream. An IDR picture with nal_unit_type equal to IDR_W_RADL does not have an associated RASL picture present in the bitstream, but may have an associated RADL picture in the bitstream. The value of nal_unit_type shall be the same for all VCL NAL units of a subpicture. A subpicture is referred to as having the same NAL unit type as the VCL NAL units of the subpicture. For the VCL NAL units of any particular picture, the following applies: -If pps_mixed_nalu_types_in_pic_flag is equal to 0, the value of nal_unit_type shall be the same for all VCL NAL units of a picture, and the picture or PU is referenced as having the same NAL unit type as the coded slice NAL units of the picture or PU. - Otherwise (pps_mixed_nalu_types_in_pic_flag is equal to 1), all of the following constraints apply. -A picture shall have at least two sub-pictures. - A VCL NAL unit of a picture shall have two or more distinct nal_unit_type values. - There shall be no VCL NAL unit of the picture with nal_unit_type equal to GDR_NUT. When a VCL NAL unit of a picture has nal_unit_type equal to nalUnitTypeA equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT, all other VCL NAL units of the picture shall have nal_unit_type equal to nalUnitTypeA or TRAIL_NUT. The value of nal_unit_type shall be the same for all pictures in an IRAP or GDR AU. When sps_video_parameter_set_id is greater than 0, vps_max_tid_il_ref_pics_plus1[i][j] is equal to 0 for j equal to GeneralLayerIdx[nuh_layer_id] and for any value of i in the range of j+1 to vps_max_layers_minus1, inclusive, pps_mixed_nalu_types_in_pic_flag is equal to 1, and the value of nal_unit_type is not equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT. The requirements for bitstream conformance are that the following constraints apply: When a picture is a leading picture of an IRAP picture, it shall be a RADL or RASL picture. When a subpicture is the leading subpicture of an IRAP subpicture, it shall be a RADL or RASL subpicture. - If a picture is not the leading picture of an IRAP picture, it shall not be a RADL or RASL picture. - If a subpicture is not the leading subpicture of an IRAP subpicture, it shall not be a RADL or RASL subpicture. - There shall be no RASL pictures associated with an IDR picture in the bitstream. - No RASL sub-pictures associated with an IDR sub-picture shall be present in the bitstream. - There shall be no RADL pictures associated with an IDR picture with nal_unit_type equal to IDR_N_LP in the bitstream. NOTE - Performing random access at the position of an IRAP AU by discarding all PUs before the IRAP AU (and correctly decoding non-RASL pictures in the IRAP AU and all following AUs in decoding order) is possible, provided that each parameter set is available (in the bitstream or by external means not specified in this specification) at the time it is referenced. - There shall be no RADL sub-pictures associated with an IDR sub-picture with nal_unit_type equal to IDR_N_LP in the bitstream. -Any picture with nuh_layer_id equal to the particular value layerId that precedes in decoding order an IRAP picture with nuh_layer_id equal to layerId shall precede the IRAP picture in output order and shall precede any RADL picture associated with the IRAP picture in output order. - Any subpicture with nuh_layer_id equal to the specified value layerId and subpicIdx subpicture index equal to the specified value that precedes in decoding order an IRAP subpicture with nuh_layer_id equal to layerId and subpicture index equal to subpicIdx shall precede in output order the IRAP subpicture and all its associated RADL subpictures. Any picture with nuh_layer_id equal to the specified value layerId that precedes in decoding order the recovery point picture with nuh_layer_id equal to -layerId shall precede the recovery point picture in output order. Any subpicture with nuh_layer_id equal to a specific value layerId and subpicIdx equal to a specific value that precedes in decoding order a subpicture with nuh_layer_id equal to layerId and subpicIdx equal to a specific value in the recovery point picture shall precede that subpicture in the recovery point picture in output order. - Any RASL picture associated with a CRA picture shall precede any RADL picture associated with the CRA picture in output order. Any RASL sub-picture associated with a CRA sub-picture shall precede any RADL sub-picture associated with the CRA sub-picture in output order. - Any RASL picture having nuh_layer_id equal to a particular value layerId and associated with a CRA picture shall follow in output order any IRAP or GDR picture having nuh_layer_id equal to layerId that precedes the CRA picture in decoding order. - Any RASL subpicture associated with a CRA subpicture and having nuh_layer_id equal to a particular value layerId and a subpicture index equal to a particular value subpicIdx shall follow in output order any IRAP or GDR subpicture with nuh_layer_id equal to layerId and a subpicture index equal to subpicIdx that precedes the CRA subpicture in decoding order. - if sps_field_seq_flag is equal to 0, the following applies: if the current picture with nuh_layer_id equal to a particular value layerId is a leading picture associated with an IRAP picture, it shall precede, in decoding order, all non-leading pictures associated with the same IRAP picture. Otherwise (sps_field_seq_flag is equal to 1), let picA and picB be the first and last leading pictures associated with an IRAP picture, respectively, in decoding order, there shall be at most one non-leading picture with nuh_layer_id equal to layerId preceding picA in decoding order, and there shall be no non-leading picture with nuh_layer_id equal to layerId between picA and picB in decoding order. - if sps_field_seq_flag is equal to 0, the following applies: if the current subpicture with nuh_layer_id equal to a particular value layerId and subpicture index equal to a particular value subpicIdx is a leading subpicture associated with an IRAP subpicture, then it shall precede, in decoding order, all non-leading subpictures associated with the same IRAP subpicture. Otherwise (sps_field_seq_flag is equal to 1), let subpicA and subpicB be the first and last leading subpictures associated with an IRAP subpicture, respectively, in decoding order, then there shall be at most one non-leading subpicture with nuh_layer_id equal to layerId and subpicture index equal to subpicIdx preceding subpicA in decoding order, and there shall be no non-leading pictures with nuh_layer_id equal to layerId and subpicture index equal to subpicIdx between picA and picB in decoding order.
[0049] A NAL unit may include a supplemental enhancement information (SEI) syntax structure, as provided in Table 2. Tables 3 and 4 show the supplemental enhancement information (SEI) syntax structure defined in JVET-T2001.
[0050] [Table 5]
[0051] [Table 6]
[0052] With respect to Tables 3 and 4, JVET-T2001 specifies the following semantics: Each SEI message consists of variables that specify the type, payloadType, and size, payloadSize, of the SEI message payload. The SEI message payload is specified. The derived SEI message payload size, payloadSize, is specified in bytes and shall be equal to the number of RBSP bytes in the SEI message payload. NOTE - A NAL unit byte sequence containing an SEI message may contain one or more emulation prevention bytes (represented by the emulation_prevention_three_byte syntax element). Since the payload size of an SEI message is specified in RBSP bytes, the amount of emulation prevention bytes is not included in the size of the SEI payload, payloadSize. payload_type_byte is the payload type byte of the SEI message. payload_size_byte is the payload size in bytes of the SEI message.
[0053] It should be noted that JVET-T2001 defines payload types, and "Additional SEI messages for VSEI (Draft 6)" 25th Meeting of ISO / IEC JTC1 / SC29 / WG11 12-21 January 2022, Teleconference, document JVET-Y2006-v1, incorporated herein by reference and referred to as JVET-Y2006, defines additional payload types. Table 5 shows the sei_payload() syntax structure in a schematic manner. That is, Table 5 shows the sei_payload() syntax structure, but for brevity, not all possible types of payloads are included in Table 5.
[0054] [Table 7]
[0055] With respect to Table 5, JVET-T2001 specifies the following semantics: sei_reserved_payload_extension_data SHALL NOT be present in bitstreams that conform to this version of this specification. However, decoders that conform to this version of this specification SHALL ignore the presence and value of sei_reserved_payload_extension_data. When present, the length in bits of sei_reserved_payload_extension_data shall be 8 * Equal to payloadSize-nEarlierBits-nPayloadZeroBits-1, where nEarlierBits is the number of bits in the sei_payload() syntax structure preceding the sei_reserved_payload_extension_data syntax element and nPayloadZeroBits is the number of sei_payload_bit_equal_to_zero syntax elements at the end of the sei_payload() syntax structure. If more_data_in_payload() is true after parsing an SEI message syntax structure (for example, the buffering_period() syntax structure), and nPayloadZeroBits is not equal to 7, then PayloadBits is set to 8. * payloadSize-n is set equal to PayloadZeroBits-1, otherwise PayloadBits is set equal to 8 * Set equal to payloadSize. payload_bit_equal_to_one shall be equal to 1. payload_bit_equal_to_zero shall be equal to 0. NOTE - SEI messages with the same value of payloadType are conceptually the same SEI message, regardless of whether they are included in a prefix or suffix SEI NAL unit. NOTE - For SEI messages specified in this specification and the VSEI specification (ITU-T H.274|ISO / IEC 23002-7), payloadType values are aligned with similar SEI messages specified in AVC (Rec. ITU-T H.264|ISO / IEC 14496-10) and HEVC (Rec. ITU-T H.265|ISO / IEC 23008-2). The semantics and duration for each SEI message are specified in the semantics specification for each particular SEI message. NOTE – The persistence information in the SEI messages is summarized for informational purposes.
[0056] JVET-T2001 further specifies the following: SEI messages having the syntax structure identified in [Table 5] specified in Rec. ITU-T H.274 | ISO / IEC 23002-7 may be used with bitstreams specified by this specification. When any particular Rec. ITU-T H.274|ISO / IEC 23002-7 SEI message is included in a bitstream specified by this specification, the SEI payload syntax shall be contained in the sei_payload() syntax structure specified in [Table 5] and shall use the payloadType value specified in [Table 5], plus any SEI message-specific constraints specified in this Annex for that particular SEI message shall apply. The value of PayloadBits is passed to the parser for the SEI message syntax structure specified in Rec. ITU-T H.274 | ISO / IEC 23002-7, as specified above.
[0057] As mentioned above, JVET-AB2006 defines NN postfilter supplemental enhancement information messages. In particular, JVET-AB2006 defines a neural network postfilter characteristics SEI message (payloadType==210) and a neural network postfilter activation SEI message (payloadType==211). Table 6 shows the syntax of the neural network postfilter characteristics SEI message defined in JVET-AB2006. Note that the neural network postfilter characteristics SEI message is sometimes referred to as the NNPFC SEI.
[0058] [Table 8-1]
[0059] [Table 8-2]
[0060] With respect to Table 6, JVET-AB2006 specifies the following semantics: The neural-network post-filter characteristics (NNPFC) SEI message specifies neural networks that can be used as post-processing filters. The use of a specified post-processing filter for a particular picture is indicated in the neural-network post-filter activation SEI message. Use of this SEI message requires the definition of the following variables: The width and height of the cropped decoded output picture in units of luma samples, denoted herein by CroppedWidth and CroppedHeight, respectively. - When present, the luma sample array CroppedYPic[idx] and the chroma sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] of the cropped decoded output picture with idx in the range of 0 to numInputPics-1, inclusive, that are used as input for the post-processing filters. -BitDepth - the bit depth of the luma sample array of the cropped decoded output picture Y . -BitDepth - the bit depth of the chroma sample array, if any, of the cropped decoded output picture C . A chroma format indicator, denoted herein by ChromaFormatIdc. When -nnpfc_auxiliary_inp_idc is equal to 1, the filtering strength control value StrengthControlVal shall be a real number in the range of 0 to 1, inclusive. The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc as specified by Table 7.
[0061] [Table 9] NOTE - There may be more than one NNPFC SEI message for the same picture. When more than one NNPFC SEI message with different values of nnpfc_id is present or activated for the same picture, they may have the same or different values of nnpfc_purpose and nnpfc_mode_idc. nnpfc_id contains an identification number that can be used to identify a post-processing filter. The value of nnpfc_id is in the range 0 to 2, inclusive. 32 -2. The range is from 256 to 511, inclusive, and 256 to 511, inclusive. 31 ~2 32Values of nnpfc_id up to -2 are reserved for future use by ITU-T | ISO / IEC. 31 ~2 32 Decoders compliant with this version of this document that encounter an NNPFC SEI message with an nnpfc_id within the range of -2 shall ignore that SEI message. When the NNPFC SEI message is the first NNPFC SEI message, in decoding order, with a particular nnpfc_id value in the current CLVS, the following applies. -This SEI message specifies a basic post-processing filter. - This SEI message relates to the current decoded picture and all subsequent decoded pictures of the current layer in output order until the end of the current CLVS. When an NNPFC SEI message is a repetition of a previous NNPFC SEI message in the current CLVS in decoding order, the following semantics apply as if this SEI message was the only NNPFC SEI message with the same content within the current CLVS. When the NNPFC SEI message is not the first NNPFC SEI message, in decoding order, with a particular nnpfc_id value in the current CLVS, the following applies. -This SEI message defines an update to the preceding elementary post-processing filter in decoding order that has the same nnpfc_id value. -This SEI message relates to the current decoded picture and all subsequent decoded pictures of the current layer, in output order, until the end of the current CLVS or until the next NNPFC SEI message with that particular nnpfc_id value in output order within the current CLVS. nnpfc_mode_idc equal to 0 indicates that this SEI message contains an ISO / IEC 15938-17 bitstream that specifies a basic post-processing filter or is an update to a basic post-processing filter with the same nnpfc_id value. When the NNPFC SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc_id value in the current CLVS, an nnpfc_mode_idc equal to 1 specifies that the basic post-processing filter associated with the nnpfc_id value is a neural network identified by the URI indicated by nnpfc_uri having the format identified by the tag URI nnpfc_tag_uri. When the NNPFC SEI message is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value in the current CLVS, nnpfc_mode_idc equal to 1 specifies that updates to basic post-processing filters with the same nnpfc_id value are defined by the URI indicated by nnpfc_uri having the format identified by the tag URI nnpfc_tag_uri. Values of nnpfc_mode_idc shall be in the range 0 to 1, inclusive, in bitstreams conforming to this version of this document. Values of nnpfc_mode_idc between 2 and 255, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_mode_idc in the range 2 to 255, inclusive. Values of nnpfc_mode_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use. When this SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc_id value in the current CLVS, the post-processing filter PostProcessingFilter() is assigned to be the same as the base post-processing filter. When this SEI message is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value in the current CLVS, the post-processing filter PostProcessingFilter() is obtained by applying the updates defined by this SEI message to the basic post-processing filter. The updates are not cumulative; rather, each update is applied to the base post-processing filter, which is the post-processing filter specified by the first NNPFC SEI message, in decoding order, with a particular nnpfc_id value in the current CLVS. nnpfc_reserved_zero_bit_a shall be equal to 0 in bitstreams conforming to this version of this document. Decoders shall ignore NPFC SEI messages in which nnpfc_reserved_zero_bit_a is not equal to 0. The nnpfc_tag_uri contains a tag URI with syntax and semantics specified in IETF RFC 4151 that identifies the format and associated information about a neural network to be used as a base post-processing filter or an update to a base post-processing filter with the same nnpfc_id value specified by the nnpfc_uri. NOTE - The nnpfc_tag_uri makes it possible to uniquely identify the format of the neural network data specified by the nnrpf_uri without the need for a central registry. An nnpfc_tag_uri equal to "tag:iso.org,2023:15938-17" indicates that the neural network data identified by the nnpfc_uri conforms to ISO / IEC 15938-17. The nnpfc_uri contains a URI with syntax and semantics specified in IETF Internet Standard 66 that identifies a neural network to be used as a base post-processing filter or an update to a base post-processing filter with the same nnpfc_id value. nnpfc_formatting_and_purpose_flag equal to 1 specifies that syntax elements related to filter purpose, input formatting, output formatting, and complexity are present. nnpfc_formatting_and_purpose_flag equal to 0 specifies that syntax elements related to filter purpose, input formatting, output formatting, and complexity are not present. When this SEI message is the first NNPFC SEI message, in decoding order, with a particular nnpfc_id value in the current CLVS, nnpfc_formatting_and_purpose_flag shall be equal to 1. When this SEI message is not the first NNPFC SEI message, in decoding order, with a particular nnpfc_id value in the current CLVS, nnpfc_formatting_and_purpose_flag shall be equal to 0. nnpfc_purpose indicates the purpose of the post-processing filter as specified in Table 8. Values of nnpfc_purpose shall be in the range 0 to 5, inclusive, in bitstreams conforming to this version of this document. Values of nnpfc_purpose between 6 and 1023, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoding devices conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_purpose in the range 6 to 1203, inclusive. Values of nnpfc_purpose greater than 1023 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use.
[0062] [Table 10] NOTE – When reserved values of nnpfc_purpose are used in the future by ITU-T|ISO / IEC, the syntax of this SEI message may be extended with syntax elements whose presence is conditioned by nnpfc_purpose being equal to that value. When SubWidthC is equal to 1 and SubHeightC is equal to 1, nnpfc_purpose shall not be equal to 2 or 4. nnpfc_out_sub_c_flag equal to 1 specifies that outSubWidthC is equal to 1 and outSubHeightC is equal to 1. nnpfc_out_sub_c_flag equal to 0 specifies that outSubWidthC is equal to 2 and outSubHeightC is equal to 1. When nnpfc_out_sub_c_flag is not present, outSubWidthC is inferred to be equal to SubWidthC and outSubHeightC is inferred to be equal to SubHeightC. When ChromaFormatIdc is equal to 2 and nnpfc_out_sub_c_flag is present, the value of nnpfc_out_sub_c_flag shall be equal to 1. nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples specify the width and height, respectively, of the picture's luma sample array resulting from applying the post-processing filter identified by nnpfc_id to the cropped decoded output picture. When nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples are not present, they are inferred to be equal to CroppedWidth and CroppedHeight, respectively. The value of nnpfc_pic_width_in_luma_samples ranges from CroppedWidth to CroppedWidth, inclusive. * The value of nnpfc_pic_height_in_luma_samples is in the range CroppedHeight~CroppedHeight, inclusive.* It shall be within the range of 16-1. nnpfc_num_input_pics_minus2 plus 2 specifies the number of decoded output pictures used as input for the post-processing filter. nnpfc_interpolated_pics[i] specifies the number of interpolated pictures generated by the post-processing filter between the i-th picture and the (i+1)-th picture that are used as input for the post-processing filter. The variables numInputPics, which specify the number of pictures used as input for the post-processing filter, and numOutputPics, which specify the total number of pictures resulting from the post-processing filter, are derived as follows:
[0063] [Table 11] nnpfc_component_last_flag equal to 1 indicates that the last dimension in the input tensor inputTensor to the post-processing filter and the output tensor outputTensor resulting from the post-processing filter is used for the current channel. nnpfc_component_last_flag equal to 0 indicates that the third dimension in the input tensor inputTensor to the post-processing filter and the output tensor outputTensor resulting from the post-processing filter is used for the current channel. NOTE - The first dimension in the input and output tensors is used for the batch index, which is the practice in some neural network frameworks. Although the formula in the semantics of this SEI message uses the batch size corresponding to a batch index equal to 0, it is up to the post-processing implementation to determine the batch size used as input to the neural network inference. NOTE - For example, when nnpfc_inp_order_idc is equal to 3 and nnpfc_auxiliary_inp_idc is equal to 1, the input tensor has 7 channels including 4 luma matrices, 2 chroma matrices, and 1 auxiliary input matrix. In this case, the process DeriveInputTensors() derives each of these 7 channels of the input tensor one by one, and when a particular one of these channels is processed, that channel is called the current channel in the process. nnpfc_inp_format_idc indicates how to convert the sample values of the cropped decoded output picture into input values to the post-processing filter. When nnpfc_inp_format_idc is equal to 0, the input values to the post-processing filter are real numbers, and the functions InpY() and InpC() are specified as follows: InpY(x) = x÷((1< <BitDepthY)-1) InpC(x) = x ÷ ((1< <BitDepthC)-1) When nnpfc_inp_format_idc is equal to 1, the input values to the post-processing filter are unsigned integers and the functions InpY() and InpC() are specified as follows:
[0064] [Table 12] The variable inpTensorBitDepth is derived from the syntax element nnpfc_inp_tensor_bitdepth_minus8, as specified below. Values of nnpfc_inp_format_idc greater than 1 are reserved for future specification by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages which contain reserved values of nnpfc_inp_format_idc. nnpfc_inp_tensor_bitdepth_minus8 plus 8 specifies the bit depth of the luma sample values in the input integer tensor. The value of inpTensorBitDepth is derived as follows: inpTensorBitDepth=nnpfc_inp_tensor_bitdepth_minus8+8 It is a bitstream conformance requirement that the value of nnpfc_inp_tensor_bitdepth_minus8 be in the range 0 to 24, inclusive. nnpfc_inp_order_idc indicates how to order the sample array of the cropped decoded output picture as one of the input pictures to the post-processing filter. Values of nnpfc_inp_order_idc shall be in the range 0 to 3, inclusive, in bitstreams conforming to this version of this document. Values of nnpfc_inp_order_idc between 4 and 255, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range 4 to 255, inclusive. Values of nnpfc_inp_order_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use. When ChromaFormatIdc is not equal to 1, nnpfc_inp_order_idc shall be equal to 3. Table 9 contains useful descriptions of the nnpfc_inp_order_idc values.
[0065] [Table 13] A patch is a rectangular array of samples from a component (eg, a luma component or a chroma component) of a picture. An nnpfc_auxiliary_inp_idc greater than 0 indicates that the auxiliary input data is present in the input tensor of the neural network postfilter. An nnpfc_auxiliary_inp_idc equal to 0 indicates that the auxiliary input data is not present in the input tensor. An nnpfc_auxiliary_inp_idc equal to 1 specifies that the auxiliary input data is derived as specified in Equation 82. Values of nnpfc_auxiliary_inp_idc shall be in the range 0 to 1, inclusive, in bitstreams conforming to this version of this document. Values of nnpfc_inp_order_idc between 2 and 255, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range 2 to 255, inclusive. Values of nnpfc_inp_order_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use. The process DeriveInputTensors() for deriving an input tensor inputTensor given vertical sample coordinates cTop and horizontal sample coordinates cLeft that specify the top-left sample location of a patch of samples contained in the input tensor is specified as follows:
[0066] [Table 14-1]
[0067] [Table 14-2]
[0068] [Table 14-3]
[0069] [Table 14-4] nnfpc_separate_colour_description_present_flag equal to 1 indicates that the separate combination of colour primaries, transfer characteristics and matrix coefficients for the picture resulting from the post-processing filter is specified in the SEI message syntax structure. nnpfc_separate_colour_description_present_flag equal to 0 indicates that the combination of colour primaries, transfer characteristics and matrix coefficients for the picture resulting from the post-processing filter is the same as that indicated in the VUI parameters for CLVS. nnpfc_colour_primaries has the same semantics as specified for the vui_colour_primaries syntax element, which is as follows: vui_colour_primaries indicates the chromaticity coordinates of the source primaries. Its semantics are as specified for MatrixCoefficients in Rec. ITU-T H.273|ISO / IEC 23091-2. When the vui_colour_primaries syntax element is not present, the value of vui_colour_primaries is inferred to be equal to 2 (chromaticity is unknown, unspecified, or determined by other means not specified in this specification). Values of vui_transfer_characteristics identified as reserved for future use in Rec. ITU-T H.273|ISO / IEC 23091-2 shall not be present in bitstreams conforming to this version of this specification. Decoding devices shall interpret reserved values of vui_colour_primaries as equal to the value 2. However, the following applies: -nnpfc_colour_primaries specifies the picture primaries resulting from applying the neural network postfilter specified in the SEI message, rather than the primaries used for CLVS. When -nnpfc_colour_primaries is not present in the NNPFC SEI message, the value of nnpfc_colour_primaries is inferred to be equal to vui_colour_primaries. nnpfc_transfer_characteristics has the same semantics as specified for the vui_transfer_characteristics syntax element, which is as follows: vui_transfer_characteristics indicates the transfer characteristic function of the color representation. Its semantics are as specified for MatrixCoefficients in Rec. ITU-T H.273|ISO / IEC 23091-2. When the vui_transfer_characteristics syntax element is not present, the value of vui_transfer_characteristics is inferred to be equal to 2 (the transfer characteristics are unknown, unspecified, or determined by other means not specified in this specification). Values of vui_transfer_characteristics identified as reserved for future use in Rec. ITU-T H.273|ISO / IEC 23091-2 shall not be present in bitstreams conforming to this version of this specification. Decoders shall interpret reserved values of vui_transfer_characteristics as equal to the value 2. However, the following applies: -nnpfc_transfer_characteristics specifies the transfer characteristics of the picture resulting from applying the neural network postfilter specified in the SEI message, instead of the transfer characteristics used for CLVS. When -nnpfc_transfer_characteristics is not present in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is inferred to be equal to vui_transfer_characteristics. nnpfc_matrix_coeffs has the same semantics as specified for the vui_matrix_coeffs syntax element, which is as follows: vui_matrix_coeffs describes the formulas used in deriving the luma and chroma signals from the green, blue, and red, or Y, Z, and X primaries. The semantics are as specified for MatrixCoefficients in Rec. ITU-T H.273|ISO / IEC 23091-2. However, the following applies: -nnpfc_matrix_coeffs specifies the matrix coefficients of the picture resulting from applying the neural network postfilter specified in the SEI message, not the matrix coefficients used for CLVS. When -nnpfc_matrix_coeffs is not present in the NNPFC SEI message, the value of nnpfc_matrix_coeffs is inferred to be equal to vui_matrix_coeffs. The allowed values for -nnpfc_matrix_coeffs are not constrained by the chroma format of the decoded video picture indicated by the value of ChromaFormatIdc for the semantics of the VUI parameter. - when nnpfc_matrix_coeffs is equal to 0, nnpfc_out_order_idc shall not be equal to 1 or 3. An nnpfc_out_format_idc equal to 0 indicates that the sample values output by the post - processing filter are real numbers, and the value range of 0 to 1 including both end - values is linearly mapped to an unsigned integer value range of 0 to (1<<bitDepth) - 1 including both end - values for any desired bit depth bitDepth for subsequent post - processing or display. An nnpfc_out_format_flag equal to 1 indicates that the sample values output by the post - processing filter are unsigned integers within the range of 0 to (1<<(nnpfc_out_tensor_bitdepth_minus8 + 8)) - 1 including both end - values. Values of nnpfc_out_format_idc greater than 1 are reserved for future specifications by ITU - T|ISO / IEC and shall not be present in the bitstream conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages containing reserved values of nnpfc_out_format_idc. nnpfc_out_tensor_bitdepth_minus8 plus 8 specifies the bit depth of the sample values in the output integer tensor. The value of nnpfc_out_tensor_bitdepth_minus8 shall be within the range of 0 to 24 including both end - values. nnpfc_out_order_idc indicates the output order of samples obtained from the post - processing filter. Values of nnpfc_out_order_idc shall be in the range 0 to 3, inclusive, in bitstreams conforming to this version of this document. Values of nnpfc_out_order_idc between 4 and 255, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_out_order_idc in the range 4 to 255, inclusive. Values of nnpfc_out_order_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use. When nnpfc_purpose is equal to 2 or 4, nnpfc_out_order_idc shall not be equal to 3. Table 10 contains useful descriptions of the nnpfc_out_order_idc values.
[0070] [Table 15] The process StoreOutputTensors() for deriving sample values in filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor given vertical sample coordinates cTop and horizontal sample coordinates cLeft that specify the top left sample location of a patch of samples contained in the input tensor is specified as follows:
[0071] [Table 16-1]
[0072] [Table 16-2]
[0073] [Table 16-3] nnpfc_constant_patch_size_flag equal to 1 indicates that the post-processing filter accepts as input exactly the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. nnpfc_constant_patch_size_flag equal to 0 indicates that the post-processing filter accepts as input any patch size that is a positive integer multiple of the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. nnpfc_patch_width_minus1+1 indicates the horizontal sample count of the patch size required for input to the post-processing filter when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_width_minus1 shall be in the range 0 to Min(32766,CroppedWidth-1), inclusive. nnpfc_patch_height_minus1+1 indicates the vertical sample count of the patch size required for input to the post-processing filter when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_height_minus1 shall be in the range 0 to Min(32766,CroppedHeight-1), inclusive. The variables inpPatchWidth and inpPatchHeight are the patch size width and patch size height, respectively. If nnpfc_constant_patch_size_flag is equal to 0, the following applies: The values of -inpPatchWidth and inpPatchHeight are either provided by external means not specified in this document, or are set by the post-processor itself. - The value of inpPatchWidth shall be a positive integer multiple of nnpfc_patch_width_minus1+1 and shall be less than or equal to CroppedWidth. The value of inpPatchHeight shall be a positive integer multiple of nnpfc_patch_height_minus1+1 and shall be less than or equal to CroppedHeight. Otherwise (nnpfc_constant_patch_size_flag is equal to 1), the value of inpPatchWidth is set equal to nnpfc_patch_width_minus1+1, and the value of inpPatchHeight is set equal to nnpfc_patch_height_minus1+1. nnpfc_overlap indicates the overlapping horizontal and vertical sample counts of adjacent input tensors of the post-processing filter. The value of nnpfc_overlap must be in the range of 0 to 16383, inclusive. The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows:
[0074] [Table 17] outPatchWidth * CroppedWidth is nnpfc_pic_width_in_luma_samples * shall be equal to inpPatchWidth, and outPatchHeight * CroppedHeight is nnpfc_pic_height_in_luma_samples * It is a bitstream conformance requirement that it shall be equal to inpPatchHeight. nnpfc_padding_type indicates the padding process when referring to sample positions outside the boundary of the cropped decoded output picture as described in Table 11. The value of nnpfc_padding_type shall be in the range of 0 to 15, inclusive.
[0075] [Table 18] nnpfc_luma_padding_val indicates the luma value used for padding when nnpfc_padding_type is equal to 4. nnpfc_cb_padding_val indicates the Cb value used for padding when nnpfc_padding_type is equal to 4. nnpfc_cr_padding_val indicates the Cr value used for padding when nnpfc_padding_type is equal to 4. The function InpSampleVal(y, x, picHeight, picWidth, croppedPic) whose inputs are vertical sample position y, horizontal sample position x, picture height picHeight, picture width picWidth, and sample array croppedPic returns the value of sampleVal derived as follows: NOTE - For input to function InpSampleVal(), the vertical positions are listed before the horizontal positions for compatibility with the input tensor conventions of some inference engines.
[0076] [Table 19] The following example process may be used to filter the cropped decoded output picture in patches using a post-processing filter PostProcessingFilter() to generate a filtered picture including Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated by nnpfc_out_order_idc.
[0077] [Table 20] nnpfc_complexity_info_present_flag equal to 1 specifies that there are one or more syntax elements indicating the complexity of the post-processing filter associated with nnpfc_id. nnpfc_complexity_info_present_flag equal to 0 specifies that there are no syntax elements indicating the complexity of the post-processing filter associated with nnpfc_id. nnpfc_parameter_type_idc equal to 0 indicates that the neural network uses integer parameters only. nnpfc_parameter_type_flag equal to 1 indicates that the neural network may use floating point or integer parameters. nnpfc_parameter_type_idc equal to 2 indicates that the neural network uses binary parameters only. nnpfc_parameter_type_idc equal to 3 is reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_parameter_type_idc equal to 3. nnpfc_log2_parameter_bit_length_minus3 equal to 0, 1, 2, and 3 indicates that the neural network shall not use parameters with bit lengths greater than 8, 16, 32, and 64, respectively. When nnpfc_parameter_type_idc is present and nnpfc_log2_parameter_bit_length_minus3 is not present, the neural network shall not use parameters with bit lengths greater than 1. nnpfc_num_parameters_idc indicates the maximum number of neural network parameters for the post-processing filters, in units of powers of 2048. nnpfc_num_parameters_idc equal to 0 indicates that the maximum number of neural network parameters is unknown. The value nnpfc_num_parameters_idc shall be in the range 0 to 52, inclusive. Values of nnpfc_num_parameters_idc greater than 52 are reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoding devices conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_num_parameters_idc greater than 52. If the value of nnpfc_num_parameters_idc is greater than 0, the variable maxNumParameters is derived as follows: maxNumParameters=(2048< <nnpfc_num_parameters_idc)-1 It is a bitstream conformance requirement that the number of neural network parameters in the post-processing filter shall be less than or equal to maxNumParameters. nnpfc_num_kmac_operations_idc greater than 0 specifies the maximum number of multiply-accumulate operations per sample for the post-processing filter. *nnpfc_num_kmac_operations_idc equal to 0 indicates that the maximum number of multiply-add operations in the network is unknown. The value of nnpfc_num_kmac_operations_idc can range from 0 to 2, inclusive. 32 It should be in the range of -1. nnpfc_total_kilobyte_size greater than 0 indicates the total size in kilobytes required to store the uncompressed parameters for the neural network. The total size in bits is equal to or greater than the sum of the bits used to store each parameter. nnpfc_total_kilobyte_size is the total size in bits divided by 8000 and rounded up. nnpfc_total_kilobyte_size equal to 0 indicates that the total size required to store the parameters of the neural network is unknown. The value of nnpfc_total_kilobyte_size can be between 0 and 2, inclusive. 32 It should be in the range of -1. nnpfc_reserved_zero_bit_b shall be equal to 0 in bitstreams conforming to this version of this document. Decoders shall ignore NPFC SEI messages in which nnpfc_reserved_zero_bit_b is not equal to 0. Let nnpfc_payload_byte[i] contain the i-th byte of a bitstream that conforms to ISO / IEC 15938-17. The byte sequence nnpfc_payload_byte[i] for all present values of i shall be a complete bitstream that conforms to ISO / IEC 15938-17.
[0078] Table 12 shows the syntax of the neural network post-filter activation SEI message defined in JVET-AB2006.
[0079] [Table 21]
[0080] With respect to Table 12, JVET-AB2006 specifies the following semantics: The neural-network post-filter activation (NNPFA) SEI message activates or deactivates the possible use of a target neural-network post-processing filter, identified by nnpfa_target_id, for post-processing filtering of a set of pictures. NOTE − There may be several NNPFA SEI messages for the same picture, for example when post-processing filters are for different purposes or filter different color components. nnpfa_target_id indicates a target neural network post-processing filter, which is specified by one or more neural network post-processing filter characteristics SEI messages associated with the current picture and having nnpfc_id equal to nnfpa_target_id. The value of nnpfa_target_id is between 0 and 2, inclusive. 32 -2. The range is from 256 to 511, inclusive, and 256 to 511, inclusive. 31 ~2 32 Values of nnpfa_target_id up to -2 are reserved for future use by ITU-T | ISO / IEC. 31 ~2 32 Decoders compliant with this version of this document that encounter an NNPFA SEI message with an nnpfa_target_id within the range of -2 shall ignore that SEI message. An NNPFA SEI message with a particular value of nnpfa_target_id shall not be present in the current PU unless one or both of the following conditions are true: - There is an NNPFC SEI message in the current CLVS with nnpfc_id equal to a particular value of nnpfa_target_id present in a PU preceding the current PU in decoding order. - There is an NNPFC SEI message in the current PU with nnpfc_id equal to a particular value of nnpfa_target_id. When a PU contains both an NNPFC SEI message with a particular value of nnpfc_id and an NNPFA SEI message with nnpfa_target_id equal to a particular value of nnpfc_id, the NNPFC SEI message shall precede the NNPFA SEI message in decoding order. nnpfa_cancel_flag equal to 1 indicates that the persistence of the target neural network post-processing filter established by any previous NNPFA SEI message with the same nnpfa_target_id as the current SEI message is canceled, i.e., the target neural network post-processing filter will no longer be used unless activated by another NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 0. nnpfa_cancel_flag equal to 0 indicates that nnpfa_persistence_flag follows. nnpfa_persistence_flag specifies the persistence of the target neural network post-processing filter for the current layer. nnpfa_persistence_flag equal to 0 specifies that the target neural network post-processing filter may be used only for post-processing filtering for the current picture. nnpfa_persistence_flag equal to 1 specifies that the target neural network post-processing filter may be used for post-processing filtering for the current picture and all subsequent pictures in the current layer in output order until one or more of the following conditions become true: -A new CLVS for the current layer begins. - The bitstream ends. - The picture in the current layer associated with an NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1 is the output following the current picture in output order. NOTE - The target neural network post-processing filter is not applied to this subsequent picture in the current layer associated with an NNPFA SEI message that has the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1.
[0081] The neural network postfilter characteristics SEI message specified in JVET-AB2006 may be less than ideal. In particular, for example, the signaling in JVET-AB2006 may be insufficient to support frame rate upsampling per sublayer. In accordance with the techniques described herein, syntax and semantics are provided to provide better support for frame rate upsampling for temporal scalability.
[0082] FIG. 1 is a block diagram illustrating an example of a system that may be configured to code (i.e., encode and / or decode) video data in accordance with one or more techniques of this disclosure. System 100 represents an example of a system that may encapsulate video data in accordance with one or more techniques of this disclosure. As shown in FIG. 1, system 100 includes a source device 102, a communication medium 110, and a destination device 120. In the example shown in FIG. 1, source device 102 may include any device configured to encode video data and transmit the encoded video data to communication medium 110. Destination device 120 may include any device configured to receive the encoded video data via communication medium 110 and decode the encoded video data. Source device 102 and / or destination device 120 may include computing devices equipped for wired and / or wireless communication, and may include, for example, set-top boxes, digital video recorders, televisions, desktop, laptop, or tablet computers, gaming consoles, medical imaging devices, and mobile devices including, for example, smartphones, cellular telephones, and personal gaming devices.
[0083] The communication medium 110 may include any combination of wireless and wired communication media and / or storage devices. The communication medium 110 may include coaxial cables, fiber optic cables, twisted pair cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other equipment that may be useful for facilitating communication between various devices and sites. The communication medium 110 may include one or more networks. For example, the communication medium 110 may include a network configured to enable access to the World Wide Web, e.g., the Internet. The network may operate according to a combination of one or more telecommunications protocols. The telecommunications protocols may include proprietary aspects and / or standardized telecommunications protocols. Examples of standardized telecommunications protocols include Digital Video Broadcasting (DVB) standards, Advanced Television Systems Committee (ATSC) standards, Integrated Services Digital Broadcasting (ISDB) standards, Data Over Cable Service Interface Specification (DOCSIS) standards, Global System Mobile Communications (GSM) standards, code division multiple access (CDMA) standards, 3rd Generation Partnership Project (3GPP) standards, European Telecommunications Standards Institute (ETSI) standards, Internet Protocol (IP) standards, Wireless Application Protocol (WAP) standards, and Institute of Electrical and Electronics Engineers (IEEE) standards.
[0084] A storage device may include any type of device or storage medium capable of storing data. A storage medium may include a tangible or non-transitory computer-readable medium. A computer-readable medium may include an optical disk, a flash memory, a magnetic memory, or any other suitable digital storage medium. In some examples, a memory device or a portion thereof may be described as a non-volatile memory, and in other examples, a portion of a memory device may be described as a volatile memory. Examples of volatile memory may include random access memory (RAM), dynamic random access memory (DRAM), and static random access memory (SRAM). Examples of non-volatile memory may include a magnetic hard disk, an optical disk, a floppy disk, a flash memory, or a form of electrically programmable memory (EPROM) or electrically erasable and programmable memory (EEPROM). The storage device(s) may include memory cards (e.g., Secure Digital (SD) memory cards), internal / external hard disk drives, and / or internal / external solid state drives. Data may be stored on the storage device according to a defined file format.
[0085] FIG. 4 is a conceptual diagram illustrating an example of components that may be included in one implementation of system 100. In the exemplary implementation shown in FIG. 4, system 100 includes one or more computing devices 402A-402N, a television service network 404, a television service provider site 406, a wide area network 408, a local area network 410, and one or more content provider sites 412A-412N. The implementation shown in FIG. 4 represents an example of a system that may be configured to enable digital media content, such as movies, live sporting events, and data and applications and their associated media presentations, to be distributed to and accessed by multiple computing devices, such as computing devices 402A-402N. In the example shown in FIG. 4, computing devices 402A-402N may include any device configured to receive data from one or more of television service network 404, wide area network 408, and / or local area network 410. For example, the computing devices 402A-402N may be equipped for wired and / or wireless communication and may be configured to receive services over one or more data channels, and may include televisions, including so-called smart televisions, set-top boxes, and digital video recorders. Additionally, the computing devices 402A-402N may include desktop, laptop, or tablet computers, gaming consoles, mobile devices, including, for example, "smart" phones, cellular phones, and personal gaming devices.
[0086] The television service network 404 is an example of a network configured to enable delivery of digital media content, which may include television services. For example, the television service network 404 may include a public terrestrial television network, a public or subscription-based satellite television service provider network, and a public or subscription-based cable television provider network and / or an over the top service provider or an Internet service provider. It should be noted that, in some examples, the television service network 404 may be primarily used to enable the provision of television services, but the television service network 404 may also enable the provision of other types of data and services based on any combination of telecommunication protocols described herein. Furthermore, it should be noted that in some examples, the television service network 404 may enable bidirectional communication between the television service provider site 406 and one or more of the computing devices 402A-402N. The television service network 404 may include any combination of wireless communication media and / or wired communication media. The television service network 404 may include coaxial cables, fiber optic cables, twisted pair cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other equipment that may be useful for facilitating communication between various devices and sites. The television service network 404 may operate according to a combination of one or more telecommunications protocols. The telecommunications protocols may include proprietary aspects and / or standardized telecommunications protocols. Examples of standardized telecommunications protocols include the DVB standard, the ATSC standard, the ISDB standard, the DTMB standard, the DMB standard, the Data Over Cable Service Interface Specification (DOCSIS) standard, the HbbTV standard, the W3C standard, and the UPnP standard.
[0087] Referring again to FIG. 4, the television service provider site 406 may be configured to distribute television services over the television service network 404. For example, the television service provider site 406 may include one or more broadcast stations, cable television providers, or satellite television providers, or Internet-based television providers. For example, the television service provider site 406 may be configured to receive transmissions including television programs over satellite uplinks / downlinks. Additionally, as shown in FIG. 4, the television service provider site 406 may be in communication with a wide area network 408 and configured to receive data from content provider sites 412A-412N. It should be noted that in some examples, the television service provider site 406 may include a television studio from which content may originate.
[0088] The wide area network 408 may include a packet-based network and may operate according to a combination of one or more telecommunications protocols. The telecommunications protocols may include proprietary aspects and / or standardized telecommunications protocols. Examples of standardized telecommunications protocols include the Global System Mobile Communications (GSM) standard, the code division multiple access (CDMA) standard, the Third Generation Partnership Project (3GPP) standard, and the IEEE 802.11 standard. rdExamples of standards that may be used include 3GPP (3rd Generation Partnership Project) standards, European Telecommunications Standards Institute (ETSI) standards, European Standards (EN), IP standards, Wireless Application Protocol (WAP) standards, and Institute of Electrical and Electronics Engineers (IEEE) standards, such as one or more of the IEEE 802 standards (e.g., Wi-Fi). Wide area network 408 may include any combination of wireless communication media and / or wired communication media. Wide area network 408 may include coaxial cables, fiber optic cables, twisted pair cables, Ethernet cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other equipment that may be useful for facilitating communication between various devices and sites. In one embodiment, wide area network 408 may include the Internet. Local area network 410 may include packet-based networks and operate according to a combination of one or more telecommunications protocols. The local area networks 410 may be distinguished from the wide area networks 408 based on the level of access and / or physical infrastructure. For example, the local area networks 410 may include a secure home network.
[0089] Referring again to FIG. 4, the content provider sites 412A-412N represent examples of sites that may provide multimedia content to the television service provider site 406 and / or the computing devices 402A-402N. For example, the content provider sites may include studios having one or more studio content servers configured to provide multimedia files and / or streams to the television service provider site 406. In one embodiment, the content provider sites 412A-412N may be configured to provide multimedia content using an IP suite. For example, the content provider sites may be configured to provide multimedia content to receiving devices according to the Real Time Streaming Protocol (RTSP), HTTP, or the like. Additionally, the content provider sites 412A-412N may be configured to provide data, including hypertext-based content, or the like, over the wide area network 408 to one or more of the receiving devices, the computing devices 402A-402N, and / or the television service provider site 406. The content provider sites 412A-412N may include one or more web servers. The data provided by the content provider sites 412A-412N may be defined according to a data format.
[0090] Referring again to FIG. 1, source device 102 includes video source 104, video encoder 106, data encapsulator 107, and interface 108. Video source 104 may include any device configured to capture and / or store video data. For example, video source 104 may include a video camera and a storage device operatively coupled thereto. Video encoder 106 may include any device configured to receive video data and generate a compliant bitstream representing the video data. A compliant bitstream may refer to a bitstream that a video decoder device can receive and from which the video data can be regenerated. The aspects of a compliant bitstream may be defined according to a video encoding standard. When generating a compliant bitstream, video encoder 106 may compress the video data. The compression may be lossy (perceptible or imperceptible to a viewer) or lossless. FIG. 5 is a block diagram illustrating an example of a video encoder 500 that may implement techniques for encoding video data described herein. It should be noted that while the exemplary video encoding device 500 is illustrated as having distinct functional blocks, such illustration is for purposes of explanation and is not intended to limit the video encoding device 500 and / or its subcomponents to any particular hardware or software architecture. The functionality of the video encoding device 500 may be realized using any combination of hardware, firmware, and / or software implementations.
[0091] The video encoding device 500 may perform intra-predictive and inter-predictive encoding of picture portions, and may therefore be referred to as a hybrid video encoding device. In the example shown in FIG. 5, the video encoding device 500 receives a source video block. In some examples, the source video block may include a portion of a picture that has been partitioned according to a coding structure. For example, the source video data may include macroblocks, CTUs, CBs, subdivisions thereof, and / or other equivalent coding units. In some examples, the video encoding device 500 may be configured to perform additional subdivisions of the source video block. It should be noted that the techniques described herein are generally applicable to video encoding, regardless of how the source video data is partitioned before and / or during encoding. In the example shown in Figure 5, the video encoding device 500 includes an adder 502, a transform coefficient generating unit 504, a coefficient quantization unit 506, an inverse quantization and transform coefficient processing unit 508, an adder 510, an intra prediction processing unit 512, an inter prediction processing unit 514, a filter unit 516, and an entropy encoding unit 518. As shown in Figure 5, the video encoding device 500 receives source video blocks and outputs a bitstream.
[0092] In the example shown in FIG. 5, the video encoding device 500 may generate residual data by subtracting a predictive video block from a source video block. Selection of the predictive video block is described in more detail below. The adder 502 represents a component configured to perform this subtraction operation. In one embodiment, the subtraction of the video blocks is performed in the pixel domain. The transform coefficient generator 504 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block or a subdivision thereof (e.g., four 8×8 transforms may be applied to a 16×16 array of residual values) to generate a set of residual transform coefficients. The transform coefficient generator 504 may be configured to perform any and all combinations of transforms included in the family of discrete triangular transforms, including approximations of the discrete triangular transform. The transform coefficient generator 504 may output the transform coefficients to a coefficient quantizer 506. The coefficient quantizer 506 may be configured to perform quantization of the transform coefficients. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may change the rate distortion (i.e., video bit rate vs. quality) of the coded video data. The degree of quantization may be changed by adjusting a quantization parameter (QP). The quantization parameter may be determined based on slice level values and / or CU level values (e.g., CU delta QP value). QP data may include any data used to determine a QP for quantizing a particular set of transform coefficients. As shown in FIG. 5, the quantized transform coefficients (which may be referred to as level values) are output to an inverse quantization and transform coefficient processing unit 508. The inverse quantization and transform coefficient processing unit 508 may be configured to apply inverse quantization and inverse transform to generate reconstructed residual data. As shown in FIG. 5, the reconstructed residual data may be added to a prediction video block at an adder 510. In this manner, the coded video block may be reconstructed, and the resulting reconstructed video block may be used to evaluate the encoding quality for a given prediction, transform, and / or quantization.The video encoding device 500 may be configured to perform multiple encoding passes (e.g., performing encoding while varying one or more of prediction, transformation parameters, and quantization parameters). The rate-distortion or other system parameters of the bitstream may be optimized based on evaluation of the reconstructed video blocks. Additionally, the reconstructed video blocks may be stored and used as references for predicting subsequent blocks.
[0093] Referring to FIG. 5, the intra-predictor 512 may be configured to select an intra-prediction mode for a video block to be coded. The intra-predictor 512 may be configured to evaluate a frame and determine an intra-prediction mode to use to code a current block. As described above, possible intra-prediction modes may include a planar prediction mode, a DC prediction mode, and an angular prediction mode. Additionally, it should be noted that in some examples, a prediction mode for a chroma component may be inferred from a prediction mode for a luma prediction mode. The intra-predictor 512 may select an intra-prediction mode after performing one or more coding passes. Additionally, in one embodiment, the intra-predictor 512 may select a prediction mode based on a rate-distortion analysis. As shown in FIG. 5, the intra-predictor 512 outputs intra-prediction data (e.g., syntax elements) to the entropy encoder 518 and the transform coefficient generator 504. As described above, the transform performed on the residual data may be mode-dependent (e.g., a secondary transform matrix may be determined based on the prediction mode).
[0094] Referring again to FIG. 5, the inter prediction processor 514 may be configured to perform inter prediction coding on the current video block. The inter prediction processor 514 may be configured to receive a source video block and calculate a motion vector for a PU of the video block. The motion vector may indicate a displacement of a prediction unit of the video block in the current video frame relative to a prediction block in a reference frame. The inter prediction coding may use one or more reference pictures. Furthermore, the motion prediction may be uni-predictive (using one motion vector) or bi-predictive (using two motion vectors). The inter prediction processor 514 may be configured to select a prediction block by calculating pixel differences determined by, for example, sum of absolute difference (SAD), sum of square difference (SSD), or other difference measures. As described above, a motion vector may be determined and determined according to the motion vector prediction. The inter prediction processor 514 may be configured to perform motion vector prediction as described above. The inter prediction processor 514 may be configured to generate a prediction block using the motion prediction data. For example, the inter prediction processor 514 may place the prediction video block in a frame buffer (not shown in FIG. 5). Note that the inter prediction processor 514 may be further configured to apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values for use in motion prediction. The inter prediction processor 514 may output motion prediction data for the calculated motion vectors to the entropy encoder 518.
[0095] Referring back to FIG. 5, the filter unit 516 receives the reconstructed video blocks and the coding parameters and outputs modified reconstructed video data. The filter unit 516 may be configured to perform deblocking and / or sample adaptive offset (SAO) filtering. SAO filtering is a non-linear amplitude mapping that may be used to improve reconstruction by adding an offset to the reconstructed video data. It should be noted that, as shown in FIG. 5, the intra-prediction unit 512 and the inter-prediction unit 514 may receive the modified reconstructed video blocks via the filter unit 216. The entropy encoder 518 receives the quantized transform coefficients and prediction syntax data (i.e., intra-prediction data and motion prediction data). It should be noted that in some examples, the coefficient quantizer 506 may perform a scan of a matrix including the quantized transform coefficients before the coefficients are output to the entropy encoder 518. In other examples, the entropy encoder 518 may perform a scan. The entropy encoder 518 may be configured to perform entropy encoding according to one or more of the techniques described herein. Thus, video encoding apparatus 500 represents one example of a device configured to generate encoded video data in accordance with one or more techniques of this disclosure.
[0096] Referring again to FIG. 1, data encapsulator 107 may receive encoded video data and generate a compliant bitstream, such as a series of NAL units, according to a defined data structure. A device receiving the compliant bitstream may regenerate the video data therefrom. Additionally, as described above, sub-bitstream extraction may refer to a process in which a device receiving a compliant bitstream forms a new compliant bitstream by discarding and / or modifying data in the received bitstream. Note that the term conforming bitstream may be used instead of the term compliant bitstream. In one embodiment, data encapsulator 107 may be configured to generate syntax according to one or more techniques described herein. Note that data encapsulator 107 need not be located in the same physical device as video encoder 106. For example, the functions described as being performed by video encoder 106 and data encapsulator 107 may be distributed between the devices shown in FIG. 4.
[0097] As explained above, the signaling in JVET-AB2006 may be insufficient to support frame rate upsampling per sublayer. In accordance with the techniques herein, a syntax is provided to specify the number of interpolated pictures generated by a post-processing filter when a temporal sublayer is dropped. That is, in one embodiment, in accordance with the techniques herein, instead of the syntax element nnpfc_interpolated_pics[i][j], the syntax element nnpfc_interpolated_pics[i] is used to provide better temporal scalability support.
[0098] As an illustrative example, the original bitstream may be a 60 frames per second (fps) bitstream and may include two temporal sublayers, namely, the 0th sublayer and the 1st sublayer, where the 0th sublayer is at 30 fps and the 0th and 1st sublayers together are at 60 fps. Furthermore, an NNPFC message according to JVET-AB2006 may signal nnpfc_num_input_pics_minus2 equal to 0 and nnpfc_interpolated_pics[0] equal to 1 for the original 60 fps bitstream, such that the interpolated bitstream and the original bitstream together result in a set of pictures with 120 fps. In this case, if the top temporal sublayer is dropped from the original bitstream such that the original bitstream becomes a 30 fps bitstream, the value of nnpfc_interpolated_pics[0] must be changed to 3 to achieve 120 fps. JVET-AB2006 does not provide a mechanism for changing the value of nnpfc_interpolated_pics[0] in such a case. In accordance with the techniques herein, an additional loop would allow for signaling nnpfc_interpolated_pics[0][1]=1 and nnpfc_interpolated_pics[0][0]=3 in this case.
[0099] Tables 13A and 13B show the relevant syntax of an example Neural Network Post Filter Characteristics SEI message for specifying the number of interpolated pictures generated by a post-processing filter when a temporal sub-layer is dropped, according to the techniques herein. With respect to Table 13A, it is noted that JVET-T2001 provides where the syntax element sps_max_sublayers_minus1 is included in the sequence parameter set syntax, seq_parameter_set_rbsp(), and has the following semantics: sps_max_sublayers_minus1 plus 1 specifies the maximum number of temporal sublayers that may exist in each CLVS that references an SPS. If sps_video_parameter_set_id is greater than 0, the value of sps_max_sublayers_minus1 shall be in the range of 0 to vps_max_sublayers_minus1, inclusive. Otherwise (sps_video_parameter_set_id is equal to 0), the following applies: The value of -sps_max_sublayers_minus1 must be in the range of 0 to 6, inclusive. The value of -vps_max_sublayers_minus1 is inferred to be equal to sps_max_sublayers_minus1. The value of -NumSubLayersInLayerInOLS[0][0] is inferred to be equal to sps_max_sublayers_minus1+1. - The value of vps_ols_ptl_idx[0] is inferred to be equal to 0, and the value of vps_ptl_max_tid[vps_ols_ptl_idx[0]], i.e. vps_ptl_max_tid[0], is inferred to be equal to sps_max_sublayers_minus1. Where: vps_max_sublayers_minus1 plus 1 specifies the maximum number of temporal sublayers that may exist in a layer specified by the VPS. The value of vps_max_sublayers_minus1 shall be in the range 0 to 6, inclusive.
[0100] [Table 22]
[0101] [Table 23]
[0102] With respect to Tables 13A and 13B, the semantics may be based on those provided above and the following: The neural-network post-filter characteristics (NNPFC) SEI message specifies neural networks that can be used as post-processing filters. The use of a specified post-processing filter for a particular picture is indicated in the neural-network post-filter activation SEI message. Use of this SEI message requires the definition of the following variables: The width and height of the cropped decoded output picture in units of luma samples, denoted herein by CroppedWidth and CroppedHeight, respectively. - When present, the luma sample array CroppedYPic[idx] and the chroma sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] of the cropped decoded output picture with idx in the range of 0 to numInputPics-1, inclusive, that are used as input for the post-processing filters. - the highest decoded temporal sublayer HighestTid (or, in one embodiment, the highest temporal sublayer present HighestTid) -BitDepth - the bit depth of the luma sample array of the cropped decoded output picture Y . -BitDepth - the bit depth of the chroma sample array, if any, of the cropped decoded output picture C . A chroma format indicator, denoted herein by ChromaFormatIdc. When -nnpfc_auxiliary_inp_idc is equal to 1, the filtering strength control value StrengthControlVal shall be a real number in the range of 0 to 1, inclusive. nnpfc_interpolated_pics[i][j] specifies the number of interpolated pictures generated by the post-processing filter between the i-th picture and the (i+1)-th picture used as input for the post-processing filter, where j is the highest temporal sublayer present in the bitstream or decoded. The variables numInputPics, which specify the number of pictures used as input for the post-processing filter, and numOutputPics, which specify the total number of pictures resulting from the post-processing filter, are derived as follows: Let Htid be the top-level Tid.
[0103] [Table 24] Now, about JVET-T2001: The variable Htid, which identifies the highest temporal sub-layer being decoded, is derived as follows. In the first AU of the bitstream the following applies: - If some external means not specified in this specification are available for setting the Htid, the Htid is set by that external means. Otherwise, if opi_htid_plus1 is present in an OPI NAL unit in the first AU of the bitstream, then Htid is set equal to ((opi_htid_plus1>0)? opi_htid_plus1-1:0). - Otherwise, Htid is set equal to vps_ptl_max_tid[vps_ols_ptl_idx[TargetOlsIdx]]. NOTE - When sps_video_parameter_set_id is equal to 0, Htid is set equal to sps_max_sublayers_minus1. In another embodiment, the code above can use a variable HighestTid to indicate the highest Tid, and the code fragment above can be written as: numOutputPics += nnpfc_interpolated_pics[i][Htid] from numOutputPics += nnpfc_interpolated_pics[i][HighestTid] can be changed to. Here, the syntax element opi_htid_plus1 is included in the operating point information syntax operating_point_information_rbsp() and has the following semantics: opi_htid_plus1 equal to 0 specifies that all pictures in the current CVS and all pictures in the next CVS up to but not including the next CVS for which opi_htid_plus1 is provided in decoding order in the OPI NAL unit are IRAP or GDR pictures with ph_recovery_poc_cnt equal to 0. opi_htid_plus1 greater than 0 specifies that all pictures in the current CVS and all pictures in the next CVS up to but not including the next CVS for which opi_htid_plus1 is provided in decoding order in the OPI NAL unit have TemporalId less than opi_htid_plus1. Here, the syntax element vps_ptl_max is included in the video parameter set syntax video_parameter_set_rbsp() and has the following semantics: vps_ptl_max_tid[i] specifies the TemporalId of the highest sublayer representation for which level information exists in the i-th profile_tier_level() syntax structure in the VPS and the TemporalId of the highest sublayer representation present in the OLS with OLS index olsIdx, such that vps_ols_ptl_idx[olsIdx] is equal to i. The value of vps_ptl_max_tid[i] shall be in the range 0 to vps_max_sublayers_minus1, inclusive. When vps_default_ptl_dpb_hrd_max_tid_flag is equal to 1, the value of vps_ptl_max_tid[i] is inferred to be equal to vps_max_sublayers_minus1.
[0104] With regard to Table 13B, note that the input variable NumSubLayerMinus1, which provides the number of temporal sublayers minus 1, may alternatively be referred to as MaxNumSubLayersMinus1 or NumSubLayersInLayerInOLS, where MaxNumSubLayersMinus1 or NumSubLayersInLayerInOLS may be based on the definitions provided in JVET-T2001.
[0105] According to JVET-AB2006, when nnpfc_constant_patch_size_flag is equal to 0 (i.e., when the post-processing filter accepts any patch size that is a positive integer multiple of the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1), the overlapping horizontal sample count and vertical sample count of adjacent input tensors of the post-processing filter need to be considered with respect to the variables inpPatchWidth and inpPatchHeight representing the patch size width and patch size height respectively. According to JVET-AB2006, the actual input size to the post-processing filter defined by the process DeriveInputTensors() is calculated by a for loop defined for yP values in the range of yP = -overlapSize ~ yP < inpPatchHeight + overlapSize and xP values in the range of xP = -overlapSize ~ xP < inpPatchWidth + overlapSize. This can result in the actual input size to the post-processing filter not being a multiple of a specific value (e.g., 8) as might be required by a typical neural network post-filter.
[0106] For example, according to JVET-AB2006, there may be a case where nnpfc_patch_width_minus1 + 1 = 8, nnpfc_patch_height_minus1 + 1 = 8, overlapSize = 3, and nnpfc_constant_patch_size_flag = 0. In this case, the filter accepts inputs of sizes that are multiples of (8,8). Further, InpPatchWidth = (nnpfc_patch_width_minus1 + 1) * k (i.e., 8 * k) and InpPatchHeight = (nnpfc_patch_height_minus1 + 1) * k (i.e., 8 * k), the input to the post-processing filter is (8 * k + 2 *overlapSize,8 * k+2 * overlapsize) (i.e. (8 * k+6,8 * k+6). Value 8 * Note that k+6 is not necessarily divisible by 8, which may lead to unexpected and / or erroneous filtering results.
[0107] In accordance with the techniques herein, an exemplary neural network post-filter characteristics SEI message may include syntax elements having semantics that allow for ensuring that the actual input size to a post-processing filter is a multiple of a particular value that may be required by a typical neural network post-filter. In one embodiment, in accordance with the techniques herein, an exemplary neural network post-filter characteristics SEI message may be based on the syntax and semantics provided in the examples of Table 6 and / or Tables 13A-13B, as well as the following: nnpfc_constant_patch_size_flag equal to 1 indicates that the post-processing filter accepts as input exactly the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. nnpfc_constant_patch_size_flag equal to 0 indicates that the post-processing filter accepts as input any patch size that is a positive integer multiple of the patch size indicated by nnpfc_patch_width_minus1, nnpfc_patch_height_minus1, and nnpfc_overlap. nnpfc_patch_height_minus1+1 indicates the vertical sample count of the patch size required for input to the post-processing filter when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_height_minus1 shall be in the range 0 to Min(32766,CroppedHeight-1), inclusive. The variables inpPatchWidth and inpPatchHeight are the patch size width and patch size height, respectively. If nnpfc_constant_patch_size_flag is equal to 0, the following applies: The values of -inpPatchWidth and inpPatchHeight are either provided by external means not specified in this document, or are set by the post-processor itself. -inpPatchWidth+2 * The value of overlapSize shall be a positive integer multiple of nnpfc_patch_width_minus1+1 and shall be less than or equal to CroppedWidth. * The value of overlapSize shall be a positive integer multiple of nnpfc_patch_height_minus1+1 and less than or equal to CroppedHeight. Otherwise (nnpfc_constant_patch_size_flag is equal to 1), the value of inpPatchWidth is set equal to nnpfc_patch_width_minus1+1, and the value of inpPatchHeight is set equal to nnpfc_patch_height_minus1+1.
[0108] In one embodiment, nnpfc_patch_height_minus1 may be based on the following: nnpfc_patch_height_minus1+1 indicates the vertical sample count of the patch size required for input to the post-processing filter when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_height_minus1 shall be in the range 0 to Min(32766,CroppedHeight-1), inclusive. The variables inpPatchWidth and inpPatchHeight are the patch size width and patch size height, respectively. If nnpfc_constant_patch_size_flag is equal to 0, the following applies: The values of -inpPatchWidth and inpPatchHeight are either provided by external means not specified in this document, or are set by the post-processor itself. -inpPatchWidth+2 * The value of overlapSize is a positive integer multiple of nnpfc_patch_width_minus1+1. In addition, inpPatchWidth is equal to or less than CroppedWidth. -inpPatchHeight+2 * The value of overlapSize shall be a positive integer multiple of nnpfc_patch_height_minus1+1. In addition, inpPatchHeight shall be less than or equal to CroppedHeight. Otherwise (nnpfc_constant_patch_size_flag is equal to 1), the value of inpPatchWidth is set equal to nnpfc_patch_width_minus1+1, and the value of inpPatchHeight is set equal to nnpfc_patch_height_minus1+1.
[0109] In one embodiment, nnpfc_patch_height_minus1 may be based on the following: nnpfc_patch_height_minus1+1 indicates the vertical sample count of the patch size required for input to the post-processing filter when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_height_minus1 shall be in the range 0 to Min(32766,CroppedHeight-1), inclusive. The variables inpPatchWidth and inpPatchHeight are the patch size width and patch size height, respectively. If nnpfc_constant_patch_size_flag is equal to 0, the following applies: The values of -inpPatchWidth and inpPatchHeight are either provided by external means not specified in this document, or are set by the post-processor itself. -The value of inpPatchWidth is nnpfc_patch_width_minus1+1-2 * It is a positive integer multiple of overlapSize and is less than or equal to CroppedWidth. The value of inpPatchHeight is nnpfc_patch_height_minus1+1-2 * It must be a positive integer multiple of overlapSize and less than or equal to CroppedHeight. Otherwise (nnpfc_constant_patch_size_flag is equal to 1), the value of inpPatchWidth is set equal to nnpfc_patch_width_minus1+1, and the value of inpPatchHeight is set equal to nnpfc_patch_height_minus1+1.
[0110] In one embodiment, in accordance with the techniques herein, an exemplary neural network post-filter characteristics SEI message may be based on the syntax and semantics provided in the examples of Table 6 and / or Tables 13A-13B, as well as the following: nnpfc_constant_patch_size_flag equal to 1 indicates that the post-processing filter accepts as input exactly the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. nnpfc_constant_patch_size_flag equal to 0 indicates that the post-processing filter accepts as input any patch size that is a positive integer power of the patch size indicated by nnpfc_patch_width_minus1, nnpfc_patch_height_minus1, and nnpfc_overlap. nnpfc_patch_height_minus1+1 indicates the vertical sample count of the patch size required for input to the post-processing filter when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_height_minus1 shall be in the range 0 to Min(32766,CroppedHeight-1), inclusive. The variables inpPatchWidth and inpPatchHeight are the patch size width and patch size height, respectively. If nnpfc_constant_patch_size_flag is equal to 0, the following applies: The values of -inpPatchWidth and inpPatchHeight are either provided by external means not specified in this document, or are set by the post-processor itself. -inpPatchWidth+2 * The value of overlapSize shall be a positive integer power of nnpfc_patch_width_minus1+1 and shall be less than or equal to CroppedWidth. * The value of overlapSize shall be a positive integer power of nnpfc_patch_height_minus1+1 and shall be less than or equal to CroppedHeight. Otherwise (nnpfc_constant_patch_size_flag is equal to 1), the value of inpPatchWidth is set equal to nnpfc_patch_width_minus1+1, and the value of inpPatchHeight is set equal to nnpfc_patch_height_minus1+1.
[0111] In this manner, video encoding apparatus 500 represents an example of a device configured to signal a neural network postfilter characteristic message and to signal, in the neural network postfilter characteristic message, a syntax element that specifies, for each of a number of sublayers associated with the neural network postfilter characteristic message, the number of interpolated pictures generated by a post-processing filter.
[0112] 1, interface 108 may include any device configured to receive data generated by data encapsulator 107 and transmit and / or store the data on a communications medium. Interface 108 may include a network interface card, such as an Ethernet card, an optical transceiver, a radio frequency transceiver, or any other type of device capable of transmitting and / or receiving information. Additionally, interface 108 may include a computer system interface that may allow files to be stored on a storage device. For example, interface 108 may include a Peripheral Component Interconnect (PCI) and Peripheral Component Interconnect Express (PCIe) bus protocols, a proprietary bus protocol, a Universal Serial Bus (USB) protocol, an I / O interface, or any other type of device ... 2 C, or any other logical and physical structure that may be used to interconnect peer devices.
[0113] 1, destination device 120 includes interface 122, data decapsulator 123, video decoder 124, and display 126. Interface 122 may include any device configured to receive data from a communication medium. Interface 122 may include a network interface card, such as an Ethernet card, an optical transceiver, a radio frequency transceiver, or any other type of device capable of receiving and / or transmitting information. Additionally, interface 122 may include an interface for a computer system that allows a compliant video bitstream to be obtained from a storage device. For example, interface 122 may include interfaces for PCI and PCIe bus protocols, proprietary bus protocols, USB protocols, I / O protocols, and the like. 2 C, or any other logical and physical structures that may be used to interconnect peer devices. The data decapsulator 123 may be configured to receive and parse any of the example syntax structures described herein.
[0114] Video decoder 124 may include any device configured to receive a bitstream (e.g., a sub-bitstream extract) and / or an acceptable variant thereof and regenerate video data therefrom. Display 126 may include any device configured to display video data. Display 126 may include one of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display. Display 126 may include a high-definition display or an ultra-high-definition display. Although in the example shown in FIG. 1, video decoder 124 is described as outputting data to display 126, it should be noted that video decoder 124 may be configured to output video data to various types of devices and / or subcomponents thereof. For example, video decoder 124 may be configured to output video data to any communication medium as described herein.
[0115] FIG. 6 is a block diagram illustrating an example of a video decoding device that may be configured to decode video data according to one or more techniques of this disclosure (e.g., the reference picture list creation decoding process described above). In one embodiment, video decoding device 600 may be configured to decode transform data and recover residual data from transform coefficients based on the decoded transform data. Video decoding device 600 may be configured to perform intra-prediction decoding and inter-prediction decoding and may therefore be referred to as a hybrid decoding device. Video decoding device 600 may be configured to parse any combination of the syntax elements described above in Tables 1-13B. Video decoding device 600 may decode video based on or in accordance with the above process and further based on the parsed values in Tables 1-13B.
[0116] In the example shown in FIG. 6, the video decoding device 600 includes an entropy decoding unit 602, an inverse quantization unit 604, an inverse transform coefficient processing unit 606, an intra prediction processing unit 608, an inter prediction processing unit 610, an adder 612, a post filter unit 614, and a reference buffer 616. The video decoding device 600 may be configured to decode video data in a manner consistent with a video encoding system. It should be noted that while the exemplary video decoding device 600 is shown having separate functional blocks, such illustration is for illustrative purposes and does not limit the video decoding device 600 and / or its subcomponents to a particular hardware or software architecture. The functionality of the video decoding device 600 may be realized using any combination of hardware, firmware, and / or software implementations.
[0117] As shown in FIG. 6, the entropy decoding unit 602 receives an entropy coded bitstream. The entropy decoding unit 602 may be configured to decode syntax elements and quantized coefficients from the bitstream according to a reverse process of the entropy coding process. The entropy decoding unit 602 may be configured to perform entropy decoding according to any of the entropy coding techniques described above. The entropy decoding unit 602 may determine values of syntax elements in the coded bitstream in accordance with a video coding standard. As shown in FIG. 6, the entropy decoding unit 602 may determine quantization parameters, values of quantized coefficients, transform data, and prediction data from the bitstream. In the example shown in FIG. 6, the inverse quantization unit 604 and the inverse transform coefficient processing unit 606 receive values of quantized coefficients from the entropy decoding unit 602 and output reconstructed residual data.
[0118] Referring back to FIG. 6, the reconstructed residual data may be provided to an adder 612. The adder 612 may add the reconstructed residual data to a predictive video block to generate reconstructed video data. The predictive video block may be determined according to a predictive video technique (i.e., intra prediction and inter-frame prediction). The intra prediction processor 608 may be configured to receive the intra prediction syntax element and retrieve the predictive video block from a reference buffer 616. The reference buffer 616 may include a memory device configured to store one or more frames of video data. The intra prediction syntax element may identify an intra prediction mode, such as the intra prediction modes described above. The inter prediction processor 610 may receive the inter prediction syntax element and generate a motion vector to identify a predictive block in one or more reference frames stored in the reference buffer 616. The inter prediction processor 610 may perform interpolation, possibly based on an interpolation filter, to generate a motion compensated block. The syntax element may include an identifier of an interpolation filter to be used for motion prediction having sub-pixel accuracy. The inter-prediction processor 610 may use an interpolation filter to calculate interpolated values for sub-integer pixels of the reference block. The post-filter unit 614 may be configured to perform filtering on the reconstructed video data. For example, the post-filter unit 614 may be configured to perform deblocking and / or sample adaptive offset (SAO) filtering, for example, based on parameters specified in the bitstream. It should be noted that in some examples, the post-filter unit 614 may also be configured to perform its own arbitrary filtering (e.g., visual enhancement such as mosquito noise reduction). As shown in FIG. 6, the reconstructed video blocks may be output by the video decoding device 600.In this manner, video decoding apparatus 600 represents an example of a device configured to receive a neural network post-filter characteristic message and parse from the neural network post-filter characteristic message a syntax element that specifies the number of interpolated pictures generated by a post-processing filter for each of a number of sub-layers associated with the neural network post-filter characteristic message.
[0119] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions may be stored or transmitted as one or more instructions or code on a computer-readable medium and executed by a hardware-based processor. A computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium such as a data storage medium or a communication medium, including any medium that facilitates the transfer of a computer program from one place to another, for example according to a communication protocol. In this manner, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0120] By way of example, and without limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, other magnetic storage devices, flash memory, or any other medium, i.e., any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. However, it should be understood that computer-readable media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory tangible storage media. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks and Blu-ray discs, where disks typically reproduce data magnetically and discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer readable media.
[0121] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, as used herein, the term "processor" may refer to any of the foregoing structures, or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured to encode and decode, or incorporated into a composite codec. Also, the techniques may be implemented entirely in one or more circuit or logic elements.
[0122] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units are illustrated in this disclosure to highlight functional aspects of devices configured to perform the disclosed techniques, but need not necessarily be realized by different hardware units. Rather, as previously described, the various units may be combined into a codec hardware unit or may be provided by a collection of interoperating hardware units, including one or more processors as previously described, along with suitable software and / or firmware.
[0123] Furthermore, each functional block or various features of the base station device and the terminal device used in each of the above implementations may be implemented or performed by a circuit, which is typically an integrated circuit or multiple integrated circuits. A circuit designed to perform the functions described herein may comprise a general-purpose processor, a digital signal processor (DSP), an application specific or general-purpose application integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, or individual hardware components, or a combination thereof. A general-purpose processor may be a microprocessor, or the processor may be a conventional processor, controller, microcontroller, or state machine. The general-purpose processor or each circuit described above may be composed of digital circuits or analog circuits. Furthermore, as semiconductor technology advances and integrated circuit technology emerges to replace current integrated circuits, integrated circuits using this technology may also be used.
[0124] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
1. Receive a neural network post-filter characteristic message that can be used as a post-processing filter; Parse the nnpfc_constant_patch_size_flag syntax element of the neural network post-filter characteristics message; where: The nnpfc_constant_patch_size_flag syntax element equal to 1 indicates that the post-processing filter accepts as input exactly the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1; The nnpfc_constant_patch_size_flag syntax element equal to 0 indicates that the post-processing filter accepts any patch size; the width of said arbitrary patch size is equal to inpPatchWidth+2*overlapSize, which is a positive integer multiple of the specified width; The height of the arbitrary patch size is equal to inpPatchHeight+2*overlapSize, which is a positive integer multiple of the specified height; inpPatchWidth is the patch size width, inpPatchHeight is the patch size height, The overlapSize is indicated by the nnpfc_overlap syntax element of the neural network postfilter characteristics message, The nnpfc_overlap syntax element indicates the overlapping horizontal and vertical sample counts of adjacent input tensors of the post-processing filter.
10. A video decoding device comprising: one or more processors configured to:
2. Signaling a neural network post-filter characteristic message that can be used as a post-processing filter; The neural network post-filter characteristics message includes a nnpfc_constant_patch_size_flag syntax element; where: The nnpfc_constant_patch_size_flag syntax element equal to 1 indicates that the post-processing filter accepts as input exactly the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1; The nnpfc_constant_patch_size_flag syntax element equal to 0 indicates that the post-processing filter accepts any patch size; the width of said arbitrary patch size is equal to inpPatchWidth+2*overlapSize, which is a positive integer multiple of the specified width; The height of the arbitrary patch size is equal to inpPatchHeight+2*overlapSize, which is a positive integer multiple of the specified height; inpPatchWidth is the patch size width, inpPatchHeight is the patch size height, The overlapSize is indicated by the nnpfc_overlap syntax element of the neural network postfilter characteristics message, The nnpfc_overlap syntax element indicates the overlapping horizontal and vertical sample counts of adjacent input tensors of the post-processing filter.
10. A video encoding device comprising: one or more processors configured to:
3. A computer-readable storage medium storing a program for causing a computer to process neural network filtering of video data, the program causing the computer to: Parse the nnpfc_constant_patch_size_flag syntax element of the neural network postfilter characteristics message that can be used as a post-processing filter; where: The nnpfc_constant_patch_size_flag syntax element equal to 1 indicates that the post-processing filter accepts as input exactly the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1; The nnpfc_constant_patch_size_flag syntax element equal to 0 indicates that the post-processing filter accepts any patch size; the width of said arbitrary patch size is equal to inpPatchWidth+2*overlapSize, which is a positive integer multiple of the specified width; The height of the arbitrary patch size is equal to inpPatchHeight+2*overlapSize, which is a positive integer multiple of the specified height; inpPatchWidth is the patch size width, inpPatchHeight is the patch size height, The overlapSize is indicated by the nnpfc_overlap syntax element of the neural network postfilter characteristics message, The nnpfc_overlap syntax element indicates the overlapping horizontal and vertical sample counts of adjacent input tensors of the post-processing filter. A computer-readable storage medium.
4. A method for encoding video data, comprising: signaling neural network post-filter characteristic messages that may be used as post-processing filters; The neural network post-filter characteristics message includes a nnpfc_constant_patch_size_flag syntax element; where: The nnpfc_constant_patch_size_flag syntax element equal to 1 indicates that the post-processing filter accepts as input exactly the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1; The nnpfc_constant_patch_size_flag syntax element equal to 0 indicates that the post-processing filter accepts any patch size; the width of said arbitrary patch size is equal to inpPatchWidth+2*overlapSize, which is a positive integer multiple of the specified width; The height of the arbitrary patch size is equal to inpPatchHeight+2*overlapSize, which is a positive integer multiple of the specified height; inpPatchWidth is the patch size width, inpPatchHeight is the patch size height, The overlapSize is indicated by the nnpfc_overlap syntax element of the neural network postfilter characteristics message, The nnpfc_overlap syntax element indicates the overlapping horizontal and vertical sample counts of adjacent input tensors of the post-processing filter.