Video decoding device, video encoding device and transmission method
Patent Information
- Application Number
- JP2023032906
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-19
- Filing Date
- 2023-03-03
- Publication Date
- 2026-02-12
AI Technical Summary
Existing video encoding standards lack efficient methods for signaling neural network post-filter parameter information, which is crucial for enhancing video quality through post-processing techniques.
The proposed techniques involve signaling neural network post-filter characteristics using syntax elements to indicate the purpose, input formatting, output formatting, and complexity of the post-processing filter, allowing devices to effectively utilize CNN-based post-filters for improved video quality.
Enables effective utilization of CNN-based post-filters by providing clear signaling of neural network post-filter parameters, thereby enhancing video quality and addressing limitations in existing encoding standards.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] FIELD This disclosure relates to video encoding, and more particularly, to techniques for signaling neural network postfilter parameter information for encoded video. [Background technology]
[0002] Digital video capabilities may be incorporated into a wide range of devices, including digital televisions, laptop or desktop computers, tablet computers, digital recording devices, digital media players, video gaming devices, cellular phones, including so-called smart phones, medical imaging devices, and the like. Digital video may be encoded according to a video encoding standard. A video encoding standard defines a format for a compliant bitstream that encapsulates the encoded video data. A compliant bitstream is a data structure that can be received and decoded by a video decoding device to generate recovered video data. A video encoding standard may incorporate video compression techniques. Examples of video encoding standards include ISO / IEC MPEG-4 Visual and ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) and High-Efficiency Video Coding (HEVC). HEVC is described in High Efficiency Video Coding (HEVC), Rec. ITU-T H.265 (December 2016), which is incorporated herein by reference and is referred to herein as ITU-T H.265. Extensions and improvements to ITU-T H.265 are being considered for the development of next generation video coding standards. For example, the ITU-T Video Coding Experts Group (VCEG) and ISO / IEC (Moving Picture Experts Group (MPEG), collectively known as the Joint Video Exploration Team (JVET)) have standardized video coding techniques with compression capabilities that significantly exceed those of the current HEVC standard.The Joint Exploration Model 7 (JEM 7), Algorithm Description of Joint Exploration Test Model 7 (JEM 7), ISO / IEC JTC1 / SC29 / WG11 Document: JVET-G1001, July 2017, Torino, IT, describes the coding features that have been under collaborative test model study by JVET as having the potential to improve video coding technology beyond the capabilities of ITU-T H.265, and is incorporated herein by reference. It should be noted that the coding features of JEM 7 are implemented in the JEM reference software. As used herein, the term JEM may collectively refer to the algorithms included in JEM 7 and the implementation of the JEM reference software. Additionally, in response to the "Joint Call for Proposals on Video Compression with Capabilities beyond HEVC" jointly issued by VCEG and MPEG, multiple descriptions of video coding tools have been published by various groups in the past 10 years. th The initial draft text of the video coding specification, derived from multiple descriptions of video coding tools, was proposed at the Meeting of ISO / IEC JTC1 / SC29 / WG11 16-20 April 2018, San Diego, CA. thThis development of a video coding standard by VCEG and MPEG is called the Versatile Video Coding (VVC) project. "Versatile Video Coding (Draft 10)", 20th Meeting of ISO / IEC JTC1 / SC29 / WG11 7-16 October 2020, Teleconference, document JVET-T2001-v2 represents the current iteration of the draft text of the video coding specification corresponding to the VVC project, which is incorporated herein by reference and referred to as JVET-T2001.
[0003] Video compression techniques allow for reducing data requirements for storing and transmitting video data. Video compression techniques may reduce data requirements by exploiting inherent redundancy in a video sequence. Video compression techniques may subdivide a video sequence into successively smaller portions (i.e., groups of pictures within a video sequence, pictures within groups of pictures, regions within pictures, sub-regions within regions, etc.). Intra-prediction coding techniques (e.g., spatial prediction techniques within a picture) and inter-prediction techniques (i.e., inter-picture techniques (temporal)) may be used to generate difference values between a unit of video data being coded and a reference unit of video data. The difference values may be referred to as residual data. The residual data may be coded as quantized transform coefficients. Syntax elements may be associated with the residual data and the reference coding units (e.g., intra-prediction mode index, and motion information). The residual data and syntax elements may be entropy coded. The entropy coded residual data and syntax elements may be included in a data structure that forms a compliant bitstream. Summary of the Invention
[0004] In general, this disclosure describes various techniques for encoding video data. In particular, this disclosure describes techniques for signaling neural network postfilter parameter information for encoded video data. It should be noted that although the techniques of this disclosure are described with respect to ITU-T H.264, ITU-T H.265, JEM, and JVET-T2001, the techniques of this disclosure are generally applicable to video encoding. For example, the encoding techniques described herein may be incorporated into video encoding systems (including video encoding systems based on future video encoding standards) that include video block structures, intra-prediction techniques, inter-prediction techniques, transform techniques, filtering techniques, and / or entropy encoding techniques other than those included in ITU-T H.265, JEM, and JVET-T2001. Thus, references to ITU-T H.264, ITU-T H.265, JEM, and / or JVET-T2001 are for illustrative purposes and should not be construed to limit the scope of the techniques described herein. Furthermore, it should be noted that the incorporation by reference of documents herein is for explanatory purposes and should not be construed as limiting or creating ambiguity with respect to the terms used herein. For example, if an incorporated reference provides a definition of a term that differs from that of another incorporated reference and / or as that term is used herein, that term should be construed to broadly include each corresponding definition and / or to include each specific definition instead.
[0005] In one embodiment, a method for encoding video data includes signaling a neural network post-filter characteristic message; signaling in the neural network post-filter characteristic message a first syntax element indicating a purpose of a post-processing filter; and signaling in the neural network post-filter characteristic message a second syntax element specifying whether syntax elements related to the purpose, input formatting, output formatting, and complexity of the post-processing filter are present in the neural network post-filter characteristic message, where the first syntax element precedes the second syntax element in the neural network post-filter characteristic message.
[0006] In one embodiment, a device comprises one or more processors configured to signal a neural network post-filter characteristic message, signal in the neural network post-filter characteristic message a first syntax element indicating a purpose of the post-processing filter, and signal in the neural network post-filter characteristic message a second syntax element specifying whether syntax elements related to the purpose, input formatting, output formatting, and complexity of the post-processing filter are present in the neural network post-filter characteristic message, wherein the first syntax element precedes the second syntax element in the neural network post-filter characteristic message.
[0007] In one embodiment, a non-transitory computer-readable storage medium includes instructions stored thereon that, when executed, cause one or more processors of the device to signal a neural network post-filter characteristic message, signal in the neural network post-filter characteristic message a first syntax element indicating a purpose of the post-processing filter, and signal in the neural network post-filter characteristic message a second syntax element specifying whether syntax elements related to the purpose, input formatting, output formatting, and complexity of the post-processing filter are present in the neural network post-filter characteristic message, wherein the first syntax element precedes the second syntax element in the neural network post-filter characteristic message.
[0008] In one embodiment, an apparatus comprises means for signaling a neural network post-filter characteristic message; means for signaling in the neural network post-filter characteristic message a first syntax element indicating a purpose of a post-processing filter; and means for signaling in the neural network post-filter characteristic message a second syntax element specifying whether syntax elements related to the purpose, input formatting, output formatting, and complexity of the post-processing filter are present in the neural network post-filter characteristic message, where the first syntax element precedes the second syntax element in the neural network post-filter characteristic message.
[0009] In one embodiment, a method for decoding video data includes receiving a neural network post-filter characteristic message; parsing from the neural network post-filter characteristic message a first syntax element indicating a purpose of a post-processing filter; and parsing from the neural network post-filter characteristic message a second syntax element specifying whether syntax elements related to a purpose, input formatting, output formatting, and complexity of the post-processing filter are present in the neural network post-filter characteristic message, where the first syntax element precedes the second syntax element in the neural network post-filter characteristic message.
[0010] In one embodiment, a device comprises one or more processors configured to receive a neural network post-filter characteristic message; parse from the neural network post-filter characteristic message a first syntax element indicating a purpose of the post-processing filter; and parse from the neural network post-filter characteristic message a second syntax element specifying whether syntax elements related to the purpose, input formatting, output formatting, and complexity of the post-processing filter are present in the neural network post-filter characteristic message, wherein the first syntax element precedes the second syntax element in the neural network post-filter characteristic message.
[0011] In one embodiment, a non-transitory computer-readable storage medium includes instructions stored thereon that, when executed, cause one or more processors of the device to receive a neural network post-filter characteristic message, parse from the neural network post-filter characteristic message a first syntax element indicating a purpose of the post-processing filter, and parse from the neural network post-filter characteristic message a second syntax element specifying whether syntax elements related to the purpose, input formatting, output formatting, and complexity of the post-processing filter are present in the neural network post-filter characteristic message, wherein the first syntax element precedes the second syntax element in the neural network post-filter characteristic message.
[0012] In one embodiment, an apparatus comprises means for receiving a neural network post-filter characteristic message; means for parsing from the neural network post-filter characteristic message a first syntax element indicating a purpose of a post-processing filter; and means for parsing from the neural network post-filter characteristic message a second syntax element specifying whether syntax elements related to the purpose, input formatting, output formatting, and complexity of the post-processing filter are present in the neural network post-filter characteristic message, where the first syntax element precedes the second syntax element in the neural network post-filter characteristic message.
[0013] The details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims. [Brief description of the drawings]
[0014] [Figure 1] FIG. 1 is a block diagram illustrating an example of a system that may be configured to encode and decode video data in accordance with one or more techniques of this disclosure. [Diagram 2] 1 is a conceptual diagram illustrating encoded video data and corresponding data structures in accordance with one or more techniques of this disclosure. [Diagram 3] 1 is a conceptual diagram illustrating a data structure that encapsulates encoded video data and corresponding metadata in accordance with one or more techniques of this disclosure. [Figure 4] A conceptual diagram illustrating an example of components that may be included in an implementation of a system that may be configured to encode and decode video data in accordance with one or more techniques of this disclosure. [Diagram 5] FIG. 1 is a block diagram illustrating an example of a video encoding device that may be configured to encode video data in accordance with one or more techniques of this disclosure. [Figure 6] 1 is a block diagram illustrating an example of a video decoding device that may be configured to decode video data in accordance with one or more techniques of this disclosure. [Figure 7] FIG. 2 is a conceptual diagram illustrating an example of a packed data channel for a luma component, in accordance with one or more techniques of this disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] Video content includes a video sequence consisting of a series of frames (or pictures). A series of frames may also be referred to as a group of pictures (GOP). Each video frame or picture may be divided into one or more regions. A region may be defined according to a base unit (e.g., a video block) and a set of rules that define the region. For example, the rule that defines a region may be that the region must be an integer number of video blocks arranged in a rectangle. Furthermore, the video blocks in a region may be ordered according to a scan pattern (e.g., a raster scan). As used herein, the term video block may generally refer to an area of a picture, or more specifically, may refer to a maximal array of sample values that may be predictively coded, its subdivisions, and / or corresponding structures. Furthermore, the term current video block may refer to a portion of a picture being coded or decoded. A video block may be defined as an array of sample values. It should be noted that in some cases, a pixel value may be described as including sample values for each component of the video data, which may also be referred to as color components (e.g., luma component (Y) and chroma components (Cb and Cr), or red, green, and blue color components). It should be noted that in some cases, the terms pixel value and sample value are used interchangeably. Furthermore, in some cases, a pixel or sample may be referred to as a pel. A video sampling format, which may be referred to as a chroma format, may be defined as the number of chroma samples included in a video block relative to the number of luma samples included in the video block. For example, for a 4:2:0 format, the sampling rate for the luma component is twice the sampling rate of the chroma components in both the horizontal and vertical directions.
[0016] A video coding device may perform predictive coding on video blocks and their subdivisions. The video blocks and their subdivisions may be referred to as nodes. ITU-T H.264 specifies macroblocks containing 16x16 luma samples. That is, in ITU-T H.264, a picture is divided into macroblocks. ITU-T H.265 specifies a similar coding tree unit (CTU) structure, sometimes called a largest coding unit (LCU). In ITU-T H.265, a picture is divided into CTUs. In ITU-T H.265, for a picture, the CTU size may be set to contain 16x16, 32x32, or 64x64 luma samples. In ITU-T H.265, a CTU is composed of a respective coding tree block (CTB) for each component of video data (e.g., luma (Y) and chroma (Cb and Cr). It should be noted that a video having one luma component and two corresponding chroma components may be described as having two channels, i.e., a luma channel and a chroma channel. Furthermore, in ITU-T H.265, a CTU may be partitioned according to a quad-tree (QT) partitioning structure, such that the CTB of a CTU is partitioned into coding blocks (CBs). That is, in ITU-T H.265, a CTU may be partitioned into quad-tree leaf nodes. According to ITU-T H.265, one luma CB, together with two corresponding chroma CBs and related syntax elements, is called a coding unit (CU). In ITU-T H.265, a minimum allowed size of a CB may be signaled. In H.265, the smallest allowable size of luma CB is 8x8 luma samples. In ITU-T H.265, the decision to code a picture portion using intra- or inter-prediction is made at the CU level. In ITU-T H.265, a CU is associated with a prediction unit structure that has its root in the CU. In ITU-T H.265, the prediction unit structure allows for splitting the luma CB and chroma CB for the purpose of generating corresponding reference samples.That is, in ITU-T H.265, the luma CB and the chroma CB may be divided into respective luma and chroma prediction blocks (PBs), where the PBs include blocks of sample values to which the same prediction is applied. In ITU-T H.265, the CB may be divided into 1, 2, or 4 PBs. ITU-T H.265 supports PB sizes from 64×64 samples to 4×4 samples. In ITU-T H.265, square PBs are supported for intra prediction, where the CB may form a PB, or the CB may be divided into four square PBs. In addition to square PBs, rectangular PBs are supported for inter prediction, where the CB may be bisected vertically or horizontally to form the PBs. Furthermore, ITU-T H.265 supports four asymmetric PB division for inter prediction, where the CB is divided into two PBs by one-quarter of the height (at the top or bottom) or width (at the left or right) of the CB. Using intra prediction data (e.g., intra prediction mode syntax element) or inter prediction data (e.g., motion data syntax element) corresponding to the PB, reference sample values and / or predicted sample values for the PB are generated.
[0017] JEM specifies a CTU with a maximum size of 256x256 luma samples. JEM specifies a quad-tree plus binary tree (QTBT) block structure. In JEM, the QTBT structure allows a quad-tree leaf node to be further split by a binary tree structure (BT). That is, in JEM, the binary tree structure allows a quad-tree leaf node to be split recursively vertically or horizontally. In JVET-T2001, the CTU is split according to a quad-tree plus multi-type tree (QTMT or QT+MTT) structure. QTMT in JVET-T2001 is similar to QTBT in JEM. However, in JVET-T2001, the multi-type tree, in addition to indicating a binary split, may also indicate a so-called ternary (or triple tree (TT)) split. Ternary splitting splits a block into three blocks vertically or horizontally. In the case of vertical TT division, the block is divided at a quarter of its width from the left edge and at a quarter of its width from the right edge, and in the case of horizontal TT division, the block is divided at a quarter of its height from the top edge and at a quarter of its height from the bottom edge.
[0018] As mentioned above, each video frame or picture may be divided into one or more regions. For example, according to ITU-T H.265, each video frame or picture may be divided to include one or more slices, and further divided to include one or more tiles, where each slice includes a sequence of CTUs (e.g., in raster scan order), and a tile is a sequence of CTUs corresponding to a rectangular portion of a picture. It should be noted that a slice, in ITU-T H.265, is a sequence of one or more slice segments starting with an independent slice segment and including all subsequent dependent slice segments (if any) preceding the next independent slice segment (if any). A slice segment, like a slice, is a sequence of CTUs. Thus, in some cases, the terms slice and slice segment may be used interchangeably to indicate a sequence of CTUs arranged in raster scan order. It should be further noted that in ITU-T H.265, a tile may consist of CTUs included in two or more slices, and a slice may consist of CTUs included in two or more tiles. However, ITU-T H.265 specifies that one or both of the following conditions must be met: (1) all CTUs in a slice belong to the same tile, and (2) all CTUs in a tile belong to the same slice.
[0019] For JVET-T2001, a slice is instead only required to consist of an integer number of CTUs, but is instead required to consist of an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile. Note that in JVET-T2001, the slice design does not include slice segments (i.e., there are no independent / dependent slice segments). Thus, in JVET-T2001, a picture may contain a single tile, where the single tile is contained within a single slice, or a picture may contain multiple tiles, where multiple tiles (or their CTU rows) may be contained within one or more slices. In JVET-T2001, the division of a picture into tiles is specified by specifying the height of each of the tile rows and the width of each of the tile columns. Thus, in JVET-T2001, a tile is a rectangular region of CTUs within a particular tile row and a particular tile column location. Further, it should be noted that JVET-T2001 specifies the case where a picture may be divided into sub-pictures, where a sub-picture is a rectangular region of a CTU within a picture. The top-left CTU of a sub-picture may be located at any CTU position within a picture, with the sub-picture being constrained to contain one or more slices. Thus, unlike tiles, sub-pictures are not necessarily limited to a specific row and column position. It should be noted that a sub-picture may be useful for encapsulating an area of interest within a picture, and the sub-bitstream extraction process may be used only to decode and display the specific area of interest. That is, as described in more detail below, a bitstream of coded video data includes a sequence of network abstraction layer (NAL) units, where the NAL units encapsulate the coded video data (i.e., video data corresponding to a slice of a picture) or the NAL units encapsulate metadata (e.g., parameter sets) used to decode the video data, and the sub-bitstream extraction process forms a new bitstream by removing one or more NAL units from the bitstream.
[0020] FIG. 2 is a conceptual diagram illustrating an example of a picture in a picture group divided according to tiles, slices, and subpictures. It should be noted that the techniques described herein may be applicable to tiles, slices, subpictures, their subdivisions, and / or their equivalent structures. That is, the techniques described herein may be generally applicable regardless of how a picture is divided into regions. For example, in some cases, the techniques described herein may be applicable when tiles may be divided into so-called bricks, where a brick is a rectangular region of a CTU row in a particular tile. Further, for example, in some cases, the techniques described herein may be applicable when one or more tiles may be included in a so-called tile group, where a tile group includes an integer number of adjacent tiles. In the example shown in FIG. 2, Pic3 is divided into 16 tiles (i.e., Tile 0 to Tile 1). 15 ), and three slices (i.e., Slice 0 through Slice 2). In the example shown in FIG. 2, Slice 0 contains four tiles (i.e., Tile 0 through Tile 3), Slice 1 contains eight tiles (i.e., Tile 4 through Tile 5), 11 ), Slice2 contains four tiles (i.e., Tile 12 From Tile 15), Pic3 includes two subpictures (i.e., Subpicture0 and Subpicture1), where Subpicture0 includes Slice0 and Slice1, and Subpicture1 includes Slice2. As discussed above, subpictures may be useful for encapsulating regions of interest within a picture, and the sub-bitstream extraction process may be used to selectively decode (and display) the regions of interest. For example, with reference to FIG. 2, Subpicture0 may correspond to the action portion (e.g., a view of the field) of a sporting event presentation, and Subpicture1 may correspond to a scrolling banner displayed during the sporting event presentation. By organizing pictures into subpictures in this manner, a viewer may be able to disable the display of the scrolling banner. That is, through the sub-bitstream extraction process, the Slice2 NAL unit may be removed from the bitstream (and thus may not be decoded and / or displayed), and the Slice0 NAL unit and the Slice1 NAL unit may be decoded and displayed. The encapsulation of slices of a picture into respective NAL unit data structures and sub-bitstream extraction is described in further detail below.
[0021] For intra-prediction coding, the intra-prediction mode may specify the location of the reference sample in the picture. In ITU-T H.265, the defined possible intra-prediction modes include planar (i.e., surface fitting) prediction mode, DC (i.e., monotonic global averaging) prediction mode, and 33 angular prediction modes (predMode: 2-34). In JEM, the defined possible intra-prediction modes include planar, DC, and 65 angular prediction modes. It should be noted that the planar and DC prediction modes may be referred to as non-directional prediction modes, and the angular prediction modes may be referred to as directional prediction modes. It should be noted that the techniques described herein may be generally applicable regardless of the number of defined possible prediction modes.
[0022] For inter-predictive coding, a reference picture is determined and a motion vector (MV) identifies samples in the reference picture used to generate a prediction for the current video block. For example, the current video block may be predicted using reference sample values located in one or more previously coded pictures, and the motion vector is used to indicate the position of the reference block relative to the current video block. The motion vector may be, for example, a horizontal displacement component of the motion vector (i.e., the MV x ), the vertical displacement component of the motion vector (i.e., MV y), and the resolution of the motion vectors (e.g., ¼ pixel precision, ½ pixel precision, 1 pixel precision, 2 pixel precision, 4 pixel precision). Previously decoded pictures, which may include pictures output before or after the current picture, may be organized into one or more reference picture lists and identified using reference picture index values. Furthermore, in inter-predictive coding, uni-prediction refers to generating a prediction using sample values from a single reference picture, and bi-prediction refers to generating a prediction using respective sample values from two reference pictures. That is, in uni-prediction, a single reference picture and corresponding motion vector are used to generate a prediction for the current video block, and in bi-prediction, a first reference picture and corresponding first motion vector, and a second reference picture and corresponding second motion vector are used to generate a prediction for the current video block. In bi-prediction, the respective sample values are combined (e.g., added, rounded, clipped, or averaged according to a weight) to generate a prediction. A picture and its regions may be classified based on what kind of prediction modes may be used to code its video blocks. That is, for a region having a B type (e.g., B slice), bi-prediction mode, uni-prediction mode, and intra-prediction mode may be used, for a region having a P type (e.g., P slice), uni-prediction mode and intra-prediction mode may be used, and for a region having an I type (e.g., I slice), only intra-prediction mode may be used. As mentioned above, the reference picture is identified by a reference index. For example, in a P slice, there may be a single reference picture list RefPicList0, and in a B slice, in addition to RefPicList0, there may be a second independent reference picture list RefPicList1. It should be noted that uni-prediction in a B slice may use one of RefPicList0 or RefPicList1 to generate a prediction. It should also be noted that during the decoding process, at the start of decoding a picture, a reference picture list is generated from a previously decoded picture stored in a decoded picture buffer (DPB).
[0023] Furthermore, the coding standard may support various modes of motion vector prediction. Motion vector prediction allows a value of a motion vector for a current video block to be derived based on another motion vector. For example, a set of candidate blocks with associated motion information may be derived from spatial and temporal neighboring blocks to the current video block. Furthermore, generated (or default) motion information may be used for motion vector prediction. Examples of motion vector prediction include advanced motion vector prediction (AMVP), temporal motion vector prediction (TMVP), the so-called "combined" mode, as well as "skip" and "direct" motion estimation. Furthermore, other examples of motion vector prediction include advanced temporal motion vector prediction (ATMVP) and spatial-temporal motion vector prediction (STMVP). In motion vector prediction, both the video encoding device and the video decoding device perform the same process to derive a set of candidates. Thus, for the current video block, the same set of candidates is generated during encoding and decoding.
[0024] As mentioned above, in inter-prediction coding, reference samples in previously coded pictures are used to code video blocks in a current picture. A previously coded picture that is available for use as a reference when coding a current picture is called a reference picture. It should be noted that the decoding order does not necessarily correspond to the picture output order, i.e., the temporal order of pictures in a video sequence. In ITU-T H.265, when a picture is decoded, the picture is stored in a decoded picture buffer (DPB) (which may be called a frame buffer, a reference buffer, a reference picture buffer, etc.). In ITU-T H.265, pictures stored in the DPB are removed from the DPB when they are output and are no longer needed for coding a subsequent picture. In ITU-T H.265, the decision of whether a picture should be removed from the DPB is performed once per picture after decoding the slice header, i.e., at the beginning of the decoding of the picture. For example, referring to FIG. 2, Pic2 is shown as referring to Pic1. Similarly, Pic3 is shown as referencing Pic0. With reference to FIG. 2, assuming that the picture numbers correspond to the decoding order, the DPB is populated as follows: after decoding Pic0, the DPB contains {Pic0}, at the start of decoding Pic1, the DPB contains {Pic0}, after decoding Pic1, the DPB contains {Pic0, Pic1}, and at the start of decoding Pic2, the DPB contains {Pic0, Pic1}. Then, Pic2 is decoded with reference to Pic1, and after decoding Pic2, the DPB contains {Pic0, Pic1, Pic2}. At the start of decoding Pic3, pictures Pic0 and Pic1 are marked for removal from the DPB as they are not required to decode Pic3 (or any subsequent pictures not shown), and assuming that Pic1 and Pic2 have been output, the DPB is updated to contain {Pic0}. Then, Pic3 is decoded by referencing Pic0. The process of marking pictures for removal from the DPB is sometimes called Reference Picture Set (RPS) management.
[0025] As mentioned above, the intra prediction data or the inter prediction data is used to generate reference sample values for a block of sample values. The difference between the sample values included in the current PB or another type of picture substructure and the associated reference sample (e.g., the reference sample generated using the prediction) may be referred to as residual data. The residual data may include a respective array of difference values corresponding to each component of the video data. The residual data may be in the pixel domain. A transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), an integer transform, a wavelet transform, or a conceptually similar transform, may be applied to the array of difference values to generate transform coefficients. It is noted that in ITU-T H.265 and JVET-T2001, a CU is associated with a transform tree structure with its root at the CU level. The transform tree is divided into one or more transform units (TUs). That is, the array of difference values may be partitioned (e.g., four 8×8 transforms may be applied to a 16×16 array of residual values) for purposes of generating transform coefficients. For each component of video data, such subdivision of the difference values may be referred to as a Transform Block (TB). Note that in some cases, a core transform and subsequent secondary transforms may be applied at the video encoder to generate the transform coefficients. For the video decoder, the order of the transforms is reversed.
[0026] The quantization process may be performed directly on the transform coefficients or on the residual sample values (e.g., in the case of palette coding quantization). Quantization approximates the transform coefficients with amplitudes limited to a particular set of values. Quantization essentially scales the transform coefficients to change the amount of data required to represent a group of transform coefficients. Quantization may include division of the transform coefficients (or values resulting from adding an offset value to the transform coefficients) by a quantization scale factor and any associated rounding function (e.g., rounding to the nearest integer). Quantized transform coefficients are sometimes referred to as coefficient level values. Inverse quantization (or "dequantization") may include multiplication of coefficient level values by a quantization scale factor and any mutual rounding or offset addition operations. It should be noted that as used herein, the term quantization process may refer in some cases to division by a quantization scale factor to generate level values and in some cases to multiplication by a quantization scale factor to recover transform coefficients. That is, the quantization process may refer in some cases to quantization and in some cases to inverse quantization. Furthermore, while in some of the following examples the quantization process is described with respect to arithmetic operations associated with decimal number representation, it should be noted that such description is for illustrative purposes and should not be construed as limiting. For example, the techniques described herein may be implemented in devices using binary arithmetic, etc. For example, the multiplication and division operations described herein may be implemented using bit shifting operations, etc.
[0027] The quantized transform coefficients and syntax elements (e.g., syntax elements indicating a coding structure of a video block) may be entropy coded according to an entropy coding technique. The entropy coding process includes coding the values of the syntax elements using a lossless data compression algorithm. Examples of entropy coding techniques include content adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), probability interval partitioning entropy coding (PIPE), etc. The entropy coded quantized transform coefficients and the corresponding entropy coded syntax elements may form a compliant bitstream that can be used to reproduce the video data at a video decoding device. The entropy coding process, e.g., CABAC, may include performing binarization on the syntax elements. Binarization refers to the process of converting the value of a syntax element into a series of one or more bits. These bits are sometimes called "bins." The binarization may include one or a combination of the following encoding techniques: fixed-length encoding, unary encoding, shortened unary encoding, shortened Rice encoding, Golomb encoding, k-th exponential Golomb encoding, and Golomb-Rice encoding. For example, the binarization may include representing an integer value of 5 of a syntax element as 00000101 using an 8-bit fixed-length binarization technique, or representing an integer value of 5 as 11110 using a unary encoding binarization technique. As used herein, each of the terms fixed-length encoding, unary encoding, shortened unary encoding, shortened Rice encoding, Golomb encoding, k-th exponential Golomb encoding, and Golomb-Rice encoding may refer to a general implementation of these techniques and / or a more specific implementation of these encoding techniques. For example, an implementation of Golomb-Rice encoding may be specifically defined according to a video encoding standard.In a CABAC example, for a particular bin, the context provides a most probable state (MPS) value for the bin (i.e., the MPS for the bin is one of 0 or 1) and a probability value that the bin is the MPS or least probably state (LPS). For example, the context may indicate that the MPS of the bin is 0 and the probability that the bin is 1 is 0.3. Note that the context may be determined based on values of previously coded bins, including bins in the current syntax element and previously coded syntax elements. For example, values of syntax elements associated with neighboring video blocks may be used to determine the context of the current bin.
[0028] As mentioned above, the sample values of the reconstructed block may differ from the sample values of the encoded current video block. Furthermore, it should be noted that in some cases, encoding video data block by block may result in artifacts (e.g., so-called blocking artifacts, banding artifacts, etc.). For example, blocking artifacts may cause encoded block boundaries of the reconstructed video data to be visually perceptible to a user. In this manner, the reconstructed sample values may be modified to minimize the difference between the encoded sample values of the current video block and the reconstructed block and / or to minimize artifacts introduced by the video encoding process. Such modification may be generally referred to as filtering. It should be noted that filtering may occur as part of an in-loop filtering process or a post-loop (or post-filtering) filtering process. In an in-loop filtering process, the sample values resulting from the filtering process may be used for the predicted video block (e.g., stored in a reference frame buffer for subsequent encoding in a video encoding device and subsequent decoding in a video decoding device). In a post-loop filtering process, the sample values resulting from the filtering process are simply output as part of the decoding process (e.g., not used for subsequent encoding). For example, in a video decoding device, in an in-loop filtering process, sample values resulting from filtering the reconstructed block are used for subsequent decoding (e.g., stored in a reference buffer) and output (e.g., to a display), whereas in a post-loop filtering process, the reconstructed block is used for subsequent decoding, and sample values resulting from filtering the reconstructed block are output and not used for subsequent decoding.
[0029] Deblocking (or de-blocking), deblock filtering, or applying a deblocking filter refers to a process of smoothing (i.e., making the boundary less perceptible by a viewer) the boundary of adjacent reconstructed video blocks. Smoothing the boundary of adjacent reconstructed video blocks may include modifying sample values contained in a row or column adjacent to the boundary. JVET-T2001 specifies when a deblocking filter is applied to reconstructed sample values as part of an in-loop filtering process. In addition to applying a deblocking filter as part of an in-loop filtering process, JVET-T2001 specifies when sample adaptive offset (SAO) filtering may be applied in the in-loop filtering process. In general, SAO is a process of modifying deblocked sample values in a region by conditionally adding an offset value. Another type of filtering process includes the so-called adaptive loop filter (ALF). ALF with block-based adaptation is specified in JEM. In JEM, ALF is applied after the SAO filter. It should be noted that the ALF may be applied to the reconstructed samples independently of other filtering techniques. The process for applying the ALF specified in the JEM in a video coding device may be summarized as follows: (1) each 2×2 block of the luma component of the reconstructed picture is classified according to a classification index, (2) a set of filter coefficients is derived for each classification index, (3) a filtering decision is determined for the luma component, (4) a filtering decision is determined for the chroma components, and (5) the filter parameters (e.g., coefficients and decisions) are signaled. JVET-T2001 specifies deblocking, SAO, and ALF filters that may be described as generally based on the deblocking, SAO, and ALF filters specified in ITU-T H.265 and JEM.
[0030] It is noted that JVET-T2001 is referred to as a pre-release version of ITU-T H.266 and is therefore a near-finished draft of the video coding standard resulting from the VVC project, and is therefore sometimes referred to as the first version of the VVC standard (alternatively, VVC or VVC version 1 or ITU-H.266). It is noted that during the VVC project, convolutional neural network (CNN)-based techniques that showed potential for artifact removal and objective quality improvement were investigated, but it was decided not to include such techniques in the VVC standard. However, CNN-based techniques are currently being considered for extending and / or improving VVC. Some CNN-based techniques are related to post-filtering. For example, "AHG11: Content-adaptive neural network post-filter, 26th Meeting of ISO / IEC JTC1 / SC29 / WG11 20-29 April 2022, Teleconference, document JVET-Z0082-v2" (herein referred to as JVET-Z0082) describes a content-adaptive neural network based post-filter. Note that in JVET-Z0082, content adaptation is achieved by overfitting a NN post-filter to a test video. Further note that the result of the overfitting process in JVET-Z0082 is a weight update.JVET-Z0082 describes when the weight updates are coded in ISO / IEC FDIS 15938-17. Information technology - Multimedia content description interface - Part 17: Compression of neural networks for multimedia content description and analysis and Test Model of Incremental Compression of Neural Networks for Multimedia Content Description and Analysis (INCTM), N0179. February 2022, which are sometimes collectively referred to as the MPEG NNR (Neural Network Representation) or Neural Network Coding (NNC) standards. JVET-Z0082 further describes when the coded weight updates are signaled in the video bitstream as NNR postfilter SEI messages. "AHG9: NNR post-filter SEI message", 26th Meeting of ISO / IEC JTC1 / SC29 / WG11 20-29 April 2022, Teleconference, document JVET-Z0052-v1 (referred to herein as JVET-Z0052) describes the NNR post-filter SEI message utilized by JVET-Z0082. The elements of the NN post-filter described in JVET-Z0082 and the NNR post-filter SEI message described in JVET-Z0052 were adopted in "Additional SEI messages for VSEI (Draft 2)", 27th Meeting of ISO / IEC JTC1 / SC29 / WG11 13-22 July 2022, Teleconference, document JVET-AA2006-v2 (referred to herein as JVET-AA2006). JVET-AA2006 specifies a versatile supplemental enhancement information message for coded video bitstreams (VSEI).JVET-AA2006 specifies syntax and semantics for the neural network postfilter characteristics SEI message and for the neural network postfilter activation SEI message. The neural network postfilter characteristics SEI message specifies the neural networks that may be used as post-processing filters. The use of the specified post-processing filters for a particular picture is indicated in the neural network postfilter activation SEI message. Furthermore, "Information technology - MPEG video technologies - Part 7: Versatile supplemental enhancement information messages for coded video bitstreams, AMENDMENT 1: Additional SEI messages" (28th Meeting of ISO / IEC JTC1 / SC29 / WG5 9, November 2022, Mainz, DE) documents JVET-AB2006, m61498 (referred to herein as JVET-AB2006). JVET-AB2006 is described in further detail below. The techniques described herein provide techniques for signaling neural network postfilter messages.
[0031] For formulas used herein, the following arithmetic operators may be used:
[0032] [Table 1]
[0033] Additionally, the following mathematical functions may be used: Log2(x), the base 2 logarithm of x;
[0034]
number
[0035] With respect to the example syntax used herein, the following definitions of logical operators may apply: x&&y The Boolean logic "product" of x and y x||y The Boolean "union" of x and y ! Boolean logic "no" x?y:zIf x is true or not equal to 0, then evaluate the value of y, else evaluate the value of z.
[0036] In addition, the following relational operators may be applied:
[0037] [Table 2]
[0038] Furthermore, in the syntax descriptors used herein, it should be noted that the following descriptors may apply: -b(8): A byte with an arbitrary pattern of bits (8 bits). The parsing process of this descriptor is specified by the return value of the function read_bits(8). -f(n): fixed pattern bit sequence using n bits written left bit first (left to right). The parsing process of this descriptor is specified by the return value of the function read_bits(n). -se(v): A signed integer zeroth-order Exp-Golomb encoded syntax element, left bit first. -tb(v): A truncated binary using up to maxVal bits, where maxVal is defined in the semantics of the syntax element. -tu(v): A shortened unary using up to maxVal bits, where maxVal is defined in the semantics of the syntax element. -u(n): An unsigned integer using n bits. When n is "v" in the syntax table, the number of bits varies depending on the values of other syntax elements. The parsing process of this descriptor is specified by the return value of the function read_bits(n), which is interpreted as a binary representation of an unsigned integer written most significant bit first. -ue(v): An unsigned integer zeroth-order Exp-Golomb encoded syntax element, left bit first.
[0039] As mentioned above, video content includes a video sequence consisting of a series of pictures, and each picture may be divided into one or more regions. In JVET-T2001, the coded representation of a picture includes the VCL NAL units of a particular layer in an AU, including all CTUs of the picture. For example, referring back to FIG. 2, the coded representation of Pic3 is encapsulated in three coded slice NAL units (i.e., Slice0 NAL unit, Slice1 NAL unit, and Slice2 NAL unit). It should be noted that the term video coding layer (VCL) NAL unit is used as a generic term for coded slice NAL units, i.e., VCL NAL is a generic term that includes all types of slice NAL units. As mentioned above and described in more detail below, NAL units may encapsulate metadata used to decode video data. NAL units that encapsulate metadata used to decode a video sequence are generally referred to as non-VCL NAL units. Thus, in JVET-T2001, a NAL unit can be a VCL NAL unit or a non-VCL NAL unit. Note that a VCL NAL unit contains slice header data that provides information used to decode a particular slice. Thus, in JVET-T2001, information used to decode video data, sometimes referred to as metadata in some cases, is not limited to being contained in a non-VCL NAL unit. JVET-T2001 specifies the case where a picture unit (PU) is a set of NAL units associated with each other according to a specified classification rule, consecutive in decoding order, and containing exactly one coded picture, and an access unit (AU) is a set of PUs containing coded pictures that belong to different layers and are associated at the same time for output from the DPB. JVET-T2001 further specifies the case where a layer is a set of VCL NAL units and associated non-VCL NAL units that all have a particular value of a layer identifier.Furthermore, in JVET-T2001, a PU consists of zero or one PH NAL unit, one coded picture containing one or more VCL NAL units, and zero or more other non-VCL NAL units. Furthermore, in JVET-T2001, a coded video sequence (CVS) is a sequence of AUs consisting of, in decoding order, a CVSS AU followed by zero or more AUs that are not CVSS AUs, including all subsequent AUs up to but not including any subsequent AU that is a CVSS AU, where a coded video sequence start (CVSS) AU is an AU with a PU for each layer in the CVS, and a coded picture in each existing picture unit is a coded layer video sequence start (CLVSS) picture. In JVET-T2001, a coded layer video sequence (CLVS) is a sequence of PUs in the same layer consisting, in decoding order, of a CLVSS PU followed by zero or more PUs that are not CLVSS PUs, including all subsequent PUs up to but not including any subsequent PU that is a CLVSS PU. That is, in JVET-T2001, a bitstream may be described as containing a sequence of AUs that form one or more CVSs.
[0040] Multi-layer video coding allows a video presentation to be decoded / displayed as a presentation corresponding to a base layer of video data and as one or more additional presentations corresponding to enhancement layers of the video data. For example, the base layer may allow a video presentation to be presented having a basic level of quality (e.g., high resolution rendering and / or 30 Hz frame rate), and the enhancement layer may allow a video presentation to be presented having an increased level of quality (e.g., ultra-high resolution rendering and / or 60 Hz frame rate). The enhancement layer may be coded by referencing the base layer. That is, for example, a picture in the enhancement layer may be coded (e.g., using inter-layer prediction techniques) by referencing one or more pictures in the base layer (including scaled versions thereof). It should be noted that layers may also be coded independently of each other. In this case, there may be no inter-layer prediction between two layers. Each NAL unit may include an identifier indicating the layer of video data with which the NAL unit is associated. As mentioned above, a sub-bitstream extraction process may be used to decode and display only a particular region of interest of a picture. Additionally, the sub-bitstream extraction process may be used to decode and display only a particular layer of a video. Sub-bitstream extraction may refer to a process in which a device receiving a compliant or conforming bitstream forms a new compliant or conforming bitstream by discarding and / or modifying data in the received bitstream. For example, sub-bitstream extraction may be used to form a new compliant or conforming bitstream that corresponds to a particular representation (e.g., a higher quality representation) of the video.
[0041] In JVET-T2001, each of a video sequence, a GOP, a picture, a slice, and a CTU may be associated with metadata that describes video coding properties, and some types of metadata may be encapsulated in non-VCL NAL units. JVET-T2001 defines parameter sets that may be used to describe video data and / or video coding properties. In particular, JVET-T2001 includes four types of parameter sets: video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), and adaptation parameter set (APS), where an SPS applies to zero or more entire CVSs, a PPS applies to zero or more entire coded pictures, an APS applies to zero or more slices, and a VPS may be optionally referenced by an SPS. A PPS applies to the individual coded pictures that reference it. In JVET-T2001, parameter sets may be encapsulated as non-VCL NAL units and / or signaled as messages. JVET-T2001 also includes a picture header (PH) encapsulated as a non-VCL NAL unit. In JVET-T2001, the picture header applies to all slices of a coded picture. JVET-T2001 further allows decoding capability information (DCI) and supplemental enhancement information (SEI) messages to be signaled. In JVET-T2001, the DCI and SEI messages assist processes related to decoding, display, or other purposes, but the DCI and SEI messages may not be required to create luma or chroma samples following the decoding process. In JVET-T2001, the DCI and SEI messages may be signaled in the bitstream using non-VCL NAL units. Furthermore, the DCI and SEI messages may be conveyed by some mechanism other than by being present in the bitstream (i.e., signaled out-of-band).
[0042] FIG. 3 shows an example of a bitstream including multiple CVSs, where a CVS includes AUs, and an AU includes picture units. The example shown in FIG. 3 corresponds to an example of encapsulating slice NAL units in a bitstream shown in the example of FIG. 2. In the example shown in FIG. 3, the corresponding picture unit for Pic3 includes three VCL NAL coded slice NAL units, namely, a Slice0 NAL unit, a Slice1 NAL unit, and a Slice2 NAL unit, and two non-VCL NAL units, namely, a PPS NAL unit and a PH NAL unit. It should be noted that in FIG. 3, the headers are NAL unit headers (i.e., should not be confused with slice headers). It should also be noted that in FIG. 3, other non-VCL NAL units not shown, such as an SPS NAL unit, a VPS NAL unit, an SEI message NAL unit, etc., may be included in the CVS. Further, note that in other examples, the PPS NAL unit used to decode Pic3 may be included elsewhere in the bitstream, e.g., in the picture unit corresponding to Pic0, or may be provided by an external mechanism. As described in more detail below, in JVET-T2001, the PH syntax structure may be present in the slice header of a VCL NAL unit or in the PH NAL unit of the current PU.
[0043] JVET-T2001 defines NAL unit header semantics that specify the type of raw byte sequence payload (RBSP) data structure contained in a NAL unit. Table 1 shows the syntax of the NAL unit header defined in JVET-T2001.
[0044] [Table 3]
[0045] JVET-T2001 specifies the following definitions for each syntax element shown in Table 1. forbidden_zero_bit shall be equal to 0. nuh_reserved_zero_bit shall be equal to 0. A value of 1 for nuh_reserved_zero_bit may be specified in the future by ITU-T|ISO / IEC. Although the value of nuh_reserved_zero_bit is required to be equal to 0 in this version of this specification, decoders conforming to this version of this specification shall allow a value of nuh_reserved_zero_bit equal to 1 to appear in the syntax and shall ignore (i.e., remove from the bitstream and discard) NAL units with nuh_reserved_zero_bit equal to 1. nuh_layer_id specifies the identifier of the layer to which a VCL NAL unit belongs, or the identifier of the layer to which a non-VCL NAL unit applies. Values of nuh_layer_id shall be in the range of 0 to 55, inclusive. Other values of nuh_layer_id are reserved for future use by ITU-T|ISO / IEC. Although values of nuh_layer_id are required to be in the range of 0 to 55, inclusive, in this version of this specification, decoders conforming to this version of this specification shall allow values of nuh_layer_id greater than 55 to appear in the syntax and shall ignore (i.e., remove from the bitstream and discard) NAL units with nuh_layer_id greater than 55. The value of nuh_layer_id shall be the same for all VCL NAL units of a coded picture. The value of nuh_layer_id of a coded picture or PU is the value of nuh_layer_id of the VCL NAL units of the coded picture or PU. When nal_unit_type is equal to PH_NUT or FD_NUT, nuh_layer_id shall be equal to the nuh_layer_id of the associated VCL NAL unit. When nal_unit_type is equal to EOS_NUT, nuh_layer_id shall be equal to one of the nuh_layer_id values of a layer present in CVS. NOTE – The values of nuh_layer_id in the DCI, OPI, VPS, AUD, and EOB NAL units are not constrained. nuh_temporal_id_plus1 minus 1 specifies the temporal identifier for the NAL unit. The value of nuh_temporal_id_plus1 shall not be equal to 0. The variable TemporalId is derived as follows: TemporalId=nuh_temporal_id_plus1-1 When nal_unit_type is in the range IDR_W_RADL to RSV_IRAP_11, inclusive, TemporalId shall be equal to 0. When nal_unit_type is equal to STSA_NUT and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, TemporalId shall be greater than 0. The value of TemporalId shall be the same for all VCL NAL units of an AU. The value of TemporalId of a coded picture, PU, or AU is the value of TemporalId of the VCL NAL units of the coded picture, PU, or AU. The value of TemporalId of a sublayer representation is the largest value of TemporalId of all VCL NAL units in the sublayer representation. The values of TemporalId for non-VCL NAL units are constrained as follows: - if nal_unit_type is equal to DCI_NUT, OPI_NUT, VPS_NUT, or SPS_NUT, TemporalId shall be equal to 0 and the TemporalId of the AU containing the NAL unit shall be equal to 0. Otherwise, if nal_unit_type is equal to PH_NUT, then TemporalId shall be equal to the TemporalId of the PU that contains the NAL unit. Otherwise, if nal_unit_type is equal to EOS_NUT or EOB_NUT, TemporalId shall be equal to 0. Otherwise, if nal_unit_type is equal to AUD_NUT, FD_NUT, PREFIX_SEI_NUT, or SUFFIX_SEI_NUT, TemporalId shall be equal to the TemporalId of the AU that contains the NAL unit. Otherwise, when nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, TemporalId shall be greater than or equal to the TemporalId of the PU that contains the NAL unit. NOTE - When the NAL unit is a non-VCL NAL unit, the value of TemporalId is equal to the minimum of the TemporalId values of all AUs to which the non-VCL NAL unit applies. When nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, TemporalId may be greater than or equal to the TemporalId of the containing AU, since all PPSs and APSs may be included at the beginning of the bitstream (e.g., when they are transported out-of-band and the receiver places them at the beginning of the bitstream), and the first coded picture has TemporalId equal to 0. nal_unit_type specifies the NAL unit type, i.e., the type of RBSP data structure contained in the NAL unit as specified in Table 2. NAL units with nal_unit_type in the range UNSPEC28 to UNSPEC31, inclusive, whose semantics are unspecified, SHALL NOT affect the decoding process specified in this specification. NOTE - NAL unit types in the range of UNSPEC_28 to UNSPEC_31 may be used as determined by the application. The decoding process for these values of nal_unit_type is not specified in this specification. Since different applications may use these NAL unit types for different purposes, it is expected that particular care will be taken in the design of encoders that generate NAL units with these nal_unit_type values, and in the design of decoders that interpret the content of NAL units with these nal_unit_type values. This specification does not define the management of these values. These nal_unit_type values may only be suitable for use in contexts where "collisions" of usage (i.e., different definitions of the meaning of the NAL unit content for the same nal_unit_type value) are not significant or possible, or are managed, e.g., defined or managed in a controlling application or transport specification, or by controlling the environment in which the bitstream is delivered. For purposes other than determining the amount of data in a DU of the bitstream, a decoder SHALL ignore (remove from the bitstream and discard) the content of all NAL units that use reserved values of nal_unit_type. NOTE - This requirement allows for the future definition of compatible extensions to this specification. [Table 4] NOTE - A clean random access (CRA) picture may have an associated RASL or RADL picture present in the bitstream. NOTE - An instantaneous decoding refresh (IDR) picture with nal_unit_type equal to IDR_N_LP does not have an associated leading picture present in the bitstream. An IDR picture with nal_unit_type equal to IDR_W_RADL does not have an associated RASL picture present in the bitstream, but may have an associated RADL picture in the bitstream. The value of nal_unit_type shall be the same for all VCL NAL units of a subpicture. A subpicture is referred to as having the same NAL unit type as the VCL NAL units of the subpicture. For the VCL NAL units of any particular picture, the following applies: - if pps_mixed_nalu_types_in_pic_flag is equal to 0, the value of nal_unit_type shall be the same for all VCL NAL units of a picture, A picture or PU is referred to as having the same NAL unit type as the coded slice NAL unit of the picture or PU. - Otherwise (pps_mixed_nalu_types_in_pic_flag is equal to 1), all of the following constraints apply. -A picture shall have at least two sub-pictures. - A VCL NAL unit of a picture shall have two or more distinct nal_unit_type values. - There shall be no VCL NAL unit of the picture with nal_unit_type equal to GDR_NUT. When a VCL NAL unit of a picture has nal_unit_type equal to nalUnitTypeA equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT, all other VCL NAL units of the picture shall have nal_unit_type equal to nalUnitTypeA or TRAIL_NUT. The value of nal_unit_type shall be the same for all pictures in an IRAP or GDR AU. When sps_video_parameter_set_id is greater than 0, vps_max_tid_il_ref_pics_plus1[i][j] is equal to 0 for j equal to GeneralLayerIdx[nuh_layer_id] and for any value of i in the range of j+1 to vps_max_layers_minus1, inclusive, pps_mixed_nalu_types_in_pic_flag is equal to 1, and the value of nal_unit_type is not equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT. The requirements for bitstream conformance are that the following constraints apply: When a picture is a leading picture of an IRAP picture, it shall be a RADL or RASL picture. When a subpicture is the leading subpicture of an IRAP subpicture, it shall be a RADL or RASL subpicture. - If a picture is not the leading picture of an IRAP picture, it shall not be a RADL or RASL picture. - If a subpicture is not the leading subpicture of an IRAP subpicture, it shall not be a RADL or RASL subpicture. - There shall be no RASL pictures associated with an IDR picture in the bitstream. - No RASL sub-pictures associated with an IDR sub-picture shall be present in the bitstream. - There shall be no RADL pictures associated with an IDR picture with nal_unit_type equal to IDR_N_LP in the bitstream. NOTE - Performing random access at the position of an IRAP AU by discarding all PUs before the IRAP AU (and correctly decoding non-RASL pictures in the IRAP AU and all following AUs in decoding order) is possible, provided that each parameter set is available (in the bitstream or by external means not specified in this specification) at the time it is referenced. - There shall be no RADL sub-pictures associated with an IDR sub-picture with nal_unit_type equal to IDR_N_LP in the bitstream. -Any picture with nuh_layer_id equal to the particular value layerId that precedes in decoding order an IRAP picture with nuh_layer_id equal to layerId shall precede the IRAP picture in output order and shall precede any RADL picture associated with the IRAP picture in output order. - Any subpicture with nuh_layer_id equal to the specified value layerId and subpicIdx subpicture index equal to the specified value that precedes in decoding order an IRAP subpicture with nuh_layer_id equal to layerId and subpicture index equal to subpicIdx shall precede in output order the IRAP subpicture and all its associated RADL subpictures. Any picture with nuh_layer_id equal to the specified value layerId that precedes in decoding order the recovery point picture with nuh_layer_id equal to -layerId shall precede the recovery point picture in output order. Any subpicture with nuh_layer_id equal to a specific value layerId and subpicIdx equal to a specific value that precedes in decoding order a subpicture with nuh_layer_id equal to layerId and subpicIdx equal to a specific value in the recovery point picture shall precede that subpicture in the recovery point picture in output order. - Any RASL picture associated with a CRA picture shall precede any RADL picture associated with the CRA picture in output order. Any RASL sub-picture associated with a CRA sub-picture shall precede any RADL sub-picture associated with the CRA sub-picture in output order. - Any RASL picture having nuh_layer_id equal to a particular value layerId and associated with a CRA picture shall follow in output order any IRAP or GDR picture having nuh_layer_id equal to layerId that precedes the CRA picture in decoding order. - Any RASL subpicture associated with a CRA subpicture and having nuh_layer_id equal to a particular value layerId and a subpicture index equal to a particular value subpicIdx shall follow in output order any IRAP or GDR subpicture with nuh_layer_id equal to layerId and a subpicture index equal to subpicIdx that precedes the CRA subpicture in decoding order. - if sps_field_seq_flag is equal to 0, the following applies: if the current picture with nuh_layer_id equal to a particular value layerId is a leading picture associated with an IRAP picture, it shall precede, in decoding order, all non-leading pictures associated with the same IRAP picture. Otherwise (sps_field_seq_flag is equal to 1), let picA and picB be the first and last leading pictures associated with an IRAP picture, respectively, in decoding order, there shall be at most one non-leading picture with nuh_layer_id equal to layerId preceding picA in decoding order, and there shall be no non-leading picture with nuh_layer_id equal to layerId between picA and picB in decoding order. - if sps_field_seq_flag is equal to 0, the following applies: if the current subpicture with nuh_layer_id equal to a particular value layerId and subpicture index equal to a particular value subpicIdx is a leading subpicture associated with an IRAP subpicture, then it shall precede, in decoding order, all non-leading subpictures associated with the same IRAP subpicture. Otherwise (sps_field_seq_flag is equal to 1), let subpicA and subpicB be the first and last leading subpictures associated with an IRAP subpicture, respectively, in decoding order, then there shall be at most one non-leading subpicture with nuh_layer_id equal to layerId and subpicture index equal to subpicIdx preceding subpicA in decoding order, and there shall be no non-leading pictures with nuh_layer_id equal to layerId and subpicture index equal to subpicIdx between picA and picB in decoding order.
[0046] A NAL unit may include a supplemental enhancement information (SEI) syntax structure, as provided in Table 2. Tables 3 and 4 show the supplemental enhancement information (SEI) syntax structure defined in JVET-T2001.
[0047] [Table 5]
[0048] With respect to Tables 3 and 4, JVET-T2001 specifies the following semantics: Each SEI message consists of variables that specify the type, payloadType, and size, payloadSize, of the SEI message payload. The SEI message payload is specified. The derived SEI message payload size, payloadSize, is specified in bytes and shall be equal to the number of RBSP bytes in the SEI message payload. NOTE - A NAL unit byte sequence containing an SEI message may contain one or more emulation prevention bytes (represented by the emulation_prevention_three_byte syntax element). Since the payload size of an SEI message is specified in RBSP bytes, the amount of emulation prevention bytes is not included in the size of the SEI payload, payloadSize. payload_type_byte is the payload type byte of the SEI message. payload_size_byte is the payload size in bytes of the SEI message.
[0049] It should be noted that JVET-T2001 defines payload types, and "Additional SEI messages for VSEI (Draft 6)" 25th Meeting of ISO / IEC JTC1 / SC29 / WG11 12-21 January 2022, Teleconference, document JVET-Y2006-v1, incorporated herein by reference and referred to as JVET-Y2006, defines additional payload types. Table 5 shows the sei_payload() syntax structure in a schematic manner. That is, Table 5 shows the sei_payload() syntax structure, but for brevity, not all possible types of payloads are included in Table 5.
[0050] [Table 6]
[0051] With respect to Table 5, JVET-T2001 specifies the following semantics: sei_reserved_payload_extension_data SHALL NOT be present in bitstreams that conform to this version of this specification. However, decoders that conform to this version of this specification SHALL ignore the presence and value of sei_reserved_payload_extension_data. When present, the length in bits of sei_reserved_payload_extension_data shall be 8 * Equal to payloadSize-nEarlierBits-nPayloadZeroBits-1, where nEarlierBits is the number of bits in the sei_payload() syntax structure preceding the sei_reserved_payload_extension_data syntax element and nPayloadZeroBits is the number of sei_payload_bit_equal_to_zero syntax elements at the end of the sei_payload() syntax structure. If more_data_in_payload() is true after parsing an SEI message syntax structure (for example, the buffering_period() syntax structure), and nPayloadZeroBits is not equal to 7, then PayloadBits is set to 8. * payloadSize-n is set equal to PayloadZeroBits-1, otherwise PayloadBits is set equal to 8 * Set equal to payloadSize. payload_bit_equal_to_one shall be equal to 1. payload_bit_equal_to_zero shall be equal to 0. NOTE - SEI messages with the same value of payloadType are conceptually the same SEI message, regardless of whether they are included in a prefix or suffix SEI NAL unit. NOTE - For SEI messages specified in this specification and the VSEI specification (ITU-T H.274|ISO / IEC 23002-7), payloadType values are aligned with similar SEI messages specified in AVC (Rec. ITU-T H.264|ISO / IEC 14496-10) and HEVC (Rec. ITU-T H.265|ISO / IEC 23008-2). The semantics and duration for each SEI message are specified in the semantics specification for each particular SEI message. NOTE – The persistence information in the SEI messages is summarized for informational purposes.
[0052] JVET-T2001 further specifies the following: SEI messages having the syntax structure identified in [Table 5] specified in Rec. ITU-T H.274 | ISO / IEC 23002-7 may be used with bitstreams specified by this specification. When any particular Rec. ITU-T H.274|ISO / IEC 23002-7 SEI message is included in a bitstream specified by this specification, the SEI payload syntax shall be contained in the sei_payload() syntax structure specified in [Table 5] and shall use the payloadType value specified in [Table 5], plus any SEI message-specific constraints specified in this Annex for that particular SEI message shall apply. The value of PayloadBits is passed to the parser for the SEI message syntax structure specified in Rec. ITU-T H.274 | ISO / IEC 23002-7, as specified above.
[0053] As mentioned above, JVET-AB2006 defines NN postfilter supplemental enhancement information messages. In particular, JVET-AB2006 defines a neural network postfilter characteristics SEI message (payloadType==210) and a neural network postfilter activation SEI message (payloadType==211). Table 6 shows the syntax of the neural network postfilter characteristics SEI message defined in JVET-AB2006. Note that the neural network postfilter characteristics SEI message is sometimes referred to as the NNPFC SEI.
[0054] [Table 7-1]
[0055] [Table 7-2]
[0056] [Table 7-3]
[0057] With respect to Table 6, JVET-AB2006 specifies the following semantics: The neural-network post-filter characteristics (NNPFC) SEI message specifies neural networks that can be used as post-processing filters. The use of a specified post-processing filter for a particular picture is indicated in the neural-network post-filter activation SEI message. Use of this SEI message requires the definition of the following variables: The width and height of the cropped decoded output picture in units of luma samples, denoted herein by CroppedWidth and CroppedHeight, respectively. - When present, the luma sample array CroppedYPic[idx] and the chroma sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] of the cropped decoded output picture with idx in the range of 0 to numInputPics-1, inclusive, that are used as input for the post-processing filters. -BitDepth - the bit depth of the luma sample array of the cropped decoded output picture Y . -BitDepth - the bit depth of the chroma sample array, if any, of the cropped decoded output picture C . A chroma format indicator, denoted herein by ChromaFormatIdc. When -nnpfc_auxiliary_inp_idc is equal to 1, the filtering strength control value StrengthControlVal shall be a real number in the range of 0 to 1, inclusive. The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc as specified by Table 7.
[0058] [Table 8] NOTE - There may be more than one NNPFC SEI message for the same picture. When more than one NNPFC SEI message with different values of nnpfc_id is present or activated for the same picture, they may have the same or different values of nnpfc_purpose and nnpfc_mode_idc. nnpfc_id contains an identification number that can be used to identify a post-processing filter. The value of nnpfc_id is in the range 0 to 2, inclusive. 32 -2. The range is from 256 to 511, inclusive, and 256 to 511, inclusive. 31 ~2 32Values of nnpfc_id up to -2 are reserved for future use by ITU-T | ISO / IEC. 31 ~2 32 Decoders compliant with this version of this document that encounter an NNPFC SEI message with an nnpfc_id within the range of -2 shall ignore that SEI message. When the NNPFC SEI message is the first NNPFC SEI message, in decoding order, with a particular nnpfc_id value in the current CLVS, the following applies. -This SEI message specifies a basic post-processing filter. - This SEI message relates to the current decoded picture and all subsequent decoded pictures of the current layer in output order until the end of the current CLVS. When an NNPFC SEI message is a repetition of a previous NNPFC SEI message in the current CLVS in decoding order, the following semantics apply as if this SEI message was the only NNPFC SEI message with the same content within the current CLVS. When the NNPFC SEI message is not the first NNPFC SEI message, in decoding order, with a particular nnpfc_id value in the current CLVS, the following applies. -This SEI message defines an update to the preceding elementary post-processing filter in decoding order that has the same nnpfc_id value. -This SEI message relates to the current decoded picture and all subsequent decoded pictures of the current layer in output order until the end of the current CLVS or until the next NNPFC SEI message with that particular nnpfc_id value in output order within the current CLVS. nnpfc_mode_idc equal to 0 indicates that this SEI message contains an ISO / IEC 15938-17 bitstream that specifies a basic post-processing filter or is an update to a basic post-processing filter with the same nnpfc_id value. When the NNPFC SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc_id value in the current CLVS, an nnpfc_mode_idc equal to 1 specifies that the basic post-processing filter associated with the nnpfc_id value is a neural network identified by the URI indicated by nnpfc_uri having the format identified by the tag URI nnpfc_tag_uri. When the NNPFC SEI message is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value in the current CLVS, nnpfc_mode_idc equal to 1 specifies that updates to basic post-processing filters with the same nnpfc_id value are defined by the URI indicated by nnpfc_uri having the format identified by the tag URI nnpfc_tag_uri. Values of nnpfc_mode_idc shall be in the range 0 to 1, inclusive, in bitstreams conforming to this version of this document. Values of nnpfc_mode_idc between 2 and 255, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_mode_idc in the range 2 to 255, inclusive. Values of nnpfc_mode_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use. When this SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc_id value in the current CLVS, the post-processing filter PostProcessingFilter() is assigned to be the same as the base post-processing filter. When this SEI message is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value in the current CLVS, the post-processing filter PostProcessingFilter() is obtained by applying the updates defined by this SEI message to the basic post-processing filter. The updates are not cumulative; rather, each update is applied to the base post-processing filter, which is the post-processing filter specified by the first NNPFC SEI message, in decoding order, with a particular nnpfc_id value in the current CLVS. nnpfc_reserved_zero_bit_a shall be equal to 0 in bitstreams conforming to this version of this document. Decoders shall ignore NPFC SEI messages in which nnpfc_reserved_zero_bit_a is not equal to 0. The nnpfc_tag_uri contains a tag URI with syntax and semantics specified in IETF RFC 4151 that identifies the format and associated information about a neural network to be used as a base post-processing filter or an update to a base post-processing filter with the same nnpfc_id value specified by the nnpfc_uri. NOTE - The nnpfc_tag_uri makes it possible to uniquely identify the format of the neural network data specified by the nnrpf_uri without the need for a central registry. An nnpfc_tag_uri equal to "tag:iso.org,2023:15938-17" indicates that the neural network data identified by the nnpfc_uri conforms to ISO / IEC 15938-17. The nnpfc_uri contains a URI with syntax and semantics specified in IETF Internet Standard 66 that identifies a neural network to be used as a base post-processing filter or an update to a base post-processing filter with the same nnpfc_id value. nnpfc_formatting_and_purpose_flag equal to 1 specifies that syntax elements related to filter purpose, input formatting, output formatting, and complexity are present. nnpfc_formatting_and_purpose_flag equal to 0 specifies that syntax elements related to filter purpose, input formatting, output formatting, and complexity are not present. nnpfc_formatting_and_purpose_flag shall be equal to 1 when this SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc_id value in the current CLVS. nnpfc_formatting_and_purpose_flag shall be equal to 0 when this SEI message is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value in the current CLVS. nnpfc_purpose indicates the purpose of the post-processing filter as specified in Table 8. Values of nnpfc_purpose shall be in the range 0 to 5, inclusive, in bitstreams conforming to this version of this document. Values of nnpfc_purpose between 6 and 1023, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoding devices conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_purpose in the range 6 to 1203, inclusive. Values of nnpfc_purpose greater than 1023 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use.
[0059] [Table 9] NOTE – When reserved values of nnpfc_purpose are used in the future by ITU-T|ISO / IEC, the syntax of this SEI message may be extended with syntax elements whose presence is conditioned by nnpfc_purpose being equal to that value. When SubWidthC is equal to 1 and SubHeightC is equal to 1, nnpfc_purpose shall not be equal to 2 or 4. nnpfc_out_sub_c_flag equal to 1 specifies that outSubWidthC is equal to 1 and outSubHeightC is equal to 1. nnpfc_out_sub_c_flag equal to 0 specifies that outSubWidthC is equal to 2 and outSubHeightC is equal to 1. When nnpfc_out_sub_c_flag is not present, outSubWidthC is inferred to be equal to SubWidthC and outSubHeightC is inferred to be equal to SubHeightC. When ChromaFormatIdc is equal to 2 and nnpfc_out_sub_c_flag is present, the value of nnpfc_out_sub_c_flag shall be equal to 1. nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples specify the width and height, respectively, of the picture's luma sample array resulting from applying the post-processing filter identified by nnpfc_id to the cropped decoded output picture. When nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples are not present, they are inferred to be equal to CroppedWidth and CroppedHeight, respectively. The value of nnpfc_pic_width_in_luma_samples ranges from CroppedWidth to CroppedWidth, inclusive. * The value of nnpfc_pic_height_in_luma_samples is in the range CroppedHeight~CroppedHeight, inclusive.* It shall be within the range of 16-1. nnpfc_num_input_pics_minus2 plus 2 specifies the number of decoded output pictures used as input for the post-processing filter. nnpfc_interpolated_pics[i] specifies the number of interpolated pictures generated by the post-processing filter between the i-th picture and the (i+1)-th picture that are used as input for the post-processing filter. The variables numInputPics, which specify the number of pictures used as input for the post-processing filter, and numOutputPics, which specify the total number of pictures resulting from the post-processing filter, are derived as follows:
[0060] [Table 10] nnpfc_component_last_flag equal to 1 indicates that the last dimension in the input tensor inputTensor to the post-processing filter and the output tensor outputTensor resulting from the post-processing filter is used for the current channel. nnpfc_component_last_flag equal to 0 indicates that the third dimension in the input tensor inputTensor to the post-processing filter and the output tensor outputTensor resulting from the post-processing filter is used for the current channel. NOTE - The first dimension in the input and output tensors is used for the batch index, which is the practice in some neural network frameworks. Although the formula in the semantics of this SEI message uses the batch size corresponding to a batch index equal to 0, it is up to the post-processing implementation to determine the batch size used as input to the neural network inference. NOTE - For example, when nnpfc_inp_order_idc is equal to 3 and nnpfc_auxiliary_inp_idc is equal to 1, the input tensor has 7 channels including 4 luma matrices, 2 chroma matrices, and 1 auxiliary input matrix. In this case, the process DeriveInputTensors() derives each of these 7 channels of the input tensor one by one, and when a particular one of these channels is processed, that channel is called the current channel in the process. nnpfc_inp_format_idc indicates how to convert the sample values of the cropped decoded output picture into input values to the post-processing filter. When nnpfc_inp_format_idc is equal to 0, the input values to the post-processing filter are real numbers, and the functions InpY() and InpC() are specified as follows: InpY(x)=x-((1< <BitDepth Y )-1) InpC(x)=x-((1< <BitDepth C )-1) When nnpfc_inp_format_idc is equal to 1, the input values to the post-processing filter are unsigned integers and the functions InpY() and InpC() are specified as follows:
[0061] [Table 11] The variable inpTensorBitDepth is derived from the syntax element nnpfc_inp_tensor_bitdepth_minus8, as specified below. Values of nnpfc_inp_format_idc greater than 1 are reserved for future specification by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages which contain reserved values of nnpfc_inp_format_idc. nnpfc_inp_tensor_bitdepth_minus8 plus 8 specifies the bit depth of the luma sample values in the input integer tensor. The value of inpTensorBitDepth is derived as follows: inpTensorBitDepth=nnpfc_inp_tensor_bitdepth_minus8+8 It is a bitstream conformance requirement that the value of nnpfc_inp_tensor_bitdepth_minus8 be in the range 0 to 24, inclusive. nnpfc_inp_order_idc indicates how to order the sample array of the cropped decoded output picture as one of the input pictures to the post-processing filter. Values of nnpfc_inp_order_idc shall be in the range 0 to 3, inclusive, in bitstreams conforming to this version of this document. Values of nnpfc_inp_order_idc between 4 and 255, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range 4 to 255, inclusive. Values of nnpfc_inp_order_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use. When ChromaFormatIdc is not equal to 1, nnpfc_inp_order_idc shall be equal to 3. Table 9 contains useful descriptions of the nnpfc_inp_order_idc values.
[0062] [Table 12] A patch is a rectangular array of samples from a component (eg, a luma component or a chroma component) of a picture. An nnpfc_auxiliary_inp_idc greater than 0 indicates that auxiliary input data is present in the neural network postfilter input tensor. An nnpfc_auxiliary_inp_idc equal to 0 indicates that auxiliary input data is not present in the input tensor. An nnpfc_auxiliary_inp_idc equal to 1 specifies that the auxiliary input data is derived as specified in Equation 82. Values of nnpfc_auxiliary_inp_idc shall be in the range 0 to 1, inclusive, in bitstreams conforming to this version of this document. Values of nnpfc_inp_order_idc between 2 and 255, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range 2 to 255, inclusive. Values of nnpfc_inp_order_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use. The process DeriveInputTensors() for deriving an input tensor inputTensor given vertical sample coordinates cTop and horizontal sample coordinates cLeft that specify the top-left sample location of a patch of samples contained in the input tensor is specified as follows:
[0063] [Table 13-1]
[0064] [Table 13-2]
[0065] [Table 13-3]
[0066] [Table 13-4] nnfpc_separate_colour_description_present_flag equal to 1 indicates that the separate combination of colour primaries, transfer characteristics and matrix coefficients for the picture resulting from the post-processing filter is specified in the SEI message syntax structure. nnpfc_separate_colour_description_present_flag equal to 0 indicates that the combination of colour primaries, transfer characteristics and matrix coefficients for the picture resulting from the post-processing filter is the same as that indicated in the VUI parameters for CLVS. nnpfc_colour_primaries has the same semantics as specified for the vui_colour_primaries syntax element, which is as follows: vui_colour_primaries indicates the chromaticity coordinates of the source primaries. Its semantics are as specified for MatrixCoefficients in Rec. ITU-T H.273|ISO / IEC 23091-2. When the vui_colour_primaries syntax element is not present, the value of vui_colour_primaries is inferred to be equal to 2 (chromaticity is unknown, unspecified, or determined by other means not specified in this specification). Values of vui_transfer_characteristics identified as reserved for future use in Rec. ITU-T H.273|ISO / IEC 23091-2 shall not be present in bitstreams conforming to this version of this specification. Decoding devices shall interpret reserved values of vui_colour_primaries as equal to the value 2. However, the following applies: -nnpfc_colour_primaries specifies the picture primaries resulting from applying the neural network postfilter specified in the SEI message, rather than the primaries used for CLVS. When -nnpfc_colour_primaries is not present in the NNPFC SEI message, the value of nnpfc_colour_primaries is inferred to be equal to vui_colour_primaries. nnpfc_transfer_characteristics has the same semantics as specified for the vui_transfer_characteristics syntax element, which is as follows: vui_transfer_characteristics indicates the transfer characteristic function of the color representation. Its semantics are as specified for MatrixCoefficients in Rec. ITU-T H.273 | ISO / IEC 23091-2. When the vui_transfer_characteristics syntax element is not present, the value of vui_transfer_characteristics is: vui_transfer_characteristics is inferred to be equal to the value 2 (transfer characteristics are unknown or unspecified, or determined by other means not specified in this specification). Values of vui_transfer_characteristics identified as reserved for future use in Rec. ITU-T H.273|ISO / IEC 23091-2 shall not be present in bitstreams conforming to this version of this specification. Decoding devices shall interpret reserved values of vui_transfer_characteristics as equal to the value 2. However, the following applies: -nnpfc_transfer_characteristics specifies the transfer characteristics of the picture resulting from applying the neural network postfilter specified in the SEI message, instead of the transfer characteristics used for CLVS. When -nnpfc_transfer_characteristics is not present in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is inferred to be equal to vui_transfer_characteristics. nnpfc_matrix_coeffs has the same semantics as specified for the vui_matrix_coeffs syntax element, which is: vui_matrix_coeffs describes the formulas used in deriving the luma and chroma signals from the green, blue, and red, or Y, Z, and X primaries. The semantics are as specified for MatrixCoefficients in Rec. ITU-T H.273|ISO / IEC 23091-2, with the following: -nnpfc_matrix_coeffs specifies the matrix coefficients of the picture resulting from applying the neural network postfilter specified in the SEI message, not the matrix coefficients used for CLVS. When -nnpfc_matrix_coeffs is not present in the NNPFC SEI message, the value of nnpfc_matrix_coeffs is inferred to be equal to vui_matrix_coeffs. The allowed values for -nnpfc_matrix_coeffs are not constrained by the chroma format of the decoded video picture indicated by the value of ChromaFormatIdc for the semantics of the VUI parameter. - when nnpfc_matrix_coeffs is equal to 0, nnpfc_out_order_idc shall not be equal to 1 or 3. An nnpfc_out_format_idc equal to 0 indicates that the sample values output by the post - processing filter are real numbers, and the value range of 0 to 1 including both end - points is linearly mapped to an unsigned integer value range of 0 to (1<<bitDepth)-1 including both end - points for any desired bit depth bitDepth for subsequent post - processing or display. An nnpfc_out_format_flag equal to 1 indicates that the sample values output by the post - processing filter are unsigned integers within the range of 0 to (1<<(nnpfc_out_tensor_bitdepth_minus8 + 8))-1 including both end - points. Values of nnpfc_out_format_idc greater than 1 are reserved for future specifications by ITU - T|ISO / IEC and shall not be present in bitstreams compliant with this version of this document. Decoders compliant with this version of this document shall ignore NNPFC SEI messages containing reserved values of nnpfc_out_format_idc. nnpfc_out_tensor_bitdepth_minus8 plus 8 specifies the bit depth of the sample values in the output integer tensor. The value of nnpfc_out_tensor_bitdepth_minus8 shall be within the range of 0 to 24 including both end - points. nnpfc_out_order_idc indicates the output order of samples obtained from the post - processing filter. Values of nnpfc_out_order_idc shall be in the range 0 to 3, inclusive, in bitstreams conforming to this version of this document. Values of nnpfc_out_order_idc between 4 and 255, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_out_order_idc in the range 4 to 255, inclusive. Values of nnpfc_out_order_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use. When nnpfc_purpose is equal to 2 or 4, nnpfc_out_order_idc shall not be equal to 3. Table 10 contains useful descriptions of the nnpfc_out_order_idc values.
[0067] [Table 14] The process StoreOutputTensors() for deriving sample values in filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor given vertical sample coordinates cTop and horizontal sample coordinates cLeft that specify the top left sample location of a patch of samples contained in the input tensor is specified as follows:
[0068] [Table 15-1]
[0069] [Table 15-2]
[0070] [Table 15-3] nnpfc_constant_patch_size_flag equal to 1 indicates that the post-processing filter accepts as input exactly the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. nnpfc_constant_patch_size_flag equal to 0 indicates that the post-processing filter accepts as input any patch size that is a positive integer multiple of the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. nnpfc_patch_width_minus1+1 indicates the horizontal sample count of the patch size required for input to the post-processing filter when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_width_minus1 shall be in the range 0 to Min(32766,CroppedWidth-1), inclusive. nnpfc_patch_height_minus1+1 indicates the vertical sample count of the patch size required for input to the post-processing filter when nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_height_minus1 shall be in the range 0 to Min(32766,CroppedHeight-1), inclusive. The variables inpPatchWidth and inpPatchHeight are the patch size width and patch size height, respectively. If nnpfc_constant_patch_size_flag is equal to 0, the following applies: The values of -inpPatchWidth and inpPatchHeight are either provided by external means not specified in this document, or are set by the post-processor itself. - The value of inpPatchWidth shall be a positive integer multiple of nnpfc_patch_width_minus1+1 and shall be less than or equal to CroppedWidth. The value of inpPatchHeight shall be a positive integer multiple of nnpfc_patch_height_minus1+1 and shall be less than or equal to CroppedHeight. Otherwise (nnpfc_constant_patch_size_flag is equal to 1), the value of inpPatchWidth is set equal to nnpfc_patch_width_minus1+1, and the value of inpPatchHeight is set equal to nnpfc_patch_height_minus1+1. nnpfc_overlap indicates the overlapping horizontal and vertical sample counts of adjacent input tensors of the post-processing filter. The value of nnpfc_overlap must be in the range of 0 to 16383, inclusive. The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows: outPatchWidth=(nnpfc_pic_width_in_luma_samples * inpPatchWidth) / CroppedWidth outPatchHeight=(nnpfc_pic_height_in_luma_samples * inpPatchHeight) / CroppedHeight horCScaling=SubWidthC / outSubWidthC verCScaling=SubHeightC / outSubHeightC outPatchCWidth=outPatchWidth * horCScaling outPatchCHeight=outPatchHeight *verCScaling overlapSize=nnpfc_overlap outPatchWidth * CroppedWidth is nnpfc_pic_width_in_luma_samples * shall be equal to inpPatchWidth, and outPatchHeight * CroppedHeight is nnpfc_pic_height_in_luma_samples * It is a bitstream conformance requirement that it shall be equal to inpPatchHeight. nnpfc_padding_type indicates the padding process when referring to sample positions outside the boundary of the cropped decoded output picture as described in Table 11. The value of nnpfc_padding_type shall be in the range of 0 to 15, inclusive.
[0071] [Table 16] nnpfc_luma_padding_val indicates the luma value used for padding when nnpfc_padding_type is equal to 4. nnpfc_cb_padding_val indicates the Cb value used for padding when nnpfc_padding_type is equal to 4. nnpfc_cr_padding_val indicates the Cr value used for padding when nnpfc_padding_type is equal to 4. The function InpSampleVal(y, x, picHeight, picWidth, croppedPic) whose inputs are vertical sample position y, horizontal sample position x, picture height picHeight, picture width picWidth, and sample array croppedPic returns the value of sampleVal derived as follows: NOTE - For input to function InpSampleVal(), the vertical positions are listed before the horizontal positions for compatibility with the input tensor conventions of some inference engines.
[0072] [Table 17] The following example process may be used to filter the cropped decoded output picture in patches using a post-processing filter PostProcessingFilter() to generate a filtered picture including Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated by nnpfc_out_order_idc.
[0073] [Table 18] nnpfc_complexity_info_present_flag equal to 1 specifies that there are one or more syntax elements indicating the complexity of the post-processing filter associated with nnpfc_id. nnpfc_complexity_info_present_flag equal to 0 specifies that there are no syntax elements indicating the complexity of the post-processing filter associated with nnpfc_id. nnpfc_parameter_type_idc equal to 0 indicates that the neural network uses integer parameters only. nnpfc_parameter_type_flag equal to 1 indicates that the neural network may use floating point or integer parameters. nnpfc_parameter_type_idc equal to 2 indicates that the neural network uses binary parameters only. nnpfc_parameter_type_idc equal to 3 is reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_parameter_type_idc equal to 3. nnpfc_log2_parameter_bit_length_minus3 equal to 0, 1, 2, and 3 indicates that the neural network shall not use parameters with bit lengths greater than 8, 16, 32, and 64, respectively. When nnpfc_parameter_type_idc is present and nnpfc_log2_parameter_bit_length_minus3 is not present, the neural network shall not use parameters with bit lengths greater than 1. nnpfc_num_parameters_idc indicates the maximum number of neural network parameters for the post-processing filters, in units of powers of 2048. nnpfc_num_parameters_idc equal to 0 indicates that the maximum number of neural network parameters is unknown. The value nnpfc_num_parameters_idc shall be in the range 0 to 52, inclusive. Values of nnpfc_num_parameters_idc greater than 52 are reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoding devices conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_num_parameters_idc greater than 52. If the value of nnpfc_num_parameters_idc is greater than 0, the variable maxNumParameters is derived as follows: maxNumParameters=(2048< <nnpfc_num_parameters_idc)-1 It is a bitstream conformance requirement that the number of neural network parameters in the post-processing filter shall be less than or equal to maxNumParameters. nnpfc_num_kmac_operations_idc greater than 0 specifies the maximum number of multiply-accumulate operations per sample for the post-processing filter. * nnpfc_num_kmac_operations_idc equal to 0 indicates that the maximum number of multiply-add operations in the network is unknown. The value of nnpfc_num_kmac_operations_idc can range from 0 to 2, inclusive. 32 It should be in the range of -1. nnpfc_total_kilobyte_size greater than 0 indicates the total size in kilobytes required to store the uncompressed parameters for the neural network. The total size in bits is equal to or greater than the sum of the bits used to store each parameter. nnpfc_total_kilobyte_size is the total size in bits divided by 8000 and rounded up. nnpfc_total_kilobyte_size equal to 0 indicates that the total size required to store the parameters of the neural network is unknown. The value of nnpfc_total_kilobyte_size can be between 0 and 2, inclusive. 32 It should be in the range of -1. nnpfc_reserved_zero_bit_b shall be equal to 0 in bitstreams conforming to this version of this document. Decoders shall ignore NPFC SEI messages in which nnpfc_reserved_zero_bit_b is not equal to 0. Let nnpfc_payload_byte[i] contain the i-th byte of a bitstream that conforms to ISO / IEC 15938-17. The byte sequence nnpfc_payload_byte[i] for all present values of i shall be a complete bitstream that conforms to ISO / IEC 15938-17.
[0074] Table 12 shows the syntax of the neural network post-filter activation SEI message defined in JVET-AB2006.
[0075] [Table 19]
[0076] With respect to Table 12, JVET-AB2006 specifies the following semantics: The neural-network post-filter activation (NNPFA) SEI message activates or deactivates the possible use of a target neural-network post-processing filter, identified by nnpfa_target_id, for post-processing filtering of a set of pictures. NOTE − There may be several NNPFA SEI messages for the same picture, for example when post-processing filters are for different purposes or filter different color components. nnpfa_target_id indicates a target neural network post-processing filter, which is specified by one or more neural network post-processing filter characteristics SEI messages associated with the current picture and having nnpfc_id equal to nnfpa_target_id. The value of nnpfa_target_id is between 0 and 2, inclusive. 32 -2. The range is from 256 to 511, inclusive, and 256 to 511, inclusive. 31 ~232 Values of nnpfa_target_id up to -2 are reserved for future use by ITU-T | ISO / IEC. 31 ~2 32 Decoders compliant with this version of this document that encounter an NNPFA SEI message with an nnpfa_target_id within the range of -2 shall ignore that SEI message. An NNPFA SEI message with a particular value of nnpfa_target_id shall not be present in the current PU unless one or both of the following conditions are true: - There is an NNPFC SEI message in the current CLVS with nnpfc_id equal to a particular value of nnpfa_target_id present in a PU preceding the current PU in decoding order. - There is an NNPFC SEI message in the current PU with nnpfc_id equal to a particular value of nnpfa_target_id. When a PU contains both an NNPFC SEI message with a particular value of nnpfc_id and an NNPFA SEI message with nnpfa_target_id equal to a particular value of nnpfc_id, the NNPFC SEI message shall precede the NNPFA SEI message in decoding order. nnpfa_cancel_flag equal to 1 indicates that the persistence of the target neural network post-processing filter established by any previous NNPFA SEI message with the same nnpfa_target_id as the current SEI message is canceled, i.e., the target neural network post-processing filter will no longer be used unless activated by another NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 0. nnpfa_cancel_flag equal to 0 indicates that nnpfa_persistence_flag follows. nnpfa_persistence_flag specifies the persistence of the target neural network post-processing filter for the current layer. nnpfa_persistence_flag equal to 0 specifies that the target neural network post-processing filter may be used only for post-processing filtering for the current picture. nnpfa_persistence_flag equal to 1 specifies that the target neural network post-processing filter may be used for post-processing filtering for the current picture and all subsequent pictures in the current layer in output order until one or more of the following conditions become true: -A new CLVS for the current layer begins. - The bitstream ends. - The picture in the current layer associated with an NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1 is the output following the current picture in output order. NOTE - The target neural network post-processing filter is not applied to this subsequent picture in the current layer associated with an NNPFA SEI message that has the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1.
[0077] The neural network post-filter characteristic SEI message defined in JVET-AB2006 may be less than ideal. In particular, for example, the signaling in JVET-AB2006 may be insufficient to allow a receiving device to determine whether a particular message is useful. In accordance with the techniques described herein, additional syntax and semantics are provided, and the placement of syntax elements is optimized.
[0078] FIG. 1 is a block diagram illustrating an example of a system that may be configured to code (i.e., encode and / or decode) video data in accordance with one or more techniques of this disclosure. System 100 represents an example of a system that may encapsulate video data in accordance with one or more techniques of this disclosure. As shown in FIG. 1, system 100 includes a source device 102, a communication medium 110, and a destination device 120. In the example shown in FIG. 1, source device 102 may include any device configured to encode video data and transmit the encoded video data to communication medium 110. Destination device 120 may include any device configured to receive the encoded video data via communication medium 110 and decode the encoded video data. Source device 102 and / or destination device 120 may include computing devices equipped for wired and / or wireless communication, and may include, for example, set-top boxes, digital video recorders, televisions, desktop, laptop, or tablet computers, gaming consoles, medical imaging devices, and mobile devices including, for example, smartphones, cellular telephones, and personal gaming devices.
[0079] The communication medium 110 may include any combination of wireless and wired communication media and / or storage devices. The communication medium 110 may include coaxial cables, fiber optic cables, twisted pair cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other equipment that may be useful for facilitating communication between various devices and sites. The communication medium 110 may include one or more networks. For example, the communication medium 110 may include a network configured to enable access to the World Wide Web, e.g., the Internet. The network may operate according to a combination of one or more telecommunications protocols. The telecommunications protocols may include proprietary aspects and / or standardized telecommunications protocols. Examples of standardized telecommunications protocols include Digital Video Broadcasting (DVB) standards, Advanced Television Systems Committee (ATSC) standards, Integrated Services Digital Broadcasting (ISDB) standards, Data Over Cable Service Interface Specification (DOCSIS) standards, Global System Mobile Communications (GSM) standards, code division multiple access (CDMA) standards, 3rd Generation Partnership Project (3GPP) standards, European Telecommunications Standards Institute (ETSI) standards, Internet Protocol (IP) standards, Wireless Application Protocol (WAP) standards, and Institute of Electrical and Electronics Engineers (IEEE) standards.
[0080] A storage device may include any type of device or storage medium capable of storing data. A storage medium may include a tangible or non-transitory computer-readable medium. A computer-readable medium may include an optical disk, a flash memory, a magnetic memory, or any other suitable digital storage medium. In some examples, a memory device or a portion thereof may be described as a non-volatile memory, and in other examples, a portion of a memory device may be described as a volatile memory. Examples of volatile memory may include random access memory (RAM), dynamic random access memory (DRAM), and static random access memory (SRAM). Examples of non-volatile memory may include a magnetic hard disk, an optical disk, a floppy disk, a flash memory, or a form of electrically programmable memory (EPROM) or electrically erasable and programmable memory (EEPROM). The storage device(s) may include memory cards (e.g., Secure Digital (SD) memory cards), internal / external hard disk drives, and / or internal / external solid state drives. Data may be stored on the storage device according to a defined file format.
[0081] FIG. 4 is a conceptual diagram illustrating an example of components that may be included in one implementation of system 100. In the exemplary implementation shown in FIG. 4, system 100 includes one or more computing devices 402A-402N, a television service network 404, a television service provider site 406, a wide area network 408, a local area network 410, and one or more content provider sites 412A-412N. The implementation shown in FIG. 4 represents an example of a system that may be configured to enable digital media content, such as movies, live sporting events, and data and applications and their associated media presentations, to be distributed to and accessed by multiple computing devices, such as computing devices 402A-402N. In the example shown in FIG. 4, computing devices 402A-402N may include any device configured to receive data from one or more of television service network 404, wide area network 408, and / or local area network 410. For example, the computing devices 402A-402N may be equipped for wired and / or wireless communication and may be configured to receive services over one or more data channels, and may include televisions, including so-called smart televisions, set-top boxes, and digital video recorders. Additionally, the computing devices 402A-402N may include desktop, laptop, or tablet computers, gaming consoles, mobile devices, including, for example, "smart" phones, cellular phones, and personal gaming devices.
[0082] The television service network 404 is an example of a network configured to enable delivery of digital media content, which may include television services. For example, the television service network 404 may include a public terrestrial television network, a public or subscription-based satellite television service provider network, and a public or subscription-based cable television provider network and / or an over the top service provider or an Internet service provider. It should be noted that, in some examples, the television service network 404 may be primarily used to enable the provision of television services, but the television service network 404 may also enable the provision of other types of data and services based on any combination of telecommunication protocols described herein. Furthermore, it should be noted that in some examples, the television service network 404 may enable bidirectional communication between the television service provider site 406 and one or more of the computing devices 402A-402N. The television service network 404 may include any combination of wireless communication media and / or wired communication media. The television service network 404 may include coaxial cables, fiber optic cables, twisted pair cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other equipment that may be useful for facilitating communication between various devices and sites. The television service network 404 may operate according to a combination of one or more telecommunications protocols. The telecommunications protocols may include proprietary aspects and / or standardized telecommunications protocols. Examples of standardized telecommunications protocols include the DVB standard, the ATSC standard, the ISDB standard, the DTMB standard, the DMB standard, the Data Over Cable Service Interface Specification (DOCSIS) standard, the HbbTV standard, the W3C standard, and the UPnP standard.
[0083] Referring again to FIG. 4, the television service provider site 406 may be configured to distribute television services over the television service network 404. For example, the television service provider site 406 may include one or more broadcast stations, cable television providers, or satellite television providers, or Internet-based television providers. For example, the television service provider site 406 may be configured to receive transmissions including television programs over satellite uplinks / downlinks. Additionally, as shown in FIG. 4, the television service provider site 406 may be in communication with a wide area network 408 and configured to receive data from content provider sites 412A-412N. It should be noted that in some examples, the television service provider site 406 may include a television studio from which content may originate.
[0084] The wide area network 408 may include a packet-based network and may operate according to a combination of one or more telecommunications protocols. The telecommunications protocols may include proprietary aspects and / or standardized telecommunications protocols. Examples of standardized telecommunications protocols include the Global System Mobile Communications (GSM) standard, the code division multiple access (CDMA) standard, the Third Generation Partnership Project (3GPP) standard, and the IEEE 802.11 standard. rdExamples of standards that may be used include 3GPP (International Telecommunications Standards Project), ETSI (European Telecommunications Standards Institute) standards, EN (European Standards), IP standards, Wireless Application Protocol (WAP) standards, and Institute of Electrical and Electronics Engineers (IEEE) standards, such as one or more of the IEEE 802 standards (e.g., Wi-Fi). Wide area network 408 may include any combination of wireless communication media and / or wired communication media. Wide area network 408 may include coaxial cables, fiber optic cables, twisted pair cables, Ethernet cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other equipment that may be useful for facilitating communication between various devices and sites. In one embodiment, wide area network 408 may include the Internet. Local area network 410 may include packet-based networks and operate according to a combination of one or more telecommunication protocols. Local area network 410 may be differentiated from wide area network 408 based on the level of access and / or physical infrastructure. For example, the local area network 410 may include a secure home network.
[0085] Referring again to FIG. 4, the content provider sites 412A-412N represent examples of sites that may provide multimedia content to the television service provider site 406 and / or the computing devices 402A-402N. For example, the content provider sites may include studios having one or more studio content servers configured to provide multimedia files and / or streams to the television service provider site 406. In one embodiment, the content provider sites 412A-412N may be configured to provide multimedia content using an IP suite. For example, the content provider sites may be configured to provide multimedia content to receiving devices according to the Real Time Streaming Protocol (RTSP), HTTP, or the like. Additionally, the content provider sites 412A-412N may be configured to provide data, including hypertext-based content, or the like, over the wide area network 408 to one or more of the receiving devices, the computing devices 402A-402N, and / or the television service provider site 406. The content provider sites 412A-412N may include one or more web servers. The data provided by the content provider sites 412A-412N may be defined according to a data format.
[0086] Referring again to FIG. 1, source device 102 includes video source 104, video encoder 106, data encapsulator 107, and interface 108. Video source 104 may include any device configured to capture and / or store video data. For example, video source 104 may include a video camera and a storage device operatively coupled thereto. Video encoder 106 may include any device configured to receive video data and generate a compliant bitstream representing the video data. A compliant bitstream may refer to a bitstream that a video decoder device can receive and from which the video data can be regenerated. The aspects of a compliant bitstream may be defined according to a video encoding standard. When generating a compliant bitstream, video encoder 106 may compress the video data. The compression may be lossy (perceptible or imperceptible to a viewer) or lossless. FIG. 5 is a block diagram illustrating an example of a video encoder 500 that may implement techniques for encoding video data described herein. It should be noted that while the exemplary video encoding device 500 is illustrated as having distinct functional blocks, such illustration is for purposes of explanation and is not intended to limit the video encoding device 500 and / or its subcomponents to any particular hardware or software architecture. The functionality of the video encoding device 500 may be realized using any combination of hardware, firmware, and / or software implementations.
[0087] The video encoding device 500 may perform intra-predictive and inter-predictive encoding of picture portions, and may therefore be referred to as a hybrid video encoding device. In the example shown in FIG. 5, the video encoding device 500 receives a source video block. In some examples, the source video block may include a portion of a picture that has been partitioned according to a coding structure. For example, the source video data may include macroblocks, CTUs, CBs, subdivisions thereof, and / or other equivalent coding units. In some examples, the video encoding device 500 may be configured to perform additional subdivisions of the source video block. It should be noted that the techniques described herein are generally applicable to video encoding, regardless of how the source video data is partitioned before and / or during encoding. In the example shown in Figure 5, the video encoding device 500 includes an adder 502, a transform coefficient generating unit 504, a coefficient quantization unit 506, an inverse quantization and transform coefficient processing unit 508, an adder 510, an intra prediction processing unit 512, an inter prediction processing unit 514, a filter unit 516, and an entropy encoding unit 518. As shown in Figure 5, the video encoding device 500 receives source video blocks and outputs a bitstream.
[0088] In the example shown in FIG. 5, the video encoding device 500 may generate residual data by subtracting a predictive video block from a source video block. Selection of the predictive video block is described in more detail below. The adder 502 represents a component configured to perform this subtraction operation. In one embodiment, the subtraction of the video blocks is performed in the pixel domain. The transform coefficient generator 504 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block or a subdivision thereof (e.g., four 8×8 transforms may be applied to a 16×16 array of residual values) to generate a set of residual transform coefficients. The transform coefficient generator 504 may be configured to perform any and all combinations of transforms included in the family of discrete triangular transforms, including approximations of the discrete triangular transform. The transform coefficient generator 504 may output the transform coefficients to a coefficient quantizer 506. The coefficient quantizer 506 may be configured to perform quantization of the transform coefficients. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may change the rate distortion (i.e., video bit rate vs. quality) of the coded video data. The degree of quantization may be changed by adjusting a quantization parameter (QP). The quantization parameter may be determined based on slice level values and / or CU level values (e.g., CU delta QP value). QP data may include any data used to determine a QP for quantizing a particular set of transform coefficients. As shown in FIG. 5, the quantized transform coefficients (which may be referred to as level values) are output to an inverse quantization and transform coefficient processing unit 508. The inverse quantization and transform coefficient processing unit 508 may be configured to apply inverse quantization and inverse transform to generate reconstructed residual data. As shown in FIG. 5, the reconstructed residual data may be added to a prediction video block at an adder 510. In this manner, the coded video block may be reconstructed, and the resulting reconstructed video block may be used to evaluate the encoding quality for a given prediction, transform, and / or quantization.The video encoding device 500 may be configured to perform multiple encoding passes (e.g., performing encoding while varying one or more of prediction, transformation parameters, and quantization parameters). The rate-distortion or other system parameters of the bitstream may be optimized based on evaluation of the reconstructed video blocks. Additionally, the reconstructed video blocks may be stored and used as references for predicting subsequent blocks.
[0089] Referring to FIG. 5, the intra-predictor 512 may be configured to select an intra-prediction mode for a video block to be coded. The intra-predictor 512 may be configured to evaluate a frame and determine an intra-prediction mode to use to code a current block. As described above, possible intra-prediction modes may include a planar prediction mode, a DC prediction mode, and an angular prediction mode. Additionally, it should be noted that in some examples, a prediction mode for a chroma component may be inferred from a prediction mode for a luma prediction mode. The intra-predictor 512 may select an intra-prediction mode after performing one or more coding passes. Additionally, in one embodiment, the intra-predictor 512 may select a prediction mode based on a rate-distortion analysis. As shown in FIG. 5, the intra-predictor 512 outputs intra-prediction data (e.g., syntax elements) to the entropy encoder 518 and the transform coefficient generator 504. As described above, the transform performed on the residual data may be mode-dependent (e.g., a secondary transform matrix may be determined based on the prediction mode).
[0090] Referring again to FIG. 5, the inter prediction processor 514 may be configured to perform inter prediction coding on the current video block. The inter prediction processor 514 may be configured to receive a source video block and calculate a motion vector for a PU of the video block. The motion vector may indicate a displacement of a prediction unit of the video block in the current video frame relative to a prediction block in a reference frame. The inter prediction coding may use one or more reference pictures. Furthermore, the motion prediction may be uni-predictive (using one motion vector) or bi-predictive (using two motion vectors). The inter prediction processor 514 may be configured to select a prediction block by calculating pixel differences determined by, for example, sum of absolute difference (SAD), sum of square difference (SSD), or other difference measures. As described above, a motion vector may be determined and determined according to the motion vector prediction. The inter prediction processor 514 may be configured to perform motion vector prediction as described above. The inter prediction processor 514 may be configured to generate a prediction block using the motion prediction data. For example, the inter prediction processor 514 may place the prediction video block in a frame buffer (not shown in FIG. 5). Note that the inter prediction processor 514 may be further configured to apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values for use in motion prediction. The inter prediction processor 514 may output motion prediction data for the calculated motion vectors to the entropy encoder 518.
[0091] Referring back to FIG. 5, the filter unit 516 receives the reconstructed video blocks and the coding parameters and outputs modified reconstructed video data. The filter unit 516 may be configured to perform deblocking and / or sample adaptive offset (SAO) filtering. SAO filtering is a non-linear amplitude mapping that may be used to improve reconstruction by adding an offset to the reconstructed video data. It should be noted that, as shown in FIG. 5, the intra-prediction unit 512 and the inter-prediction unit 514 may receive the modified reconstructed video blocks via the filter unit 216. The entropy encoder 518 receives the quantized transform coefficients and prediction syntax data (i.e., intra-prediction data and motion prediction data). It should be noted that in some examples, the coefficient quantizer 506 may perform a scan of a matrix including the quantized transform coefficients before the coefficients are output to the entropy encoder 518. In other examples, the entropy encoder 518 may perform a scan. The entropy encoder 518 may be configured to perform entropy encoding according to one or more of the techniques described herein. Thus, video encoding apparatus 500 represents one example of a device configured to generate encoded video data in accordance with one or more techniques of this disclosure.
[0092] Referring again to FIG. 1, data encapsulator 107 may receive encoded video data and generate a compliant bitstream, such as a series of NAL units, according to a defined data structure. A device receiving the compliant bitstream may regenerate the video data therefrom. Additionally, as described above, sub-bitstream extraction may refer to a process in which a device receiving a compliant bitstream forms a new compliant bitstream by discarding and / or modifying data in the received bitstream. Note that the term conforming bitstream may be used instead of the term compliant bitstream. In one embodiment, data encapsulator 107 may be configured to generate syntax according to one or more techniques described herein. Note that data encapsulator 107 need not be located in the same physical device as video encoder 106. For example, the functions described as being performed by video encoder 106 and data encapsulator 107 may be distributed between the devices shown in FIG. 4.
[0093] As mentioned above, the signaling defined in JVET-AB2006 may be insufficient. In accordance with the techniques herein, in one embodiment, a syntax element nnpfc_purpose is provided earlier in the NNPFC SEI message to allow the receiver device to determine whether a particular NNPFC SEI message has a purpose that is useful or interesting to the receiver device, and whether the receiver device needs to parse other data in the message based on that purpose. Furthermore, in one embodiment, in accordance with the techniques herein, nnpfc_purpose may be signaled unconditionally. Note that knowing the purpose of the NNPFC message directly, rather than indirectly by keeping track of the nnpfc_id, allows the receiver device to quickly and explicitly know the purpose information. It should be noted that according to JVET-AB2006, the first NNPFC SEI message in decoding order with a particular nnpfc_id value in the current CLVS (having nnpfc_formatting_and_purpose_flag equal to 1 and containing the nnpfc_purpose syntax element) may be lost and a second NNPFC SEI message may be signaled with the same nnpfc_id value and nnpfc_formatting_and_purpose_flag equal to 0. In this case, additional parsing is required before discovering the purpose of the second NNPFC SEI message. Currently, the additional parsing includes parsing the nnpfc_mode_idc and, if it is equal to 1, parsing any byte-aligned nnpfc_reserved_zero_bit bit, the nnpfc_tag_uri, the nnpfc_uri, and then the nnpfc_formatting_and_purpose_flag syntax element. It is argued that the techniques herein will help mitigate such scenarios by moving the nnpfc_purpose flag earlier in the NNPFC SEI message and signaling it unconditionally.Further, in accordance with the techniques herein, in one embodiment, the flag nnpfc_formatting_and_purpose_flag may be renamed to nnpfc_formatting_flag.
[0094] In one embodiment, in accordance with the techniques herein, a syntax element is provided in the neural network postfilter characteristics SEI message that indicates whether the input data is manipulated according to normalized data. Tables 13A-13D show the associated syntax of an example neural network postfilter characteristics SEI message in which the nnpfc_purpose syntax element is moved in accordance with the techniques herein.
[0095] [Table 20]
[0096] [Table 21]
[0097] [Table 22]
[0098] [Table 23]
[0099] In another embodiment, the nnpfc_purpose syntax element may be moved up as described in one of the embodiments above and still be conditionally signaled based on the nnpfc_formatting_and_purpose_flag, and thus in these cases the position of the nnpfc_formatting_and_purpose_flag may also be moved up. Tables 14A-14C show the associated syntax of an example Neural Network Post Filter Properties SEI message in which the nnpfc_purpose syntax element is moved, in accordance with the techniques herein.
[0100] [Table 24]
[0101] [Table 25]
[0102] [Table 26]
[0103] For Tables 13A-14C, the semantics may be based on the semantics provided above.
[0104] The semantics of JVET-AB2006 assert that for the syntax element nnpfc_out_order_idc, the outermost for loop defining the process StoreOutputTensors() should use the variable numOutputPics instead of the variable numInputPics since this process should be applied to the number of output pictures. Furthermore, the semantics of JVET-AB2006 assert that the variable numOutputPics is not derived when nnpfc_purpose has a value other than the value 5. It asserts that the variable numOutputPics should be used in the formula of nnpfc_outout_order_idc defining the process StoreOutputTensors(). Thus, according to the techniques herein, the variable numOutputPics is initialized when nnpfc_purpose is not equal to 5. In one embodiment, according to the techniques herein, the semantics of nnpfc_out_order_idc and nnpfc_interpolated_pics may be as follows: nnpfc_out_order_idc indicates the output order of the samples coming from the post-processing filter. Values of nnpfc_out_order_idc shall be in the range 0 to 3, inclusive, in bitstreams conforming to this version of this document. Values of nnpfc_out_order_idc between 4 and 255, inclusive, are reserved for future use by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_out_order_idc in the range 4 to 255, inclusive. Values of nnpfc_out_order_idc greater than 255 shall not be present in bitstreams conforming to this version of this document and are not reserved for future use. When nnpfc_purpose is equal to 2 or 4, nnpfc_out_order_idc shall not be equal to 3. Table 10 contains useful descriptions of the nnpfc_out_order_idc values. The process StoreOutputTensors() for deriving sample values in filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor given vertical sample coordinates cTop and horizontal sample coordinates cLeft that specify the top left sample location of a patch of samples contained in the input tensor is specified as follows:
[0105] [Table 27-1]
[0106] [Table 27-2]
[0107] [Table 27-3] nnpfc_interpolated_pics[i] specifies the number of interpolated pictures generated by the post-processing filter between the i-th picture and the (i+1)-th picture that are used as input for the post-processing filter. The variables numInputPics, which specify the number of pictures used as input for the post-processing filter, and numOutputPics, which specify the total number of pictures resulting from the post-processing filter, are derived as follows:
[0108] [Table 28]
[0109] In another embodiment, the above modifications to the for loop and to initializing numOutputPic to be equal to 1 when nnpfc_purpose is not equal to 5 may be included in the semantics of other syntax elements of the NNPFC SEI message.
[0110] In one embodiment, in accordance with the techniques herein, valid value ranges may be defined for two ue(v) coding syntax elements nnpfc_num_input_pics_minus2 and nnpfc_interpolated_pics[i]. In JVET-AB2006, there are no ranges specified for these. This may make it difficult for a decoder device to allocate suitable buffers. In one embodiment, in accordance with the techniques herein, the semantics of nnpfc_num_input_pics_minus2 and nnpfc_interpolated_pics[i] may be as follows: nnpfc_num_input_pics_minus2 plus 2 specifies the number of decoded output pictures used as input for the post-processing filter. The value of nnpfc_num_input_pics_minus2 shall be in the range 0 to 14, inclusive, in bitstreams conforming to this version of this document. In alternative embodiments, values different from 14 may be used as the upper limit and for defining the range of the syntax element nnpfc_num_input_pics_minus2, for example values of 30 or 62 or 126 or some other value may be used. nnpfc_interpolated_pics[i] specifies the number of interpolated pictures generated by the post-processing filter between the i-th picture and the (i+1)-th picture that are used as input for the post-processing filter. The value of nnpfc_interpolated_pics[i] shall be in the range 0 to 64, inclusive, in bitstreams conforming to this version of this document. In alternative embodiments, values different from 64 may be used as the upper limit and for defining the range of the syntax element nnpfc_interpolated_pics[i], for example values of 32 or 96 or 128 or 256 or some other value may be used.
[0111] According to JVET-AB2006, the NNPFC SEI message takes separate bit depths for luma and chroma as input. Y and BitDepth C ) and defines two separate functions InpY(x) and InpC(x) for luma and chroma samples. However, in the semantics, InpC(x) for chroma is calculated using the inpTensorBitDepth variable, which is calculated based on the syntax element nnpfc_inp_tensor_bitdepth_minus8. However, the semantics of nnpfc_inp_tensor_bitdepth_minus8 specifies that it plus 8 specifies the bit depth of the luma sample values in the input integer tensor. In one embodiment, in accordance with the techniques herein, an additional syntax element nnpfc_inp_tensor_bitdepth_chroma_minus8 may be defined to specify the bit depth of the chroma sample values in the input integer tensor. The NNPFC SEI message in JVET-AB2006 defines separate bit value depths BitDepth for the luma sample array and the chroma sample array of the cropped decoded output picture. Y and BitDepth Cand because the computation in the semantics for nnpfc_inp_format_idc defines separate functions InpY() and InpC(), it is argued that it becomes more flexible to define a separate syntax element nnpfc_inp_tensor_bitdepth_chroma_minus8 for specifying a separate bit-depth of the chroma sample values in the input integer tensor that may differ from the bit-depth of the luma sample values in the input integer tensor. Table 15 shows the relevant syntax of an example Neural Network Post Filter Properties SEI message with syntax element nnpfc_inp_tensor_bitdepth_chroma_minus8 in accordance with the techniques herein.
[0112] [Table 29]
[0113] With respect to Table 15, the semantics may be based on those provided above and the following: nnpfc_inp_format_idc indicates how to convert the sample values of the cropped decoded output picture into input values to the post-processing filter. When nnpfc_inp_format_idc is equal to 0, the input values to the post-processing filter are real numbers, and the functions InpY() and InpC() are specified as follows: InpY(x) = x÷((1< <BitDepth Y )-1) InpC(x) = x ÷ ((1< <BitDepth C )-1) When nnpfc_inp_format_idc is equal to 1, the input values to the post-processing filter are unsigned integers and the functions InpY() and InpC() are specified as follows:
[0114] [Table 30] The variable inpTensorBitDepth is derived from the syntax element nnpfc_inp_tensor_bitdepth_luma_minus8, as specified below. The variable inpTensorBitDepthChroma is derived from the syntax element nnpfc_inp_tensor_bitdepth_chroma_minus8, as specified below. Values of nnpfc_inp_format_idc greater than 1 are reserved for future specification by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages which contain reserved values of nnpfc_inp_format_idc. nnpfc_inp_tensor_bitdepth_luma_minus8 plus 8 specifies the bit depth of the luma sample values in the input integer tensor. The value of inpTensorBitDepth is derived as follows: inpTensorBitDepth=nnpfc_inp_tensor_bitdepth_luma_minus8+8 It is a bitstream conformance requirement that the value of nnpfc_inp_tensor_bitdepth_luma_minus8 be in the range 0 to 24, inclusive. nnpfc_inp_tensor_bitdepth_chroma_minus8 plus 8 specifies the bit depth of the chroma sample values in the input integer tensor. The value of inpTensorBitDepth is derived as follows: inpTensorBitDepthChroma=nnpfc_inp_tensor_bitdepth_chroma_minus8+8 It is a bitstream conformance requirement that the value of nnpfc_inp_tensor_bitdepth_chroma_minus8 be in the range 0 to 24, inclusive.
[0115] In one embodiment, the variable inpTensorBitDepth may be renamed to inpTensorBitDepthLuma.
[0116] Note that in some embodiments, the example syntax provided in Table 15 may be combined with the example syntax provided in Tables 13A-13D and Tables 14A-14C, i.e., the syntax element nnpfc_inp_tensor_bitdepth_chroma_minus8 may be included in any of Tables 13A-13D and Tables 14A-14C.
[0117] In one embodiment, in accordance with the techniques herein, the semantics of nnpfc_inp_tensor_bitdepth_minus8 may be defined such that nnpfc_inp_tensor_bitdepth_minus8 plus 8 specifies the bit depth of the luma and chroma sample values in the input integer tensor. That is, in one embodiment, the semantics of nnpfc_inp_tensor_bitdepth_minus8 may be based on the following: nnpfc_inp_tensor_bitdepth_minus8 plus 8 specifies the bit depth of the luma and chroma sample values in the input integer tensor. The value of inpTensorBitDepth is derived as follows: inpTensorBitDepth=nnpfc_inp_tensor_bitdepth_minus8+8 It is a bitstream conformance requirement that the value of nnpfc_inp_tensor_bitdepth_minus8 be in the range 0 to 24, inclusive.
[0118] According to JVET-AB2006, the NNPFC SEI message supports the following update mechanisms: When the NNPFC SEI message is not the first NNPFC SEI message, in decoding order, with a particular nnpfc_id value in the current CLVS, the following applies. -This SEI message defines an update to the preceding elementary post-processing filter in decoding order that has the same nnpfc_id value. -This SEI message relates to the current decoded picture and all subsequent decoded pictures of the current layer in output order until the end of the current CLVS or until the next NNPFC SEI message with that particular nnpfc_id value in output order within the current CLVS. Also, When this SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc_id value in the current CLVS, the post-processing filter PostProcessingFilter() is assigned to be the same as the base post-processing filter. When this SEI message is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value in the current CLVS, the post-processing filter PostProcessingFilter() is obtained by applying the updates defined by this SEI message to the basic post-processing filter.
[0119] Note that according to JVET-AB2006, updates are not cumulative, rather each update applies to the base post-processing filter, which is the post-processing filter specified by the first NNPFC SEI message, in decoding order, with a particular nnpfc_id value in the current CLVS.
[0120] It is argued that in a scenario where the first NNPFC SEI message in decoding order with a particular nnpfc_id value in the current CLVS is lost, a receiver receiving a subsequent NNPFC SEI message with the same nnpfc_id value may need to parse several syntax elements before understanding whether the received message is an update or not. Also, in this case, the received message that is an update typically needs to be discarded after doing all of this parsing, since the first NNPFC SEI message that defines the basic post-processing filter is lost. In accordance with the techniques herein, this processing and parsing can be simplified by signaling a flag directly in the NNPFC SEI message that indicates whether the message is an update or the first NNPFC SEI message for the signaled nppfc_id.
[0121] In one embodiment, according to the techniques herein, a flag may be signaled to indicate whether an NNPFC SEI message is the first (or new) NNPFC SEI message in a CLVS, thereby specifying a base post-processing filter for a particular nnpfc_id, or whether it is an update of a previously signaled SEI message with the same nnpfc_id. It is argued that this can make it easier for a receiver to process an NNPFC SEI message by simply parsing this flag, instead of keeping track of the received nnpfc_id value and parsing other fields in the SEI message to identify whether it is an update or not. For example, some syntax elements in an NNPFC SEI message may need to be parsed to identify that the message contains an MPEG Neural Network Representation (NNR) update. It is argued that this can increase the processing required to understand an NNPFC message in an environment where SEI messages may be lost, which can be avoided by signaling the proposed flag.
[0122] Table 16 illustrates the associated syntax of an example neural network post-filter characteristics SEI message with a flag indicating whether the NNPFC SEI message is a new message or an update, in accordance with the techniques herein.
[0123] [Table 31]
[0124] With respect to Table 16, the semantics may be based on those provided above and the following: An nnpfc_update_flag equal to 1 specifies that the contents of this NNPFC SEI message provide an update to the previous (first) NNPFC message information with the same nnpfc_id. An nnpfc_update_flag equal to 0 specifies that the contents of this NNPFC SEI message contain standalone NNPFC message information that is not an update to the previous NNPFC message information with the same nnpfc_id. Alternatively, the semantics may be defined as follows: An nnpfc_update_flag equal to 1 specifies that the contents of this NNPFC SEI message provide an update to the previous (first) NNPFC SEI message with the same nnpfc_id. An nnpfc_update_flag equal to 0 specifies that this SEI message is the first NNPFC SEI message, in decoding order, with a particular nnpfc_id value in the current CLVS.
[0125] In another embodiment, the nnpfc_update_flag may be signaled in a different location in the NNPFC SEI message, for example, it may be signaled immediately after the nnpfc_id syntax element, or immediately after the nnpfc_mode_idc syntax element, or immediately before the nnfc_formatting_and_purpose_flag syntax element, or some other place in the NNPFC SEI message.
[0126] It should be noted that in some embodiments, the example syntax provided in Table 16 may be combined with the example syntax provided in Tables 13A-13D, 14A-14C, 15, and combinations thereof. That is, the syntax element nnpfc_update_flag may be included in any of Tables 13A-13D, 14A-14C, 15, and combinations thereof.
[0127] In accordance with the techniques herein, in some embodiments, video encoding apparatus 500 thus represents an example of a configured device that signals a neural network post-filter characteristic message, signals in the neural network post-filter characteristic message a first syntax element indicating the purpose of the post-processing filter, and signals in the neural network post-filter characteristic message a second syntax element that specifies whether syntax elements related to the purpose, input formatting, output formatting, and complexity of the post-processing filter are present in the neural network post-filter characteristic message, and the first syntax element precedes the second syntax element in the neural network post-filter characteristic message.
[0128] 1, interface 108 may include any device configured to receive data generated by data encapsulator 107 and transmit and / or store the data on a communications medium. Interface 108 may include a network interface card, such as an Ethernet card, an optical transceiver, a radio frequency transceiver, or any other type of device capable of transmitting and / or receiving information. Additionally, interface 108 may include a computer system interface that may allow files to be stored on a storage device. For example, interface 108 may include a Peripheral Component Interconnect (PCI) and Peripheral Component Interconnect Express (PCIe) bus protocols, a proprietary bus protocol, a Universal Serial Bus (USB) protocol, an I / O interface, or any other type of device ... 2 C, or any other logical and physical structure that may be used to interconnect peer devices.
[0129] 1, destination device 120 includes interface 122, data decapsulator 123, video decoder 124, and display 126. Interface 122 may include any device configured to receive data from a communication medium. Interface 122 may include a network interface card, such as an Ethernet card, an optical transceiver, a radio frequency transceiver, or any other type of device capable of receiving and / or transmitting information. Additionally, interface 122 may include an interface for a computer system that allows a compliant video bitstream to be obtained from a storage device. For example, interface 122 may include interfaces for PCI and PCIe bus protocols, proprietary bus protocols, USB protocols, I / O protocols, and the like. 2C, or any other logical and physical structures that may be used to interconnect peer devices. The data decapsulator 123 may be configured to receive and parse any of the example syntax structures described herein.
[0130] Video decoder 124 may include any device configured to receive a bitstream (e.g., a sub-bitstream extract) and / or an acceptable variant thereof and regenerate video data therefrom. Display 126 may include any device configured to display video data. Display 126 may include one of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display. Display 126 may include a high-definition display or an ultra-high-definition display. Although in the example shown in FIG. 1, video decoder 124 is described as outputting data to display 126, it should be noted that video decoder 124 may be configured to output video data to various types of devices and / or subcomponents thereof. For example, video decoder 124 may be configured to output video data to any communication medium as described herein.
[0131] FIG. 6 is a block diagram illustrating an example of a video decoding device that may be configured to decode video data according to one or more techniques of this disclosure (e.g., the reference picture list creation decoding process described above). In one embodiment, the video decoding device 600 may be configured to decode transform data and recover residual data from transform coefficients based on the decoded transform data. The video decoding device 600 may be configured to perform intra-prediction decoding and inter-prediction decoding and may therefore be referred to as a hybrid decoding device. The video decoding device 600 may be configured to parse any combination of the syntax elements described above in Tables 1-16. The video decoding device 600 may decode video based on or in accordance with the above process and further based on the parsed values in Tables 1-16.
[0132] In the example shown in FIG. 6, the video decoding device 600 includes an entropy decoding unit 602, an inverse quantization unit 604, an inverse transform coefficient processing unit 606, an intra prediction processing unit 608, an inter prediction processing unit 610, an adder 612, a post filter unit 614, and a reference buffer 616. The video decoding device 600 may be configured to decode video data in a manner consistent with a video encoding system. It should be noted that while the exemplary video decoding device 600 is shown having separate functional blocks, such illustration is for illustrative purposes and does not limit the video decoding device 600 and / or its subcomponents to a particular hardware or software architecture. The functionality of the video decoding device 600 may be realized using any combination of hardware, firmware, and / or software implementations.
[0133] As shown in FIG. 6, the entropy decoding unit 602 receives an entropy coded bitstream. The entropy decoding unit 602 may be configured to decode syntax elements and quantized coefficients from the bitstream according to a reverse process of the entropy coding process. The entropy decoding unit 602 may be configured to perform entropy decoding according to any of the entropy coding techniques described above. The entropy decoding unit 602 may determine values of syntax elements in the coded bitstream in accordance with a video coding standard. As shown in FIG. 6, the entropy decoding unit 602 may determine quantization parameters, values of quantized coefficients, transform data, and prediction data from the bitstream. In the example shown in FIG. 6, the inverse quantization unit 604 and the inverse transform coefficient processing unit 606 receive values of quantized coefficients from the entropy decoding unit 602 and output reconstructed residual data.
[0134] Referring back to FIG. 6, the reconstructed residual data may be provided to an adder 612. The adder 612 may add the reconstructed residual data to a predictive video block to generate reconstructed video data. The predictive video block may be determined according to a predictive video technique (i.e., intra prediction and inter-frame prediction). The intra prediction processor 608 may be configured to receive the intra prediction syntax element and retrieve the predictive video block from a reference buffer 616. The reference buffer 616 may include a memory device configured to store one or more frames of video data. The intra prediction syntax element may identify an intra prediction mode, such as the intra prediction modes described above. The inter prediction processor 610 may receive the inter prediction syntax element and generate a motion vector to identify a predictive block in one or more reference frames stored in the reference buffer 616. The inter prediction processor 610 may perform interpolation, possibly based on an interpolation filter, to generate a motion compensated block. The syntax element may include an identifier of an interpolation filter to be used for motion prediction having sub-pixel accuracy. The inter-prediction processor 610 may use an interpolation filter to calculate interpolated values for sub-integer pixels of the reference block. The post-filter unit 614 may be configured to perform filtering on the reconstructed video data. For example, the post-filter unit 614 may be configured to perform deblocking and / or sample adaptive offset (SAO) filtering, for example, based on parameters specified in the bitstream. It should be noted that in some examples, the post-filter unit 614 may also be configured to perform its own arbitrary filtering (e.g., visual enhancement such as mosquito noise reduction). As shown in FIG. 6, the reconstructed video blocks may be output by the video decoding device 600.In this manner, video decoding apparatus 600 represents an example of a device configured to receive a neural network post-filter characteristic message, parse from the neural network post-filter characteristic message a first syntax element indicating a purpose of the post-processing filter, and parse from the neural network post-filter characteristic message a second syntax element specifying whether syntax elements related to the purpose, input formatting, output formatting, and complexity of the post-processing filter are present in the neural network post-filter characteristic message, wherein the first syntax element precedes the second syntax element in the neural network post-filter characteristic message.
[0135] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions may be stored or transmitted as one or more instructions or code on a computer-readable medium and executed by a hardware-based processor. A computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium such as a data storage medium or a communication medium, including any medium that facilitates the transfer of a computer program from one place to another, for example according to a communication protocol. In this manner, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0136] By way of example, and without limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, other magnetic storage devices, flash memory, or any other medium, i.e., any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. However, it should be understood that computer-readable media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory tangible storage media. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks and Blu-ray discs, where disks typically reproduce data magnetically and discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer readable media.
[0137] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, as used herein, the term "processor" may refer to any of the foregoing structures, or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a composite codec. Also, the techniques may be implemented entirely in one or more circuit or logic elements.
[0138] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units are illustrated in this disclosure to highlight functional aspects of devices configured to perform the disclosed techniques, but need not necessarily be realized by different hardware units. Rather, as previously described, the various units may be combined into a codec hardware unit or may be provided by a collection of interoperating hardware units, including one or more processors as previously described, along with suitable software and / or firmware.
[0139] Furthermore, each functional block or various features of the base station device and the terminal device used in each of the above implementations may be implemented or performed by a circuit, which is typically an integrated circuit or multiple integrated circuits. A circuit designed to perform the functions described herein may comprise a general-purpose processor, a digital signal processor (DSP), an application specific or general-purpose application integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, or individual hardware components, or a combination thereof. A general-purpose processor may be a microprocessor, or the processor may be a conventional processor, controller, microcontroller, or state machine. The general-purpose processor or each circuit described above may be composed of digital circuits or analog circuits. Furthermore, as semiconductor technology advances and integrated circuit technology emerges to replace current integrated circuits, integrated circuits using this technology may also be used.
[0140] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
1. A video decoding device for receiving a bitstream, comprising: receiving the bitstream including a neural network post-filter characteristics SEI message; The neural network post-filter characteristics SEI message includes: a first syntax element indicating the purpose of the post-processing filter; a second syntax element indicating whether syntax elements related to filter purpose, input formatting, output formatting, and complexity are present; The video decoding device, wherein in the neural network post-filter characteristics SEI message, the first syntax element precedes the second syntax element.
2. A video encoding device for transmitting a bitstream, comprising: Neural network post-filter characteristics included in the SEI message: a first syntax element indicating the purpose of the post-processing filter; a second syntax element indicating whether syntax elements related to filter purpose, input formatting, output formatting, and complexity are present; an entropy encoding unit that transmits the bitstream including the neural network post-filter characteristics SEI message; The video decoding device, wherein in the neural network post-filter characteristics SEI message, the first syntax element precedes the second syntax element.
3. A transmission method for transmitting a bitstream, comprising: transmitting the bitstream including a neural network post-filter characteristics SEI message; The neural network post-filter characteristic SEI message is a first syntax element indicating the purpose of the post-processing filter; a second syntax element indicating whether syntax elements related to filter purpose, input formatting, output formatting, and complexity are present; 2. The method of claim 1, wherein in the neural network post-filter characteristics SEI message, the first syntax element precedes the second syntax element.