System and method for signaling neural network postfilter frame rate upsampling information in video encoding

JP2024057562A5Pending Publication Date: 2025-12-02SHARP KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023003991
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-12
Filing Date
2023-01-13
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing video encoding standards lack efficient methods for signaling neural network postfilter parameter information, particularly for frame rate upsampling, which is crucial for improving video quality but not adequately addressed in current standards like ITU-T H.264, ITU-T H.265, JEM, and JVET-T2001.

Method used

The proposed techniques involve signaling neural network postfilter characteristic messages that specify the number of input pictures and the manner of concatenation for picture interpolation processes, using syntax elements to enhance frame rate upsampling in video encoding.

Benefits of technology

This approach allows for improved video quality by effectively signaling neural network postfilter parameters, enabling better frame rate upsampling and enhancing the overall video encoding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a device, a system, and a method for signaling neural network postfilter parameter information for an encoded video.SOLUTION: A video decoder performs frame rate upsampling on the basis of information included in a neural network postfilter characteristic message including a syntax element that specifies the number of input pictures used as input for the neural network postfilter picture interpolation process, and a syntax element with a value that specifies the manner in which input pictures are concatenated before being input into the neural network postfilter picture interpolation process.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] FIELD This disclosure relates to video encoding, and more particularly, to techniques for signaling neural network postfilter parameter information for encoded video. [Background technology]

[0002] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, laptop or desktop computers, tablet computers, digital recording devices, digital media players, video gaming devices, cellular phones, including so-called smart phones, medical imaging devices, and the like. Digital video can be encoded according to a video encoding standard. A video encoding standard defines a format for a compliant bitstream that encapsulates the encoded video data. A compliant bitstream is a data structure that can be received and decoded by a video decoding device to generate recovered video data. A video encoding standard can incorporate video compression techniques. Examples of video encoding standards include ISO / IEC MPEG-4 Visual and ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) and High-Efficiency Video Coding (HEVC). HEVC is described in High Efficiency Video Coding (HEVC), Rec. ITU-T H.265 (December 2016), which is incorporated herein by reference and is referred to herein as ITU-T H.265. Extensions and improvements to ITU-T H.265 are being considered for the development of next generation video coding standards. For example, the ITU-T Video Coding Experts Group (VCEG) and ISO / IEC (Moving Picture Experts Group (MPEG), collectively known as the Joint Video Exploration Team (JVET)) have standardized video coding techniques with compression capabilities that significantly exceed those of the current HEVC standard.The Joint Exploration Model 7 (JEM 7), Algorithm Description of Joint Exploration Test Model 7 (JEM 7), ISO / IEC JTC1 / SC29 / WG11 Document: JVET-G1001, July 2017, Torino, IT, describes the coding features that have been under collaborative test model study by JVET as having the potential to improve video coding technology beyond the capabilities of ITU-T H.265, and is incorporated herein by reference. It should be noted that the coding features of JEM 7 are implemented in the JEM reference software. As used herein, the term JEM may collectively refer to the algorithms included in JEM 7 and the implementation of the JEM reference software. Additionally, in response to the "Joint Call for Proposals on Video Compression with Capabilities beyond HEVC" jointly issued by VCEG and MPEG, multiple descriptions of video coding tools have been published by various groups in the past 10 years. th The initial draft text of the video coding specification, derived from multiple descriptions of video coding tools, was proposed at the Meeting of ISO / IEC JTC1 / SC29 / WG11 16-20 April 2018, San Diego, CA. thThis development of a video coding standard by VCEG and MPEG is called the Versatile Video Coding (VVC) project. "Versatile Video Coding (Draft 10)", 20th Meeting of ISO / IEC JTC1 / SC29 / WG11 7-16 October 2020, Teleconference, document JVET-T2001-v2 represents the current iteration of the draft text of the video coding specification corresponding to the VVC project, which is incorporated herein by reference and referred to as JVET-T2001.

[0003] Video compression techniques allow for reducing data requirements for storing and transmitting video data. Video compression techniques can reduce data requirements by exploiting inherent redundancy in a video sequence. Video compression techniques may subdivide a video sequence into successively smaller portions (i.e., groups of pictures in a video sequence, pictures in groups of pictures, regions in pictures, sub-regions in regions, etc.). Intra-prediction coding techniques (e.g., spatial prediction techniques within a picture) and inter-prediction techniques (i.e., inter-picture techniques (temporal)) can be used to generate difference values ​​between a unit of video data being coded and a reference unit of video data. The difference values ​​are sometimes referred to as residual data. The residual data can be coded as quantized transform coefficients. Syntax elements can associate the residual data with the reference coding units (e.g., intra-prediction mode index, and motion information). The residual data and syntax elements can be entropy coded. The entropy coded residual data and syntax elements can be included in a data structure that forms a compliant bitstream. Summary of the Invention

[0004] In general, this disclosure describes various techniques for encoding video data. In particular, this disclosure describes techniques for signaling neural network postfilter parameter information for encoded video data. It should be noted that although the techniques of this disclosure are described with respect to ITU-T H.264, ITU-T H.265, JEM, and JVET-T2001, the techniques of this disclosure are generally applicable to video encoding. For example, the encoding techniques described herein can be incorporated into video encoding systems (including video encoding systems based on future video encoding standards) that include video block structures, intra-prediction techniques, inter-prediction techniques, transform techniques, filtering techniques, and / or entropy encoding techniques other than those included in ITU-T H.265, JEM, and JVET-T2001. Thus, references to ITU-T H.264, ITU-T H.265, JEM, and / or JVET-T2001 are for illustrative purposes and should not be construed to limit the scope of the techniques described herein. Furthermore, it should be noted that the incorporation by reference of documents herein is for explanatory purposes and should not be construed as limiting or creating ambiguity with respect to the terms used herein. For example, if an incorporated reference provides a definition of a term that differs from that of another incorporated reference and / or as that term is used herein, that term should be construed to broadly include each corresponding definition and / or to include each specific definition instead.

[0005] In one example, a method of encoding video data includes signaling a neural network postfilter characteristic message, signaling a first syntax element in the neural network postfilter characteristic message that specifies a number of input pictures to be used as input for the neural network postfilter picture interpolation process, and signaling a second syntax element in the neural network postfilter characteristic message having a value that specifies a manner in which the input pictures are concatenated before being input to the neural network postfilter picture interpolation process.

[0006] In one example, a device comprises one or more processors configured to signal a neural network postfilter characteristic message, signal a first syntax element in the neural network postfilter characteristic message that specifies a number of input pictures to be used as input for the neural network postfilter picture interpolation process, and signal a second syntax element in the neural network postfilter characteristic message having a value that specifies a manner in which the input pictures are concatenated before being input to the neural network postfilter picture interpolation process.

[0007] In one example, a non-transitory computer-readable storage medium includes instructions stored on the non-transitory computer-readable storage medium that, when executed, cause one or more processors of the device to signal a neural network postfilter characteristic message, signal a first syntax element in the neural network postfilter characteristic message that specifies a number of input pictures to be used as input for the neural network postfilter picture interpolation process, and signal a second syntax element in the neural network postfilter characteristic message having a value that specifies a manner in which the input pictures are concatenated before being input to the neural network postfilter picture interpolation process.

[0008] In one example, an apparatus comprises means for signaling in a neural network postfilter property message a first syntax element that specifies a number of input pictures to be used as input for a neural network postfilter picture interpolation process, and means for signaling in the neural network postfilter property message a second syntax element having a value that specifies a manner in which the input pictures are concatenated before being input to the neural network postfilter picture interpolation process.

[0009] In one example, a method for decoding video data includes receiving a neural network postfilter characteristic message, parsing a first syntax element from the neural network postfilter characteristic message that specifies a number of input pictures to be used as input for a neural network postfilter picture interpolation process, parsing a second syntax element from the neural network postfilter characteristic message having a value that specifies a manner in which the input pictures are concatenated before being input to the neural network postfilter picture interpolation process, and concatenating the input pictures based on the value of the second syntax element.

[0010] In one example, a device comprises one or more processors configured to receive a neural network postfilter characteristic message, parse a first syntax element from the neural network postfilter characteristic message that specifies a number of input pictures to be used as input for the neural network postfilter picture interpolation process, parse a second syntax element from the neural network postfilter characteristic message having a value that specifies a manner in which the input pictures are concatenated before being input to the neural network postfilter picture interpolation process, and concatenate the input pictures based on the value of the second syntax element.

[0011] In one example, a non-transitory computer-readable storage medium includes instructions on the non-transitory computer-readable storage medium that, when executed, cause one or more processors of the device to receive a neural network postfilter characteristic message, parse a first syntax element from the neural network postfilter characteristic message that specifies a number of input pictures to be used as input for the neural network postfilter picture interpolation process, parse a second syntax element from the neural network postfilter characteristic message having a value that specifies a manner in which the input pictures are concatenated before being input to the neural network postfilter picture interpolation process, and concatenate the input pictures based on the value of the second syntax element.

[0012] In one example, an apparatus comprises means for receiving a neural network postfilter characteristic message, means for parsing a first syntax element from the neural network postfilter characteristic message that specifies a number of input pictures to be used as input for the neural network postfilter picture interpolation process, means for parsing a second syntax element from the neural network postfilter characteristic message having a value that specifies a manner in which the input pictures are concatenated before being input to the neural network postfilter picture interpolation process, and means for concatenating the input pictures based on the value of the second syntax element.

[0013] The details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims. [Brief description of the drawings]

[0014] [Figure 1] FIG. 1 is a block diagram illustrating an example of a system that can be configured to encode and decode video data in accordance with one or more techniques of this disclosure. [Diagram 2]1 is a conceptual diagram illustrating encoded video data and corresponding data structures in accordance with one or more techniques of this disclosure. [Diagram 3] 1 is a conceptual diagram illustrating a data structure that encapsulates encoded video data and corresponding metadata in accordance with one or more techniques of this disclosure. [Figure 4] A conceptual diagram illustrating an example of components that may be included in an implementation of a system that may be configured to encode and decode video data in accordance with one or more techniques of this disclosure. [Diagram 5] 1 is a block diagram illustrating an example of a video encoding device that can be configured to encode video data in accordance with one or more techniques of this disclosure. [Figure 6] 1 is a block diagram illustrating an example of a video decoding device that can be configured to decode video data in accordance with one or more techniques of this disclosure. [Figure 7] FIG. 2 is a conceptual diagram illustrating an example of a packed data channel for a luma component, in accordance with one or more techniques of this disclosure.

[0015] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Video content includes a video sequence consisting of a series of frames (or pictures). A series of frames may also be referred to as a group of pictures (GOP). Each video frame or picture may be divided into one or more regions. A region may be defined according to a base unit (e.g., a video block) and a set of rules that define the region. For example, the rule that defines a region may be that the region must be an integer number of video blocks arranged in a rectangle. Furthermore, the video blocks in a region may be ordered according to a scan pattern (e.g., raster scan). As used herein, the term video block may generally refer to a portion of a picture, or more specifically, may refer to a maximum array of sample values ​​that can be predictively encoded, its subdivision, and / or corresponding structure. Furthermore, the term current video block may refer to a portion of a picture being encoded or decoded. A video block may be defined as an array of sample values. It should be noted that in some cases, a pixel value may be described as including sample values ​​for each component of video data, which may also be referred to as color components (e.g., luma component (Y) and chroma components (Cb and Cr), or red, green, and blue color components). It should be noted that in some cases, the terms pixel value and sample value are used interchangeably. Furthermore, in some cases, a pixel or a sample may be referred to as a pel. A video sampling format, sometimes referred to as a chroma format, may define the number of chroma samples included in a video block relative to the number of luma samples included in the video block. For example, for a 4:2:0 format, the sampling rate for the luma component is twice the sampling rate of the chroma components in both the horizontal and vertical directions.

[0016] A video coding device may perform predictive coding on video blocks and their subdivisions. The video blocks and their subdivisions may be referred to as nodes. ITU-T H.264 specifies macroblocks containing 16x16 luma samples. That is, in ITU-T H.264, a picture is divided into macroblocks. ITU-T H.265 specifies a similar coding tree unit (CTU) structure, sometimes called a largest coding unit (LCU). In ITU-T H.265, a picture is divided into CTUs. In ITU-T H.265, for a picture, the CTU size may be set to contain 16x16, 32x32, or 64x64 luma samples. In ITU-T H.265, a CTU is composed of a respective coding tree block (CTB) for each component of video data (e.g., luma (Y) and chroma (Cb and Cr). It should be noted that a video having one luma component and two corresponding chroma components can be described as having two channels, i.e., a luma channel and a chroma channel. Furthermore, in ITU-T H.265, a CTU can be divided according to a quad-tree (QT) partitioning structure, so that the CTB of a CTU is divided into coding blocks (CBs). That is, in ITU-T H.265, a CTU can be divided into quad-tree leaf nodes. According to ITU-T H.265, one luma CB, together with two corresponding chroma CBs and related syntax elements, is called a coding unit (CU). In ITU-T H.265, the minimum allowed size of a CB can be signaled. In H.265, the smallest allowed size of luma CB is 8x8 luma samples. In ITU-T H.265, the decision to code a picture portion using intra or inter prediction is made at the CU level.

[0017] In ITU-T H.265, a CU is associated with a prediction unit structure with root in the CU. In ITU-T H.265, the prediction unit structure allows for splitting the luma CB and chroma CB for the purpose of generating corresponding reference samples. That is, in ITU-T H.265, the luma CB and chroma CB can be split into respective luma and chroma prediction blocks (PBs), where a PB includes a block of sample values ​​to which the same prediction is applied. In ITU-T H.265, the CB can be split into 1, 2, or 4 PBs. ITU-T H.265 supports PB sizes from 64x64 samples to 4x4 samples. In ITU-T H.265, square PBs are supported for intra prediction, where the CB can form a PB or the CB can be split into 4 square PBs. In addition to square PB, ITU-T H.265 supports rectangular PB for inter prediction, where CB can be bisected vertically or horizontally to form PB. In addition, ITU-T H.265 supports four asymmetric PB division for inter prediction, where CB is divided into two PBs by one-quarter of CB height (at top or bottom) or width (at left or right). Using intra prediction data (e.g., intra prediction mode syntax element) or inter prediction data (e.g., motion data syntax element) corresponding to PB, reference sample value and / or predicted sample value for PB are generated.

[0018] JEM specifies a CTU with a maximum size of 256x256 luma samples. JEM specifies a quad tree plus binary tree (QTBT) block structure. In JEM, the QTBT structure allows a quad tree leaf node to be further split by a binary tree structure (BT). That is, in JEM, the binary tree structure allows a quad tree leaf node to be split recursively vertically or horizontally. In JVET-T2001, the CTU is split according to a quad tree plus multitype tree (QTMT or QT+MTT) structure. QTMT in JVET-T2001 is similar to QTBT in JEM. However, in JVET-T2001, in addition to exhibiting a binary split, the multitype tree may exhibit a so-called ternary (or triple tree (TT)) split. The ternary split splits a block into three blocks vertically or horizontally. In the case of vertical TT division, the block is divided at a quarter of its width from the left edge and at a quarter of its width from the right edge, and in the case of horizontal TT division, the block is divided at a quarter of its height from the top edge and at a quarter of its height from the bottom edge.

[0019] As mentioned above, each video frame or picture may be divided into one or more regions. For example, according to ITU-T H.265, each video frame or picture may be divided to include one or more slices, and further divided to include one or more tiles, where each slice includes a sequence of CTUs (e.g., in raster scan order), and a tile is a sequence of CTUs corresponding to a rectangular portion of a picture. It should be noted that a slice, in ITU-T H.265, is a sequence of one or more slice segments starting with an independent slice segment and including all subsequent dependent slice segments (if any) preceding the next independent slice segment (if any). A slice segment, like a slice, is a sequence of CTUs. Thus, in some cases, the terms slice and slice segment may be used interchangeably to indicate a sequence of CTUs arranged in raster scan order. It should be further noted that in ITU-T H.265, a tile may consist of CTUs included in two or more slices, and a slice may consist of CTUs included in two or more tiles. However, ITU-T H.265 specifies that one or both of the following conditions must be met: (1) all CTUs in a slice belong to the same tile, and (2) all CTUs in a tile belong to the same slice.

[0020] For JVET-T2001, a slice is instead only required to consist of an integer number of CTUs, but is instead required to consist of an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile. Note that in JVET-T2001, the slice design does not include slice segments (i.e., there are no independent / dependent slice segments). Thus, in JVET-T2001, a picture may contain a single tile, where the single tile is contained within a single slice, or a picture may contain multiple tiles, where multiple tiles (or their CTU rows) may be contained within one or more slices. In JVET-T2001, the division of a picture into tiles is specified by specifying the height of each of the tile rows and the width of each of the tile columns. Thus, in JVET-T2001, a tile is a rectangular region of CTUs within a particular tile row and a particular tile column location. Further, it should be noted that JVET-T2001 specifies the case where a picture may be divided into sub-pictures, where a sub-picture is a rectangular region of a CTU within a picture. The top-left CTU of a sub-picture may be located at any CTU position within a picture, with the sub-picture being constrained to contain one or more slices. Thus, unlike tiles, sub-pictures are not necessarily limited to a specific row and column position. It should be noted that a sub-picture may be useful for encapsulating an area of ​​interest within a picture, and the sub-bitstream extraction process may be used only to decode and display the specific area of ​​interest. That is, as described in more detail below, a bitstream of coded video data includes a sequence of network abstraction layer (NAL) units, where the NAL units encapsulate the coded video data (i.e., video data corresponding to a slice of a picture) or the NAL units encapsulate metadata (e.g., parameter sets) used to decode the video data, and the sub-bitstream extraction process forms a new bitstream by removing one or more NAL units from the bitstream.

[0021] FIG. 2 is a conceptual diagram illustrating an example of a picture in a picture group divided according to tiles, slices, and subpictures. It should be noted that the techniques described herein may be applicable to tiles, slices, subpictures, their subdivisions, and / or their equivalent structures. That is, the techniques described herein may be generally applicable regardless of how a picture is divided into regions. For example, in some cases, the techniques described herein may be applicable when tiles may be divided into so-called bricks, where a brick is a rectangular region of a CTU row in a particular tile. Further, for example, in some cases, the techniques described herein may be applicable when one or more tiles may be included in a so-called tile group, where a tile group includes an integer number of adjacent tiles. In the example shown in FIG. 2, Pic3 is shown as including 16 tiles (i.e., Tile0 to Tile15) and three slices (i.e., Slice0 to Slice2). In the example shown in FIG. 2, Slice 0 includes four tiles (i.e., Tile 0 to Tile 3), Slice 1 includes eight tiles (i.e., Tile 4 to Tile 11), and Slice 2 includes four tiles (i.e., Tile 12 to Tile 15). Further, as shown in the example of FIG. 2, Pic 3 is shown as including two subpictures (i.e., Subpicture 0 and Subpicture 1), where Subpicture 0 includes Slice 0 and Slice 1, and Subpicture 1 includes Slice 2. As mentioned above, subpictures can be useful for encapsulating regions of interest within a picture, and the sub-bitstream extraction process can be used to selectively decode (and display) the regions of interest. For example, referring to FIG. 2, Subpicture 0 may correspond to the action portion of a sporting event presentation (e.g., a view of the field), and Subpicture 1 may correspond to a scrolling banner displayed during the sporting event presentation. By organizing pictures into subpictures in this manner, a viewer can disable the display of the scrolling banner.That is, through the sub-bitstream extraction process, the Slice2 NAL unit may be removed from the bitstream (and thus may not be decoded and / or displayed), and the Slice0 and Slice1 NAL units may be decoded and displayed. The encapsulation of slices of a picture into respective NAL unit data structures and the sub-bitstream extraction are described in further detail below.

[0022] For intra-prediction coding, the intra-prediction mode may specify the location of the reference sample in the picture. In ITU-T H.265, the defined possible intra-prediction modes include planar (i.e., surface fitting) prediction mode, DC (i.e., monotonic global averaging) prediction mode, and 33 angular prediction modes (predMode: 2-34). In JEM, the defined possible intra-prediction modes include planar, DC, and 65 angular prediction modes. It should be noted that the planar and DC prediction modes may be referred to as non-directional prediction modes, and the angular prediction modes may be referred to as directional prediction modes. It should be noted that the techniques described herein may be generally applicable regardless of the number of defined possible prediction modes.

[0023] For inter-predictive coding, a reference picture is determined, and a motion vector (MV) identifies samples in the reference picture used to generate a prediction for the current video block. For example, the current video block may be predicted using reference sample values ​​located in one or more previously encoded picture(s), and the motion vector is used to indicate the location of the reference block relative to the current video block. The motion vector may, for example, describe a horizontal displacement component (i.e., MVx) of the motion vector, a vertical displacement component (i.e., MVy) of the motion vector, and a resolution (e.g., ¼ pixel precision, ½ pixel precision, 1 pixel precision, 2 pixel precision, 4 pixel precision) of the motion vector. Previously decoded pictures, which may include pictures output before or after the current picture, may be organized into one or more reference picture lists and identified using reference picture index values. Furthermore, in inter-predictive coding, uni-prediction refers to generating a prediction using sample values ​​from a single reference picture, and bi-prediction refers to generating a prediction using respective sample values ​​from two reference pictures. That is, in uni-prediction, a single reference picture and corresponding motion vector are used to generate a prediction for the current video block, and in bi-prediction, a first reference picture and corresponding first motion vector and a second reference picture and corresponding second motion vector are used to generate a prediction for the current video block. In bi-prediction, the respective sample values ​​are combined (e.g., added according to weights, rounded, clipped, or averaged) to generate a prediction. A picture and its regions may be classified based on what kind of prediction mode may be used to encode its video block. That is, for a region having a B type (e.g., B slice), bi-prediction mode, uni-prediction mode, and intra-prediction mode may be used, for a region having a P type (e.g., P slice), uni-prediction mode and intra-prediction mode may be used, and for a region having an I type (e.g., I slice), only intra-prediction mode may be used. As described above, the reference picture is identified by a reference index.For example, in a P slice, there may be a single reference picture list RefPicList0, and in a B slice, in addition to RefPicList0, there may be a second independent reference picture list RefPicList1. Note that for uni-prediction in a B slice, one of RefPicList0 or RefPicList1 may be used to generate the prediction. Furthermore, note that during the decoding process, at the start of decoding a picture, the reference picture list(s) are generated from previously decoded pictures stored in the decoded picture buffer (DPB).

[0024] Furthermore, the coding standard may support various modes of motion vector prediction. Motion vector prediction allows a value of a motion vector for a current video block to be derived based on another motion vector. For example, a set of candidate blocks with associated motion information may be derived from spatial and temporal neighboring blocks to the current video block. Furthermore, generated (or default) motion information may be used for motion vector prediction. Examples of motion vector prediction include advanced motion vector prediction (AMVP), temporal motion vector prediction (TMVP), the so-called "combined" mode, as well as "skip" and "direct" motion estimation. Furthermore, other examples of motion vector prediction include advanced temporal motion vector prediction (ATMVP) and spatial-temporal motion vector prediction (STMVP). In motion vector prediction, both the video encoding device and the video decoding device perform the same process to derive a set of candidates. Thus, the same set of candidates is generated for the current video block during encoding and decoding.

[0025] As mentioned above, in inter-prediction coding, reference samples in previously coded pictures are used to code video blocks in a current picture. A previously coded picture that is available for use as a reference when coding a current picture is called a reference picture. It should be noted that the decoding order does not necessarily correspond to the picture output order, i.e., the temporal order of pictures in a video sequence. In ITU-T H.265, when a picture is decoded, the picture is stored in a decoded picture buffer (DPB) (which may be called a frame buffer, a reference buffer, a reference picture buffer, etc.). In ITU-T H.265, pictures stored in the DPB are removed from the DPB when they are output and are no longer needed for coding a subsequent picture. In ITU-T H.265, the decision of whether a picture should be removed from the DPB is performed once per picture after decoding the slice header, i.e., at the beginning of the decoding of the picture. For example, referring to FIG. 2, Pic2 is shown as referring to Pic1. Similarly, Pic3 is shown as referencing Pic0. With reference to FIG. 2, assuming that the picture numbers correspond to the decoding order, the DPB is populated as follows: after decoding Pic0, the DPB contains {Pic0}, at the start of decoding Pic1, the DPB contains {Pic0}, after decoding Pic1, the DPB contains {Pic0, Pic1}, and at the start of decoding Pic2, the DPB contains {Pic0, Pic1}. Pic2 is then decoded with reference to Pic1, and after decoding Pic2, the DPB contains {Pic0, Pic1, Pic2}. At the start of decoding Pic3, pictures Pic0 and Pic1 are marked for removal from the DPB as they are not required to decode Pic3 (or any subsequent pictures not shown), and assuming that Pic1 and Pic2 have been output, the DPB is updated to contain {Pic0}. Pic3 is then decoded by referencing Pic0. The process of marking pictures for removal from the DPB is sometimes called Reference Picture Set (RPS) management.

[0026] As mentioned above, the intra prediction data or the inter prediction data is used to generate reference sample values ​​for a block of sample values. The difference between the sample values ​​included in the current PB or another type of picture substructure and the associated reference sample (e.g., the reference sample generated using the prediction) may be referred to as residual data. The residual data may include respective arrays of difference values ​​corresponding to respective components of the video data. The residual data may be in the pixel domain. A transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), an integer transform, a wavelet transform, or a conceptually similar transform, may be applied to the array of difference values ​​to generate transform coefficients. It is noted that in ITU-T H.265 and JVET-T2001, a CU is associated with a transform tree structure with the root at the CU level. The transform tree is divided into one or more transform units (TUs). That is, for purposes of generating transform coefficients, the array of difference values ​​may be partitioned (e.g., four 8×8 transforms may be applied to a 16×16 array of residual values). For each component of video data, such subdivision of the difference values ​​may be referred to as a Transform Block (TB). Note that in some cases, a core transform and subsequent secondary transforms may be applied at the video encoder to generate the transform coefficients. For the video decoder, the order of the transforms is reversed.

[0027] The quantization process may be performed directly on the transform coefficients or on the residual sample values ​​(e.g., in the case of palette coding quantization). Quantization approximates the transform coefficients with amplitudes limited to a particular set of values. Quantization essentially scales the transform coefficients to change the amount of data required to represent a group of transform coefficients. Quantization may include division of the transform coefficients (or values ​​resulting from adding an offset value to the transform coefficients) by a quantization scale factor and any associated rounding function (e.g., rounding to the nearest integer). The quantized transform coefficients are sometimes referred to as coefficient level values. Inverse quantization (or "dequantization") may include multiplication of the coefficient level values ​​by a quantization scale factor and any mutual rounding or offset addition operations. It should be noted that as used herein, the term quantization process may refer in some cases to division by a quantization scale factor to generate level values ​​and in some cases to multiplication by a quantization scale factor to recover the transform coefficients. That is, the quantization process may refer in some cases to quantization and in some cases to inverse quantization. Furthermore, while in some of the following examples the quantization process is described with respect to arithmetic operations associated with decimal number representation, it should be noted that such description is for illustrative purposes and should not be construed as limiting. For example, the techniques described herein may be implemented in devices using binary arithmetic, etc. For example, the multiplication and division operations described herein may be implemented using bit shifting operations, etc.

[0028] The quantized transform coefficients and syntax elements (e.g., syntax elements indicating a coding structure of a video block) may be entropy coded according to an entropy coding technique. The entropy coding process includes coding the values ​​of the syntax elements using a lossless data compression algorithm. Examples of entropy coding techniques include content adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), probability interval partitioning entropy coding (PIPE), etc. The entropy coded quantized transform coefficients and the corresponding entropy coded syntax elements may form an adapted bitstream that may be used to reproduce the video data at a video decoding device. The entropy coding process, e.g., CABAC, may include performing binarization on the syntax elements. Binarization refers to the process of converting the value of the syntax elements into a series of one or more bits. These bits may be referred to as "bins." The binarization may include one or a combination of the following encoding techniques: fixed-length encoding, unary encoding, shortened unary encoding, shortened Rice encoding, Golomb encoding, k-th exponential Golomb encoding, and Golomb-Rice encoding. For example, the binarization may include representing an integer value of 5 of a syntax element as 00000101 using an 8-bit fixed-length binarization technique, or representing an integer value of 5 as 11110 using a unary encoding binarization technique. As used herein, each of the terms fixed-length encoding, unary encoding, shortened unary encoding, shortened Rice encoding, Golomb encoding, k-th exponential Golomb encoding, and Golomb-Rice encoding may refer to a general implementation of these techniques and / or a more specific implementation of these encoding techniques. For example, an implementation of Golomb-Rice encoding may be specifically defined according to a video encoding standard.In a CABAC example, for a particular bin, the context provides a most probable state (MPS) value for the bin (i.e., the MPS for the bin is one of 0 or 1) and a probability value that the bin is the MPS or least probably state (LPS). For example, the context may indicate that the MPS of the bin is 0 and the probability that the bin is 1 is 0.3. Note that the context may be determined based on values ​​of previously coded bins, including bins in the current syntax element and previously coded syntax elements. For example, values ​​of syntax elements associated with neighboring video blocks may be used to determine the context of the current bin.

[0029] As mentioned above, the sample values ​​of the reconstructed block may differ from the sample values ​​of the encoded current video block. Furthermore, it should be noted that in some cases, encoding video data block by block may result in artifacts (e.g., so-called blocking artifacts, banding artifacts, etc.). For example, blocking artifacts may cause encoded block boundaries of the reconstructed video data to be visually perceptible to a user. In this manner, the reconstructed sample values ​​may be modified to minimize the difference between the encoded sample values ​​of the current video block and the reconstructed block and / or to minimize artifacts introduced by the video encoding process. Such modification may be generally referred to as filtering. It should be noted that filtering may occur as part of an in-loop filtering process or a post-loop (or post-filtering) filtering process. In an in-loop filtering process, the sample values ​​resulting from the filtering process may be used for the predicted video block (e.g., stored in a reference frame buffer for subsequent encoding in a video encoding device and subsequent decoding in a video decoding device). In a post-loop filtering process, the sample values ​​resulting from the filtering process are simply output as part of the decoding process (e.g., not used for subsequent encoding). For example, in a video decoding device, in an in-loop filtering process, sample values ​​resulting from filtering the reconstructed block are used for subsequent decoding (e.g., stored in a reference buffer) and output (e.g., to a display), whereas in a post-loop filtering process, the reconstructed block is used for subsequent decoding, and sample values ​​resulting from filtering the reconstructed block are output and not used for subsequent decoding.

[0030] Deblocking (or de-blocking), deblock filtering, or applying a deblocking filter refers to a process of smoothing (i.e., making the boundary less perceptible by a viewer) the boundary of adjacent reconstructed video blocks. Smoothing the boundary of adjacent reconstructed video blocks may include modifying sample values ​​contained in rows or columns adjacent to the boundary. JVET-T2001 specifies when a deblocking filter is applied to reconstructed sample values ​​as part of an in-loop filtering process. In addition to applying a deblocking filter as part of an in-loop filtering process, JVET-T2001 specifies when sample adaptive offset (SAO) filtering may be applied in the in-loop filtering process. In general, SAO is a process of modifying deblocked sample values ​​in a region by conditionally adding an offset value. Another type of filtering process includes the so-called adaptive loop filter (ALF). ALF with block-based adaptation is specified in JEM. In JEM, ALF is applied after the SAO filter. It should be noted that the ALF may be applied to the reconstructed samples independently of other filtering techniques. The process for applying the ALF specified in the JEM in a video coding device may be summarized as follows: (1) each 2×2 block of the luma component of the reconstructed picture is classified according to a classification index, (2) a set of filter coefficients is derived for each classification index, (3) a filtering decision is determined for the luma component, (4) a filtering decision is determined for the chroma components, and (5) the filter parameters (e.g., coefficients and decisions) are signaled. JVET-T2001 specifies deblocking, SAO, and ALF filters that may be described as generally based on the deblocking, SAO, and ALF filters specified in ITU-T H.265 and JEM.

[0031] It is noted that JVET-T2001 is referred to as a pre-release version of ITU-T H.266 and is therefore a near-finished draft of the video coding standard resulting from the VVC project, and is therefore sometimes referred to as the first version of the VVC standard (alternatively, VVC or VVC version 1 or ITU-H.266). It is noted that during the VVC project, convolutional neural network (CNN)-based techniques that showed potential for artifact removal and objective quality improvement were investigated, but it was decided not to include such techniques in the VVC standard. However, CNN-based techniques are currently being considered for extending and / or improving VVC. Some CNN-based techniques are related to post-filtering. For example, "AHG11: Content-adaptive neural network post-filter", 26th Meeting of ISO / IEC JTC1 / SC29 / WG11 20-29 April 2022, Teleconference, document JVET-Z0082-v2 (referred to herein as JVET-Z0082) describes a content-adaptive neural network based post-filter. Note that in JVET-Z0082, content adaptation is achieved by overfitting a NN post-filter to a test video. Further note that the result of the overfitting process in JVET-Z0082 is a weight update.JVET-Z0082 describes when the weight updates are coded in ISO / IEC FDIS 15938-17. Information technology-Multimedia content description interface-Part 17: Compression of neural networks for multimedia content description and analysis and Test Model of Incremental Compression of Neural Networks for Multimedia Content Description and Analysis (INCTM), N0179. February 2022, which are sometimes collectively referred to as the MPEG NNR (Neural Network Representation) or Neural Network Coding (NNC) standards. JVET-Z0082 further describes when the coded weight updates are signaled within the video bitstream as NNR postfilter SEI messages. "AHG9: NNR post-filter SEI message", 26th Meeting of ISO / IEC JTC1 / SC29 / WG11 20-29 April 2022, Teleconference, document JVET-Z0052-v1 (referred to herein as JVET-Z0052) describes the NNR post-filter SEI message utilized by JVET-Z0082. The elements of the NN post-filter described in JVET-Z0082 and the NNR post-filter SEI message described in JVET-Z0052 were adopted in "Additional SEI messages for VSEI (Draft 2)", 27th Meeting of ISO / IEC JTC1 / SC29 / WG11 13-22 July 2022, Teleconference, document JVET-AA2006-v2 (referred to herein as JVET-AA2006). JVET-AA2006 specifies a versatile supplemental enhancement information message for coded video bitstreams (VSEI).JVET-AA2006 specifies syntax and semantics for the neural network postfilter properties SEI message and for the neural network postfilter activation SEI message. The neural network postfilter properties SEI message specifies neural networks that may be used as post-processing filters. The use of a specified post-processing filter for a particular picture is indicated in the neural network postfilter activation SEI message. JVET-AA2006 is described in further detail below. The techniques described herein provide techniques for signaling neural network postfilter messages.

[0032] For the formulas used herein, the following arithmetic operators may be used:

[0033] [Table 1]

[0034] In addition, the following mathematical functions can be used: Log2(x), the base 2 logarithm of x;

number

[0035] With respect to the example syntax used herein, the following definitions of logical operators may apply: x && y The Boolean logic "product" of x and y x||y The Boolean "union" of x and y ! Boolean logic "no" x?y:zIf x is true or not equal to 0, then evaluate the value of y, else evaluate the value of z.

[0036] In addition, the following relational operators may be applied:

[0037] [Table 2]

[0038] Furthermore, in the syntax descriptors used herein, it should be noted that the following descriptors may apply: -b(8): A byte with an arbitrary pattern of bits (8 bits). The parsing process of this descriptor is specified by the return value of the function read_bits(8). -f(n): fixed pattern bit sequence using n bits written (left to right), left bit first. The parsing process of this descriptor is specified by the return value of the function read_bits(n). -se(v): A signed integer zeroth-order Exp-Golomb encoded syntax element, left bit first. -tb(v): A shortened binary using up to maxVal bits, where maxVal is defined in the semantics of the syntax element. -tu(v): A shorthand unary that uses up to maxVal bits, where maxVal is defined in the semantics of the syntax element. -u(n): An unsigned integer using n bits. If n is "v" in the syntax table, the number of bits varies depending on the values ​​of other syntax elements. The parsing process of this descriptor is specified by the return value of the function read_bits(n), which is interpreted as a binary representation of an unsigned integer written most significant bit first. -ue(v): An unsigned integer zeroth-order Exp-Golomb encoded syntax element, left bit first.

[0039] As mentioned above, video content includes a video sequence consisting of a series of pictures, and each picture may be divided into one or more regions. In JVET-T2001, the coded representation of a picture includes the VCL NAL units of a particular layer in an AU, including all CTUs of the picture. For example, referring back to FIG. 2, the coded representation of Pic3 is encapsulated in three coded slice NAL units (i.e., Slice0 NAL unit, Slice1 NAL unit, and Slice2 NAL unit). It should be noted that the term video coding layer (VCL) NAL unit is used as a generic term for coded slice NAL units, i.e., VCL NAL is a generic term that includes all types of slice NAL units. As mentioned above and described in more detail below, NAL units may encapsulate metadata used to decode video data. NAL units that encapsulate metadata used to decode a video sequence are generally referred to as non-VCL NAL units. Thus, in JVET-T2001, a NAL unit can be a VCL NAL unit or a non-VCL NAL unit. Note that a VCL NAL unit contains slice header data that provides information used to decode a particular slice. Thus, in JVET-T2001, information used to decode video data, sometimes referred to as metadata in some cases, is not limited to being contained in a non-VCL NAL unit. JVET-T2001 specifies the case where a picture unit (PU) is a set of NAL units associated with each other according to a specified classification rule, consecutive in decoding order, and containing exactly one coded picture, and an access unit (AU) is a set of PUs containing coded pictures that belong to different layers and are associated at the same time for output from the DPB. JVET-T2001 further specifies the case where a layer is a set of VCL NAL units and associated non-VCL NAL units that all have a particular value of a layer identifier.Furthermore, in JVET-T2001, a PU consists of zero or one PH NAL unit, one coded picture containing one or more VCL NAL units, and zero or more other non-VCL NAL units. Furthermore, in JVET-T2001, a coded video sequence (CVS) is a sequence of AUs consisting of, in decoding order, a CVSS AU followed by zero or more AUs that are not CVSS AUs, including all subsequent AUs up to but not including any subsequent AU that is a CVSS AU, where a coded video sequence start (CVSS) AU is an AU with a PU for each layer in the CVS, and a coded picture in each existing picture unit is a coded layer video sequence start (CLVSS) picture. In JVET-T2001, a coded layer video sequence (CLVS) is a sequence of PUs in the same layer consisting, in decoding order, of a CLVSS PU followed by zero or more PUs that are not CLVSS PUs, including all subsequent PUs up to but not including any subsequent PU that is a CLVSS PU. That is, in JVET-T2001, a bitstream may be described as containing a sequence of AUs that form one or more CVSs.

[0040] Multi-layer video coding allows a video presentation to be decoded / displayed as a presentation corresponding to a base layer of video data and as one or more additional presentations corresponding to enhancement layers of the video data. For example, the base layer may allow a video presentation to be presented having a basic level of quality (e.g., high resolution rendering and / or 30 Hz frame rate), and the enhancement layer may allow a video presentation to be presented having an increased level of quality (e.g., ultra-high resolution rendering and / or 60 Hz frame rate). The enhancement layer may be coded by referencing the base layer. That is, for example, a picture in the enhancement layer may be coded (e.g., using inter-layer prediction techniques) by referencing one or more pictures in the base layer (including scaled versions thereof). It should be noted that layers may also be coded independently of each other. In this case, there may be no inter-layer prediction between two layers. Each NAL unit may include an identifier indicating the layer of video data with which the NAL unit is associated. As mentioned above, a sub-bitstream extraction process may be used to decode and display only a particular region of interest of a picture. Additionally, a sub-bitstream extraction process may be used to decode and display only a particular layer of a video. Sub-bitstream extraction may refer to a process in which a device receiving a compliant or conforming bitstream forms a new compliant or conforming bitstream by discarding and / or modifying data in the received bitstream. For example, sub-bitstream extraction may be used to form a new compliant or conforming bitstream that corresponds to a particular representation (e.g., a higher quality representation) of the video.

[0041] In JVET-T2001, each of a video sequence, a GOP, a picture, a slice, and a CTU may be associated with metadata that describes video coding properties, and some types of metadata may be encapsulated in non-VCL NAL units. JVET-T2001 defines parameter sets that may be used to describe video data and / or video coding properties. In particular, JVET-T2001 includes four types of parameter sets: video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), and adaptation parameter set (APS), where an SPS applies to zero or more entire CVSs, a PPS applies to zero or more entire coded pictures, an APS applies to zero or more slices, and a VPS may be optionally referenced by an SPS. A PPS applies to the individual coded pictures that reference it. In JVET-T2001, parameter sets may be encapsulated as non-VCL NAL units and / or signaled as messages. JVET-T2001 also includes a picture header (PH) encapsulated as a non-VCL NAL unit. In JVET-T2001, the picture header applies to all slices of a coded picture. JVET-T2001 further allows decoding capability information (DCI) and supplemental enhancement information (SEI) messages to be signaled. In JVET-T2001, the DCI and SEI messages assist processes related to decoding, display, or other purposes, but the DCI and SEI messages may not be required to create luma or chroma samples following the decoding process. In JVET-T2001, the DCI and SEI messages may be signaled in the bitstream using non-VCL NAL units. Furthermore, the DCI and SEI messages may be conveyed by some mechanism other than by being present in the bitstream (i.e., signaled out-of-band).

[0042] FIG. 3 shows an example of a bitstream including multiple CVSs, where a CVS includes AUs, and an AU includes picture units. The example shown in FIG. 3 corresponds to an example of encapsulating slice NAL units in a bitstream shown in the example of FIG. 2. In the example shown in FIG. 3, the corresponding picture unit for Pic3 includes three VCL NAL coded slice NAL units, namely, a Slice0 NAL unit, a Slice1 NAL unit, and a Slice2 NAL unit, and two non-VCL NAL units, namely, a PPS NAL unit and a PH NAL unit. It should be noted that in FIG. 3, the headers are NAL unit headers (i.e., should not be confused with slice headers). It should also be noted that in FIG. 3, other non-VCL NAL units not shown, such as an SPS NAL unit, a VPS NAL unit, an SEI message NAL unit, etc., may be included in the CVS. Further, note that in other examples, the PPS NAL unit used to decode Pic3 may be included elsewhere in the bitstream, e.g., in the picture unit corresponding to Pic0, or may be provided by an external mechanism. As described in more detail below, in JVET-T2001, the PH syntax structure may be present in the slice header of a VCL NAL unit or in the PH NAL unit of the current PU.

[0043] JVET-T2001 defines NAL unit header semantics that specify the type of raw byte sequence payload (RBSP) data structure contained in a NAL unit. Table 1 shows the syntax of the NAL unit header defined in JVET-T2001.

[0044] [Table 3]

[0045] JVET-T2001 specifies the following definitions for each syntax element shown in Table 1: forbidden_zero_bit shall be equal to 0. nuh_reserved_zero_bit shall be equal to 0. A value of 1 for nuh_reserved_zero_bit may be specified in ITU-T|ISO / IEC. Although the value of nuh_reserved_zero_bit is required to be equal to 0 in this version of this specification, decoders conforming to this version of this specification shall allow a value of nuh_reserved_zero_bit equal to 1 to appear in the syntax and shall ignore (i.e., remove from the bitstream and discard) NAL units with nuh_reserved_zero_bit equal to 1. nuh_layer_id specifies the identifier of the layer to which a VCL NAL unit belongs, or the identifier of the layer to which a non-VCL NAL unit applies. Values ​​of nuh_layer_id shall be in the range of 0 to 55, inclusive. Other values ​​of nuh_layer_id are reserved for future use by ITU-T | ISO / IEC. Although values ​​of nuh_layer_id are required to be in the range of 0 to 55, inclusive, in this version of this specification, decoders conforming to this version of this specification shall allow values ​​of nuh_layer_id greater than 55 to appear in the syntax and shall ignore (i.e., remove from the bitstream and discard) NAL units with nuh_layer_id greater than 55. The value of nuh_layer_id shall be the same for all VCL NAL units of a coded picture. The value of nuh_layer_id of a coded picture or PU is the value of nuh_layer_id of the VCL NAL units of the coded picture or PU. When nal_unit_type is equal to PH_NUT or FD_NUT, nuh_layer_id shall be equal to the nuh_layer_id of the associated VCL NAL unit. When nal_unit_type is equal to EOS_NUT, nuh_layer_id shall be equal to one of the nuh_layer_id values ​​of a layer present in CVS. NOTE – The values ​​of nuh_layer_id in the DCI, OPI, VPS, AUD, and EOB NAL units are not constrained. nuh_temporal_id_plus1 minus 1 specifies the temporal identifier for the NAL unit. The value of nuh_temporal_id_plus1 must not be equal to 0. The variable TemporalId is derived as follows: TemporalId=nuh_temporal_id_plus1-1 When nal_unit_type is in the range IDR_W_RADL to RSV_IRAP_11, inclusive, TemporalId shall be equal to 0. When nal_unit_type is equal to STSA_NUT and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, TemporalId shall be greater than 0. The value of TemporalId shall be the same for all VCL NAL units of an AU. The value of TemporalId of a coded picture, PU, ​​or AU is the value of TemporalId of the VCL NAL units of the coded picture, PU, ​​or AU. The value of TemporalId of a sublayer representation is the largest value of TemporalId of all VCL NAL units in the sublayer representation. The values ​​of TemporalId for non-VCL NAL units are constrained as follows: - if nal_unit_type is equal to DCI_NUT, OPI_NUT, VPS_NUT, or SPS_NUT, TemporalId shall be equal to 0 and the TemporalId of the AU containing the NAL unit shall be equal to 0. Otherwise, if nal_unit_type is equal to PH_NUT, then TemporalId shall be equal to the TemporalId of the PU that contains the NAL unit. Otherwise, if nal_unit_type is equal to EOS_NUT or EOB_NUT, TemporalId shall be equal to 0. Otherwise, if nal_unit_type is equal to AUD_NUT, FD_NUT, PREFIX_SEI_NUT, or SUFFIX_SEI_NUT, TemporalId shall be equal to the TemporalId of the AU that contains the NAL unit. Otherwise, when nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, TemporalId shall be greater than or equal to the TemporalId of the PU that contains the NAL unit. NOTE - When the NAL unit is a non-VCL NAL unit, the value of TemporalId is equal to the minimum of the TemporalId values ​​of all AUs to which the non-VCL NAL unit applies. When nal_unit_type is equal to PPS_NUT, PREFIX_APS_NUT, or SUFFIX_APS_NUT, TemporalId may be greater than or equal to the TemporalId of the containing AU, since all PPSs and APSs may be included at the beginning of the bitstream (e.g., when they are transported out-of-band and the receiver places them at the beginning of the bitstream), and the first coded picture has TemporalId equal to 0. nal_unit_type specifies the NAL unit type, i.e., the type of RBSP data structure contained in the NAL unit as specified in Table 2. NAL units with nal_unit_type in the range UNSPEC28 to UNSPEC31, inclusive, whose semantics are unspecified, SHALL NOT affect the decoding process specified in this specification. NOTE - NAL unit types in the range of UNSPEC_28 to UNSPEC_31 may be used as determined by the application. The decoding process for these values ​​of nal_unit_type is not specified in this specification. Since different applications may use these NAL unit types for different purposes, it is expected that particular care will be taken in the design of encoders that generate NAL units with these nal_unit_type values, and in the design of decoders that interpret the content of NAL units with these nal_unit_type values. This specification does not define the management of these values. These nal_unit_type values ​​may only be suitable for use in contexts where "collisions" of usage (i.e., different definitions of the meaning of the NAL unit content for the same nal_unit_type value) are not significant or possible, or are managed, e.g., defined or managed in a controlling application or transport specification, or by controlling the environment in which the bitstream is delivered. For purposes other than determining the amount of data in a DU of the bitstream, a decoder SHALL ignore (remove from the bitstream and discard) the content of all NAL units that use reserved values ​​of nal_unit_type. NOTE - This requirement allows for the future definition of compatible extensions to this specification.

[0046] [Table 4-1]

[0047] [Table 4-2] NOTE - A Clean Random Access (CRA) picture may have an associated RASL or RADL picture present in the bitstream. NOTE - An Instantaneous Decoding Refresh (IDR) picture with nal_unit_type equal to IDR_N_LP does not have an associated leading picture present in the bitstream. An IDR picture with nal_unit_type equal to IDR_W_RADL does not have an associated RASL picture present in the bitstream, but may have an associated RADL picture in the bitstream. The value of nal_unit_type shall be the same for all VCL NAL units of a subpicture. A subpicture is referred to as having the same NAL unit type as the VCL NAL units of the subpicture. For the VCL NAL units of any particular picture, the following applies: -If pps_mixed_nalu_types_in_pic_flag is equal to 0, the value of nal_unit_type shall be the same for all VCL NAL units of a picture, and the picture or PU is referenced as having the same NAL unit type as the coded slice NAL units of the picture or PU. - Otherwise (pps_mixed_nalu_types_in_pic_flag is equal to 1), all of the following constraints apply. -A picture shall have at least two sub-pictures. - A VCL NAL unit of a picture shall have two or more distinct nal_unit_type values. - There shall be no VCL NAL unit of the picture with nal_unit_type equal to GDR_NUT. When a VCL NAL unit of a picture has nal_unit_type equal to nalUnitTypeA equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT, all other VCL NAL units of the picture shall have nal_unit_type equal to nalUnitTypeA or TRAIL_NUT. The value of nal_unit_type shall be the same for all pictures in an IRAP or GDR AU. When sps_video_parameter_set_id is greater than 0, vps_max_tid_il_ref_pics_plus1[i][j] is equal to 0 for j equal to GeneralLayerIdx[nuh_layer_id] and for any value of i in the range of j+1 to vps_max_layers_minus1, inclusive, pps_mixed_nalu_types_in_pic_flag is equal to 1, and the value of nal_unit_type is not equal to IDR_W_RADL, IDR_N_LP, or CRA_NUT. The requirements for bitstream conformance are that the following constraints apply: When a picture is a leading picture of an IRAP picture, it shall be a RADL or RASL picture. When a subpicture is the leading subpicture of an IRAP subpicture, it shall be a RADL or RASL subpicture. - If a picture is not the leading picture of an IRAP picture, it shall not be a RADL or RASL picture. - If a subpicture is not the leading subpicture of an IRAP subpicture, it shall not be a RADL or RASL subpicture. - There shall be no RASL pictures associated with an IDR picture in the bitstream. - No RASL sub-pictures associated with an IDR sub-picture shall be present in the bitstream. - There shall be no RADL pictures associated with an IDR picture with nal_unit_type equal to IDR_N_LP in the bitstream. NOTE - Performing random access at the position of an IRAP AU by discarding all PUs before the IRAP AU (and correctly decoding non-RASL pictures in the IRAP AU and all following AUs in decoding order) is possible, provided that each parameter set is available (in the bitstream or by external means not specified in this specification) at the time it is referenced. - There shall be no RADL sub-pictures associated with an IDR sub-picture with nal_unit_type equal to IDR_N_LP in the bitstream. -Any picture with nuh_layer_id equal to the particular value layerId that precedes in decoding order an IRAP picture with nuh_layer_id equal to layerId shall precede the IRAP picture in output order and shall precede any RADL picture associated with the IRAP picture in output order. - Any subpicture with nuh_layer_id equal to the specified value layerId and subpicIdx subpicture index equal to the specified value that precedes in decoding order an IRAP subpicture with nuh_layer_id equal to layerId and subpicture index equal to subpicIdx shall precede in output order the IRAP subpicture and all its associated RADL subpictures. Any picture with nuh_layer_id equal to the specified value layerId that precedes in decoding order the recovery point picture with nuh_layer_id equal to -layerId shall precede the recovery point picture in output order. Any subpicture with nuh_layer_id equal to a specific value layerId and subpicIdx equal to a specific value that precedes in decoding order a subpicture with nuh_layer_id equal to layerId and subpicIdx equal to a specific value in the recovery point picture shall precede that subpicture in the recovery point picture in output order. - Any RASL picture associated with a CRA picture shall precede any RADL picture associated with the CRA picture in output order. Any RASL sub-picture associated with a CRA sub-picture shall precede any RADL sub-picture associated with the CRA sub-picture in output order. - Any RASL picture with nuh_layer_id equal to a particular value layerId associated with a CRA picture shall follow in output order any IRAP or GDR picture with nuh_layer_id equal to layerId that precedes the CRA picture in decoding order. - Any RASL subpicture associated with a CRA subpicture and having nuh_layer_id equal to a particular value layerId and a subpicture index equal to a particular value subpicIdx shall follow in output order any IRAP or GDR subpicture with nuh_layer_id equal to layerId and a subpicture index equal to subpicIdx that precedes the CRA subpicture in decoding order. - if sps_field_seq_flag is equal to 0, the following applies: if the current picture with nuh_layer_id equal to a particular value layerId is a leading picture associated with an IRAP picture, it shall precede, in decoding order, all non-leading pictures associated with the same IRAP picture. Otherwise (sps_field_seq_flag is equal to 1), let picA and picB be the first and last leading pictures associated with an IRAP picture, respectively, in decoding order, there shall be at most one non-leading picture with nuh_layer_id equal to layerId preceding picA in decoding order, and there shall be no non-leading picture with nuh_layer_id equal to layerId between picA and picB in decoding order. - if sps_field_seq_flag is equal to 0, the following applies: if the current subpicture with nuh_layer_id equal to a particular value layerId and subpicture index equal to a particular value subpicIdx is a leading picture associated with an IRAP subpicture, then it shall precede, in decoding order, all non-leading subpictures associated with the same IRAP subpicture. Otherwise (sps_field_seq_flag is equal to 1), let subpicA and subpicB be the first and last leading subpictures associated with an IRAP subpicture, respectively, in decoding order, then there shall be at most one non-leading subpicture with nuh_layer_id equal to layerId and subpicture index equal to subpicIdx preceding subpicA in decoding order, and there shall be no non-leading pictures with nuh_layer_id equal to layerId and subpicture index equal to subpicIdx between picA and picB in decoding order.

[0048] A NAL unit may include a supplemental enhancement information (SEI) syntax structure, as provided in Table 2. Tables 3 and 4 show the supplemental enhancement information (SEI) syntax structure defined in JVET-T2001.

[0049] [Table 5]

[0050] [Table 6]

[0051] With respect to Tables 3 and 4, JVET-T2001 specifies the following semantics: Each SEI message consists of variables that specify the type, payloadType, and size, payloadSize, of the SEI message payload. The SEI message payload is specified. The derived SEI message payload size, payloadSize, is specified in bytes and shall be equal to the number of RBSP bytes in the SEI message payload. NOTE - A NAL unit byte sequence containing an SEI message may contain one or more emulation prevention bytes (represented by the emulation_prevention_three_byte syntax element). Since the payload size of an SEI message is specified in RBSP bytes, the amount of emulation prevention bytes is not included in the size of the SEI payload, payloadSize. payload_type_byte is the payload type byte of the SEI message. payload_size_byte is the payload size in bytes of the SEI message.

[0052] It should be noted that JVET-T2001 defines payload types, and "Additional SEI messages for VSEI (Draft 6)" 25th Meeting of ISO / IEC JTC1 / SC29 / WG11 12-21 January 2022, Teleconference, document JVET-Y2006-v1, incorporated herein by reference and referred to as JVET-Y2006, defines additional payload types. Table 5 shows the sei_payload() syntax structure in a schematic manner. That is, Table 5 shows the sei_payload() syntax structure, but for brevity, not all possible types of payloads are included in Table 5.

[0053] [Table 7]

[0054] With respect to Table 5, JVET-T2001 specifies the following semantics: sei_reserved_payload_extension_data SHALL NOT be present in bitstreams that conform to this version of this specification. However, decoders that conform to this version of this specification SHALL ignore the presence and value of sei_reserved_payload_extension_data. When present, the length in bits of sei_reserved_payload_extension_data shall be 8 * Equal to payloadSize-nEarlierBits-nPayloadZeroBits-1, where nEarlierBits is the number of bits in the sei_payload() syntax structure preceding the sei_reserved_payload_extension_data syntax element and nPayloadZeroBits is the number of sei_payload_bit_equal_to_zero syntax elements at the end of the sei_payload() syntax structure. If more_data_in_payload() is true after parsing an SEI message syntax structure (for example, the buffering_period() syntax structure), and nPayloadZeroBits is not equal to 7, then PayloadBits is set to 8. * payloadSize-n is set equal to PayloadZeroBits-1, otherwise PayloadBits is set equal to 8 * Set equal to payloadSize. payload_bit_equal_to_one shall be equal to 1. payload_bit_equal_to_zero shall be equal to 0. NOTE - SEI messages with the same value of payloadType are conceptually the same SEI message, regardless of whether they are included in a prefix or suffix SEI NAL unit. NOTE - For SEI messages specified in this specification and the VSEI specification (ITU-T H.274|ISO / IEC 23002-7), payloadType values ​​are aligned with similar SEI messages specified in AVC (Rec. ITU-T H.264|ISO / IEC 14496-10) and HEVC (Rec. ITU-T H.265|ISO / IEC 23008-2). The semantics and duration for each SEI message are specified in the semantics specification for each particular SEI message. NOTE – The persistence information in SEI messages is summarized in Table 142 for informational purposes.

[0055] JVET-T2001 further specifies the following: SEI messages having the syntax structure identified in [Table 5] specified in Rec. ITU-T H.274 | ISO / IEC 23002-7 may be used with bitstreams specified by this specification. When any particular Rec. ITU-T H.274|ISO / IEC 23002-7 SEI message is included in a bitstream specified by this specification, the SEI payload syntax shall be contained in the sei_payload() syntax structure specified in [Table 5] and shall use the payloadType value specified in [Table 5], plus any SEI message-specific constraints specified in this Annex for that particular SEI message shall apply. The value of PayloadBits is passed to a parser for the SEI message syntax structure specified in Rec. ITU-T H.274 | ISO / IEC 23002-7, as specified above.

[0056] As mentioned above, JVET-AA2006 defines NN postfilter supplemental enhancement information messages. In particular, JVET-AA2006 defines a neural network postfilter characteristics SEI message (payloadType==210) and a neural network postfilter activation SEI message (payloadType==211). Tables 6 and 7 show the syntax of the neural network postfilter characteristics SEI message defined in JVET-AA2006. It should be noted that the neural network postfilter characteristics SEI message is sometimes referred to as the NNPFC SEI.

[0057] [Table 8-1]

[0058] [Table 8-2]

[0059] [Table 8-3]

[0060] [Table 9]

[0061] With respect to Tables 6 and 7, JVET-AA2006 specifies the following semantics: This SEI message specifies a neural network that can be used as a post-processing filter. The use of a specified post-processing filter for a particular picture is indicated in a Neural Network Post Filter Activation SEI message. Use of this SEI message requires the definition of the following variables: The width and height of the cropped decoded output picture in units of luma samples, denoted herein by CroppedWidth and CroppedHeight, respectively. - If present, the luma sample array CroppedYPic and the chroma sample arrays CroppedCbPic and CroppedCrPic of the cropped decoded output picture for the vertical coordinate y and horizontal coordinate x, where the top-left corner of the sample arrays has coordinate y equal to 0 and coordinate x equal to 0. -BitDepth - the bit depth of the luma sample array of the cropped decoded output picture Y . -BitDepth - the bit depth of the chroma sample array, if any, of the cropped decoded output picture C . A chroma format indicator, denoted herein by ChromaFormatIdc. - Quantized strength value StrengthControlVal when nnpfc_auxiliary_inp_idc is equal to 1. When this SEI message specifies a neural network that can be used as a post-processing filter, the semantics specify the derivation of the luma sample array FilteredYPic[x][y] and chroma sample arrays FilteredCbPic[x][y] and FilteredCrPic[x][y], as indicated by the value of nnpfc_out_order_idc, that contain the output of the post-processing filter. The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc as specified by Table 8.

[0062] [Table 10] nnpfc_id contains an identification number that can be used to identify a post-processing filter. The value of nnpfc_id is in the range 0 to 2, inclusive. 32 It should be in the range of -2. 256 to 511, inclusive, and 2 31 ~2 32 Values ​​of nnpfc_id up to -2 are reserved for future use by ITU-T | ISO / IEC. 31 ~2 32 A decoder encountering a value of nnpfc_id within the range of -2 shall ignore it. An nnpfc_mode_idc equal to 0 specifies that the post-processing filter associated with the nnpfc_id value is determined by external means not specified in this specification. nnpfc_mode_idc equal to 1 specifies that the post-processing filter associated with the nnpfc_id value is a neural network represented by the ISO / IEC 15938-17 bit stream included in this SEI message. An nnpfc_mode_idc equal to 2 specifies that the post-processing filter associated with the nnpfc_id value is a neural network identified by the specified tag uniform resource identifier (URI) (nnpfc_uri_tag[i]) and neural network information URI (nnpfc_uri[i]). The value of nnpfc_mode_idc shall be in the range 0 to 255, inclusive. Values ​​of nnpfc_mode_idc greater than 2 are reserved for future specification by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification shall ignore SEI messages containing reserved values ​​of nnpfc_mode_idc. nnpfc_purpose_and_formatting_flag equal to 0 specifies that syntax elements related to filter purpose, input formatting, output formatting, and complexity are not present. nnpfc_purpose_and_formatting_flag equal to 1 specifies that syntax elements related to filter purpose, input formatting, output formatting, and complexity are present. When nnpfc_mode_idc is equal to 1 and the current CLVS does not contain a preceding neural network postfilter characteristics SEI message in decoding order that has a value of nnpfc_id equal to the value of nnpfc_id in this SEI message, nnpfc_purpose_and_formatting_flag shall be equal to 1. When the current CLVS contains a preceding neural network postfilter characteristics SEI message in decoding order that has the same value of nnpfc_id equal to the value of nnpfc_id in this SEI message, at least one of the following conditions shall apply: -This SEI message has nnpfc_mode_idc equal to 1 and nnpfc_purpose_and_formatting_flag equal to 0 to provide neural network updates. This SEI message has the same content as the preceding Neural Network Post-Filter Characteristics SEI message. When this SEI message is the first neural network post-filter characteristics SEI message in decoding order with a particular nnpfc_id value in the current CLVS, it specifies a base post-processing filter that pertains to the current decoded picture and all subsequent decoded pictures of the current layer in output order until the end of the current CLVS. When this SEI message is not the first neural network post-filter characteristics SEI message in decoding order with a particular nnpfc_id value in the current CLVS, this SEI message pertains to the current decoded picture and all subsequent decoded pictures of the current layer in output order until the end of the current CLVS or the next neural network post-filter characteristics SEI message in output order with that particular nnpfc_id value in the current CLVS. nnpfc_purpose indicates the purpose of the post-processing filter as specified in Table 9. The value of nnpfc_purpose is between 0 and 2, inclusive. 32 nnpfc_purpose shall be in the range -2. Values ​​of nnpfc_purpose not appearing in Table 9 are reserved for future specification by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification shall ignore SEI messages containing reserved values ​​of nnpfc_purpose.

[0063] [Table 11] NOTE – When reserved values ​​of nnpfc_purpose are used in the future by ITU-T|ISO / IEC, the syntax of this SEI message may be extended with syntax elements whose presence is conditioned by nnpfc_purpose being equal to that value. When SubWidthC is equal to 1 and SubHeightC is equal to 1, nnpfc_purpose shall not be equal to 2 or 4. nnpfc_out_sub_c_flag equal to 1 specifies that outSubWidthC is equal to 1 and outSubHeightC is equal to 1. nnpfc_out_sub_c_flag equal to 0 specifies that outSubWidthC is equal to 2 and outSubHeightC is equal to 1. When nnpfc_out_sub_c_flag is not present, outSubWidthC is inferred to be equal to SubWidthC and outSubHeightC is inferred to be equal to SubHeightC. When SubWidthC is equal to 2 and SubHeightC is equal to 1, nnpfc_out_sub_c_flag shall not be equal to 0. nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples specify the width and height, respectively, of the picture luma sample array obtained by applying the post-processing filter identified by nnpfc_id to the cropped decoded output picture. When nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples are not present, they are inferred to be equal to CroppedWidth and CroppedHeight, respectively. nnpfc_component_last_flag equal to 0 specifies that the second dimension in the input tensor inputTensor to the post-processing filter and the output tensor outputTensor resulting from the post-processing filter are used for the channels. nnpfc_component_last_flag equal to 1 specifies that the last dimension in the input tensor iinputTensor to the post-processing filter and the output tensor outputTensor resulting from the post-processing filter are used for the channels. NOTE - The first dimension in the input and output tensors is used for the batch index, which is the practice in some neural network frameworks. Although the semantics of this SEI message use a batch size equal to 1, it is up to the post-processing implementation to determine the batch size used as input to the neural network inference. NOTE − A color component is an example of a channel. nnpfc_inp_format_flag indicates how to convert the sample values ​​of the cropped decoded output picture into input values ​​to the post-processing filter. When nnpfc_inp_format_flag is equal to 0, the input values ​​to the post-processing filter are real numbers, and the functions InpY() and InpC() are specified as follows: InpY(x)=x-((1< <BitDepthY)-1) InpC(x)=x-((1< <BitDepthc)-1) When nnpfc_inp_format_flag is equal to 1, the input values ​​to the post-processing filter are unsigned integers and the functions InpY() and InpC() are specified as follows:

[0064] [Table 12] The variable inpTensorBitDepth is derived from the syntax element nnpfc_inp_tensor_bitdepth_minus8, as specified below. nnpfc_inp_tensor_bitdepth_minus8 plus 8 specifies the bit depth of the luma sample values ​​in the input integer tensor. The value of inpTensorBitDepth is derived as follows: inpTensorBitDepth=nnpfc_inp_tensor_bitdepth_minus8+8 It is a bitstream conformance requirement that the value of nnpfc_inp_tensor_bitdepth_minus8 be in the range 0 to 24, inclusive. nnpfc_auxiliary_inp_idc not equal to 0 specifies that the auxiliary input data is present in the input tensor of the neural network postfilter. nnpfc_auxiliary_inp_idc equal to 0 indicates that the auxiliary input data is not present in the input tensor. nnpfc_auxiliary_inp_idc equal to 1 specifies that the auxiliary input data is derived as specified in Table 14. Values ​​of nnpfc_auxiliary_inp_idc shall be in the range 0 to 255, inclusive. Values ​​of nnpfc_auxiliary_inp_idc greater than 1 are reserved for future specification by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification shall ignore SEI messages that contain reserved values ​​of nnpfc_auxiliary_inp_idc. nnfpc_separate_colour_description_present_flag equal to 1 indicates that the separate combination of colour primaries, transfer characteristics and matrix coefficients for the picture resulting from the post-processing filter is specified in the SEI message syntax structure. nnpfc_separate_colour_description_present_flag equal to 0 indicates that the combination of colour primaries, transfer characteristics and matrix coefficients for the picture resulting from the post-processing filter is the same as that indicated in the VUI parameters for CLVS. nnpfc_colour_primaries has the same semantics as specified for the vui_colour_primaries syntax element, which is as follows: vui_colour_primaries indicates the chromaticity coordinates of the source primaries. Its semantics are as specified for the ColourPrimaries parameter in Rec. ITU-T H.273|ISO / IEC 23091-2. When the vui_colour_primaries syntax element is not present, the value of vui_colour_primaries is inferred to be equal to 2 (chromaticity is unknown, unspecified, or determined by other means not specified in this specification). Values ​​of vui_colour_primaries identified as reserved for future use in Rec. ITU-T H.273|ISO / IEC 23091-2 shall not be present in bitstreams conforming to this version of this specification. Decoding devices shall interpret reserved values ​​of vui_colour_primaries as equal to the value 2. However, the following applies: -nnpfc_colour_primaries specifies the picture primaries resulting from applying the neural network postfilter specified in the SEI message, rather than the primaries used for CLVS. When -nnpfc_colour_primaries is not present in the Neural Network Postfilter Features SEI message, the value of nnpfc_colour_primaries is inferred to be equal to vui_colour_primaries. nnpfc_transfer_characteristics has the same semantics as specified for the vui_transfer_characteristics syntax element, which is as follows: vui_transfer_characteristics indicates the transfer characteristic function of the color representation. Its semantics are as specified for the TransferCharacteristics parameter in Rec. ITU-T H.273|ISO / IEC 23091-2. When the vui_transfer_characteristics syntax element is not present, the value of vui_transfer_characteristics is inferred to be equal to 2 (the transfer characteristics are unknown, unspecified, or determined by other means not specified in this specification). Values ​​of vui_transfer_characteristics identified as reserved for future use in Rec. ITU-T H.273|ISO / IEC 23091-2 shall not be present in bitstreams conforming to this version of this specification. Decoding devices shall interpret reserved values ​​of vui_transfer_characteristics as equal to the value 2. However, the following applies: -nnpfc_transfer_characteristics specifies the transfer characteristics of the picture resulting from applying the neural network postfilter specified in the SEI message, instead of the transfer characteristics used for CLVS. When -nnpfc_transfer_characteristics is not present in the Neural Network Postfilter Characteristics SEI message, the value of nnpfc_transfer_characteristics is inferred to be equal to vui_transfer_characteristics. nnpfc_matrix_coeffs has the same semantics as specified for the vui_matrix_coeffs syntax element, which is: vui_matrix_coeffs describes the formulas used in deriving the luma and chroma signals from the green, blue, and red, or Y, Z, and X primaries. The semantics are as specified for MatrixCoefficients in Rec. ITU-T H.273|ISO / IEC 23091-2. vui_matrix_coeffs shall not be equal to 0 unless both of the following conditions are true: -BitDepthC is equal to BitDepthY. -ChromaFormatIdc equals 3 (4:4:4 chroma format). The use of vui_matrix_coeffs equal to 0 under all other conditions is reserved for future use by ITU-T|ISO / IEC. vui_matrix_coeffs shall not be equal to 8 unless one of the following conditions is true: -BitDepthC is equal to BitDepthY, -BitDepthC is equal to BitDepthY+1 and ChromaFormatIdc is equal to 3 (4:4:4 chroma format). The specification of the use of vui_matrix_coeffs equal to 8 under all other conditions is reserved for future use by ITU-T|ISO / IEC. When the vui_matrix_coeffs syntax element is not present, the value of vui_matrix_coeffs is inferred to be equal to 2 (unknown, unspecified, or determined by other means not specified in this specification). However, the following applies: -nnpfc_matrix_coeffs specifies the matrix coefficients of the picture resulting from applying the neural network postfilter specified in the SEI message, not the matrix coefficients used for CLVS. When -nnpfc_matrix_coeffs is not present in the Neural Network Postfilter Properties SEI message, the value of nnpfc_matrix_coeffs is inferred to be equal to vui_matrix_coeffs. The allowed values ​​for -nnpfc_matrix_coeffs are not constrained by the chroma format of the decoded video picture indicated by the value of ChromaFormatIdc for the semantics of the VUI parameter. - when nnpfc_matrix_coeffs is equal to 0, nnpfc_out_order_idc shall not be equal to 1 or 3. nnpfc_inp_order_idc indicates how the sample arrays of the cropped decoded output picture are ordered as input to the post-processing filters. Table 10 contains informative descriptions of nnpfc_inp_order_idc values. The semantics of nnpfc_inp_order_idc, in the range 0 to 3, inclusive, are specified in Table 12, which specifies the process for deriving the input tensor inputTensor for different values ​​of nnpfc_inp_order_idc and for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft that specify the top-left sample location of a patch of samples contained in the input tensor. When the chroma format of the cropped decoded output picture is not 4:2:0, nnpfc_inp_order_idc shall not be equal to 3. Values ​​of nnpfc_inp_order_idc shall be in the range 0 to 255, inclusive. Values ​​of nnpfc_inp_order_idc greater than 3 are reserved for future specification by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification shall ignore SEI messages containing reserved values ​​of nnpfc_inp_order_idc.

[0065] [Table 13] A patch is a rectangular array of samples from a component (eg, a luma component or a chroma component) of a picture. nnpfc_constant_patch_size_flag equal to 0 specifies that the post-processing filter accepts as input any patch size that is a positive integer multiple of the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. When nnpfc_constant_patch_size_flag equal to 0, the patch size width shall be less than or equal to CroppedWidth. When nnpfc_constant_patch_size_flag equal to 0, the patch size height shall be less than or equal to CroppedHeight. nnpfc_constant_patch_size_flag equal to 1 specifies that the post-processing filter accepts as input exactly the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. nnpfc_patch_width_minus1+1 specifies the horizontal sample count of the patch size required for input to the post-processing filter when nnpfc_constant_patch_size_flag is equal to 1. When nnpfc_constant_patch_size_flag is equal to 0, any positive integer multiple of (nnpfc_patch_width_minus1+1) may be used as the horizontal sample count of the patch size used for input to the post-processing filter. The value of nnpfc_patch_width_minus1 shall be in the range of 0 to Min(32766,CroppedWidth-1), inclusive. nnpfc_patch_height_minus1+1 specifies the vertical sample count of the patch size required for input to the post-processing filter when nnpfc_constant_patch_size_flag is equal to 1. When nnpfc_constant_patch_size_flag is equal to 0, any positive integer multiple of (nnpfc_patch_height_minus1+1) may be used as the vertical sample count of the patch size used for input to the post-processing filter. The value of nnpfc_patch_height_minus1 shall be in the range of 0 to Min(32766,CroppedHeight-1), inclusive. nnpfc_overlap specifies the horizontal and vertical sample count of overlap of adjacent input tensors for the post-processing filter. The value of nnpfc_overlap must be in the range 0 to 16383, inclusive. The variables inpPatchWidth, inpPatchHeight, outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows:

[0066] [Table 14] outPatchWidth * CroppedWidth is nnpfc_pic_width_in_luma_samples * shall be equal to inpPatchWidth, and outPatchHeight * CroppedHeight is nnpfc_pic_height_in_luma_samples * It is a bitstream conformance requirement that it shall be equal to inpPatchHeight. nnpfc_padding_type specifies the padding process when referring to sample positions outside the cropped decoded output picture boundary, as described in Table 11. The value of nnpfc_padding_type shall be in the range of 0 to 15, inclusive.

[0067] [Table 15] nnpfc_luma_padding_val specifies the luma value used for padding when nnpfc_padding_type is equal to 4. nnpfc_cb_padding_val specifies the Cb value used for padding when nnpfc_padding_type is equal to 4. nnpfc_cr_padding_val specifies the Cr value used for padding when nnpfc_padding_type is equal to 4. The function InpSampleVal(y, x, picHeight, picWidth, croppedPic) whose inputs are vertical sample position y, horizontal sample position x, picture height picHeight, picture width picWidth, and sample array croppedPic returns the value of sampleVal derived as follows:

[0068] [Table 16]

[0069] [Table 17-1]

[0070] [Table 17-2]

[0071] [Table 17-3] An nnpfc_complexity_idc greater than 0 specifies that there may be one or more syntax elements indicating the complexity of the post-processing filter associated with nnpfc_id. An nnpfc_complexity_idc equal to 0 specifies that there are no syntax elements indicating the complexity of the post-processing filter associated with nnpfc_id. The value nnpfc_complexity_idc shall be in the range 0 to 255, inclusive. Values ​​of nnpfc_complexity_idc greater than 1 are reserved for future specification by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification shall ignore SEI messages containing reserved values ​​of nnpfc_complexity_idc. nnpfc_out_format_flag equal to 0 indicates that the sample values ​​output by the post-processing filter are real numbers, and the functions OutY() and OutC() for converting the luma and chroma sample values, respectively, output by the post-processing filter to integer values ​​at bit depths BitDepthY and BitDepthC, respectively, are specified as follows: OutY(x)=Clip3(0,(1< <BitDepthY)-1,Round(x * ((1< <BitDepthY)-1))) OutC(x) = Clip3(0,(1< <BitDepthC)-1,Round(x * ((1< <BitDepthC)-1))) nnpfc_out_format_flag equal to 1 indicates that the sample values ​​output by the post-processing filter are unsigned integers and the functions OutY() and OutC() are specified as follows:

[0072] [Table 18] The variable outTensorBitDepth is derived from the syntax element nnpfc_out_tensor_bitdepth_minus8, as described below. nnpfc_out_tensor_bitdepth_minus8 plus 8 specifies the bit depth of the sample values ​​in the output integer tensor. The value of outTensorBitDepth is derived as follows: outTensorBitDepth=nnpfc_out_tensor_bitdepth_minus8+8 It is a bitstream conformance requirement that the value of nnpfc_out_tensor_bitdepth_minus8 be in the range 0 to 24, inclusive. nnpfc_out_order_idc indicates the output order of samples resulting from the post-processing filter. Table 13 contains an informative description of nnpfc_out_order_idc values. The semantics of nnpfc_out_order_idc, in the range 0 to 3, inclusive, are specified in Table 14, which specifies the process for deriving sample values ​​in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor for different values ​​of nnpfc_out_order_idc and for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft that specify the top-left sample location of a patch of samples contained in the input tensor. When nnpfc_purpose is equal to 2 or 4, nnpfc_out_order_idc shall not be equal to 3. Values ​​of nnpfc_out_order_idc shall be in the range 0 to 255, inclusive. Values ​​of nnpfc_out_order_idc greater than 3 are reserved for future specification by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification shall ignore SEI messages containing reserved values ​​of nnpfc_out_order_idc.

[0073] [Table 19]

[0074] [Table 20-1]

[0075] [Table 20-2] The base post-processing filter for the cropped decoded output picture picA is the filter identified by the first Neural Network Post-Filter Characteristics SEI message in decoding order with a particular nnpfc_id value in the CLVS. If there is another neural network post-filter characteristics SEI message related to picture picA, with the same nnpfc_id value, with nnpfc_mode_idc equal to 1, and with different content than the neural network post-processing filter defining the base post-processing filter, then the base post-processing filter is updated by decoding the ISO / IEC 15938-17 bitstream in that neural network post-filter characteristics SEI message to obtain the post-processing filter PostProcessingFilter(). Otherwise, the post-processing filter PostProcessingFilter() is assigned to be the same as the base post-processing filter. The following process is used to filter the cropped decoded output picture using the post-processing filter PostProcessingFilter() to generate a filtered picture that includes Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated by nnpfc_out_order_idc.

[0076] [Table 21] nnpfc_reserved_zero_bit shall be equal to 0. nnpfc_uri_tag[i] contains a null-terminated UTF-8 string that specifies a tag URI. The UTF-8 string shall contain a URI with syntax and semantics as specified in IETF RFC 4151 that uniquely identifies the format and associated information about the neural network to be used as the post-processing filter specified by the nnrpf_uri[i] value. NOTE - The nnrpf_uri_tag[i] elements represent "tag" URIs that allow the format of the neural network data specified by the nnrpf_uri[i] values ​​to be uniquely identified without the need for a central registration authority. nnpfc_uri[i] shall contain a null-terminated UTF-8 string as specified in ISO / IEC 10646. The UTF-8 string shall contain a URI with syntax and semantics as specified in IETF Internet Standard 66 that identifies the neural network information (e.g., data representation) to be used as a post-processing filter. Let nnpfc_payload_byte[i] contain the i-th byte of a bitstream that conforms to ISO / IEC 15938-17. The byte sequence nnpfc_payload_byte[i] for all present values ​​of i shall be a complete bitstream that conforms to ISO / IEC 15938-17. nnpfc_parameter_type_idc equal to 0 indicates that the neural network uses integer parameters only. nnpfc_parameter_type_flag equal to 1 indicates that the neural network may use floating point or integer parameters. nnpfc_parameter_type_idc equal to 2 indicates that the neural network uses binary parameters only. nnpfc_parameter_type_idc equal to 3 is reserved for future specification by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification shall ignore SEI messages containing reserved values ​​of nnpfc_parameter_type_idc. nnpfc_log2_parameter_bit_length_minus3 equal to 0, 1, 2, and 3 indicates that the neural network shall not use parameters with bit lengths greater than 8, 16, 32, and 64, respectively. When nnpfc_parameter_type_idc is present and nnpfc_log2_parameter_bit_length_minus3 is not present, the neural network shall not use parameters with bit lengths greater than 1. nnpfc_num_parameters_idc indicates the maximum number of neural network parameters for the post-processing filters, in units of powers of 2048. nnpfc_num_parameters_idc equal to 0 indicates that the maximum number of neural network parameters is unspecified. The value nnpfc_num_parameters_idc shall be in the range 0 to 52, inclusive. Values ​​of nnpfc_num_parameters_idc greater than 52 are reserved for future specification by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification shall ignore SEI messages containing reserved values ​​of nnpfc_num_parameters_idc. If the value of nnpfc_num_parameters_idc is greater than 0, the variable maxNumParameters is derived as follows: maxNumParameters=(2048< <nnpfc_num_parameters_idc)-1 It is a bitstream conformance requirement that the number of neural network parameters in the post-processing filter shall be less than or equal to maxNumParameters. nnpfc_num_kmac_operations_idc greater than 0 specifies the maximum number of multiply-accumulate operations per sample for the post-processing filter. * Specifies that the maximum number of network multiply-add operations is not specified. nnpfc_num_kmac_operations_idc equal to 0 specifies that the maximum number of network multiply-add operations is not specified. The value of nnpfc_num_kmac_operations_idc can range from 0 to 2, inclusive. 32 It should be in the range of -1.

[0077] Table 15 shows the syntax of the neural network post-filter activation SEI message defined in JVET-AA2006.

[0078] [Table 22]

[0079] With respect to Table 15, JVET-AA2006 specifies the following semantics: This SEI message specifies the neural network post-processing filters that can be used for post-processing filtering for the current picture. The neural network post-processing filter activation SEI message only persists for the current picture. NOTE - There may be several neural network post-processing filter activation SEI messages for the same picture, for example when the post-processing filters are for different purposes or filter different color components. nnpfa_id relates to the current picture and specifies that the neural network post-processing filters specified by one or more Neural Network Post-Processing Filter Characteristics SEI messages with nnpfc_id equal to nnfpa_id may be used for post-processing filtering for the current picture. Furthermore, in some cases, with respect to the use of the NN postfilter characteristics SEI message, please note the following: For interpretation of the Neural Network Post Filter Characteristics SEI message, the following variables are specified: -InpPicWidthInLumaSamples pps_pic_width_in_luma_samples-SubWidthC * It is set equal to (pps_conf_win_left_offset + pps_conf_win_right_offset). -InpPicHeightInLumaSamples pps_pic_height_in_luma_samples-SubHeightC * It is set equal to (pps_conf_win_top_offset + pps_conf_win_bottom_offset). - the variable CroppedYPic[y][x] and the chroma sample array CroppedCbPic[y][x] and CroppedCrPic[y][x], when present, is the cropped decoded output picture to which the neural network postfilter characteristic SEI message is applied, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 29, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55 番目 , 1 番目 , and 2 番目 The input signal is set to be a two-dimensional array of decoded sample values ​​of the components of - BitDepthY and BitDepthC are both set equal to BitDepth. -InpSubWidthC is set equal to SubWidthC. - InpSubHeightC is set equal to SubHeightC. - SliceQPY is set equal to SliceQpY. When neural network postfilter features SEI messages with the same nnpfc_id and different contents are present in the same picture unit, both neural network postfilter features SEI messages shall be present in the same SEI NAL unit.

[0080] The neural network postfilter characteristics SEI message defined in JVET-AA2006 may be less than ideal. In particular, for example, the signaling in JVET-AA2006 may be insufficient for signaling neural network postfilter parameters related to picture interpolation for frame rate upsampling. In accordance with the techniques described herein, additional syntax and semantics are provided for indicating neural network postfilter parameters, including for generating interpolated pictures for frame rate upsampling.

[0081] FIG. 1 is a block diagram illustrating an example of a system that may be configured to code (i.e., encode and / or decode) video data in accordance with one or more techniques of this disclosure. System 100 represents an example of a system that may encapsulate video data in accordance with one or more techniques of this disclosure. As shown in FIG. 1, system 100 includes a source device 102, a communication medium 110, and a destination device 120. In the example shown in FIG. 1, source device 102 may include any device configured to encode video data and transmit the encoded video data to communication medium 110. Destination device 120 may include any device configured to receive the encoded video data via communication medium 110 and decode the encoded video data. Source device 102 and / or destination device 120 may include computing devices equipped for wired and / or wireless communication, and may include, for example, set-top boxes, digital video recorders, televisions, desktop, laptop, or tablet computers, gaming consoles, medical imaging devices, and mobile devices including, for example, smartphones, cellular phones, personal gaming devices.

[0082] The communication medium 110 may include any combination of wireless and wired communication media and / or storage devices. The communication medium 110 may include coaxial cables, fiber optic cables, twisted pair cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other equipment that may be useful for facilitating communication between various devices and sites. The communication medium 110 may include one or more networks. For example, the communication medium 110 may include a network configured to enable access to the World Wide Web, e.g., the Internet. The network may operate according to a combination of one or more telecommunications protocols. The telecommunications protocols may include proprietary aspects and / or standardized telecommunications protocols. Examples of standardized telecommunications protocols include Digital Video Broadcasting (DVB) standards, Advanced Television Systems Committee (ATSC) standards, Integrated Services Digital Broadcasting (ISDB) standards, Data Over Cable Service Interface Specification (DOCSIS) standards, Global System Mobile Communications (GSM) standards, code division multiple access (CDMA) standards, 3rd Generation Partnership Project (3GPP) standards, European Telecommunications Standards Institute (ETSI) standards, Internet Protocol (IP) standards, Wireless Application Protocol (WAP) standards, and Institute of Electrical and Electronics Engineers (IEEE) standards.

[0083] A storage device may include any type of device or storage medium capable of storing data. A storage medium may include a tangible or non-transitory computer-readable medium. A computer-readable medium may include an optical disk, a flash memory, a magnetic memory, or any other suitable digital storage medium. In some examples, a memory device or a portion thereof may be described as a non-volatile memory, and in other examples, a portion of a memory device may be described as a volatile memory. Examples of volatile memory may include random access memory (RAM), dynamic random access memory (DRAM), and static random access memory (SRAM). Examples of non-volatile memory may include a magnetic hard disk, an optical disk, a floppy disk, a flash memory, or a form of electrically programmable memory (EPROM) or electrically erasable and programmable memory (EEPROM). The storage device(s) may include memory cards (e.g., Secure Digital (SD) memory cards), internal / external hard disk drives, and / or internal / external solid state drives. Data may be stored on the storage device according to a defined file format.

[0084] FIG. 4 is a conceptual diagram illustrating an example of components that may be included in one implementation of system 100. In the exemplary implementation shown in FIG. 4, system 100 includes one or more computing devices 402A-402N, a television service network 404, a television service provider site 406, a wide area network 408, a local area network 410, and one or more content provider sites 412A-412N. The implementation shown in FIG. 4 represents an example of a system that may be configured to enable digital media content, such as movies, live sporting events, as well as data and applications and their associated media presentations, to be distributed to and accessed by multiple computing devices, such as computing devices 402A-402N. In the example shown in FIG. 4, computing devices 402A-402N may include any device configured to receive data from one or more of television service network 404, wide area network 408, and / or local area network 410. For example, the computing devices 402A-402N may be equipped for wired and / or wireless communication and may be configured to receive services over one or more data channels, and may include televisions, including so-called smart televisions, set-top boxes, and digital video recorders. Additionally, the computing devices 402A-402N may include desktop, laptop, or tablet computers, gaming consoles, mobile devices, including, for example, "smart" phones, cellular phones, and personal gaming devices.

[0085] The television service network 404 is an example of a network configured to enable the delivery of digital media content, which may include television services. For example, the television service network 404 may include a public terrestrial television network, a public or subscription-based satellite television service provider network, and a public or subscription-based cable television provider network and / or an over-the-top service provider or an Internet service provider. It should be noted that, in some examples, the television service network 404 may be primarily used to enable the provision of television services, but the television service network 404 may also enable the provision of other types of data and services based on any combination of telecommunication protocols described herein. Furthermore, it should be noted that in some examples, the television service network 404 may enable bidirectional communication between the television service provider site 406 and one or more of the computing devices 402A-402N. The television service network 404 may include any combination of wireless communication media and / or wired communication media. The television service network 404 may include coaxial cables, fiber optic cables, twisted pair cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other equipment that may be useful for facilitating communication between various devices and sites. The television service network 404 may operate according to a combination of one or more telecommunications protocols. The telecommunications protocols may include proprietary aspects and / or standardized telecommunications protocols. Examples of standardized telecommunications protocols include the DVB standard, the ATSC standard, the ISDB standard, the DTMB standard, the DMB standard, the Data Over Cable Service Interface Specification (DOCSIS) standard, the HbbTV standard, the W3C standard, and the UPnP standard.

[0086] Referring again to FIG. 4, the television service provider site 406 can be configured to distribute television services over the television service network 404. For example, the television service provider site 406 can include one or more broadcast stations, cable television providers, or satellite television providers, or Internet-based television providers. For example, the television service provider site 406 can be configured to receive transmissions including television programs over satellite uplinks / downlinks. Additionally, as shown in FIG. 4, the television service provider site 406 can be in communication with a wide area network 408 and can be configured to receive data from content provider sites 412A-412N. It should be noted that in some examples, the television service provider site 406 can include a television studio from which content can originate.

[0087] The wide area network 408 may include a packet-based network and may operate according to a combination of one or more telecommunications protocols. The telecommunications protocols may include proprietary aspects and / or standardized telecommunications protocols. Examples of standardized telecommunications protocols include the Global System Mobile Communications (GSM) standard, the code division multiple access (CDMA) standard, the Third Generation Partnership Project (3GPP) standard, and the IEEE 802.11 standard. rdExamples of standards that may be used include 3GPP (Global Positioning System) standards, ETSI (European Telecommunications Standards Institute) standards, EN (European Standards), IP standards, Wireless Application Protocol (WAP) standards, and Institute of Electrical and Electronics Engineers (IEEE) standards, such as one or more of the IEEE 802 standards (e.g., Wi-Fi). The wide area network 408 may include any combination of wireless communication media and / or wired communication media. The wide area network 408 may include coaxial cables, fiber optic cables, twisted pair cables, Ethernet cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other equipment that may be useful for facilitating communication between various devices and sites. In one example, the wide area network 408 may include the Internet. The local area network 410 may include packet-based networks and operate according to a combination of one or more telecommunication protocols. The local area network 410 may be distinguished from the wide area network 408 based on the level of access and / or physical infrastructure. For example, the local area network 410 may include a secure home network.

[0088] Referring again to FIG. 4, the content provider sites 412A-412N represent examples of sites that may provide multimedia content to the television service provider site 406 and / or the computing devices 402A-402N. For example, the content provider sites may include a studio having one or more studio content servers configured to provide multimedia files and / or streams to the television service provider site 406. In one example, the content provider sites 412A-412N may be configured to provide multimedia content using an IP suite. For example, the content provider sites may be configured to provide multimedia content to receiving devices according to the Real Time Streaming Protocol (RTSP), HTTP, or the like. Additionally, the content provider sites 412A-412N may be configured to provide data, including hypertext-based content, or the like, over the wide area network 408 to one or more of the receiving devices, the computing devices 402A-402N, and / or the television service provider site 406. The content provider sites 412A-412N may include one or more web servers. The data provided by the data provider sites 412A-412N may be defined according to a data format.

[0089] Referring again to FIG. 1, source device 102 includes video source 104, video encoder 106, data encapsulator 107, and interface 108. Video source 104 may include any device configured to capture and / or store video data. For example, video source 104 may include a video camera and a storage device operatively coupled thereto. Video encoder 106 may include any device configured to receive video data and generate a compliant bitstream representing the video data. A compliant bitstream may refer to a bitstream that a video decoder may receive and from which the video data can be regenerated. Aspects of a compliant bitstream may be defined according to a video encoding standard. When generating a compliant bitstream, video encoder 106 may compress the video data. The compression may be lossy (perceptible or imperceptible to a viewer) or lossless. FIG. 5 is a block diagram illustrating an example of a video encoder 500 that may implement techniques for encoding video data described herein. It should be noted that while the exemplary video encoding device 500 is shown as having distinct functional blocks, such illustration is for purposes of explanation and is not intended to limit the video encoding device 500 and / or its subcomponents to any particular hardware or software architecture. The functionality of the video encoding device 500 may be realized using any combination of hardware, firmware, and / or software implementations.

[0090] The video encoding device 500 may perform intra-prediction and inter-prediction coding of a picture portion, and may therefore be referred to as a hybrid video encoding device. In the example shown in FIG. 5, the video encoding device 500 receives a source video block. In some examples, the source video block may include a portion of a picture that is partitioned according to a coding structure. For example, the source video data may include a macroblock, a CTU, a CB, a subdivision thereof, and / or another equivalent coding unit. In some examples, the video encoding device 500 may be configured to perform additional subdivision of the source video block. It should be noted that the techniques described herein are generally applicable to video encoding, regardless of how the source video data is partitioned before and / or during encoding. In the example shown in FIG. 5, the video encoding device 500 includes an adder 502, a transform coefficient generator 504, a coefficient quantizer 506, an inverse quantization and transform coefficient processor 508, an adder 510, an intra-prediction processor 512, an inter-prediction processor 514, a filter unit 516, and an entropy encoder 518. As shown in FIG. 5, a video encoding device 500 receives source video blocks and outputs a bitstream.

[0091] In the example shown in FIG. 5, the video encoding device 500 may generate residual data by subtracting a predictive video block from a source video block. Selection of the predictive video block is described in more detail below. The adder 502 represents a component configured to perform this subtraction operation. In one example, the subtraction of the video blocks is performed in the pixel domain. The transform coefficient generator 504 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block or a subdivision thereof (e.g., four 8×8 transforms may be applied to a 16×16 array of residual values) to generate a set of residual transform coefficients. The transform coefficient generator 504 may be configured to perform any and all combinations of transforms included in the family of discrete triangular transforms, including approximations of the discrete triangular transform. The transform coefficient generator 504 may output the transform coefficients to a coefficient quantizer 506. The coefficient quantizer 506 may be configured to perform quantization of the transform coefficients. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may change the rate distortion (i.e., video bit rate vs. quality) of the coded video data. The degree of quantization may be modified by adjusting a quantization parameter (QP). The quantization parameter may be determined based on slice level values ​​and / or CU level values ​​(e.g., CU delta QP value). The QP data may include any data used to determine a QP for quantizing a particular set of transform coefficients. As shown in FIG. 5, the quantized transform coefficients (which may also be referred to as level values) are output to an inverse quantization and transform coefficient processing unit 508. The inverse quantization and transform coefficient processing unit 508 may be configured to apply inverse quantization and inverse transform to generate reconstructed residual data. As shown in FIG. 5, the reconstructed residual data may be added to a prediction video block at an adder 510. In this manner, the coded video block may be reconstructed, and the resulting reconstructed video block may be used to evaluate the encoding quality for a given prediction, transform, and / or quantization.The video encoding device 500 may be configured to perform multiple encoding passes (e.g., performing encoding while varying one or more of prediction, transformation parameters, and quantization parameters). The rate-distortion or other system parameters of the bitstream may be optimized based on evaluation of the reconstructed video blocks. Furthermore, the reconstructed video blocks may be stored and used as references for predicting subsequent blocks.

[0092] Referring to FIG. 5, the intra-predictor 512 may be configured to select an intra-prediction mode for a video block to be coded. The intra-predictor 512 may be configured to evaluate a frame and determine an intra-prediction mode to use to code a current block. As described above, possible intra-prediction modes may include a planar prediction mode, a DC prediction mode, and an angular prediction mode. Additionally, it should be noted that in some examples, a prediction mode for a chroma component may be inferred from a prediction mode for a luma prediction mode. The intra-predictor 512 may select an intra-prediction mode after performing one or more coding passes. Additionally, in one example, the intra-predictor 512 may select a prediction mode based on a rate-distortion analysis. As shown in FIG. 5, the intra-predictor 512 outputs intra-prediction data (e.g., syntax elements) to the entropy encoder 518 and the transform coefficient generator 504. As described above, a transform performed on the residual data may be mode-dependent (e.g., a secondary transform matrix may be determined based on the prediction mode).

[0093] Referring again to FIG. 5, the inter prediction processor 514 may be configured to perform inter prediction coding on the current video block. The inter prediction processor 514 may be configured to receive a source video block and calculate a motion vector for a PU of the video block. The motion vector may indicate a displacement of a prediction unit of the video block in the current video frame relative to a prediction block in a reference frame. The inter prediction coding may use one or more reference pictures. Furthermore, the motion prediction may be uni-predictive (using one motion vector) or bi-predictive (using two motion vectors). The inter prediction processor 514 may be configured to select a prediction block by calculating pixel differences determined, for example, by sum of absolute difference (SAD), sum of square difference (SSD), or other difference measures. As described above, a motion vector may be determined and determined according to the motion vector prediction. The inter prediction processor 514 may be configured to perform motion vector prediction as described above. The inter prediction processor 514 may be configured to generate a prediction block using the motion prediction data. For example, the inter prediction processor 514 may place the prediction video block in a frame buffer (not shown in FIG. 5). Note that the inter prediction processor 514 may be further configured to apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values ​​for use in motion prediction. The inter prediction processor 514 may output motion prediction data for the calculated motion vectors to the entropy encoder 518.

[0094] Referring back to FIG. 5, the filter unit 516 receives the reconstructed video blocks and the coding parameters and outputs modified reconstructed video data. The filter unit 516 may be configured to perform deblocking and / or sample adaptive offset (SAO) filtering. SAO filtering is a non-linear amplitude mapping that may be used to improve reconstruction by adding an offset to the reconstructed video data. It should be noted that, as shown in FIG. 5, the intra-prediction unit 512 and the inter-prediction unit 514 may receive the modified reconstructed video blocks via the filter unit 516. The entropy encoder 518 receives the quantized transform coefficients and prediction syntax data (i.e., intra-prediction data and motion prediction data). It should be noted that in some examples, the coefficient quantizer 506 may perform a scan of a matrix including the quantized transform coefficients before the coefficients are output to the entropy encoder 518. In other examples, the entropy encoder 518 may perform a scan. The entropy encoder 518 may be configured to perform entropy encoding according to one or more of the techniques described herein. Thus, video encoding apparatus 500 represents one example of a device configured to generate encoded video data in accordance with one or more techniques of this disclosure.

[0095] Referring again to FIG. 1, data encapsulator 107 may receive encoded video data and generate a compliant bitstream, such as a series of NAL units, according to a defined data structure. A device receiving the compliant bitstream may regenerate the video data therefrom. Additionally, as discussed above, sub-bitstream extraction may refer to a process in which a device receiving a compliant bitstream forms a new compliant bitstream by discarding and / or modifying data in the received bitstream. Note that the term conforming bitstream may be used in place of the term compliant bitstream. In one example, data encapsulator 107 may be configured to generate syntax according to one or more techniques described herein. Note that data encapsulator 107 need not be located within the same physical device as video encoder 106. For example, the functions described as being performed by video encoder 106 and data encapsulator 107 may be distributed between the devices shown in FIG. 4.

[0096] As mentioned above, the signaling specified in JVET-AA2006 may be insufficient. In particular, JVET-AA2006 does not specify sufficient signaling for picture interpolation. In accordance with the technology herein, signaling for picture interpolation and frame rate upsampling is provided. In one example, in accordance with the technology herein, a syntax element is provided in a neural network postfilter characteristics SEI message indicating how multiple pictures are concatenated. In one example, in accordance with the technology herein, a syntax element is provided in a neural network postfilter characteristics SEI message indicating whether normalization is performed. In one example, in accordance with the technology herein, a syntax element is provided in a neural network postfilter characteristics SEI message indicating whether input data is manipulated according to normalized data. Tables 16, 17, and 18 show the syntax of an exemplary neural network postfilter characteristics SEI message in accordance with the technology herein.

[0097] [Table 23-1]

[0098] [Table 23-2]

[0099] [Table 24-1]

[0100] [Table 24-2]

[0101] [Table 24-3]

[0102] [Table 25-1]

[0103] [Table 25-2]

[0104] [Table 25-3]

[0105] With respect to Tables 16, 17, and 18, in one example, one or more of the syntax elements nnpfc_input_extra_dimension and / or nnpfc_output_extra_dimension may be signaled unconditionally, rather than only if nnpfc_purpose==5. Further, in one example, the condition under which the mean and standard deviation value syntax elements (e.g., nnpfc_luma_mean_val and nnpfc_luma_std_val) are signaled may be described as if((nnpfc_inp_format_flag==0)&&(nnpfc_inp_standardization_flag==1)). Further, with respect to Tables 16, 17, and 18, in some examples, the syntax elements nnpfc_in_standardization_flag (and the syntax elements conditionally signaled therein) and nnpfc_out_standardization_flag may not be included in the syntax.

[0106] With respect to Table 16, Table 17, and Table 18, the semantics are based on the semantics provided above and may be based on the following: nnpfc_purpose indicates the purpose of the post-processing filter as specified in Table 19. The value of nnpfc_purpose is between 0 and 2, inclusive. 32nnpfc_purpose shall be in the range -2. Values ​​of nnpfc_purpose not appearing in Table 19 are reserved for future specification by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification shall ignore SEI messages containing reserved values ​​of nnpfc_purpose.

[0107] [Table 26] NOTE – When reserved values ​​of nnpfc_purpose are used in the future by ITU-T|ISO / IEC, the syntax of this SEI message may be extended with syntax elements whose presence is conditioned by nnpfc_purpose being equal to that value. When SubWidthC is equal to 1 and SubHeightC is equal to 1, nnpfc_purpose shall not be equal to 2 or 4. nnpfc_number_of_input_pictures_minus1 plus 1 specifies the number of decoded output pictures used as input for the post-processing filter. nnpfc_number_of_interpolated_pictures_minus1 plus 1 specifies the number of interpolated pictures generated by the post-processing filter. nnpfc_input_extra_dimension specifies the method used to concatenate multiple input pictures before passing them as input to a post-processing filter. A value of 0 specifies that the pictures are concatenated in the channel dimension. In this case, the extra input dimension is not used, and n pictures are concatenated in n * The image is treated as one picture with c channels. This case is typically used when the network consists of classical 2D convolution operations and the goal is not to significantly increase the computational cost. A value of 1 specifies that an extra dimension is added and the input pictures are concatenated in the first dimension (i.e., as a list of pictures). This case is typically used when a post-processing filter first processes the pictures independently before concatenation. Each picture is passed through a separate sub-neural network to extract deep features, and then these deep features are concatenated and passed through another sub-neural network. A value of 2 specifies that an extra dimension is added, and the input pictures are concatenated in the second dimension of the input tensor. This case typically arises when the video input is treated as a time series, and the output of the network at some time t depends on both the current picture and the previous picture. A value of 3 specifies that an extra dimension is added before the height dimension in the input tensor. This case is typically used when post-processing consists of a 3D convolution operation. In this case, the dimension indicating the number of pictures (or the extra dimension) is treated as the third spatial dimension, and the height and width are treated as the other two. The 3D kernel is convolved with the 3D data. nnpfc_output_extra_dimension indicates how the post-processing filter concatenates multi-picture outputs. It has values ​​from 0 to 3. Values ​​0 to 3 have the same semantics as described for nnpfc_input_extra_dimension. Knowing the nnpfc_output_extra_dimension value is necessary to apply the StoreOutputTensor() routine defined later. nnpfc_inp_format_flag indicates how to convert the sample values ​​of the cropped decoded output picture to the input values ​​to the post-processing filter. When nnpfc_inp_format_flag is equal to 0, the input values ​​to the post-processing filter are real numbers. When nnpfc_inp_format_flag is equal to 1, the input values ​​to the post-processing filter are unsigned integers. nnpfc_inp_normalization_range_flag equal to 1 specifies that InpY and InpC normalization is to the range [0,1]. nnpfc_inp_normalization_range_flag equal to 0 specifies that InpY and InpC normalization is to the range [-1,1]. Normalization range of [-1,1]: Neural network-based methods normalize input images within a small range to stabilize the training procedure. Two ranges are common for normalization: [-1,1] and [0,1]. The dominance of one over the other is highly debated. However, a significant number of research works in the field of image enhancement normalize data to the [-1,1] range. The proposed signaling provides support for normalization in both ranges. The functions InpY and InpC are specified as follows:

[0108] [Table 27] nnpfc_inp_standardization_flag equal to 1 specifies that the input to the post-processing filter is normalized. nnpfc_inp_standardization_flag equal to 0 specifies that no normalization is applied to the input to the post-filter. nnpfc_luma_mean_val specifies the mean of the luma samples used for normalization. nnpfc_cb_mean_val specifies the mean of the chroma Cb samples used for normalization. nnpfc_cr_mean_val specifies the mean of the chroma Cr samples used for normalization. nnpfc_luma_std_val specifies the standard deviation of the luma samples used for normalization. nnpfc_cb_std_val specifies the standard deviation of the chroma Cb samples used for normalization. nnpfc_cr_std_val specifies the standard deviation of the chroma Cr samples used for normalization.

[0109] [Table 28] nnpfc_inp_order_idc can be based on any one of the following: nnpfc_inp_order_idc indicates how the sample arrays of the cropped decoded output picture are ordered as input to the post-processing filter. Table 20A includes an informative description of nnpfc_inp_order_idc values ​​for an example. Table 20B includes an informative description of nnpfc_inp_order_idc values ​​for an example. The semantics of nnpfc_inp_order_idc, in the range 0 to 3 inclusive, are specified in an example Table 12, and in an example Tables 22A and 22B, which specify the process for deriving the input tensor inputTensor for different values ​​of nnpfc_inp_order_idc and for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft that specify the top-left sample location of a patch of samples contained in the input tensor. When the chroma format of the cropped decoded output picture is not 4:2:0, nnpfc_inp_order_idc shall not be equal to 3. Values ​​of nnpfc_inp_order_idc shall be in the range 0 to 255, inclusive. Values ​​of nnpfc_inp_order_idc greater than 3 are reserved for future specification by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification shall ignore SEI messages containing reserved values ​​of nnpfc_inp_order_idc.

[0110] [Table 29]

[0111] [Table 30] nnpfc_inp_order_idc indicates how the sample arrays of the cropped decoded output picture are ordered as input to the post-processing filter. Table 21A includes an informative description of nnpfc_inp_order_idc values ​​for an example. Table 21B includes an informative description of nnpfc_inp_order_idc values ​​for an example. The semantics of nnpfc_inp_order_idc, in the range 0 to 7 inclusive, are specified in an example Table 23A and an example Table 23B, which specify the process for deriving the input tensor inputTensor for different values ​​of nnpfc_inp_order_idc and for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft that specify the top-left sample location of a patch of samples contained in the input tensor. When the chroma format of the cropped decoded output picture is not 4:2:0, nnpfc_inp_order_idc shall not be equal to 3. Values ​​of nnpfc_inp_order_idc shall be in the range 0 to 255, inclusive. Values ​​of nnpfc_inp_order_idc greater than 7 are reserved for future specification by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification shall ignore SEI messages containing reserved values ​​of nnpfc_inp_order_idc.

[0112] [Table 31-1]

[0113] [Table 31-2]

[0114] [Table 32-1]

[0115] [Table 32-2] nnpfc_overlap specifies the horizontal and vertical sample count of overlap of adjacent input tensors for the post-processing filter. The value of nnpfc_overlap must be in the range 0 to 16383, inclusive. The variables inpPatchWidth, inpPatchHeight, outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows:

[0116] [Table 33]

[0117] [Table 34-1]

[0118] [Table 34-2]

[0119] [Table 34-3]

[0120] [Table 34-4] [Table 34-5]

[0121] [Table 34-6]

[0122]

Table 34-7

[0123]

Table 35-1

[0124]

Table 35-2

[0125]

Table 35-3

[0126]

Table 36-1

[0127]

Table 36-2

[0128]

Table 36-3

[0129]

Table 36-4

[0130]

Table 36-5

[0131]

Table 36-6

[0132] [Table 36-7]

[0133] [Table 36-8]

[0134] [Table 36-9]

[0135] [Table 37-1]

[0136] [Table 37-2]

[0137] [Table 37-3]

[0138] [Table 37-4] nnpfc_out_normalization_range_flag equal to 1 specifies that the output of the post-processing filter is in the range [0,1]. nnpfc_out_normalization_range_flag equal to 0 specifies that the output of the post-processing filter is in the range [-1,1].

[0139] [Table 38] nnpfc_out_format_flag equal to 1 indicates that the sample values ​​output by the post-processing filter are unsigned integers and the functions OutY and OutC are specified as follows:

[0140] [Table 39] The variable outTensorBitDepth is derived from the syntax element nnpfc_out_tensor_bitdepth_minus8. nnpfc_out_standardization_flag equal to 1 specifies that the output of post-processing should be denormalized. nnpfc_out_standardization_flag equal to 0 specifies that denormalization is not required.

[0141] [Table 40] The nnpfc_out_order_idc can be based on any one of the following: nnpfc_out_order_idc indicates the output order of samples resulting from the post-processing filter. Table 13 contains an informative description of nnpfc_out_order_idc values. The semantics of nnpfc_out_order_idc, in the range of 0 to 3, inclusive, are specified in Table 14, in one example, and in Tables 24A and 24B, in one example, which specify the process for deriving sample values ​​in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor for different values ​​of nnpfc_out_order_idc and for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft that specify the top-left sample location of a patch of samples contained in the input tensor. When nnpfc_purpose is equal to 2 or 4, nnpfc_out_order_idc shall not be equal to 3. Values ​​of nnpfc_out_order_idc shall be in the range of 0 to 255, inclusive. Values ​​of nnpfc_out_order_idc greater than 3 are reserved for future specification by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification shall ignore SEI messages containing reserved values ​​of nnpfc_out_order_idc.

[0142] [Table 41-1]

[0143] [Table 41-2]

[0144] [Table 41-3]

[0145] [Table 41-4]

[0146] [Table 42-1]

[0147] [Table 42-2] The base post-processing filter for the cropped decoded output picture picA is the filter identified by the first neural network post-filter characteristics SEI message in decoding order with a particular nnpfc_id value in the CLVS. If there is another neural network post-filter characteristics SEI message related to picture picA with the same nnpfc_id value, nnpfc_mode_idc equal to 1, and with different content than the neural network post-processing filter defining the base post-processing filter, the base post-processing filter is updated by decoding the ISO / IEC 15938-17 bitstream in that neural network post-filter characteristics SEI message to obtain the post-processing filter PostProcessingFilter(). Otherwise, the post-processing filter PostProcessingFilter() is assigned to be the same as the base post-processing filter. The following process is used to filter the cropped decoded output picture using the post-processing filter PostProcessingFilter() to generate a filtered picture that includes Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated by nnpfc_out_order_idc.

[0148] [Table 43] nnpfc_out_order_idc indicates the output order of samples resulting from the post-processing filter. Table 25 contains an informative description of nnpfc_out_order_idc values. The semantics of nnpfc_out_order_idc, which ranges from 0 to 7 inclusive, are specified in Table 26A in one example and Table 26B in another example, which specify the process for deriving sample values ​​in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor for different values ​​of nnpfc_out_order_idc and for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft that specify the top-left sample location of a patch of samples contained in the input tensor. When nnpfc_purpose is equal to 2 or 4, nnpfc_out_order_idc shall not be equal to 3. Values ​​of nnpfc_out_order_idc shall be in the range of 0 to 255 inclusive. Values ​​of nnpfc_out_order_idc greater than 3 are reserved for future specification by ITU-T|ISO / IEC and shall not be present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification shall ignore SEI messages containing reserved values ​​of nnpfc_out_order_idc.

[0149] [Table 44]

[0150] [Table 45-1]

[0151] [Table 45-2]

[0152] [Table 45-3]

[0153] [Table 45-4]

[0154] [Table 45-5]

[0155] [Table 46-1]

[0156] [Table 46-2]

[0157] [Table 46-3] The base post-processing filter for the cropped decoded output picture picA is the filter identified by the first Neural Network Post-Filter Characteristics SEI message in decoding order with a particular nnpfc_id value in the CLVS. If there is another neural network post-filter characteristics SEI message related to picture picA, with the same nnpfc_id value, with nnpfc_mode_idc equal to 1, and with different content than the neural network post-processing filter defining the base post-processing filter, then the base post-processing filter is updated by decoding the ISO / IEC 15938-17 bitstream in that neural network post-filter characteristics SEI message to obtain the post-processing filter PostProcessingFilter(). Otherwise, the post-processing filter PostProcessingFilter() is assigned to be the same as the base post-processing filter. The following process is used to filter the cropped decoded output picture using the post-processing filter PostProcessingFilter() to generate a filtered picture that includes Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated by nnpfc_out_order_idc.

[0158] [Table 47]

[0159] Note that according to the process provided in Tables 22A and 22B, the pseudocode for nnpfc_inp_order_idc equal to 0, 1, 2, and 3 is modified to be general and work for both single and multiple pictures. Note that according to the process provided in Table 23A, values ​​5-7 are defined for nnpfc_inp_order_idc to handle more than one input image (e.g., in case of frame rate upsampling). Note that according to the process provided in Table 23B, multiple input pictures are always passed to the post-processing filter as a list, and the post-processing filter is responsible for the concatenation.

[0160] Note that according to the process provided in Tables 24A and 24B, the pseudocode for nnpfc_out_order_idc equal to 0, 1, 2, and 3 is generalized to handle the multiple picture output case. Note that according to the process provided in Table 25A, values ​​5 through 7 are defined for nnpfc_out_order_idc to handle the multiple picture output case. Note that according to the process provided in Table 25B, the output of the post-processing filter is always received as a list.

[0161] With respect to Tables 16, 17, and 18, in some examples, the number of input pictures may be indicated using minus-two signaling. That is, in one example, the following syntax in Tables 16, 17, and 18:

[0162] [Table 48] can be replaced with the following syntax:

[0163] [Table 49] Or you can replace it with the following syntax:

[0164] [Table 50] During the ceremony, nnpfc_number_of_input_pictures_minus2 plus 2 specifies the number of decoded output pictures used as input for the post-processing filter. nnpfc_interpolated_pictures[i] specifies the number of interpolated pictures generated by the post-processing filter between the i-th picture and the (i+1)-th picture that are used as input for the post-processing filter. nnpfc_output_extra_dimension[i] indicates how the post-processing filter concatenates multi-picture outputs. It has values ​​from 0 to 3. Values ​​0 to 3 have the same semantics as described for nnpfc_input_extra_dimension. Knowing the nnpfc_output_extra_dimension value is necessary to apply the StoreOutputTensor() routine defined later. The variables numInputImages and numOutputImages are derived as follows:

[0165] [Table 51]

[0166] In one example, in accordance with the techniques herein, the routines DeriveInputTensors() and StoreOutputTensors() may be applied to each picture in a video. That is, in one example, the following process may be used to filter the cropped decoded output picture with a post-processing filter PostProcessingFilter() to generate a filtered picture that includes Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated by nnpfc_out_order_idc.

[0167] [Table 52-1]

[0168] [Table 52-2]

[0169] [Table 52-3]

[0170] With respect to the above process, in one example, the process ConcatenateInputTensors() may be based on the process provided in Table 27A or Table 27B. It should be noted that according to the process in Table 27A, pictures are always concatenated in the first dimension, and according to the process in Table 27B, the input tensors are concatenated based on the value of nnpfc_input_extra_dimension. Furthermore, with respect to the above process, in one example, the process SeparateOutputTensors() may be based on the process provided in Table 28A or Table 28B. It should be noted that according to the process in Table 28A, the pictures output by the post-processing filter are separated, and according to the process in Table 28B, the output tensors of the post-processing filter are separated based on the value of nnpfc_output_extra_dimension.

[0171] [Table 53-1]

[0172] [Table 53-2]

[0173] [Table 54-1]

[0174] [Table 54-2]

[0175] [Table 54-3]

[0176] [Table 54-4]

[0177]

Table 54-5

[0178]

Table 54-6

[0179]

Table 54-7

[0180]

Table 55-1

[0181]

Table 55-2

[0182]

Table 56-1

[0183]

Table 56-2

[0184]

Table 56-3

[0185]

Table 56-4

[0186] In accordance with the techniques herein, in some examples, thus, video encoding apparatus 500 represents an example of a device configured to signal a neural network postfilter characteristic message, signal a first syntax element in the neural network postfilter characteristic message that specifies the number of input pictures to be used as input for the neural network postfilter picture interpolation process, and signal a second syntax element in the neural network postfilter characteristic message having a value that specifies the manner in which the input pictures are concatenated before being input to the neural network postfilter picture interpolation process.

[0187] 1, interface 108 may include any device configured to receive data generated by data encapsulator 107 and transmit and / or store the data on a communications medium. Interface 108 may include a network interface card, such as an Ethernet card, and may include an optical transceiver, a radio frequency transceiver, or any other type of device capable of transmitting and / or receiving information. Additionally, interface 108 may include a computer system interface that may allow files to be stored on a storage device. For example, interface 108 may include a Peripheral Component Interconnect (PCI) and Peripheral Component Interconnect Express (PCIe) bus protocols, a proprietary bus protocol, a Universal Serial Bus (USB) protocol, an I / O protocol, or any other type of device that may be configured to transmit and / or receive information. 2 C, or any other logical and physical structure that can be used to interconnect peer devices.

[0188] 1, destination device 120 includes interface 122, data decapsulator 123, video decoder 124, and display 126. Interface 122 may include any device configured to receive data from a communication medium. Interface 122 may include a network interface card, such as an Ethernet card, and may include an optical transceiver, a radio frequency transceiver, or any other type of device capable of receiving and / or transmitting information. Additionally, interface 122 may include an interface for a computer system that allows a compliant video bitstream to be obtained from a storage device. For example, interface 122 may include interfaces for PCI and PCIe bus protocols, proprietary bus protocols, USB protocols, I / O protocols, and the like. 2 C, or any other logical and physical structures that may be used to interconnect peer devices. The data decapsulator 123 may be configured to receive and parse any of the example syntax structures described herein.

[0189] Video decoder 124 may include any device configured to receive a bitstream (e.g., a sub-bitstream extract) and / or an acceptable variant thereof and regenerate video data therefrom. Display 126 may include any device configured to display video data. Display 126 may include one of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display. Display 126 may include a high-definition display or an ultra-high-definition display. Although in the example shown in FIG. 1, video decoder 124 is described as outputting data to display 126, it should be noted that video decoder 124 may be configured to output video data to various types of devices and / or subcomponents thereof. For example, video decoder 124 may be configured to output video data to any communication medium as described herein.

[0190] FIG. 6 is a block diagram illustrating an example of a video decoding device that may be configured to decode video data according to one or more techniques of this disclosure (e.g., the reference picture list creation decoding process described above). In one example, video decoding device 600 may be configured to decode transform data and recover residual data from transform coefficients based on the decoded transform data. Video decoding device 600 may be configured to perform intra-prediction and inter-prediction decoding, and may therefore be referred to as a hybrid decoding device. Video decoding device 600 may be configured to parse any combination of the syntax elements described above in Tables 1-28B. Video decoding device 600 may decode video based on or in accordance with the process described above, and further based on the parsed values ​​in Tables 1-28B.

[0191] In the example shown in FIG. 6, the video decoding device 600 includes an entropy decoding unit 602, an inverse quantization unit 604, an inverse transform coefficient processing unit 606, an intra prediction processing unit 608, an inter prediction processing unit 610, an adder 612, a post filter unit 614, and a reference buffer 616. The video decoding device 600 may be configured to decode video data in a manner consistent with a video encoding system. It should be noted that while the exemplary video decoding device 600 is shown having separate functional blocks, such illustration is for illustrative purposes and does not limit the video decoding device 600 and / or its subcomponents to a particular hardware or software architecture. The functionality of the video decoding device 600 may be realized using any combination of hardware, firmware, and / or software implementations.

[0192] As shown in FIG. 6, the entropy decoding unit 602 receives an entropy coded bitstream. The entropy decoding unit 602 may be configured to decode syntax elements and quantized coefficients from the bitstream according to a reverse process of the entropy coding process. The entropy decoding unit 602 may be configured to perform entropy decoding according to any of the entropy coding techniques described above. The entropy decoding unit 602 may determine values ​​of syntax elements in the coded bitstream in accordance with a video coding standard. As shown in FIG. 6, the entropy decoding unit 602 may determine quantization parameters, values ​​of quantized coefficients, transform data, and prediction data from the bitstream. In the example shown in FIG. 6, the inverse quantization unit 604 and the inverse transform coefficient processing unit 606 receive values ​​of quantized coefficients from the entropy decoding unit 602 and output reconstructed residual data.

[0193] Referring back to FIG. 6, the reconstructed residual data may be provided to an adder 612. The adder 612 may add the reconstructed residual data to a predictive video block to generate reconstructed video data. The predictive video block may be determined according to a predictive video technique (i.e., intra prediction and inter-frame prediction). The intra prediction processor 608 may be configured to receive an intra prediction syntax element and retrieve a predictive video block from a reference buffer 616. The reference buffer 616 may include a memory device configured to store one or more frames of video data. The intra prediction syntax element may identify an intra prediction mode, such as the intra prediction modes described above. The inter prediction processor 610 may receive the inter prediction syntax element and generate a motion vector to identify a predictive block in one or more reference frames stored in the reference buffer 616. The inter prediction processor 610 may possibly perform interpolation based on an interpolation filter to generate a motion compensated block. The syntax element may include an identifier of an interpolation filter to be used for motion prediction having sub-pixel accuracy. The inter-prediction processing unit 610 may use an interpolation filter to calculate interpolated values ​​for sub-integer pixels of the reference block. The post-filter unit 614 may be configured to perform filtering on the reconstructed video data. For example, the post-filter unit 614 may be configured to perform deblocking and / or sample adaptive offset (SAO) filtering, for example, based on parameters specified in the bitstream. It should be noted that in addition, in some examples, the post-filter unit 614 may be configured to perform any filtering of its own (e.g., visual enhancement such as mosquito noise reduction). As shown in FIG. 6, the reconstructed video blocks may be output by the video decoding device 600.In this manner, video decoding apparatus 600 represents an example of a device configured to receive a neural network postfilter characteristic message, parse a first syntax element from the neural network postfilter characteristic message that specifies the number of input pictures to be used as input for the neural network postfilter picture interpolation process, parse a second syntax element from the neural network postfilter characteristic message having a value that specifies the manner in which the input pictures are to be concatenated before being input to the neural network postfilter picture interpolation process, and concatenate the input pictures based on the value of the second syntax element.

[0194] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions may be stored or transmitted as one or more instructions or code on a computer-readable medium and executed by a hardware-based processor. A computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium, such as a data storage medium or a communication medium, including any medium that facilitates the transfer of a computer program from one place to another, for example according to a communication protocol. In this manner, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium, such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0195] By way of example, and without limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, other magnetic storage, flash memory, or any other medium, i.e., any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. However, it should be understood that computer-readable media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory tangible storage media. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks and Blu-ray discs, where disks typically reproduce data magnetically and discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer readable media.

[0196] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, as used herein, the term "processor" may refer to any of the foregoing structures, or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided in dedicated hardware and / or software modules configured to encode and decode, or incorporated into a composite codec. Also, the techniques may be implemented entirely in one or more circuit or logic elements.

[0197] The techniques of this disclosure can be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units are illustrated in this disclosure to highlight functional aspects of devices configured to perform the disclosed techniques, but need not necessarily be realized by different hardware units. Rather, as previously described, the various units may be combined into a codec hardware unit or may be provided by a collection of interoperating hardware units, including one or more processors as previously described, along with suitable software and / or firmware.

[0198] Moreover, each functional block or various features of the base station device and the terminal device used in each of the above implementations may be implemented or performed by a circuit, typically an integrated circuit or multiple integrated circuits. The circuit designed to perform the functions described herein may comprise a general purpose processor, a digital signal processor (DSP), an application specific or general purpose integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, or individual hardware components, or a combination thereof. The general purpose processor may be a microprocessor, or the processor may be a conventional processor, controller, microcontroller, or state machine. The general purpose processor or each circuit described above may be composed of digital circuits or analog circuits. Furthermore, as semiconductor technology advances and integrated circuit technology emerges to replace current integrated circuits, integrated circuits using this technology may also be used.

[0199] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. 1. A device comprising one or more processors, the one or more processors comprising: receiving a neural network post-filter characteristic message; parsing a first syntax element in the neural network post-filter characteristics message, the value of the first syntax element plus one specifying the number of input pictures to be used as input to the neural network post-filter; Parsing a second syntax element in the neural network post-filter characteristics message used to calculate the number of output pictures; deriving a multi-dimensional input tensor, one dimension of the multi-dimensional input tensor corresponding to the number of the input pictures; generating a multi-dimensional output tensor, a first dimension of the multi-dimensional output tensor corresponding to the number of output pictures; The device is configured as follows:

2. The device of claim 1 , wherein the device comprises a video decoder.

3. 1. A device comprising one or more processors, the one or more processors comprising: Signaling neural network post-filter characteristic messages, signaling a first syntax element in the neural network post-filter characteristics message, the value of the first syntax element plus one specifying the number of input pictures to be used as input to the neural network post-filter; signaling in the neural network post-filter characteristics message a second syntax element used to calculate the number of output pictures; deriving a multi-dimensional input tensor, one dimension of the multi-dimensional input tensor corresponding to the number of the input pictures; generating a multi-dimensional output tensor, a first dimension of the multi-dimensional output tensor corresponding to the number of output pictures; The device is configured as follows:

4. The device of claim 3 , wherein the device comprises a video encoder.

5. A computer-readable storage medium storing a program for causing a computer to execute: receiving a neural network post-filter characteristic message; Parse a first syntax element in the neural network post-filter characteristics message, where the value of the first syntax element plus one specifies the number of input pictures to be used as input to the neural network post-filter; Parsing a second syntax element in the neural network post-filter characteristics message used to calculate a number of output pictures; deriving a multi-dimensional input tensor, one dimension of the multi-dimensional input tensor corresponding to the number of the input pictures; A multi-dimensional output tensor is generated, a first dimension of the multi-dimensional output tensor corresponding to the number of output pictures.