System and method for signaling picture information in video coding
By receiving and parsing the image parameter set identifier and sequential counting elements in the slice header syntax structure, the problem of inefficient image information transmission in the existing video encoding standards is solved, and more efficient video encoding and decoding performance is achieved, supporting the expansion of the next generation of video encoding standards.
Patent Information
- Application Number
- CN202510227022.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-24
- Filing Date
- 2020-08-20
- Publication Date
- 2025-07-22
AI Technical Summary
When sending signals to the picture information of encoded videos, existing video encoding standards have problems such as inefficiency and insufficient information transmission. Especially in the development of next-generation video encoding standards, how to efficiently parse and transmit image parameter set identifiers and sequential counting information has not been effectively solved.
By receiving the slice header syntax structure, determining whether it contains the picture header syntax structure, and parsing the specified picture parameter set identifier and the second syntax element of the picture order count when the structure is included, implementing conditional syntax element delivery.
It improves the parsing efficiency and accuracy of image information during video encoding, supports the expansion and improvement of next-generation video encoding standards, and improves the encoding and decoding performance of video data.
Smart Images

Figure CN120358358A_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese Patent Application "System and Method for Signaling Picture Information in Video Coding" (Application No.: 202080058277.3) with an application date of August 20, 2020. Technical Field
[0002] The present disclosure relates to video coding and, more particularly, to techniques for signaling picture information of an encoded video. Background Art
[0003] Digital video capabilities can be incorporated into a variety of devices, including digital televisions, laptop or desktop computers, tablet computers, digital recording devices, digital media players, video game devices, cellular telephones (including so-called smart phones), medical imaging devices, and the like. Digital video can be encoded according to video coding standards. A video coding standard defines the format of a compliant bitstream that encapsulates the encoded video data. A compliant bitstream is a data structure that can be received and decoded by a video decoding device to generate reconstructed video data. Video coding standards can incorporate video compression techniques. Examples of video coding standards include ISO / IEC MPEG-4 Visual and ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) and High Efficiency Video Coding (HEVC). HEVC is described in the ITU-T Recommendation H.265, High Efficiency Video Coding (HEVC), of December 2016, which is incorporated herein by reference and is referred to herein as ITU-T H.265. Extensions and improvements to ITU-T H.265 are currently being considered to develop the next generation of video coding standards. For example, the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) (collectively referred to as the Joint Video Exploration Team (JVET)) are working on standardizing video coding technologies with significantly greater compression capabilities than the current HEVC standard. The Joint Exploration Model 7 (JEM7), the algorithm description of the Joint Exploration Test Model 7 (JEM 7), and the ISO / IEC JTC1 / SC29 / WG11 document: JVET-G1001 (July 2017, Turin, Italy), which are incorporated herein by reference, describe the coding characteristics under the Joint Test Model studies by JVET, and this technology is a potential enhanced video coding technology that exceeds the capabilities of ITU-T H.265. It should be noted that the coding characteristics of JEM 7 are implemented in the JEM reference software. As used herein, the term JEM can collectively refer to the algorithms in JEM 7 and the specific implementation of the JEM reference software. In addition, in response to the "Joint Call for Proposals on Video Compression with Capabilities beyond HEVC" jointly issued by VCEG and MPEG, at the 10th meeting of ISO / IEC JTC1 / SC29 / WG11 held in San Diego, CA from April 16 to 20, 2018, various groups presented a variety of descriptions of video coding tools.According to various descriptions of video coding tools, the final initial draft text of the video coding specification is described in "Versatile Video Coding (Draft 1)", i.e., document JVET-J1001-v2, in the 10th meeting of ISO / IEC JTC1 / SC29 / WG11 held in San Diego, California from April 16th to 20th, 2018. This document is incorporated herein by reference and is referred to as JVET-J1001. The current development of the next-generation video coding standards by JVET and MPEG is called the Versatile Video Coding (VVC) project. "Versatile Video Coding (Draft6)" (document JVET-O2001-vE, which is incorporated herein by reference and is called JVET-O2001) in the 15th meeting of ISO / IEC JTC1 / SC29 / WG11 held in Gothenburg, Sweden from July 3rd to 12th, 2019 represents the current version of the draft text of the video coding specification corresponding to the VVC project.
[0004] Video compression techniques can reduce the data requirements for storing and transmitting video data. Video compression techniques can reduce data requirements by exploiting the redundancy inherent in a video sequence. Video compression techniques can further divide a video sequence into successive smaller parts (i.e., a group of pictures within a video sequence, pictures within a group of pictures, regions within a picture, sub-regions within a region, etc.). Intra prediction coding techniques (e.g., spatial prediction techniques within a picture) and inter prediction techniques (i.e., techniques between pictures (temporal)) can be used to generate the difference between a unit of video data to be encoded and a reference unit of video data. This difference can be referred to as residual data. The residual data can be encoded as quantized transform coefficients. Syntax elements can relate to the residual data and reference coding units (e.g., intra prediction mode index and motion information). The residual data and syntax elements can be entropy encoded. The entropy-encoded residual data and syntax elements can be included in a data structure that forms a compliant bitstream. Summary of the Invention
[0005] In one example, a method for decoding picture information for decoding video data is provided, the method comprising: receiving a slice header syntax structure; determining whether the slice header syntax structure includes a picture header syntax structure; and in the case where the slice header syntax structure includes the picture header syntax structure, parsing, from the picture header syntax structure, a first syntax element specifying a picture parameter set identifier and a second syntax element specifying a picture order count.
[0006] In one example, a device is provided that includes one or more processors configured to: receive a slice header syntax structure; determine whether the slice header syntax structure includes a picture header syntax structure; and, in the case where the slice header syntax structure includes the picture header syntax structure, parse a first syntax element specifying a picture parameter set identifier from the picture header syntax structure and parse a second syntax element specifying a picture order count from the picture header syntax structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 FIG. is a block diagram illustrating an example of a system configurable to encode and decode video data in accordance with one or more techniques of the present disclosure.
[0008] Figure 2 FIG. is a conceptual diagram illustrating encoded video data and corresponding data structures in accordance with one or more techniques of the present disclosure.
[0009] Figure 3 FIG. is a conceptual diagram illustrating a data structure encapsulating encoded video data and corresponding metadata in accordance with one or more techniques of the present disclosure.
[0010] Figure 4 FIG. is a conceptual diagram illustrating an example of components that may be included in an implementation of a system configurable to encode and decode video data in accordance with one or more techniques of the present disclosure.
[0011] Figure 5 FIG. is a block diagram illustrating an example of a video encoder configurable to encode video data in accordance with one or more techniques of the present disclosure.
[0012] Figure 6 FIG. is a block diagram illustrating an example of a video decoder configurable to decode video data in accordance with one or more techniques of the present disclosure. DETAILED DESCRIPTION
[0013] In general, the present disclosure describes various techniques for encoding video data. Specifically, the present disclosure describes techniques for signaling picture information of encoded video data. It should be noted that although the techniques of the present disclosure are described with respect to ITU-T H.264, ITU-T H.265, JEM, and JVET-O2001, the techniques of the present disclosure can be generally applied to video coding. For example, in addition to those techniques included in ITU-T H.265, JEM, and JVET-O2001, the encoding techniques described herein can be incorporated into video coding systems (including video coding systems based on future video coding standards), including video block structures, intra prediction techniques, inter prediction techniques, transform techniques, filtering techniques, and / or other entropy coding techniques. Therefore, the references to ITU-T H.264, ITU-T H.265, JEM, and / or JVET-O2001 are for descriptive purposes and should not be construed as limiting the scope of the techniques described herein. Additionally, it should be noted that the incorporation of documents by reference herein is for descriptive purposes and should not be construed as limiting or creating ambiguity regarding the terms used herein. For example, in the case where the definition of a certain term provided in one incorporated reference is different from another incorporated reference and / or the term as used herein, the term should be interpreted in a manner that broadly includes each corresponding definition and / or in a manner that includes each specific definition in the alternatives.
[0014] In one example, a method of decoding video data includes: signaling a flag in a sequence parameter set, the flag having a value indicating whether each picture that references the sequence parameter set exactly includes one slice; and conditionally signaling one or more syntax elements in a slice header based on the value of the flag.
[0015] In one example, an apparatus includes one or more processors configured to: signal a flag in a sequence parameter set, the flag having a value indicating whether each picture that references the sequence parameter set exactly includes one slice; and conditionally signal one or more syntax elements in a slice header based on the value of the flag.
[0016] In one example, a non-transitory computer-readable storage medium includes instructions stored thereon that, when executed, cause one or more processors of a device to: signal a flag in a sequence parameter set, the flag having a value indicating whether each picture that references the sequence parameter set exactly includes one slice; and conditionally signal one or more syntax elements in a slice header based on the value of the flag.
[0017] In one example, an apparatus includes: means for signaling a flag in a sequence parameter set, the flag having a value indicating whether each picture that refers to the sequence parameter set exactly includes one slice; and means for conditionally signaling one or more syntax elements in a slice header based on the value of the flag.
[0018] In one example, a method of decoding video data includes: parsing a flag in a sequence parameter set, the flag having a value indicating whether each picture that refers to the sequence parameter set exactly includes one slice; and conditionally parsing one or more syntax elements in a slice header based on the value of the flag.
[0019] In one example, a device includes one or more processors configured to: parse a flag in a sequence parameter set, the flag having a value indicating whether each picture that refers to the sequence parameter set exactly includes one slice; and conditionally parse one or more syntax elements in a slice header based on the value of the flag.
[0020] In one example, a non-transitory computer-readable storage medium includes instructions stored thereon that, when executed, cause one or more processors of a device to: parse a flag in a sequence parameter set, the flag having a value indicating whether each picture that refers to the sequence parameter set exactly includes one slice; and conditionally parse one or more syntax elements in a slice header based on the value of the flag.
[0021] In one example, an apparatus includes: means for parsing a flag in a sequence parameter set, the flag having a value indicating whether each picture that refers to the sequence parameter set exactly includes one slice; and means for conditionally parsing one or more syntax elements in a slice header based on the value of the flag.
[0022] Details of one or more examples are set forth in the following figures and description. Other features, objects, and advantages will be apparent from the description, the figures, and the claims.
[0023] The video content includes a video sequence consisting of a series of frames (or pictures). A series of frames may also be referred to as a group of pictures (GOP). Each video frame or picture can be divided into one or more regions. Regions can be defined according to a basic unit (e.g., a video block) and a set of rules defining the regions. For example, the rule for defining a region can be: the region must be an integer number of video blocks arranged in a rectangle. In addition, the video blocks in a region can be sorted according to a scanning pattern (e.g., raster scan). As used herein, the term "video block" generally can refer to a region of a picture, or can more specifically refer to the largest array of sample values that can be predictively encoded, its sub-partitions, and / or corresponding structures. In addition, the term "current video block" can refer to the region of a picture that is being encoded or decoded. A video block can be defined as an array of sample values. It should be noted that in some cases, pixel values can be described as including sample values of the corresponding components of video data, which can also be referred to as color components (e.g., luminance (Y) and chrominance (Cb and Cr) components or red, green, and blue components). It should be noted that in some cases, the terms "pixel value" and "sample value" can be used interchangeably. In addition, in some cases, a pixel or a sample can be referred to as a pel. The video sampling format (which can also be referred to as the chrominance format) can define the number of chrominance samples included in a video block relative to the number of luminance samples included in the video block. For example, for the 4:2:0 sampling format, the sampling rate of the luminance component is twice the sampling rate of the chrominance components in both the horizontal and vertical directions.
[0024] A video encoder may perform predictive coding on video blocks and their sub - partitions. The video blocks and their sub - partitions may be referred to as nodes. ITU - T H.264 specifies macroblocks that include 16×16 luma samples. That is, in ITU - T H.264, a picture is segmented into macroblocks. ITU - T H.265 specifies a similar Coding Tree Unit (CTU) structure (which may be referred to as the Largest Coding Unit (LCU)). In ITU - T H.265, a picture is segmented into CTUs. In ITU - T H.265, for a picture, the CTU size may be set to include 16×16, 32×32, or 64×64 luma samples. In ITU - T H.265, a CTU consists of corresponding Coding Tree Blocks (CTBs) for each component of the video data (e.g., luma (Y) and chroma (Cb and Cr)). It should be noted that video with one luma component and two corresponding chroma components can be described as having two channels, i.e., the luma channel and the chroma channel. Additionally, in ITU - T H.265, a CTU may be partitioned according to a quadtree (QT) partitioning structure, which causes the CTBs of the CTU to be divided into Coding Blocks (CBs). That is, in ITU - T H.265, a CTU may be divided into quadtree leaf nodes. According to ITU - T H.265, a luma CB together with two corresponding chroma CBs and associated syntax elements is called a Coding Unit (CU). In ITU - T H.265, the minimum allowed size of a CB may be signaled. In ITU - T H.265, the minimum allowed size of a luma CB is 8×8 luma samples. In ITU - T H.265, the decision to encode a picture region using intra - prediction or inter - prediction is made at the CU level.
[0025] In ITU-T H.265, a coding unit (CU) is associated with a prediction unit (PU) structure having its root at the CU. In ITU-T H.265, the PU structure allows splitting of the luma coding block (CB) and chroma CBs to generate corresponding reference samples. That is, in ITU-T H.265, the luma CB and chroma CBs can be split into respective luma prediction blocks and chroma prediction blocks (PBs), where a PB includes a block of sample values to which the same prediction is applied. In ITU-T H.265, a CB can be divided into 1, 2, or 4 PBs. ITU-T H.265 supports PB sizes from 64×64 samples down to 4×4 samples. In ITU-T H.265, square PBs are supported for intra prediction, where a CB can form a PB or a CB can be split into four square PBs. In ITU-T H.265, in addition to square PBs, rectangular PBs are supported for inter prediction, where a CB can be halved vertically or horizontally to form a PB. Further, it should be noted that in ITU-T H.265, for inter prediction, four asymmetric PB partitions are supported, where a CB is divided into two PBs at a quarter of the height (top or bottom) or width (left or right) of the CB. Intra prediction data (e.g., intra prediction mode syntax elements) or inter prediction data (e.g., motion data syntax elements) corresponding to a PB are used to generate reference and / or predicted sample values of the PB.
[0026] JEM specifies a coding tree unit (CTU) of 256×256 luma samples of the maximum size. JEM specifies a quadtree plus binary tree (QTBT) block structure. In JEM, the QTBT structure allows further partitioning of quadtree leaf nodes by a binary tree (BT) structure. That is, in JEM, the binary tree structure allows vertical or horizontal recursive partitioning of quadtree leaf nodes. In JVET-O2001, a CTU is partitioned according to a quadtree plus multi-type tree (QTMT or QT+MTT) structure. QTMT in JVET-O2001 is similar to QTBT in JEM. However, in JVET-O2001, in addition to indicating binary splitting, the multi-type tree can also indicate a so-called ternary (or trinary tree (TT)) splitting. Ternary splitting divides a block vertically or horizontally into three blocks. In the case of vertical TT splitting, the block is split at a quarter of its width from the left edge and at a quarter of its width from the right edge, and in the case of horizontal TT splitting, the block is split at a quarter of its height from the top edge and at a quarter of its height from the bottom edge.
[0027] As described above, each video frame or picture can be divided into one or more regions. For example, according to ITU-T H.265, each video frame or picture can be partitioned into including one or more slices, and further partitioned into including one or more tiles, where each slice includes a sequence of CTUs (e.g., arranged in raster scan order), and where a tile is a sequence of CTUs corresponding to a rectangular region of the picture. It should be noted that in ITU-T H.265, a slice starts from an independent slice segment and includes all subsequent dependent slice segments (if any) before the next independent slice segment (if any), which is a sequence of one or more slice segments. A slice segment (like a slice) is a sequence of CTUs. Thus, in some cases, the terms "slice" and "slice segment" can be used interchangeably to indicate a sequence of CTUs arranged in raster scan order. Additionally, it should be noted that in ITU-T H.265, a tile can be composed of CTUs included in more than one slice, and a slice can be composed of CTUs included in more than one tile. However, ITU-T H.265 stipulates that one or both of the following conditions should be met: (1) all CTUs in a slice belong to the same tile; and (2) all CTUs in a tile belong to the same slice.
[0028] Regarding JVET-O2001, a slice needs to be composed of an integer number of tiles, rather than just an integer number of CTUs. It should be noted that in JVET-O2001, the slice design does not include slice segments (i.e., there are no independent / dependent slice segments). In JVET-O2001, a tile is a rectangular CTU row region within a specific tile in the picture. Additionally, in JVET-O2001, a tile can be divided into multiple tiles, each tile composed of one or more CTU rows within the tile. A tile that is not divided into multiple tiles is also referred to as a tile. However, a tile that is a proper subset of a tile is not called a tile. Thus, in some video coding techniques, slices including a set of CTUs that do not form a rectangular region of the picture may or may not be supported. Additionally, it should be noted that in some cases, a slice may need to be composed of an integer number of complete tiles, and in such cases, the slice is called a tile group. The techniques described herein can be applied to tiles, slices, tiles, and / or tile groups. Figure 2 is a conceptual diagram showing an example of a group of pictures including slices. In Figure 2 the example shown, Pic3 is shown as including two slices (i.e., slice 0 and slice 1). In Figure 2 the example shown, slice 0 includes one tile, i.e., tile 0, and slice 1 includes two tiles, i.e., tile 1 and tile 2. It should be noted that in some cases, slice 0 and slice 1 may meet the requirements of a tile and / or a tile group and be classified as a tile and / or a tile group.
[0029] For intra prediction coding, the intra prediction mode may specify the positions of reference samples within a picture. In ITU-T H.265, the possible intra prediction modes that have been defined include the planar (i.e., surface fitting) prediction mode, the DC (i.e., flat overall average) prediction mode, and 33 angular prediction modes (predMode: 2 - 34). In JEM, the possible intra prediction modes that have been defined include the planar prediction mode, the DC prediction mode, and 65 angular prediction modes. It should be noted that the planar prediction mode and the DC prediction mode may be referred to as non-directional prediction modes, and the angular prediction modes may be referred to as directional prediction modes. It should be noted that regardless of the number of possible prediction modes that have been defined, the techniques described herein may be generally applicable.
[0030] For inter prediction coding, a reference picture is determined, and a motion vector (MV) identifies the samples in that reference picture that are used to generate a prediction of the current video block. For example, reference sample values located in one or more previously encoded pictures may be used to predict the current video block, and the motion vector is used to indicate the position of the reference block relative to the current video block. The motion vector may describe, for example, the horizontal displacement component of the motion vector (i.e., MV x ), the vertical displacement component of the motion vector (i.e., MV y), and the resolution of the motion vector (e.g., quarter-pixel accuracy, half-pixel accuracy, one-pixel accuracy, two-pixel accuracy, four-pixel accuracy). The previously decoded pictures (which may include pictures output before or after the current picture) can be organized into one or more reference picture lists and are identified using reference picture index values. In addition, in inter-prediction coding, single prediction means generating a prediction using sample values from a single reference picture, and dual prediction means generating a prediction using corresponding sample values from two reference pictures. That is, in single prediction, a single reference picture and the corresponding motion vector are used to generate a prediction for the current video block, while in dual prediction, the first reference picture and the corresponding first motion vector and the second reference picture and the corresponding second motion vector are used to generate a prediction for the current video block. In dual prediction, the corresponding sample values are combined (e.g., added, rounded and clipped, or averaged according to weights) to generate a prediction. Pictures and their regions can be classified based on which types of prediction modes can be used to encode their video blocks. That is, for regions of type B (e.g., B slices), dual prediction, single prediction, and intra prediction modes can be utilized, for regions of type P (e.g., P slices), single prediction and intra prediction modes can be utilized, and for regions of type I (e.g., I slices), only the intra prediction mode can be utilized. As described above, reference pictures are identified by reference indices. For example, for P slices, there can be a single reference picture list RefPicList0, and for B slices, in addition to RefPicList0, there can be a second independent reference picture list RefPicList1. It should be noted that for single prediction in B slices, either RefPicList0 or RefPicList1 can be used to generate a prediction. In addition, it should be noted that during the decoding process, at the start of decoding a picture, a reference picture list is generated from the previously decoded pictures stored in the decoded picture buffer (DPB).
[0031] In addition, the coding standard may support various motion vector prediction modes. Motion vector prediction enables the value of a motion vector for a current video block to be derived based on another motion vector. For example, a set of candidate blocks with associated motion information can be derived from spatially adjacent blocks and temporally adjacent blocks of the current video block. In addition, the generated (or default) motion information can be used for motion vector prediction. Examples of motion vector prediction include advanced motion vector prediction (AMVP), temporal motion vector prediction (TMVP), the so-called "merge" mode, and "skip" and "direct" motion inference. In addition, other examples of motion vector prediction include advanced temporal motion vector prediction (ATMVP) and spatio-temporal motion vector prediction (STMVP). For motion vector prediction, both the video encoder and the video decoder perform the same process to derive a set of candidates. Thus, for a current video block, the same set of candidates is generated during encoding and decoding.
[0032] As described above, for inter-prediction coding, reference samples in previously encoded pictures are used to encode video blocks in the current picture. The previously encoded pictures that are available as references when encoding the current picture are called reference pictures. It should be noted that the decoding order does not necessarily correspond to the picture output order, i.e., the temporal order of pictures in the video sequence. In ITU-T H.265, when a picture is decoded, it is stored in a decoded picture buffer (DPB) (which may be referred to as a frame buffer, reference buffer, reference picture buffer, etc.). In ITU-T H.265, the picture stored in the DPB is removed from the DPB when it is output and is no longer needed for encoding subsequent pictures. In ITU-T H.265, a determination of whether a picture should be removed from the DPB is called once for each picture after the slice header is decoded, i.e., at the start of decoding the picture. For example, referring Figure 2 , Pic3 is shown as referring to Pic2. Similarly, Pic4 is shown as referring to Pic1. Regarding Figure 2, assuming that the picture order corresponds to the decoding order, the DPB will be filled as follows: After decoding Pic1, the DPB will include {Pic1}; at the start of decoding Pic2, the DPB will include {Pic1}; after decoding Pic2, the DPB will include {Pic1, Pic2}; at the start of decoding Pic3, the DPB will include {Pic1, Pic2}. Then, Pic3 will be decoded with reference to Pic2, and after decoding Pic3, the DPB will include {Pic1, Pic2, Pic3}. At the start of decoding Pic4, pictures Pic2 and Pic3 will be marked for removal from the DPB because they are not required for decoding Pic4 (or any subsequent pictures, not shown), and assuming that Pic2 and Pic3 have been output, the DPB will be updated to include {Pic1}. Then Pic4 will be decoded with reference to Pic1. The process of marking pictures for removal from the DPB can be referred to as reference picture set (RPS) management.
[0033] As described above, intra prediction data or inter prediction data is used to generate reference sample values for blocks of sample values. The difference between the sample values included in the current PB or another type of picture region structure and the associated reference samples (e.g., those generated using prediction) can be referred to as residual data. The residual data can include corresponding difference arrays for each component of the video data. The residual data may be in the pixel domain. A transform such as a discrete cosine transform (DCT), discrete sine transform (DST), integer transform, wavelet transform, or a conceptually similar transform may be applied to the difference array to generate transform coefficients. It should be noted that in ITU-T H.265 and JVET-O2001, a CU is associated with a transform unit (TU) structure having its root at the CU level. That is, in order to generate transform coefficients, the array of differences can be partitioned (e.g., four 8×8 transforms can be applied to a 16×16 residual value array). Such subdivision of the differences for each component of the video data can be referred to as a transform block (TB). It should be noted that in some cases, a core transform and a subsequent secondary transform can be applied (in the video encoder) to generate transform coefficients. For a video decoder, the order of the transforms is reversed.
[0034] The quantization process can be directly performed on transform coefficients or residual sample values (e.g., in the case of palette coding quantization). Quantization approximates the transform coefficients by restricting the amplitude to a set of specified values. Quantization essentially scales the transform coefficients in order to change the amount of data required to represent a set of transform coefficients. Quantization can include dividing the transform coefficients (or the values obtained by adding an offset value to the transform coefficients) by a quantization scaling factor and any associated rounding function (e.g., rounding to the nearest integer). The quantized transform coefficients can be referred to as coefficient level values. Inverse quantization (or "dequantization") can include multiplying the coefficient level values by the quantization scaling factor, and any inverse rounding or offset addition operations. It should be noted that, as used herein, the term quantization process may, in some cases, refer to dividing by a scaling factor to generate level values, and in some cases may refer to multiplying by a scaling factor to recover the transform coefficients. That is, the quantization process can refer to quantization in some cases and inverse quantization in some cases. In addition, it should be noted that although the quantization process is described in terms of arithmetic operations related to decimal notation in some of the examples below, such a description is for illustrative purposes and should not be construed as limiting. For example, the techniques described herein can be implemented in devices that use binary operations, etc. For example, the multiplication and division operations described herein can be implemented using shift operations, etc.
[0035] Entropy coding can be performed on the quantized transform coefficients and syntax elements (e.g., syntax elements indicating the coding structure of a video block) according to entropy coding techniques. The entropy coding process includes encoding the syntax element values using a lossless data compression algorithm. Examples of entropy coding techniques include context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), probability interval partitioning entropy coding (PIPE), etc. The entropy-coded quantized transform coefficients and the corresponding entropy-coded syntax elements can form a compliant bitstream that can be used to reproduce video data at a video decoder. The entropy coding process, such as CABAC, can include binarizing the syntax elements. Binarization refers to the process of converting the value of a syntax element into a sequence of one or more bits. These bits can be referred to as "bins". Binarization can include one or a combination of the following coding techniques: fixed-length coding, unary coding, truncated unary coding, truncated Rice coding, Golomb coding, k-th order exponential Golomb coding, and Golomb-Rice coding. For example, binarization can include representing the integer value 5 of a syntax element as 00000101 using an 8-bit fixed-length binarization technique, or representing the integer value 5 as 11110 using a unary coding binarization technique. As used herein, each of the terms fixed-length coding, unary coding, truncated unary coding, truncated Rice coding, Golomb coding, k-th order exponential Golomb coding, and Golomb-Rice coding can refer to the general implementation of these techniques and / or more specific implementations of these coding techniques. For example, the Golomb-Rice coding implementation can be specifically defined according to a video coding standard. In an example of CABAC, for a particular bin, the context provides the maximum probability state (MPS) value of the bin (i.e., the MPS of the bin is one of 0 or 1), and the probability value that the bin is the MPS or the minimum probability state (LPS). For example, the context can indicate that the MPS of the bin is 0, and the probability that the bin is 1 is 0.3. It should be noted that the context can be determined based on the previously encoded bin values of the bins included in the current syntax element and the previously encoded syntax elements. For example, the values of the syntax elements associated with adjacent video blocks can be used to determine the context of the current bin.
[0036] Regarding the formulas used herein, the following arithmetic operators can be used:
[0037] In addition, the following mathematical functions can be used:
[0038] Log2(x), the base-2 logarithm of x;
[0039]
[0040] Ceil(x) the smallest integer greater than or equal to x.
[0041] Regarding the exemplary grammar used in this document, the following definitions of logical operators can be applied:
[0042] x&&y The Boolean logical "AND" of x and y
[0043] x||y The Boolean logical "OR" of x and y
[0044] ! Boolean logical "NOT"
[0045] x? y : z Evaluates to y if x is TRUE or not equal to 0; otherwise, evaluates to z.
[0046] In addition, the following relational operators can be applied:
[0047] In addition, it should be noted that in the grammar descriptors used in this document, the following descriptors can be applied:
[0048] - b(8): A byte (8 bits) with any bit string pattern. The parsing process of this descriptor is specified by the return value of the function read_bit(8).
[0049] - f(n): A fixed pattern bit string written using n bits (from left to right) starting from the leftmost bit. The parsing process of this descriptor is specified by the return value of the function read_bit(n).
[0050] - se(v): A syntax element of a signed integer 0th order Exp-Golomb code, starting from the leftmost bit.
[0051] - tb(v): A truncated binary code using at most maxVal bits, where maxVal is defined in the semantics of the syntax element.
[0052] - tu(v): A truncated unary code using at most maxVal bits, where maxVal is defined in the semantics of the syntax element.
[0053] - u(n): An unsigned integer using n bits. When n is "v" in the syntax table, the number of bits varies in a way that depends on the value of other syntax elements. The parsing process of this descriptor is specified by the return value of the function read_bits(n), which is interpreted as the binary representation of an unsigned integer written with the most significant bit first.
[0054] - ue(v): A syntax element of an unsigned integer 0th order Exp-Golomb code, starting from the leftmost bit.
[0055] As described above, video content includes a video sequence consisting of a series of frames (or pictures), and each video frame or picture can be divided into one or more regions. An encoded video sequence (CVS) can be encapsulated (or structured) as a series of access units, where each access unit includes video data configured as a network abstraction layer (NAL) unit. It should be noted that in some cases, an access unit may be required to exactly contain one encoded picture. A bitstream can be described as including a sequence of NAL units that form one or more CVSs. It should be noted that multi-layer extensions enable video presentation to include a base layer and one or more additional enhancement layers. For example, the base layer can enable video presentation with a basic quality level (e.g., high-definition presentation and / or 30Hz frame rate), and the enhancement layer can enable video presentation with an enhanced quality level (e.g., ultra-high-definition rendering and / or 60Hz frame rate). The enhancement layer can be encoded by referring to the base layer. That is, for example, pictures in the enhancement layer can be encoded (e.g., using inter-layer prediction techniques) by referring to one or more pictures (including scaled versions thereof) in the base layer. Each NAL unit can include an identifier indicating the video data layer with which the NAL unit is associated. It should be noted that sub-bitstream extraction can refer to the process by which a device receiving a compliant or conforming bitstream forms a new compliant or conforming bitstream by discarding and / or modifying data in the received bitstream. For example, sub-bitstream extraction can be used to form a new compliant or conforming bitstream corresponding to a specific video representation (e.g., high-quality representation). Layers can also be encoded independently of each other. In this case, there may be no inter-layer prediction between the two layers.
[0056] Reference Figure 2In the example shown in Pic3, each video data slice included in Pic3 (i.e., slice 0 slice 1) is shown to be encapsulated in a NAL unit. In JVET-O2001, each of a video sequence, a GOP, a picture, a slice, and a CTU can be associated with metadata that describes video coding attributes. JVET-O2001 defines parameter sets that can be used to describe video data and / or video coding attributes. Specifically, JVET-O2001 includes the following five types of parameter sets: a decoding parameter set (DPS), a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), and an adaptive parameter set (APS). In JVET-O2001, a parameter set can be encapsulated as a special type of NAL unit or can be signaled as a message. A NAL unit that includes encoded video data (e.g., a slice) can be referred to as a VCL (Video Coding Layer) NAL unit, and a NAL unit that includes metadata (e.g., a parameter set) can be referred to as a non-VCL NAL unit. In addition, JVET-O2001 enables supplementary enhancement information (SEI) messages to be signaled. In JVET-O2001, an SEI message assists processes related to decoding, display, or other purposes. However, an SEI message may not be required to construct luminance or chrominance samples through a decoding process. In JVET-O2001, an SEI message can be signaled in a bitstream using a non-VCL NAL unit. In addition, an SEI message can be conveyed in some way other than by being present in a bitstream (i.e., signaled out-of-band).
[0057] Figure 3 An example of a bitstream including a plurality of CVSs is shown, where the CVSs are represented by NAL units included in corresponding access units. In Figure 3 the example shown, the non-VCL NAL units include corresponding parameter set NAL units (i.e., sequence parameter set (SPS) and picture parameter set (PPS) NAL units), SEI message NAL units, and access unit delimiter NAL units. It should be noted that in Figure 3 HEADER is the NAL unit header. JVET-O2001 defines NAL unit header semantics that specify the type of the raw byte sequence payload (RBSP) data structure included in the NAL unit. Table 1 shows the syntax of the NAL unit header provided in JVET-O2001.
[0058]
[0059] JVET-02001 provides the following definitions for the corresponding syntax elements shown in Table 1.
[0060] The forbidden_zero_bit shall be equal to 0.
[0061] The nuh_reserved_zero_bit shall be equal to "0". The value of nuh_reserved_zero_bit may be specified by ITU-TIISO / IEC in the future as 1. The decoder shall ignore (i.e., remove and discard from the bitstream) NAL units for which nuh_reserved_zero_bit is equal to "1".
[0062] The nuh_layer_id specifies the identifier of the layer to which a VCL NAL unit belongs or the identifier of the layer applicable to a non-VCL NAL unit.
[0063] For all VCL NAL units of an encoded picture, the value of nah_layer_id shall be the same. The value of nah_layer_id of an encoded picture or layer access unit is the value of nah_layer_id of the VCL NAL units of the encoded picture or layer access unit.
[0064] nuh_temporal_id_plus1 minus 1 specifies the temporal identifier of the NAL unit. The value of nuh_temporalid_plus1 shall not be equal to 0.
[0065] The variable Temporalid is derived as follows:
[0066] TemporalId = nuh_temporal_id_plus1 - 1
[0067] When nal_unit_type is in the range from IDR_W_RADL to RSV IRAP VCL13 (including the end values), TemporalId shall be equal to 0.
[0068] When nal_unit_type is equal to STSA_NUT, TemporalId shall be equal to 0.
[0069] The values of TemporalId for all VCL NAL units of a layer access unit shall be the same. The value of TemporalId of an encoded picture or layer access unit is the value of Temporalid of the VCL NAL units of the encoded picture or layer access unit. The value of TemporalId of a sublayer representation is the maximum value of TemporalId of all VCL NAL units in the sublayer representation.
[0070] The values of TemporalId for non-VCL NAL units are constrained as follows:
[0071] - If nal_unit_type is equal to DPS_NUT, VPS_NUT or SPS_NUT, then Temporalld shall be equal to 0, and the TemporalId of the layer access unit containing the NAL unit shall be equal to 0.
[0072] - Otherwise, when nal_unit_type is not equal to EOS_NUT and not equal to EOB_NUT, TemporalId shall be greater than or equal to the TemporalId of the layer access unit containing the NAL unit.
[0073] Note - When the NAL unit is a non-VCL NAL unit, the value of TemporalId is equal to the minimum of the TemporalId values of all layer access units to which the non-VCL NAL unit applies. When nal_unit_type is equal to PPS_NUT or APS_NUT, TemporalId may be greater than or equal to the TemporalId of the containing layer access unit, because all PPSs and APSs may be included at the start of the bitstream where the first coded picture has a TemporalId equal to 0. When nal_unit_type is equal to PREFIX_SEI_NUT or SUFFIX_SEI NUT, TemporalId may be greater than or equal to the TemporalId of the containing layer access unit, because the SEI NAL unit may contain information applicable to a subset of the bitstream that includes layer access units whose TemporalId values are greater than the TemporalId of the layer access unit containing the SEI NAL unit.
[0074] nal_unit_type specifies the NAL unit type, i.e., the type of the RBSP data structure contained in the NAL unit as specified in Table 2.
[0075] NAL units with nal_unit_type in the range of UNSPEC28…UNSPEC31 (including the end values) whose semantics are not specified shall not affect the decoding process specified in this specification.
[0076] Note - NAL unit types in the range UNSPEC28…UNSPEC31 may be used as determined by the application. The decoding process for these values of nal_unit_type is not specified in this specification. Since different applications may use these NAL unit types for different purposes, special care must be taken when designing an encoder that generates NAL units with these nal_unit_type values and when designing a decoder that interprets the content of NAL units with these nal_unit_type values. This specification does not define any management of these values. These nal_unit_type values may only apply in contexts where "conflicts" (i.e., different definitions of the meaning of the content of NAL units with the same nal_unit_type value) are unimportant, or impossible, or are managed, e.g., defined or managed in a control application or transport specification, or managed by the environment that distributes the bitstream.
[0077] For purposes other than determining the amount of data in the decoding units of the bitstream, the decoder shall ignore (remove and discard from the bitstream) the content of all NAL units that use reserved values of nal_unit_type.
[0078] Note - This requirement allows for future definition of compatible extensions to this specification.
[0079]
[0080]
[0081] Note - A clean random access (CRA) picture may have associated RASL or RADL pictures present in the bitstream.
[0082] Note - An instantaneous decoding refresh (IDR) picture with a nal_unit_type equal to IDR_N_LP does not have associated leading pictures present in the bitstream. An IDR picture with a nal_unit_type equal to IDR_W_RADL does not have associated RASL pictures present in the bitstream, but may have associated RADL pictures in the bitstream.
[0083] For all coded NAL units of a picture, the value of nal_unit_type shall be the same. A picture or layer access unit is said to have the same NAL unit type as the coded slice NAL units of the picture or layer access unit.
[0084] For single-layer bitstreams, the following constraints apply:
[0085] - Every picture other than the first picture in decoding order in the bitstream is considered to be associated with the previous IRAP picture in decoding order.
[0086] - When a picture is a leading picture of an IRAP picture, the picture shall be a RADL or RASL picture.
[0087] - When a picture is a trailing picture of an IRAP picture, the picture shall not be a RADL or RASL picture.
[0088] - RASL pictures shall not be present in the bitstream that are associated with IDR pictures.
[0089] - RADL pictures shall not be present in the bitstream that are associated with IDR pictures having a nal_unit_type equal to IDR_N_LP.
[0090] Note - Random access at the location of an IRAP access unit can be performed by discarding all access units prior to the IRAP access unit (and correctly decoding the IRAP picture and all subsequent non-RASL pictures in decoding order), provided that each parameter set (in the bitstream or externally) is available when referenced.
[0091] - Any picture before an IRAP picture in decoding order shall be before the TRAP picture in output order and shall be before any RADL picture associated with the IRAP picture in output order.
[0092] - Any RASL picture associated with a CRA picture shall be before any RADL picture associated with the CRA picture in output order.
[0093] - Any RASL picture associated with a CRA picture shall be after any IRAP picture before the CRA picture in decoding order in output order.
[0094] - If field_seq_flag is equal to 0 and the current picture is a leading picture associated with an IRAP picture, the current picture shall be before all non-leading pictures associated with the same TRAP picture in decoding order. Otherwise, let picA and picB be the first leading picture and the last leading picture associated with the TRAP picture in decoding order respectively. There shall be at most one non-leading picture before picA in decoding order and no non-leading pictures between picA and picB in decoding order.
[0095] It should be noted that, generally speaking, for example, with respect to ITU-T H.265, an IRAP picture is a picture that does not reference any pictures other than itself during its decoding process for prediction. Usually, the first picture in the decoding order in the bitstream must be an IRAP picture. In ITU-T H.265, an IRAP picture can be a broken-link access (BLA) picture, a clean random access (CRA) picture, or an instantaneous decoding refresh (IDR) picture. ITU-T H.265 describes the concept of a leading picture, which is a picture that is before the associated IRAP picture in the output order. ITU-T H.265 also describes the concept of a trailing picture, which is a non-IRAP picture that is after the associated IRAP picture in the output order. The trailing pictures associated with an IRAP picture are also after the IRAP picture in the decoding order. For an IDR picture, there are no trailing pictures that need to reference pictures decoded before the IDR picture. ITU-T H.265 stipulates that a CRA picture can have a leading picture that is after the CRA picture in the decoding order and contains inter-picture prediction that references pictures decoded before the CRA picture. Therefore, when a CRA picture is used as a random access point, these leading pictures may not be decodable and are identified as random access skipped leading (RASL) pictures. A BLA picture can also be after an RASL picture. When these RASL pictures are not decodable, that is, when the decoder starts its decoding process at a CRA point, for both BLA pictures and CRA pictures, these RASL pictures are always discarded. Another type of picture that can be after an IRAP picture in the decoding order and before the IRAP picture in the output order is a random access decodable leading (RADL) picture, which cannot contain a reference to any picture that is before the IRAP picture in the decoding order.
[0096] As provided in Table 2, a NAL unit can include a sequence parameter set (SPS). Table 3 shows the syntax of the SPS provided in JVET-O2001.
[0097]
[0098]
[0099]
[0100]
[0101]
[0102]
[0103] Regarding Table 3, JVET-02001 provides the following semantics:
[0104] The SPS RBSP shall be available for the decoding process before it is referenced, including being provided in at least one access unit with a TemporalId equal to 0 or by an external means, and the SPS NAL unit containing the SPS RBSP shall have a nuh_layer_id equal to that of the PPS NAL unit that references it.
[0105] All SPS NAL units with a specific value of sps_seq_parameter_set_id in the CVS shall have the same content.
[0106] When sps_decoding_parameter_set_id is greater than 0, it specifies the value of dps_decoding_parameter_set_id of the DPS referenced by the SPS. When sps_decoding_parameter_set_id is equal to 0, the SPS does not reference the DPS, and when decoding each CVS that references the SPS, the DPS is not referenced. In all SPSs referenced by coded pictures in the bitstream, the value of sps_decoding_parameter_set_id shall be the same.
[0107] When sps_video_parameter_set_id is greater than 0, it specifies the value of vps_video_parameter_set_id of the VPS referenced by the SPS. When sps_video_parameter_set_id is equal to 0, the SPS does not reference the VPS, and when decoding each CVS that references the SPS, the VPS is not referenced.
[0108] sps_max_sub_layers_minus1 plus 1 specifies the maximum number of temporal sub-layers that may exist in each CVS that references the SPS. The value of sps_max_sub_layers_minus1 shall be in the range of 0 to 6 (including the end values).
[0109] In the bitstream conforming to this version of this specification, sps_reserved_zero_5bits shall be equal to 0. Other values of sps_reserved_zero_5bits are reserved for future use by ITU-T | ISO / IEC.
[0110] The gdr_enabled_flag being equal to 1 specifies that GDR pictures may exist in the CVS that references the SPS. The gdr_enabled_flag being equal to 0 specifies that GDR pictures do not exist in the CVS that references the SPS.
[0111] The sps_seq_parameter_set_id provides the identifier of the SPS for other syntax elements to reference. The value of sps_seq_parameter_set_id shall be in the range of 0 to 15, inclusive.
[0112] As specified, the chroma_format_idc specifies the chrominance sampling relative to the luma sampling. The value of chroma_format_idc shall be in the range of 0 to 3, inclusive.
[0113] Separate_colour_plane_flag being equal to 1 specifies that the three color components of the 4:4:4 chrominance format are encoded separately. Separate_colour_plane_flag being equal to 0 specifies that the color components are not encoded separately. When Separate_colour_plane_flag is absent, it is inferred to be equal to 0. When Separate_colour_plane_flag is equal to 1, the coded picture consists of three separate components, each consisting of the coded samples of a color plane (Y, Cb, or Cr) and using the monochrome coding syntax. In this case, each color plane is associated with a specific colour_plane_id value.
[0114] Note - There is no correlation in the decoding process between color planes with different colour_plane_id values. For example, the decoding process of a monochrome picture with one value of colour_plane_id does not use any data from a monochrome picture with a different colour_plane_id value for inter prediction.
[0115] According to the value of Separate_colour_plane_flag, the value of the variable ChromaArrayType is specified as follows:
[0116] - If Separate_colour_plane_flag is equal to 0, then ChromaArrayType is set to be equal to chroma_format_idc.
[0117] - Otherwise (Separate_colour_plane_flag is equal to 1), ChromaArrayType is set to be equal to 0.
[0118] pic_width_x_in_luma_samples specifies the maximum width of each coded picture that refers to the SPS, in luma samples. pic_width_max_in_luma_samples shall be equal to 0 and shall be an integer multiple of MinCbSizeY.
[0119] pic_height_max_in_luma_samples specifies the maximum height of each coded picture that refers to the SPS, in luma samples. pic_height_max_in_luma_samples shall be equal to 0 and shall be an integer multiple of MinCbSizeY.
[0120] subpics_present_flag being equal to 1 indicates that sub-picture parameters are present in the SPS RBSP syntax. subpics_present_flag being equal to 0 indicates that sub-picture parameters are not present in the SPS RBSP syntax.
[0121] Note - When the bitstream is the result of a sub-bitstream extraction process and contains only a subset of the sub-pictures of the input bitstream to the sub-bitstream extraction process, it may be necessary to set the value of subpics_present_flag to 1 in the RBSP of the SPS.
[0122] max_subpics_minus1 plus 1 specifies the maximum number of sub-pictures that may be present in the CVS. max_subpics_minus1 shall be in the range of 0 to 254. The value 255 is reserved for future use by ITU-T | ISO / IEC.
[0123] subpic_grid_col_width_minus1 plus 1 specifies the width of each element of the sub-picture identifier grid, in units of 4 samples. The length of the syntax element is Ceil(Log2(pic_width_max_in_luma_samples / 4)) bits.
[0124] The variable NumSubPicGridCols is derived as follows:
[0125] NumSubPicGridCols =
[0126] (pic_width_max_in_luma_samples + subpic_grid_colwidth_mi nus 1 * 4 +3) / (subpic_grid_col_width_minus1 * 4 + 4)
[0127] subpic_grid_row_height_minus1 plus 1 specifies the height of each element of the sub-picture identifier grid, in units of 4 samples. The length of the syntax element is Ceil(Log2(pic_height_max_in_luma_samples / 4)) bits.
[0128] The variable NumSubPicGridRows is derived as follows:
[0129] NumSubPicGridRows =
[0130] (pic_height maxin_luma_samples + subpic_grid_row_height_minus1 * 4 + 3) / (subpic_grid_row_height_minus1 * 4 + 4)
[0131] subpic_grid_idx[i][j] specifies the sub-picture index at the grid position (i, j). The length of the syntax element is Ceil(Log2(max subpics minus1 + 1)) bits.
[0132] The variables SubPicTop[subpic_grid_idx[i][j]], SubPicLeft[subpic_grid_idx[i][j]], SubPicWidth[subpic_grid_idx[i][j]], SubPicHeight[subpic_grid_idx[i][j]] and NumSubPics are derived as follows:
[0133]
[0134] subpic_treated_as_pic_flag[i] being equal to 1 specifies that the i-th sub-picture of each coded picture in the CVS is treated as a picture in the decoding process excluding in-loop filtering operations. subpic_treated_as_pic_flag[i] being equal to 0 specifies that the i-th sub-picture of each coded picture in the CVS is not treated as a picture in the decoding process excluding in-loop filtering operations. When not present, the value of subpic_treated_as_pic_flag[i] is inferred to be 0.
[0135] loop_filter_across_subpic_enabled_flag[ i ] being equal to 1 specifies that loop filtering operations can be performed across the boundaries of the i-th sub-picture of each coded picture in the CVS. loop_filter_across_subpic_enabled_flag[ i ] being equal to 0 specifies that loop filtering operations are not performed across the boundaries of the i-th sub-picture of each coded picture in the CVS. When not present, it is inferred that the value of loop_filter_across_subpic_enabled_pic_flag[ i ] is equal to 1.
[0136] Bitstream conformance requires the following constraints to apply:
[0137] - For any two sub-pictures subpicA and subpicB, when the index of subpicA is less than the index of subpicB, any coded NAL unit of subPicA shall follow any coded NAL unit of subPicB in the decoding order.
[0138] - The shape of the sub-picture shall be such that each sub-picture, when decoded, shall have its entire left boundary and entire top boundary consisting of the picture boundary or the boundary of the previously decoded sub-picture.
[0139] bit_depth_luma_minus8 specifies the bit depth of the samples of the luma array BitDepthy and the range offset QpBdOffset of the luma quantization parameter as follows Y value:
[0140] BitDepthy = 8 + bit_depth_luma_minus8
[0141] QpBdOffsety = 6 * bit_depth_luma_minus8
[0142] bit_depth_luma_minus8 shall be in the range of 0 to 8 (including the end values).
[0143] bit_depth_chroma_minus8 specifies the bit depth of the samples of the chroma array BitDepthc and the range offset QpBdOffset of the chroma quantization parameter as follows C value:
[0144] BitDepth C = 8 + bit_dcpth_chroma_minus8
[0145] QpBdOffsetC = 6 * bit_depth_chroma_minus8
[0146] bit_depth_chroma_minus8 shall be in the range of 0 to 8, inclusive of the end values.
[0147] min_qp_prime_ts_minus4 specifies the minimum allowable quantization parameter for the transform skip mode as follows:
[0148] QpPrimeTsMin = 4 + min_qp_prime_ts_minus4
[0149] log2_max_pic_order_cnt_lsb_minus4 specifies the value of the variable MaxPicOrderCntLsb used during the decoding of the picture order count as follows:
[0150] MaxPicOrderCntLsb = 2 (log2_max_pic_order_cnt_lsb_minus4 + 4)
[0151] The value of log2_max_pic_order_cnt_lsb_minus4 shall be in the range of 0 to 12, inclusive of the end values.
[0152] sps_sub_layer_ordering_info_present_flag being equal to 1 specifies that sps_max_dec_pic_buffering_minus1[ i ], sps_max_num_reorder_pics[ i ], and sps_max_latency_increase_plus1[ i ] exist for sps_max_sub_layers_minus1 + 1 sub - layers. sps_sub_layer_ordering_info_present_flag being equal to 0 specifies that the values of sps_max_dec_pic_buffering_minus1 [ sps_max_sub_layers_minus1], sps_max_num_reorder_pics[sps_max_sub_layers_minus 1 ], and sps_max_latency_increase_plus1 [ sps_max_sub_layers_minus1] apply to all sub - layers. When not present, sps_sub_layer_ordering_info_present_flag is inferred to be 0.
[0153] sps_max_dec_pic_buffering_minus1[i] plus 1 specifies the maximum required size of the decoded picture buffer for CVS when HighestTid is equal to i, in units of picture storage buffers. The value of sps_max_dec_pic_buffering_minus1[i] shall be in the range of 0 to MaxDpbSize - 1 (including the end values), where MaxDpbSize is specified elsewhere. When i is greater than 0, sps_max_dec_pic_buffering_minus1[1] shall be greater than or equal to sps_max_dec_pic_buffering_minus1[i - 1]. When there is no sps_max_dec_pic_buffering_minus1 i] for i in the range of 0 to sps_max_sub_layers_minus1 - 1 (including the end values), since sps_sub_layer_ordering_info_present_flag is equal to 0, it is inferred to be equal to sps_max_dec_pic_buffering_minus1[sps_max_sub_layers_minus1].
[0154] sps_max_num_reorder_pics[ i ] indicates the maximum allowed number of pictures that can be before any picture in CVS in the decoding order and after that picture in the output order when HighestTid is equal to i. The value of sps_max_num_reorder_pics[i] shall be in the range of 0 to sps_max_dec_pic_buffering_minus1[ i ] (including the end values). When i is greater than 0, sps_max_num_reorder_pics[i] shall be greater than or equal to sps_max_num_reorder_pics[ i - 1 ]. When there is no sps_max_num_reorder_pics[i] for i in the range of 0 to sps_max sub_layers_minus1 - 1 (including the end values), since sps_sub_layer_ordering_info_present_flag is equal to 0, it is inferred to be equal to sps_max_num_reorder_pics[ sps_max_sub_layers_minus1].
[0155] sps_max_latency_increase_plus1[i] that is not equal to 0 is used to calculate the value of SpsMaxLatencyPictures[ i ], which specifies the maximum number of pictures that may be in the output order before any picture in the CVS and after that picture in the decoding order when HighestTid is equal to i.
[0156] When sps_max_latency_increase_plus1[ i ] is not equal to 0, the value of SpsMaxLatencyPictures[ i ] is specified as follows:
[0157] SpsMaxLatencyPictures[ i =
[0158] sps_max_num_reorder_pics[ i ]+ sps_max_latency_increase_plus1[i]-1
[0159] When sps_max_latency_increase_plusi1[ i ] is equal to 0, it does not represent the corresponding restriction.
[0160] The value of sps_max_latency_increase_plus1[ i] shall be in the range of 0 to 2 32 -2 (including the end values). When there is no sps_max_latency_increase_plus1[ i ] for i in the range of 0 to sps_max_sub_layers_minus1 - 1 (including the end values), since sps_sub_layer_ordering_info_present_flag is equal to 0, it is inferred to be equal to sps_max_latency_increase_plus 1 [ sps_max_sub_layers_minus1].
[0161] long_term_ref_pics_flag being equal to 0 specifies that no LTRP is used for inter prediction of any coded picture in the CVS. long_term_ref_pics_fiag being equal to 1 specifies that LTRP can be used for inter prediction of one or more coded pictures in the CVS.
[0162] The inter_layer_ref_pics_present_flag being equal to 0 specifies that no LTRP is used for inter prediction of any coded pictures in the CVS. The inter_layerref pics_flag being equal to 1 specifies that the LTRP can be used for inter prediction of one or more coded pictures in the CVS. When sps_video_parameter_set_id is equal to 0, it is inferred that the value of the inter_layer_ref_pics_present_flag is equal to 0.
[0163] The sps_idr_rp1_present_flag being equal to 1 specifies that the reference picture list syntax element is present in the slice header of an IDR picture. The sps_idr_rp1_present_flag being equal to 0 specifies that the reference picture list syntax element is not present in the slice header of an IDR picture.
[0164] The rpl1_same_as_rpl0_flag being equal to 1 specifies that the syntax structures num_ref pic_lists_in_sps[1] and ref pic_list_struct(1, rplsIdx) do not exist, and the following applies:
[0165] - It is inferred that the value of num_ref_pic_lists_in sps[1] is equal to the value of num_ref_pic_lists_in_sps[0].
[0166] - It is inferred that the value of each syntax element in the syntax elements of ref_pic_list_struct(1, rplsIdx) is equal to the value of the corresponding syntax element in ref_pic_list_stnict(0, rplsIdx) for rplsIdx in the range of 0 to num_ref_pic_lists_in_sps[0] - 1.
[0167] num_ref_pic_lists_in_sps[i] specifies the number of ref_pic_list_struct(listIdx,rplsIdx) syntax structures included in the SPS, where listIdx is equal to i. The value of num_ref_pic_lists_in_sps[i] shall be in the range of 0 to 64 (inclusive of the end values).
[0168] Note - For each value of listIdx (equal to 0 or 1), the decoder shall allocate memory for the ref_pic_list_struct(listIdx, rplsIdx) syntax structure with a total of num_ref_pic_lists_in_sps[ i ] + 1, since there may be one ref_pic_list_struct(listIdx, rplsIdx) syntax structure signaled directly in the slice header of the current picture.
[0169] When qtbtt_dual_tree_intra_flag is equal to 1, it specifies that for I slices, each CTU is partitioned into coding units with 64×64 luma samples using implicit quadtree partitioning, and these coding units are the roots of two separate coding_tree syntax structures for luma and chroma. When qtbtt_dual_tree_intra_flag is equal to 0, it specifies that the separate coding_tree syntax structures are not used for I slices. When qtbtt_dual_tree_intra_flag is absent, it is inferred to be equal to 0.
[0170] log2_ctu_size_minus5 plus 5 specifies the luma coding tree block size of each CTU. Bitstream conformance requires that the value of log2_ctu_size_minus5 be less than or equal to 2.
[0171] log2_min_luma_coding_block_size_minus2 plus 2 specifies the minimum luma coding block size.
[0172] The variables CtbLog2SizeY, CtbSizeY, MinCbLog2SizeY, MinCbSizeY, IbcBufWidthY, IbcBufWidthC, and Vsize are derived as follows:
[0173] CtbLog2SizeY = log2_ctu_size_minus5 + 5
[0174] CtbSizeY = 1 << CtbLog2SizeY
[0175] MinCbLog2SizeY = log2_min_luma_coding_block_sizeminus2 + 2
[0176] MinCbSizeY = 1 << MinCbLog2SizeY
[0177] IbcBufWidthY = 128 * 128 / CtbSizeY
[0178] IbcBufWidthC = IbcBufWidthY / SubWidthC
[0179] VSize = Min(64, CtbSizeY)
[0180] The variables CtbWidthC and CtbHeightC respectively specify the width and height of the array of each chroma CTB, and these two variables are derived as follows:
[0181] - If chroma_format_idc is equal to 0 (monochrome) or separate_colour_plane_flag is equal to 1, then both CtbWidthC and CtbHeightC are equal to 0.
[0182] - Otherwise, CtbWidthC and CtbHeightC are derived as follows:
[0183] CtbWidthC = CtbSizeY / SubWidthC
[0184] CtbHeightC = CtbSizeY / SubHeightC
[0185] For log2BlockWidth in the range from 0 to 4 and log2BlockHeight in the range from 0 to 4 (including the end values), the specified upper-right diagonal and raster scan order array initialization processes are called with 1 << log2BlockWidth and 1 << log2BlockHeight as inputs, and the outputs are assigned to DiagScanOrder[log2BlockWidth][log2BlockHeight] and Raster2DiagScanPos[log2BlockWidth][log2BlockHeight].
[0186] For log2BlockWidth in the range of 0 to 6 and log2BlockHeight in the range of 0 to 6 (including the end values), call the specified horizontal and vertical traversal scan order array initialization process with 1 << log2BlockWidth and 1 << log2BlockHeight as inputs, and assign the outputs to HorTravScanOrder[log2BlockWidth ][ log2BlockHeight ] and VerTravScanOrder[ log2BlockWidth ][log2BlockHeight ].
[0187] The partition_constraints_override_enabled_flag being equal to 1 specifies the presence of the partition_constraints_override_flag in the slice header of a slice that references the SPS. The partition_constraints_override_enabled_flag being equal to 0 specifies the absence of the partition_constraints_override_flag in the slice header of a slice that references the SPS.
[0188] sps_log2_diff_min_qt_min_cb_intra_slice_luma specifies the default difference between the base-2 logarithm of the minimum size in the luma samples of the luma leaf blocks resulting from quadtree partitioning of a CTU and the base-2 logarithm of the minimum coded block size in the luma samples of the luma CUs in a slice with slice_type equal to 2 (I) that references the SPS. When the partition_constraints_override_flag is equal to 1, this default difference can be overridden by slice_log2_diff_min_qt_min_cb_luma present in the slice header of a slice that references the SPS. The value of sps_log2_diff_min_qt_min_cb_intra_slice_luma shall be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY (including the end values). The base-2 logarithm of the minimum size in the luma samples of the luma leaf blocks resulting from quadtree partitioning of a CTU is derived as follows:
[0189] MinQtLog2SizeIntraY = sps_log2_diff_min_qt_min_cb_intra_slice_luma + MinCbLog2SizeY
[0190] sps_log2_diff_min_qt_min_cb_inter_slice specifies the default difference between the log2 of the minimum size in the luma samples of the luma leaf blocks resulting from the quadtree partitioning of a CTU and the log2 of the minimum luma coding block size in the luma samples of the luma CUs in a slice that references an SPS with slice_type equal to 0 (B) or 1 (P). When partition_constraints_override_flag is equal to 1, this default difference can be overridden by slice_log2_diff_min_qt_min_cb_luma present in the slice header of the slice that references the SPS. The value of sps_log2_diff_min_qt_min_cb_inter_slice shall be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive. The log2 of the minimum size in the luma samples of the luma leaf blocks resulting from the quadtree partitioning of a CTU is derived as follows:
[0191] MinQtLog2SizelnterY = sps_log2_diff_min_qt_min_cb_interslice +MinCbLog2SizeY
[0192] sps_max_mtt_hierarchy_depth_inter_slice specifies the default maximum hierarchical structure depth of the coding units resulting from the multi-type tree partitioning of the quadtree leaves in a slice that references an SPS with slice_type equal to 0 (B) or 1 (P). When partition_constraints_override_flag is equal to 1, the default maximum hierarchical structure depth can be overridden by slice_max_mtt_hierarchy_depth_luma present in the slice header of the slice that references the SPS. The value of sps_max_mtt_hierarchy_depth_inter_slice shall be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive.
[0193] sps_max_mtt_hierarchy_depth_intra_slice_luma specifies the default maximum hierarchical structure depth of coding units generated by multi-type tree splitting of quadtree leaves in slices with slice_type equal to 2 (I) that reference the SPS. When partition_constraints_override_flag is equal to 1, the default maximum hierarchical structure depth can be overridden by slice_max_mtt_hierarchy_depth_luma in the slice header of the slice that references the SPS. The value of sps_max_mtt_hierarchy_depth_intra_slice_luma shall be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY (including the end values).
[0194] sps_log2_diff_max_bt_min_qt_intra_slice_luma specifies the default difference between the base-2 logarithm of the maximum size (width or height) of the luminance samples in a luminance coding block that can be split using binary splitting and the base-2 logarithm of the minimum size (width or height) of the luminance samples in a luminance leaf block generated by quadtree splitting of a CTU in a slice with slice_type equal to 2 (I) that references the SPS. When partition_constraints_override_flag is equal to 1, this default difference can be overridden by slice_log2_diff_max_bt_min_qt_luma present in the slice header of the slice that references the SPS. The value of sps_log2_diff_max_bt_min_qt_intra_slice_luma shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeIntraY (including the end values). When sps_log2_diff_max_bt_min_qt_intra_slice_luma does not exist, it is inferred that the value of sps_log2_diff_max_bt_min_qt_intra_slice_luma is equal to 0.
[0195] sps_log2_diff_max_tt_min_qt_intra_slice_luma specifies the default difference between the base-2 logarithm of the maximum size (width or height) in the luma samples of a luma coded block that can use ternary partitioning and the base-2 logarithm of the minimum size (width or height) in the luma samples of a luma leaf block resulting from quadtree partitioning of a CTU in a slice with slice_type equal to 2 (I) in the referenced SPS. When partition_constraints_override_flag is equal to 1, this default difference can be overridden by slice_log2_diff_max_tt_min_qt_luma present in the slice header of the slice in the referenced SPS. The value of sps_log2_diff_max_tt_min_qt_intra_slice_luma shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeIntraY, inclusive. When sps_log2_diff_max_tt_min_qt_intra_slice_luma is not present, it is inferred that the value of sps_log2_diff_max_tt_min_qt_intra_slice_luma is equal to 0.
[0196] sps_log2_diff_max_bt_min_qt_inter_slice specifies the default difference between the base-2 logarithm of the maximum size (width or height) in the luma samples of a luma coded block that can use binary partitioning and the base-2 logarithm of the minimum size (width or height) in the luma samples of a luma leaf block resulting from quadtree partitioning of a CTU in a slice with slice_type equal to 0 (B) or 1 (P) in the referenced SPS. When partition_constraints_override_flag is equal to 1, this default difference can be overridden by slice_log2_diff_max_bt_min_qt_luma present in the slice header of the slice in the referenced SPS. The value of sps_log2_diff_max_bt_min_qt_inter_slice shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizelnterY, inclusive. When sps_log2_diff_max_bt_min_qt_inter_slice is not present, it is inferred that the value of sps_log2_diff_max_bt_min_qt_inter_slice is equal to 0.
[0197] sps_log2_diff_max_tt_min_qt_inter_slice specifies the default difference between the log2 of the maximum size (width or height) of the luma samples in a luma coded block that can use ternary splitting and the log2 of the minimum size (width or height) of the luma samples in a luma leaf block resulting from quadtree splitting of a CTU in a slice with slice_type equal to 0 (B) or 1 (P) that references the SPS. When partition_constraints_override_flag is equal to 1, this default difference can be overridden by slice_log2_diff_max_tt_min_qt_luma present in the slice header of the slice that references the SPS. The value of sps_log2_diff_inax_tt_min_qt_inter_slice shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeInterY, inclusive. When sps_log2_diff_max_tt_min_qt_inter_slice is not present, it is inferred that the value of sps_log2_diff_max_tt_min_qt_inter_slice is equal to 0.
[0198] sps_log2_diff_min_qt_min_cb_intra_slice_chroma specifies the default difference between the log2 of the minimum size of the luma samples in the chroma leaf blocks resulting from quadtree splitting of chroma CTUs with treeType equal to DUAL_TREE_CHROMA and the log2 of the minimum coded block size of the luma samples in the chroma CUs with treeType equal to DUAL_TREE_CHROMA in slices with slice_type equal to 2 (I) that reference the SPS. When partition_constraints_override_flag is equal to 1, this default difference can be overridden by slice_log2_diff_min_qt_min_cb_chroma present in the slice header of the slice that references the SPS. The value of sps_log2_diff_min_qt_min_cb_intra_slice_chroma shall be in the range of 0 to CtbLog2SizeY - MinChLog2SizeY, inclusive. When not present, the value of sps_log2_diff_min_qt_min_cb_intra_slice_chroma is inferred to be equal to 0. The log2 of the minimum size of the luma samples in the chroma leaf blocks resulting from quadtree splitting of CTUs with treeType equal to DUAL_TREE_CHROMA is derived as follows:
[0199] MinQtLog2SizeIntraC = sps_log2_diff_min_qt_min_cb_intra_slice_chroma + MinCbLog2SizeY
[0200] sps_max_mtt_hierarchy_depth_intra_slice_chroma specifies the default maximum hierarchical structure depth of chroma coding units generated by multi-type tree splitting of chroma quad-tree leaves with treeType equal to DUAL_TREE_CHROMA in slices with slice_type equal to 2 (I) that reference the SPS. When partition_constraints_override_flag is equal to 1, the default maximum hierarchical structure depth can be overridden by slice_max_mtt_hierarchy_depth_chroma in the slice header of the slice that references the SPS. The value of sps_max_mtt_hierarchy_depth_intra_slice_chroma shall be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY (including the end values). When it does not exist, the value of sps_max_mtt_hierarchy_depth_intra_slice_chroma is inferred to be equal to 0.
[0201] sps_log2_diff_max_bt_min_qt_intra_slice_chroma specifies the default difference between the base-2 logarithm of the maximum size (width or height) in the luma samples of a chroma coding block that can be split using binary splitting and the base-2 logarithm of the minimum size (width or height) in the luma samples of a chroma leaf block generated by quadtree splitting of a chroma CTU with treeType equal to DUAL_TREE_CHROMA in slices with slice_type equal to 2 (I) that reference the SPS. When partition_constraints_override_flag is equal to 1, this default difference can be overridden by slice_log2_diff_max_bt_min_qt_chroma present in the slice header of the slice that references the SPS. The value of sps_log2_diff_max_bt_min_qt_intra_slice_chroma shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeIntraC (including the end values). When sps_log2_diff_max_bt_min_qt_intra_slice_chroma does not exist, the value of sps_log2_diff_max_bt_min_qt_intra_slice_chroma is inferred to be equal to 0.
[0202] sps_log2_diff_max_tt_min_qt_intra_slice_chroma specifies the default difference between the base-2 logarithm of the maximum size (width or height) in the luma samples of a chroma coded block that can use ternary splitting and the base-2 logarithm of the minimum size (width or height) in the luma samples of a chroma leaf block resulting from a quadtree split of a chroma CTU with treeType equal to DUAL_TREE_CHROMA in a slice with slice_type equal to 2 (I) that references the SPS. When partition_constraints_override_flag is equal to 1, this default difference can be overridden by slice_log2_diff_max_tt_min_qt_chroma present in the slice header of a slice that references the SPS. The value of sps_log2_diff_max_tt_min_qt_intra_slice_chroma shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeIntraC, inclusive. When sps_log2_diff_max_tt_min_qt_intra_slice_chroma is not present, it is inferred that the value of sps_log2_diff_max_tt_min_qt_intra_slice_chroma is equal to 0.
[0203] sps_max_lumatransforrn_size_64_flag being equal to 1 specifies that the maximum transform size in the luma samples is equal to 64. sps_max_luma_transform_size_64_flag being equal to 0 specifies that the maximum transform size in the luma samples is equal to 32. When CtbSizeY is less than 64, the value of sps_max_luma_transform_size_64_flag shall be equal to 0.
[0204] The variables MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, and MaxTbSizeY are derived as follows:
[0205] MinTbLog2SizeY = 2
[0206] MaxTbLog2SizeY = sps_max_luma_transform_size_64_flag? 6: 5
[0207] MinTbSizeY = 1 << MinTbLog2SizeY
[0208] MaxTbSizeY = 1 << MaxTbLog2SizeY
[0209] same_qp_table_for_chroma being equal to 1 specifies that only one chroma QP mapping table is signaled and that this table applies to both the Cb and Cr residuals and the joint Cb - Cr residuals. same_qp_table_for_chroma being equal to 0 specifies that three chroma QP mapping tables are signaled in the SPS. When same_qp_table_for_chroma does not exist in the bitstream, it is inferred that the value of same_qp_table_for_chroma is equal to 1.
[0210] num_points_in_qp_table_rainus1[ i ] + 1 specifies the number of points used to describe the i-th chroma QP mapping table. The value of num_points_in_qp table minus1[ i ] shall be in the range of 0 to 63 + QpBdOffsetc (including the end values). When num_points_in_qp_table_minus1[ 0 ] does not exist in the bitstream, it is inferred that the value of num_points_in_qp_table_minus1[ 0 ] is equal to 0.
[0211] delta_qp_in_val_minua1[ i ][ j ] specifies the incremental value for the input coordinate used to derive the j-th pivot point of the i-th chroma QP mapping table. When delta_qp_in_val_minus1[ 0 ][ j ] does not exist in the bitstream, it is inferred that the value of delta_qp in delta_qp_in_val_minus1[ 0 ][ j ] is equal to 0.
[0212] delta_qp_out_val[ i ][ j ] specifies the incremental value for the output coordinate used to derive the j-th pivot point of the i-th chroma QP mapping table. When delta_qp_out_val[ 0 ][ j] does not exist in the bitstream, it is inferred that the value of delta_qp_out_val[ 0 ][ j ] is equal to 0.
[0213] The i-th chroma QP mapping table ChromaQpTable[ i ] is derived as follows, where i = 0..same_qp_table_for_chroma? 0 : 2:
[0214]
[0215] When same_qp_table_for_chroma is equal to 1, ChromaQpTable[ 1 ][ k ] and ChromaQpTable[ 2 ][ k ] are set to be equal to ChromaQpTable[ 0 ][ k ], where k = -QpBdOffsetC..63.
[0216] For bitstream conformance requirements, the values of qpInVal[ i ][ j ] and qpOutVal[ i ][ j ] shall be in the range of -QpBdOffsetC to 63 (including the end values), where i = i = 0..same_qp_table_for_chroma? 02, and j = 0..num_points_in_qp_table_minus1[ i ].
[0217] When sps_weighted_pred_flag is equal to 1, it specifies that weighted prediction can be applied to P slices that reference the SPS. When sps_weighted_pred_flag is equal to 0, it specifies that weighted prediction is not applied to P slices that reference the SPS.
[0218] When sps_weighted_bipred_flag is equal to 1, it specifies that explicit weighted prediction can be applied to B slices that reference the SPS. When sps_weighted_bipred_flag is equal to 0, it specifies that explicit weighted prediction is not applied to B slices that reference the SPS.
[0219] When sps_sao_enabled_flag is equal to 1, it specifies that the sample adaptive offset process is applied to the reconstructed picture after the deblocking filtering process. When sps_sao_enabled_flag is equal to 0, it specifies that the sample adaptive offset process is not applied to the reconstructed picture after the deblocking filter process.
[0220] When sps_alf_enabled_flag is equal to 0, it specifies that the adaptive loop filter is disabled.
[0221] When sps_alf enabled_flag is equal to 1, it specifies that the adaptive loop filter is enabled.
[0222] When sps_transform_skip_enabled_flag is equal to 1, it specifies that transform_skip_flag may exist in the transform unit syntax. When tsps_transform_skip_enabled_flag is equal to 0, it specifies that transform_skip_flag does not exist in the transform unit syntax.
[0223] When sps_bdpcm_enabled_flag equals 1, it specifies that intra_bdpcm_flag may exist in the coding unit syntax for intra-coded units. When sps_bdpcm_enabled_flag equals 0, it specifies that intra_bdpcm_flag does not exist in the coding unit syntax for intra-coded units. When it does not exist, it is inferred that the value of sps_bdpcm_enabled_flag equals 0.
[0224] When sps_joint_cbcr_enabled_flag equals 0, it specifies that the joint coding of chrominance residuals is disabled. When sps_joint_cbcr_enabled_flag equals 1, it specifies that the joint coding of chrominance residuals is enabled.
[0225] When sps_ref_wraparound_enabled_flag equals 1, it specifies that horizontal wraparound motion compensation is applied in inter prediction. When sps_ref wraparound_enabled_flag equals 0, it specifies that horizontal wraparound motion compensation is not applied. When the value of (CtbSizeY / MinCbSizeY + 1) is less than or equal to (pic_width_in_luma_samples / MinCbSizeY - 1), where pic_width_in_luma_samples is the value of pic_width_in_luma_samples in any PPS that refers to the SPS, the value of sps_ref_wraparound_enabled_flag shall equal 0.
[0226] sps_ref_wraparound_offset_minus1 plus 1 specifies the offset used to calculate the horizontal wraparound position, in units of MinCbSizeY luma samples. The value of ref_wriparound_offset_minus1 shall be in the range of (CtbSizeY / MinCbSizeY) + 1 to (pic_width_in_luma_samples / MinCbSizeY) - 1 (inclusive), where pic_width_in_luma_samples is the value of pic_width_in_luma_samples in any PPS that refers to the SPS.
[0227] When sps_temporal_mvp_enabled_flag equals 1, it specifies that slice_temporal_mvp_enabled_flag exists in the slice header of slices in CVS where slice_type is not equal to I. When sps_temporal_mvp_enabled_flag equals 0, it specifies that slice_temporal_mvp_enabled_flag does not exist in the slice header and the temporal motion vector predictor is not used in CVS.
[0228] When sps_sbtmvp_enabled_flag equals 1, it specifies that the sub-block based temporal motion vector predictor can be used in picture decoding in CVS, where all slices have a slice_type not equal to I. When sps_sbtmvp_enabled_flag equals 0, it specifies that the sub-block based temporal motion vector predictor is not used in CVS. When sps_sbtmvp_enabled_flag does not exist, it is inferred to be equal to 0.
[0229] When sps_amvr_enabled_flag equals 1, it specifies that adaptive motion vector difference resolution is used in motion vector coding. When amvr_enabled_flag equals 0, it specifies that adaptive motion vector difference resolution is not used in motion vector coding.
[0230] When sps_bdof_enabled_flag equals 0, it specifies that bidirectional optical flow inter prediction is disabled. When sps_bdof_enabled_flag equals 1, it specifies that bidirectional optical flow inter prediction is enabled.
[0231] When sps_smvd_enabled_flag equals 1, it specifies that symmetric motion vector difference can be used in motion vector decoding. When sps_smvd_enabled_flag equals 0, it specifies that symmetric motion vector difference is not used in motion vector coding.
[0232] When sps_dmvr_enabled_flag equals 1, it specifies that inter dual prediction based on decoder motion vector modification is enabled. When sps_dmvr_enabled_flag equals 0, it specifies that inter dual prediction based on decoder motion vector modification is disabled.
[0233] When sps_bdof_dmvr_slice_present_flag equals 1, it specifies that slice_disable_bdof_dmvr_flag exists in the slice header that references the SPS.
[0234] The sps_bdof_dmvr_slice_present_flag being equal to 0 specifies that the slice_disable_bdof_dmvr_flag does not exist in the slice header that references the SPS. When the sps_bdof_dmvr_slice_present_flag does not exist, it is inferred that the value of the sps_bdof_dmvr_slice_present_flag is equal to 0.
[0235] The sps_mmvd_enabled_flag being equal to 1 specifies that the merge mode with motion vector difference is enabled. The sps_mmvd_enabled_flag being equal to 0 specifies that the merge mode with motion vector difference is disabled.
[0236] The sps_isp_enabled_flag being equal to 1 specifies that the intra prediction with sub - partitioning is enabled. The sps_isp_enabled_flag being equal to 0 specifies that the intra prediction with sub - partitioning is disabled.
[0237] The sps_mrl_enabled_flag being equal to 1 specifies that the intra prediction with multiple reference lines is enabled. The sps_mrl_enabled_flag being equal to 0 specifies that the intra prediction with multiple reference lines is disabled.
[0238] The sps_mip_enabled_flag being equal to 1 specifies that the matrix - based intra prediction is enabled. The sps_mip_enabled_flag being equal to 0 specifies that the matrix - based intra prediction is disabled.
[0239] The sps_cclm_enabled_flag being equal to 0 specifies that the cross - component linear model intra prediction from the luma component to the chroma component is disabled. The sps_cclm_enabled_flag being equal to 1 specifies that the cross - component linear model intra prediction from the luma component to the chroma component is enabled. When the sps_cclm_enabled_flag does not exist, it is inferred to be equal to 0.
[0240] The sps_cclm_colocated_chroma_flag being equal to 1 specifies that the left - upsampled luma samples in the cross - component linear model intra prediction are collocated with the top - left luma sample. The sps_cclm_colocated_chroma_flag being equal to 0 specifies that the left - upsampled luma samples in the cross - component linear model intra prediction are horizontally co - located with the top - left luma sample, but vertically shifted by 0.5 luma sample units relative to the top - left luma sample.
[0241] The sps_mts_enabled_flag being equal to 1 specifies that sps_explicit_ints_intra enabled_flag exists in the sequence parameter set RBSP syntax and sps_explicit_mts_inter_enabled_flag exists in the sequence parameter set RBSP syntax. The sps_mts_enabled_flag being equal to 0 specifies that sps_explicit_mts_intra_enabled_flag does not exist in the sequence parameter set RBSP syntax and sps_explicit_mts_inter_enabled_flag does not exist in the sequence parameter set RBSP syntax.
[0242] The sps_explicit_mts_intra_enabled_flag being equal to 1 specifies that tu_mts_idx may exist in the transform unit syntax for intra coded units. The sps_explicit_mts_intra_enabled_flag being equal to 0 specifies that tu_mts_idx does not exist in the transform unit syntax for intra coded units. When it does not exist, it is inferred that the value of sps_explicit_mts_intra_enabled_flag is equal to 0.
[0243] The sps_explicit_mts_inter_enabled_flag being equal to 1 specifies that tu_mts_idx may exist in the transform unit syntax for inter coded units. The sps_explicit_mts_inter_enabled_flag being equal to 0 specifies that tu_mts_idx does not exist in the transform unit syntax for inter coded units. When it does not exist, it is inferred that the value of sps_explicit_mts_inter_enabled_flag is equal to 0.
[0244] The sps_sbt_enabled_flag being equal to 0 specifies that sub-block transform for inter prediction CUs is disabled. The sps_sbt_enabled_flag being equal to 1 specifies that sub-block transform for inter prediction CUs is enabled.
[0245] The sps_sbt_max_size_64_flag being equal to 0 specifies that the maximum CU width and height for allowing sub-block transform is 32 luma samples. The sps_sbt_max_size_64_flag being equal to 1 specifies that the maximum CU width and height for allowing sub-block transform is 64 luma samples.
[0246] MaxSbiSize = Min(MaxTbSizeY, sps_sht_max_size_64_flag ? 64 : 32)
[0247] The sps_affine_enabled_flag specifies whether motion compensation based on the affine model can be used for inter prediction. The sps_affine_enabled_flag specifies whether motion compensation based on the affine model can be used for inter prediction. If the sps_affine_enabled_flag is equal to 0, the syntax shall be constrained such that motion compensation based on the affine model is not used in the CVS, and there are no inter_affine_flag and cu_affine_type_flag in the coding unit syntax of the CVS. Otherwise (the sps_affine_enabled_flag is equal to 1), motion compensation based on the affine model can be used in the CVS.
[0248] The sps_affine_type_flag specifies whether motion compensation based on the 6-parameter affine model can be used for inter prediction. If the sps_affine_type_flag is equal to 0, the syntax shall be constrained such that motion compensation based on the 6-parameter affine model is not used in the CVS, and there is no cu_affine_type_flag in the coding unit syntax of the CVS. Otherwise (the sps_affine_type_flag is equal to 1), motion compensation based on the 6-parameter affine model can be used in the CVS. When it does not exist, it is inferred that the value of the sps_affine_type_flag is equal to 0.
[0249] The sps_affine_amvr_enabled_flag being equal to 1 specifies the use of adaptive motion vector difference resolution in the motion vector coding of the affine inter mode. The sps_affine_amvr_enabled_flag being equal to 0 specifies the non-use of adaptive motion vector difference resolution in the motion vector coding of the affine inter mode.
[0250] The sps_affine_prof_enabled_flag specifies whether prediction correction using optical flow can be used for affine motion compensation. If the sps_affine_prof_enabled_flag is equal to 0, optical flow shall not be applied to correct affine motion compensation. Otherwise (the sps_affine_prof_enabled_flag is equal to 1), optical flow can be applied to correct affine motion compensation. When it does not exist, it is inferred that the value of the sps_affine_prof_enabled_flag is equal to 0.
[0251] The sps_palette_enabled_flag being equal to 1 specifies that pred_mode_plt_flag may be present in the coding unit syntax. The sps_palette_enabled_flag being equal to 0 specifies that pred_mode_plt_flag is not present in the coding unit syntax. When the sps_palette_enabled_flag is not present, it is inferred to be equal to 0.
[0252] The sps_bcw_enabled_flag specifies whether dual prediction with CU weights can be used for inter prediction. If the sps_bcw_enabled_flag is equal to 0, the syntax shall be constrained such that dual prediction with CU weights is not used in the CVS and bcw_idx is not present in the coding unit syntax of the CVS. Otherwise (sps_bcw_enabled_flag equal to 1), dual prediction with CU weights can be used in the CVS.
[0253] The sps_ibc_enabled_flag being equal to 1 specifies that the IBC prediction mode can be used in the decoding of pictures in the CVS. The sps_ibc_enabled_flag being equal to 0 specifies that the IBC prediction mode is not used in the CVS. When the sps_ibc_enabled_flag is not present, it is inferred to be equal to 0.
[0254] The sps_ciip_enabled_flag specifies that ciip_flag may be present in the coding unit syntax for inter coding units. The sps_ciip_enabled_flag being equal to 0 specifies that ciip_flag is not present in the coding unit syntax for inter coding units.
[0255] The sps_fpel_mmvd_enabled_flag being equal to 1 specifies that the merge mode with motion vector difference is using integer sample precision. The sps_fpel_mmvd_enabled_flag being equal to 0 specifies that the merge mode with motion vector difference can use fractional sample precision.
[0256] The sps_triangle_enabled_flag specifies whether triangle - based motion compensation can be used for inter - frame prediction. When sps_triangle_enabled_flag equals 0, the syntax shall be constrained such that triangle - based motion compensation is not used in the CVS, and merge_triangle_split_dir, merge_triangle_idx0, and merge_triangle_idxl do not exist in the coding unit syntax of the CVS. When sps_triangle_enabled_flag equals 1, triangle - based motion compensation can be used in the CVS.
[0257] When sps_lmcs_enabled_flag equals 1, luminance mapping with chroma scaling is used in the CVS. When sps_lmcs_enabled_flag equals 0, luminance mapping with chroma scaling is not used in the CVS.
[0258] When sps_lfnst_enabled_flag equals 1, lfnst_idx may exist in the residual coding syntax for intra - coded units. When sps_lfnst_enabled_flag equals 0, lfnst_idx does not exist in the residual coding syntax for intra - coded units.
[0259] When sps_ladf_enabled_flag equals 1, sps_nurn_ladf_intervals_minus2, sps_ladf_lowest_interval_qp_offset, sps_ladf_qp_offset[i], and sps_ladf_delta_threshold_minus1[i] exist in the SPS.
[0260] sps_num_ladf_intervals_minus2 plus 1 specifies the number of the sps_ladf_delta_threshold_minus1[i] and sps_ladf_qp_offset i syntax elements that exist in the SPS. The value of sps_num_ladf_intervals_minus2 shall be in the range of 0 to 3 (including the end values).
[0261] sps_ladf_lowest_interval_qp_offset specifies the offset used to derive the specified variable qP. The value of ps_ladf_lowest_interval_qp_offset shall be in the range of 0 to 63 (including the end values).
[0262] sps_ladf_qp_offset[1] specifies the offset array used to derive the specified variable qP. The value of sps_ladf_qp_offset[ i ] shall be in the range of 0 to 63, inclusive.
[0263] sps_ladf_delta_threshold_minus1[ i ] is used to calculate the value of SpsLadfIntervalLowerBound[ i ], which specifies the lower bound of the i-th luminance intensity level interval. The value of sps_ladf_delta_threshold_minus1[ i ] shall be in the range of 0 to 2 BrtDepthY - to 3, inclusive.
[0264] Set the value of SpsLadflntervalLowerBound[ 0 ] to be equal to 0.
[0265] For each value of i in the range of 0 to sps_num_ladf_intervals_minus2, inclusive, the variable SpsLadflntervalLowerBound[ i + 1 ] is derived as follows:
[0266] SpsLadflntervalLowerBound[ i + 1 ] = SpsLadflntervalLowerBound[ i] +sps_ladf_delta_threshold_minus1[ i ]+ 1
[0267] sps_scaling_list_enabled_flag being equal to 1 specifies that the scaling list is used for the scaling process of transform coefficients. sps_scaling_list_enabled_flag being equal to 0 specifies that the scaling list is not used for the scaling process of transform coefficients.
[0268] hrd_parameters_present_flag being equal to 1 specifies the presence of the syntax elements num_units_in_tick and time_scale and the syntax structure general_hrd_parameters() in the SPS RBSP syntax structure. general_hrd_parameters_present_flag being equal to 0 specifies the absence of the syntax elements num_units_in_tick and time_scale and the syntax structure general_hrd_parameters() in the SPS RBSP syntax structure.
[0269] num_units_in_tick is the number of time units of a clock operating at a frequency of time_scale Hz, which corresponds to an increment of the clock tick counter (referred to as a clock tick). num_units_in_tick shall be greater than 0. The clock tick in seconds is equal to the quotient of num_units_in_tick divided by time_scale. For example, when the picture rate of a video signal is 25 Hz, time_scale can be equal to 27,000,000 and num_units_in_tick can be equal to 1,080,000, and thus the clock tick can be equal to 0.04 seconds.
[0270] time_scale is the number of time units elapsed in one second. For example, the time_scale of a time coordinate system measuring time using a 27 MHz clock is 27,000,000. The value of time_scale shall be greater than 0.
[0271] sub_layer_cpb_parameters_present_flag being equal to 1 specifies that the syntax structure general_hrd_parameters() is present in the SPS RBSP and includes HRD parameters for sub-layer representations with TemporalId in the range from 0 to sps_max_sub_layers_minus1 (including the end values). sub_layer_cpb_parameters_present_flag being equal to 0 specifies that the syntax structure general_hrd_parameters() is present in the SPS RBSP and includes HRD parameters for a sub-layer representation with TemporalId equal to sps_max_sub_layers_minus1.
[0272] vui_parameters_present_flag being equal to 1 specifies that the syntax structure vui_parameters() is present in the SPS RBSP syntax structure. vui_parameters_present_flag being equal to 0 specifies that the syntax structure vui_parameters() is not present in the SPS RBSP syntax structure.
[0273] When sps_extension_flag equals 0, it specifies that the sps_extension_data_flag syntax structure does not exist in the SPS RBSP syntax structure. When sps_extension_flag equals 1, it specifies that the sps_extension_data_flag syntax structure exists in the SPS RBSP syntax structure.
[0274] sps_extension_data_flag can have any value. Its presence and value do not affect the decoder's compliance with the profiles specified in this version of the present specification. Decoders compliant with this version of the present specification shall ignore all sps_extension_data_flag syntax elements.
[0275] As provided in Table 2, the NAL unit may include a Picture Parameter Set (PPS). Table 4 shows the syntax of the PPS provided in JVET-O2001.
[0276]
[0277]
[0278]
[0279]
[0280]
[0281] Regarding Table 4, JVET-02001 provides the following semantics:
[0282] The PPS RBSP shall be available for the decoding process before it is referenced, including being provided in at least one access unit having a TemporalId less than or equal to the TemporalId of the PPS NAL unit or by an external means, and the PPS NAL unit containing the PPS RBSP shall have a nuh_layer_id equal to that of the coded slice NAL unit that references it.
[0283] All pps NAL units having a specific value of pps_pic_parameter_set_id within an access unit shall have the same content.
[0284] pps_pic_parameter_set_id identifies the PPS for reference by other syntax elements.
[0285] The value of pps_pic_parameter_set_id shall be in the range of 0 to 63 (including the end values).
[0286] The pps_seq_parameter_set_id specifies the value of sps_seq_parameter_set_id of the SPS. The value of pps_seq_parameter_set_id shall be in the range of 0 to 15, inclusive. In all PPSs referenced by coded pictures in the CVS, the value of pps_seq_parameter_set_id shall be the same.
[0287] The pic_width_in_luma_samples specifies the width of each decoded picture that references the PPS, in luma samples. The pic_width_in_luma_samples shall not be equal to 0, shall be an integer multiple of Max(8, MinCbSizeY), and shall be less than or equal to pic_width_max_in_luma_samples.
[0288] When subpics_present_flag is equal to 1, the value of pic_width_in_luma_samples shall be equal to pic_width_max_in_luma_samples.
[0289] The pic_height_in_luma_samples specifies the height of each decoded picture that references the PPS, in luma samples. The pic_height_in_luma_samples shall not be equal to 0 and shall be an integer multiple of Max(8, MinCbSizeY), and shall be greater than or equal to pic_height_max_in_luma_samples.
[0290] When subpics_present_flag is equal to 1, the value of pic_height_in_luma_samples shall be equal to pic_height_max_in_luma_samples.
[0291] Let refPicWidthInLumaSamples and refPicHeightInLumaSamples be the pic_width_in_luma_samples and pic_height_in_luma_samples of the reference picture of the current picture that references this PPS, respectively. Bitstream conformance requires all of the following conditions to be met:
[0292] - pic_widthinluma_samples * 2 shall be greater than or equal to refPicWidthInLumaSamples.
[0293] - The pic_height_in Jumasamples * 2 shall be greater than or equal to refPicHeightInLumaSamples.
[0294] - pic_width_in_luma_samples shall be less than or equal to refPicWidthInLumaSamples * 8.
[0295] - pic_height_in_luma_samples shall be less than or equal to refPicHeightInLumaSamples * 8.
[0296] The variables PicWidthlnCtbsY, PicHeightlnCtbsY, PicSizelnCtbsY, PicWidthlnMinCbsY, PicHeightInMinCb sY, PicSizeInMinCb sY, PicSizelnSamplesY, PicWidthlnSamplesC and PicHeightInSamplesC are derived as follows:
[0297] PicWidthlnCtbsY = Ceil(pic_width_in_luma_samples ÷ CtbSizeY)
[0298] PicHeightlnCtbsY = Ceil(pic_height_in_lumasamples ÷ CtbSizeY)
[0299] PicSizelnCtbsY PicWidthInCtbsY * PicHeightInCtbsY
[0300] PicWidthInMinCbsY = pic_width_in_luma_samples / MinCbSizeY
[0301] PicHeightInMinCbsY = pic_height_in_luma_samples / MinCbSizeY
[0302] PicSizeInMinCbsY = PicWidthInMinCbsY * PicHeightInMinCbsY
[0303] PicSizeInSamplesY = pic_width_in_luma_samples * pic_height_in_luma_samples
[0304] PicWidthInSamplesC = pic_width_in_luma_samples / SubWidthC
[0305] PicHeightInSamplesC = pic_height_in_luma_samples / SubHeightC
[0306] The conformance_window_flag being equal to 1 indicates that the conformance cropping window offset parameter immediately follows in the SPS. The conformance_window_flag being equal to 0 indicates that there is no conformance cropping window offset parameter.
[0307] conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset specify the samples of the picture output from the decoding process in the CVS according to the rectangular region specified in the coordinates of the picture to be output. When the conformance_window_flag is equal to 0, it is inferred that conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset are equal to 0.
[0308] The conformance cropping window contains the luma samples with horizontal picture coordinates from SubWidthC * conf_win_left_offset to pic_width_in_luma_samples - (SubWidthC * conf_win_right_offset + 1), and vertical picture coordinates from SubHeightC * conf_win_top_offset to pic_height_in_luma_samples - (SubHeightC * conf_win_bottom_offset + 1) (including the end values).
[0309] The value of SubWidthC * (conf_win_left_offset + conf_win_right_offset) shall be less than pic_width_in_luma_samples, and the value of SubHeightC * (conf_win_top_offset + conf_win_bottom_offset) shall be less than pic_height_in_luma_samples.
[0310] The variables PicOutputWidthL and PicOutputHeightL are derived as follows:
[0311] PicOutputWidthL = pic_width_in_luma_samples - SubWidthC * (conf_winright_offset + conf_win_left_offset)
[0312] PicOutputHeightL = pic_height_in_pic_size_units -
[0313] When ChromaArrayType is not equal to 0, the corresponding specified samples of the two chroma matrices are the samples with picture coordinates (x / SubWidthC, y / SubHeightC), where (x, y) are the picture coordinates of the specified luma sample.
[0314] Note - The conforming crop window offset parameters are only applied at output. All internal decoding processes are applied to the uncropped picture size.
[0315] Let ppsA and ppsB be any two PPSs that reference the same SPS. Bitstream conformance requires that when ppsA and ppsB have the same values of pic_width_in_luma_samples and pic_height_in_luma_samples respectively, ppsA and ppsB shall have the same values of conf_win_left_offset, conf_win_right_offset, conf_win_top_offset, and conf_win_bottom_offset respectively.
[0316] The output_flag_present_flag being equal to 1 indicates the presence of the pic_output_flag syntax element in the slice header that references the PPS. The output_flag_present_flag being equal to 0 indicates the absence of the pic_output_flag syntax element in the slice header that references the PPS.
[0317] The single_tile_in_pic_flag being equal to 1 specifies the presence of only one tile in each picture that references the PPS. The single_tile_in_pic_flag being equal to 0 specifies the presence of more than one tile in each picture that references the PPS.
[0318] Note - In the case where there is no further brick partitioning within a tile, the entire tile is referred to as a brick. When a picture contains only a single tile without further brick partitioning, the picture is referred to as a single brick.
[0319] Bitstream conformance requires that for all PPSs referenced by coded pictures within a CVS, the value of the single_tile_in_pic_flag shall be the same.
[0320] The uniform_filespacing_flag being equal to 1 specifies that tile column boundaries and the same tile row boundaries are distributed uniformly across the picture, and the tile column boundaries and tile row boundaries are signaled using the syntax elements tile_cols_width_minus1 and tile_rows_height_minus1. The uniform_tile_spacing_flag being equal to 0 specifies that tile column boundaries and the same tile row boundaries may be distributed either uniformly or non-uniformly across the picture, and the tile column boundaries and tile row boundaries are signaled using the syntax elements num_tile_columns_minus1 and num_tile_rows_minus1 and a list of syntax element pairs tile_column_width_minus1[ i ] and tile_row_height_minus1[ i ]. When absent, the value of the uniform_tile_spacing_flag is inferred to be equal to 1.
[0321] tile_cols_width_rainus1 plus 1 specifies the width of the tile columns, except for the rightmost tile column of the picture, in CTB when uniform_tile_spacing_flag is equal to 1. The value of tile_cols_width_minus1 shall be in the range of 0 to PicWidthInCtbsY - 1 (including the end values). When not present, it is inferred that the value of tile_cols_width_minus1 is equal to PicWidthlnCtbsY - 1.
[0322] tile_rows_height_minus1 plus 1 specifies the height of the tile rows, except for the bottom tile row of the picture, in CTB when uniform_tile_spacing_flag is equal to 1. The value of tile_rows_height_minus1 shall be in the range of 0 to PicHeightlnCtbsY - 1 (including the end values). When not present, it is inferred that the value of tile_rows_height_minus1 is equal to PicHeightInCtbsY - 1.
[0323] num_tile_columns_minus1 plus 1 specifies the number of tile columns that divide the picture when uniform_tile_spacing_flag is equal to 0. The value of num_tile_columns_minus1 shall be in the range of 0 to PicWidthlnCtbsY - 1 (including the end values). When single_tile_in_pic_flag is equal to 1, it is inferred that the value of num_tile_columns_minus1 is equal to 0. Otherwise, when uniform_tile_spacing_flag is equal to 1, the value of num_tile_columns_minus1 is inferred according to the specification.
[0324] num_tik_rows_minus1 plus 1 specifies the number of tile rows that divide the picture when uniform_tile_spacing_flag is equal to 0. The value of num_tile_rows_minus1 shall be in the range of 0 to PicHeightlnCtbsY - 1 (including the end values). When single_tile_in_pic_flag is equal to 1, it is inferred that the value of nuratile_rows_minus1 is equal to 0. Otherwise, when uniform_tile_spacing_flag is equal to 1, the value of num_tile_rows_minus1 is inferred according to the specification.
[0325] The variable NumTileslnPic is set to be equal to (num_tile_columns_minus1 + 1) * (num_tile_rows_minus1 + 1).
[0326] When single_tile_in_pic_flag is equal to 0, NumTileslnPic shall be greater than 1.
[0327] tile_column_width_minus1[[ i ] plus 1 specifies the width of the i-th tile column in units of CTB.
[0328] tile_row_height_mainus1[ i ] plus 1 specifies the height of the i-th tile row in units of CTB.
[0329] brick_splitting_present_flag being equal to 1 specifies that one or more tiles of a picture that references the PPS can be split into two or more bricks. brick_splitting_present_flag being equal to 0 specifies that no tiles of pictures that do not reference the PPS are split into two or more bricks.
[0330] num_tiles_in_pic_minus1 plus 1 specifies the number of slices in each picture that references the PPS. The value of num_tiles_in_pic_minus1 shall be equal to NumTileslnPic - 1. When it does not exist, it is inferred that the value of num_tiles_in_pic_minus1 is equal to NumTileslnPic - 1.
[0331] brick_split_flag[[ i ] being equal to 1 specifies that the i-th tile is split into two or more bricks. brick_split_flag[ i ] being equal to 0 specifies that the i-th tile is not split into two or more bricks. When it does not exist, it is inferred that the value of brick_split_flag[ i ] is equal to 0.
[0332] uniform_brick_spacing_flag[ i ] being equal to 1 specifies that the horizontal brick boundaries are evenly distributed over the i-th brick and signals the horizontal brick boundaries using the syntax element brick_height_minus1[ i ]. uniform_brick_spacing_flag[ i ] being equal to 0 specifies that the horizontal brick boundaries may be evenly distributed or may not be evenly distributed over the i-th tile, and signals the horizontal brick boundaries using the syntax element num_brick_rows_minus2[ i ] and a list of the syntax elements brick_row_height_minus1[ i ][ j ]. When not present, it is inferred that the value of uniform_brick_spacing_flag[ i ] is equal to 1.
[0333] brick_height_minus1[ i ] plus 1 specifies the height, in CTB, of the rows of bricks in the i-th tile other than the bottom brick when uniform_brick_spacing_flag[ i ] is equal to 1. When present, the value of brick_height_minus1 shall be in the range 0 to RowHeight[ i ] - 2, inclusive. When not present, it is inferred that the value of brick_height_minus1[ i ] is equal to RowHeight[ i ] - 1.
[0334] num_brick_rows_minus2[ i ] plus 2 specifies the number of bricks that divide the i-th tile when uniform_brick_spacing_flag[ i ] is equal to 0. When present, the value of num_brick_rows_minus2[ i ] shall be in the range 0 to RowHeight[ i ] - 2, inclusive. When brick_split_flag[ i ] is equal to 0, it is inferred that the value of num_brick_rows_minus2[ i ] is equal to -1. Otherwise, when uniform_brick_spacing_flag[ i ] is equal to 1, the value of num_brick_rows_minus2[ i ] is inferred as specified.
[0335] brick_row_height_minus1[ i ][ j ] plus 1 specifies the height, in CTB, of the j-th brick in the i-th tile when uniform_tile_spacing_flag is equal to 0.
[0336] Derive the following variables, and when uniform_tile_spacing_flag is equal to 1, infer the values of num_tile_columns_minus1 and num_tile_rows_minus1, and for each i in the range 0 to NumTilesInPic - 1 (inclusive), when uniform_brick_spacing_flag[ i ] is equal to 1, infer the value of num_brick_rows_minus2[ i ] by calling the specified CTB raster and brick scan conversion procedures:
[0337] - List RowHeight[ j ], where j is in the range 0 to num_tile_rows_minus1 (inclusive), which specifies the height of the j-th tile row in CTB units,
[0338] - List CtbAddrRsToBs[ ctbAddrRs ], where ctbAddrRs is in the range 0 to PicSizelnCtbsY - 1 (inclusive), which specifies the conversion from the CTB address in the CTB raster scan of the picture to the CTB address in the brick scan,
[0339] - List CtbAddrBsToRs[ctbAddrBs], where ctbAddrBs is in the range 0 to PicSizelnCtbsY - 1 (inclusive), which specifies the conversion from the CTB address in the brick scan to the CTB address in the CTB raster scan of the picture,
[0340] - List BrickId[ ctbAddrBs ], where ctbAddrBs is in the range 0 to PicSizelnCtbsY - 1 (inclusive), which specifies the conversion from the CTB address in the brick scan to the brick ID,
[0341] - List NumCtusInBrick[ brickIdx], where brickIdx is in the range 0 to NumBrickslnPic - 1, inclusive, which specifies the conversion from the brick index to the number of CTUs in the brick,
[0342] - List FirstCtbAddrBs[brickIdx], where brickIdx is in the range 0 to NumBricksInPic - 1 (inclusive), which specifies the conversion from the brick ID to the CTB address in the brick scan of the first CTB in the brick.
[0343] A single_brick_per_slice_flag equal to 1 specifies that each slice that references this PPS includes one brick. A single_brick_per_slice_flag equal to 0 specifies that the slices that reference this PPS may include more than one brick. When not present, it is inferred that the value of single_brick_per_slice_flag is equal to 1.
[0344] A rect_slice_flag equal to 0 specifies that the bricks within each slice are in raster scan order and no slice information is signaled in the PPS. A rect_slice_flag equal to 1 specifies that the bricks within each slice cover a rectangular region of the picture and slice information is signaled in the PPS. When brick_splitting_present_flag is equal to 1, the value of rect_slice_flag shall be equal to 1. When not present, it is inferred that rect_slice_flag is equal to 1.
[0345] num_slices_in_pic_minus1 plus 1 specifies the number of slices that reference the PPS in each picture. The value of num_slices_in_pic_minus1 shall be in the range of 0 to NumBricksInPic - 1 (inclusive). When not present and single_brick_per_slice_fiag is equal to 1, it is inferred that the value of num_slices_in_pic_minus1 is equal to NumBricksInPic - 1.
[0346] bottom_right_brick_idx_length_minus1 plus 1 specifies the number of bits used to represent the syntax element bottom_right_brick_idx_delta[ i ]. The value of bottom right_brick_idx_length_minus1 shall be in the range of 0 to Ceil(Log2(NumBricksInPic)) - 1 (inclusive).
[0347] bottom_right_brick_idx_delta[ i ] specifies the difference between the brick index of the brick at the bottom - right corner of the i - th slice and the brick index of the brick at the bottom - right corner of the (i - 1) - th slice when i > 0. bottom_right_brick_idx_delta[ 0 ] specifies the brick index of the brick at the bottom - right corner of the 0 - th slice. When single_brick_per_slice_flag equals 1, it is inferred that the value of bottom_right_brick_idx_delta[ i ] equals 1. It is inferred that the value of BottomRightBrickIdx[num_slices_in_pic_minus1] equals NumBricksInPic - 1. The syntax element of bottom_right_brick_idx_delta[ i ] has a length of bottom_right_brick_idx_length_minus1 + 1 bits.
[0348] brick_idx_delta_sign_flag[ i ] being equal to 1 indicates the positive sign of bottom_right_brick_idx_delta[i]. sign bottom_right_brick_idx_delta[ i ] being equal to 0 indicates the negative sign of bottom_right_brick_idx_delta[ i ].
[0349] Bit - stream conformance requires that a slice shall consist of multiple complete tiles or only a sequentially - consecutive complete brick of a single tile.
[0350] The variables TopLeftBrickIdx[i], BottomRightBrickIdx[ i ], NumBricksInSlice[ i ], and BricksToSliceMap[ j ], which specify the brick index of the brick at the top - left corner of the i - th slice, the brick index of the brick at the bottom - right corner of the i - th slice, the number of bricks in the i - th slice, and the mapping of bricks to slices, are derived as follows:
[0351]
[0352] The loop_filter_across_bricks_enabled_flag being equal to 1 specifies that loop filter operations can be performed across brick boundaries in pictures that reference the PPS. The loop_filter_across_bricks_enabled_flag being equal to 0 specifies that loop filter operations are not performed across brick boundaries in pictures that reference the PPS. Loop filter operations include deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When not present, it is inferred that the value of the loop_filter_across_bricks_enabled_flag is equal to 1.
[0353] The loop_filter_across_slices_enabled_flag being equal to 1 specifies that loop filter operations can be performed across slice boundaries in pictures that reference the PPS. The loop_filter_across_slice_enabled_flag being equal to 0 specifies that loop filter operations are not performed across slice boundaries in pictures that reference the PPS. Loop filter operations include deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When not present, it is inferred that the value of the loop_filter_across_slices_enabled_flag is equal to 0.
[0354] The signalled_slice_id_flag being equal to 1 specifies that the slice ID of each slice is signalled. The signalled_slice_id_flag being equal to 0 specifies that the slice ID is not signalled. When the rect_slice_flag is equal to 0, it is inferred that the value of the signalled_slice_id_flag is equal to 0.
[0355] Note - For a bitstream that is the result of sub-bitstream extraction, when each picture in the bitstream contains a proper subset of the slices included in the pictures of the "original" bitstream, and the subset of included slices does not include the top-left slice of the pictures in the "original" bitstream, the value of the signalled_slice_id_flag in the PPS of the extracted bitstream must be equal to 1.
[0356] signalled_slice_id_length_minus1 plus 1 specifies the number of bits used to represent the syntax element slice_id[ i ] (when present) and the syntax element slice_address in the slice header. The value of signalled_slice_id_length_minus1 shall be in the range of 0 to 15 (inclusive). When not present, the value of signalled_slice_id_length_minus1 is inferred to be Ceil(Log2(Max(2, num_slices_in_pic_minus1 + 1))) - 1.
[0357] slice_id[ i ] specifies the slice ID of the i-th slice. The length of the slice_id[ i ] syntax element is signalled_slice_id_length_minus1+1 bits. When not present, for each i in the range of 0 to num_slices_in_pic_minus1 (inclusive), the value of slice_id[ i ] is inferred to be i.
[0358] entropy_coding_sync_enabled_flag being equal to 1 specifies that a specific synchronization process of context variables is called before decoding the CTU of the first CTB in the CTB rows of each tile in each picture that references a PPS, and a specific storage process of context variables is called after decoding the CTU of the first CTB in the CTB rows of each tile in each picture that references a PPS. entropy_coding_sync_enabled_flag being equal to 0 specifies that a specific synchronization process of context variables is not required to be called before decoding the CTU of the first CTB in the CTB rows of each tile in each picture that references a PPS, and a specific storage process of context variables is not required to be called after decoding the CTU of the first CTB in the CTB rows of each tile in each picture that references a PPS.
[0359] For bitstream conformance, for all PPSs referenced by coded pictures within CVS, the value of entropy_coding_sync_enabled_flag shall be the same.
[0360] cabac_init_present_flag being equal to 1 specifies the presence of cabac_init_flag in the slice header that references a PPS. cabac_init_present_flag being equal to 0 specifies the absence of cabac_init_flag in the slice header that references a PPS.
[0361] Adding 1 to num_ref id.x_default_active_minus1[ i ] specifies the inferred value of variable NumRefIdxActive[0] for P slices and B slices when i equals 0, where num_ref_idx_active_override_flag equals 0, and specifies the inferred value of NumRefIdxActive[1] for B slices when i equals 1, where num_ref_idx_active_override_flag equals 0. The value of num_ref_idx_default_active_minus1[ i ] shall be in the range of 0 to 14, inclusive.
[0362] rpl1_idx_present_flag being equal to 0 specifies that ref_pic_list_sps_flag[1] and ref_pic_list_idx[1] are not present in the slice header. rpl1_idx_present_flag being equal to 1 specifies that ref_pic_list_sps_flag[1] and ref_pic_list_idx[1] may be present in the slice header.
[0363] Adding 26 to init_qp_minus26 specifies the initial value of SliceQ for each slice that references the PPS. When a non-zero value of slice_qp_delta is decoded, the SliceQ is modified at the slice layer PY of the initial value. When a non-zero value of slice_qp_delta is decoded, the SliceQ is modified at the slice layer PY of the initial value. The value of init_qp_minus26 shall be in the range of -(26 + QpBdOffsety) to +37, inclusive.
[0364] log2_transform_skip_max_size_minus2 specifies the maximum block size for transform skip and shall be in the range of 0 to 3.
[0365] When not present, the inferred value of log2_transform_skip_max_size_minus2 is equal to 0.
[0366] The variable MaxTsSize is set to be equal to 1 << (log2_transform_skip_max_size_minus2 + 2).
[0367] The cu_qp_delta_enabled_flag being equal to 1 specifies that the cu_qp_delta_subdiv syntax element exists in the PPS and the cu_qp_delta_abs may exist in the transform unit syntax. The cu_qp_delta_enabled_flag being equal to 0 specifies that the cu_qp_delta_subdiv syntax element does not exist in the PPS and the cu_qp_delta_abs does not exist in the transform unit syntax.
[0368] cu_gp_delta_subdiv specifies the maximum cbSubdiv value of the coding unit that conveys cu_qp_delta_abs and cu_qp_delta_sign_flag. The value range of cu_qp_delta_subdiv is specified as follows:
[0369] - If slice_type is equal to I, the value of cu_qp_delta_subdiv shall be in the range of 0 to 2*(log2_ctu_size_minus2 - log2_min_qt_size intraslice_minus2 + MaxMttDepthY) (including the end values).
[0370] - Otherwise (slice_type is not equal to I), the value of cu_qp_delta_subdiv shall be in the range of 0 to 2*(log2_ctu_size_minus2 - log2_min_qt_size_inter slice_minus2 + MaxMttDepthY) (including the end values).
[0371] When it does not exist, it is inferred that the value of cu_qp_delta_subdiv is equal to 0.
[0372] pps_cb_qp_offset and pps_cr_qp_offset respectively specify the offsets for deriving Qp' Cb and Qp' Cr of the luma quantization parameter Qp' Y The values of pps_cb_qp_offset and pps_cr_qp_offset shall be in the range of -12 to +12 (including the end values). When ChromaArrayType is equal to 0, pps_cb_qp_offset and pps_cr_qp_offset are not used during the decoding process, and the decoder shall ignore their values.
[0373] pps_joint_cbcr_qp_offset specifies the one associated with the Qp' used for derivationCbCr luminance quantization parameter Qp' Y Offset. The value of pps_joint_cbcr_qp_offset shall be in the range of - 12 to + 12 (including the end values). When ChromaArrayType is equal to 0 or sps_joint_cbcr_enabled_flag is equal to 0, pps_joint_cbcr_qp_offset is not used during decoding, and the decoder shall ignore its value.
[0374] pps_slice_chroma_qp_offsets_present_flag being equal to 1 indicates the presence of the slice_cb_qp_offset and slice_cr_qp_offset syntax elements in the associated slice header. pps_slice_chroma_qp_offsets present_flag being equal to 0 indicates the absence of these syntax elements in the associated slice header. When ChromaArrayType is equal to 0, pps_slice_chroma_qp_offsets_present_flag shall be equal to 0.
[0375] cu_chroma_qp_offset_enabled_flag being equal to 1 specifies that cu_chroma_qp_offset_flag may be present in the transform unit syntax.
[0376] cu_chroma_qp_offset_enabled_fiag being equal to 0 specifies that cu_chroma_qp_offset_flag is not present in the transform unit syntax. When ChromaArrayType is equal to 0, the bitstream conformance requires that the value of cu_chroma_cjp_offset_enabled_flag shall be equal to 0.
[0377] cu_chroma_qp_offset_subdiv specifies the maximum cbSubdiv value of the coding unit that conveys cu_chroma_qp_offset_flag. The value range of cu_chroma_qp_offset_subdiv is specified as follows:
[0378] - If slice_type is equal to I, the value of cu_chroma_qp_offset_subdiv shall be in the range of 0 to 2 * (log2_ctu_size_minus2 - log2_min_qt_size_intra_slice_minus2 + MaxMttDepthY) (including the end values).
[0379] - Otherwise (slice_type is not equal to I), the value of eu_chromaqp_offset_subdiv shall be in the range of 0 to 2 * (log2_ctu_size_minus2 - log2_min_qt_size_intera_slice_minus2 + MaxMaDepthY) (including the end values).
[0380] When not present, the inferred value of cu_chroma_qp_offset_subdiv is equal to 0.
[0381] chroma_qp_offset_list_len_minus1 plus 1 specifies the number of the syntax elements cb_qp_offset_list[ i ], cr_qp_offset_list[ i ], and joint_cbcr_qp_offset_list[ i ] present in the SPS. The value of chroma_qp_offset_list_len_minus1 shall be in the range of 0 to 5 (including the end values).
[0382] cb_qp_offset_list[ i ], cr_qp_offset_list[ i ], and joint_cbcr_qp_offset_list[ i ] specify the offsets used in the derivation of Qp' Cb 、Qp' Cr and Qp' CbCr respectively. The values of ch_qp_offset_list[ i ], cr_qp_offset_list[ i ], and joint_cbcr_qp_offset_list[ i ] shall be in the range of -12 to +12 (including the end values).
[0383] pps_weighted_pred_flag being equal to 0 specifies that weighted prediction is not applied to P slices that reference the SPS. pps_weighted_pred_flag being equal to 1 specifies that weighted prediction is applied to P slices that reference the SPS. When sps_weighted_pred_flag is equal to 0, the value of pps_weighted_pred_flag shall be equal to 0.
[0384] When pps_weighted_bipred_flag equals 0, it specifies that explicit weighted prediction is not applied to B slices that reference the PPS. When pps_weighted_bipred_flag equals 1, it specifies that explicit weighted prediction is applied to B slices that reference the PPS. When sps_weighted_bipred_flag equals 0, the value of pps_weighted_bipred_flag shall equal 0.
[0385] When deblocking_filter_control_present_flag equals 1, it specifies that there are deblocking filter control syntax elements in the PPS. When deblocking_filter_control_present_flag equals 0, it specifies that there are no deblocking filter control syntax elements in the PPS.
[0386] When deblocking_filter_override_enabled_flag equals 1, it specifies that deblocking_filter_override_flag exists in the slice header of the picture that references the PPS. When deblocking_filter_override_enabled_flag equals 0, it specifies that deblocking_filter_override_flag does not exist in the slice header of the picture that references the PPS. When it does not exist, it is inferred that the value of deblocking_filter_override_enabled_flag equals 0.
[0387] When pps_deblocking_filter_thsabled_flag equals 1, it specifies that the operation of the deblocking filter is not applied to slices that reference the PPS in which slice_deblocking_filter_disabled_flag does not exist. When pps_deblocking_filter_disabled_flag equals 0, it specifies that the operation of the deblocking filter is applied to slices that reference the PPS in which slice_deblocking_filter_disabled_flag does not exist. When it does not exist, it is inferred that the value of pps_deblocking_filter_disabled_flag equals 0.
[0388] pps_beta_offset_div2 and pps_te_offset_div2 specify the default deblocking parameter offsets (divided by 2) for β and tC applied to slices that reference the PPS, unless the default deblocking parameter offsets are overridden by deblocking parameter offsets present in the slice header of the slices that reference the PPS. The values of pps_beta_offset_div2 and pps_tc_offset_div2 shall each be in the range of -6 to 6, inclusive. When not present, pps_beta_offset_div2 and pps_tc_offset_div2 are inferred to have a value equal to 0.
[0389] constant_slice_header_params_enabled_flag equal to 0 specifies that pps_dep_quant_enabled_idc, pps_ref_pic_list_sps_idc[ i ], pps_temporal_mvp_enable_idc, pps_mvd_l1_zero_idc, pps_collocated_from_l0_idc, pps_six_minus_max_num_merge_cand_plus1, pps_five_minus_max_num_subblock_merge_cand_plus1, and pps_max_num_merge_cand_minus_max_num_triangle_cand_minus1 are inferred to be equal to 0. constant_slice_header_params_enabled_flag equal to 1 specifies that these syntax elements are present in the PPS.
[0390] pps_dep_quant_enabled_idc equal to 0 specifies that the syntax element dep_quant_enabled_flag is present in the slice header of the slices that reference the PPS. pps_dep_quant_enabled_idc equal to 1 or 2 specifies that the syntax element dep_quant_enabled_flag is not present in the slice header of the slices that reference the PPS. pps_dep_quant_enabled_idc equal to 3 is reserved for future use by ITU-T | ISO / IEC.
[0391] When pps_ref_pic_list_sps_idc[ i ] is equal to 0, it specifies that the syntax element ref_pic_list_sps_flag[ i ] exists in the slice header of the slice that references the PPS. When pps_ref_pic_list_sps_idc[ i ] is equal to 1 or 2, it specifies that the syntax element ref_pic_list_sps_flag[ i ] does not exist in the slice header of the slice that references the PPS. pps_ref_pic_list sps_idc[ i ] equal to 3 is reserved for future use by ITU-T | ISO / IEC.
[0392] When pps_temporal_mvp_enabled_idc is equal to 0, it specifies that the syntax element slice_temporal_mvp_enabled_flag exists in the slice header of the slice whose slice_type is not equal to I in the slice that references the PPS. When pps_temporal_mvp_enabled_idc is equal to 1 or 2, it specifies that the slice_temporal_mvp_enabled_flag does not exist in the slice header of the slice that references the PPS. pps_temporal_mvp_enabled_idc equal to 3 is reserved for future use by ITU-T | ISO / IEC.
[0393] When pps_mvd_l1_zero_ide is equal to 0, it specifies that the syntax element mvd_l1_zero_flag exists in the slice header of the slice that references the PPS. When pps_mvd_l1_zero_idc is equal to 1 or 2, it specifies that the mvd_l1_zero_flag does not exist in the slice header of the slice that references the PPS. pps_mvd_l1_zero_idc equal to 3 is reserved for future use by ITU-T | ISO / IEC.
[0394] When pps_collocated_from_l0_idc is equal to 0, it specifies that the syntax element collocated_from_l0_flag exists in the slice header of the slice that references the PPS. When pps_collocated_from_l0_idc is equal to 1 or 2, it specifies that the syntax element collocated_from_l0_flag does not exist in the slice header of the slice that references the PPS. pps_collocated_from_l0_idc equal to 3 is reserved for future use by ITU-T | ISO / IEC.
[0395] That pps_six_minus_max_num_merge_cand_plus1 equals 0 specifies that six_minus_max_num_merge_cand exists in the slice header of the slice that references the PPS. That pps_six_minus_max_num_merge_cand_plus1 is greater than 0 specifies that six_minus_max_num_merge_cand does not exist in the slice header of the slice that references the PPS. The value of pps_six_minus_max_num_merge_cand_plus1 shall be in the range of 0 to 6, inclusive of the end values.
[0396] That pps_five_minus_max_num_subblock_merge_cand_plus1 equals 0 specifies that five_minus_max_num_subblock_merge_cand exists in the slice header of the slice that references the PPS. That pps_five_minus_max_num_subblock_merge_cand_plus1 is greater than 0 specifies that five_minus_max_num_subblock_merge_cand does not exist in the slice header of the slice that references the PPS. The value of pps_five_minus_max_num_subblock_merge_cand_plus1 shall be in the range of 0 to 6, inclusive of the end values.
[0397] That pps_max_num_merge_cand_minus_max_num_triangle_cand_minus1 equals 0 specifies that max_num_merge_cand_minus_max_num_triangle_cand exists in the slice header of the slice that references the PPS. That pps_max_num_merge_cand_minus_max_num_triangle_cand_minus1 is greater than 0 specifies that max_num_merge_cand_minus_max_num_triangle_cand does not exist in the slice header of the slice that references the PPS. The value of pps_max_num_merge_cand_minus max num_triangle_cand_minus1 shall be in the range of 0 to MaxNumMergeCand - 1, inclusive of the end values.
[0398] The pps_loop_filter_across_virtual_boundaries_disabled_flag being equal to 1 specifies that the loop filter operations are disabled across virtual boundaries in pictures that reference the PPS. The pps_loop_filter_across_virtual_boundaries_disabled_flag being equal to 0 specifies that such disabling of the loop filter operations is not applied in pictures that reference the PPS. The loop filter operations include deblocking filter, sample adaptive offset filter, and adaptive loop filter operations. When absent, the value of pps_loopfilter_across_virtual_boundaries_disabled_flag is inferred to be equal to 0.
[0399] pps_num_ver_virtual_boundaries specifies the number of pps_virtual_boundaries_pos_x[ i ] syntax elements present in the PPS. When pps_virtualboundariespos_x[ i ] is absent, it is inferred to be equal to 0.
[0400] pps_virtual_boundaries_pos_x[i] is used to calculate the value of PpsVirtualBoundariesPosX[i], which specifies the position of the i-th vertical virtual boundary in units of luma samples. pps_virtual_boundaries_pos_x[i] shall be in the range of 1 to Ceil(pic_widthin_luma_samples ÷ 8) - 1, inclusive.
[0401] The position of the vertical virtual boundary PpsVirtualBoundariesPosX[ i ] is derived as follows:
[0402] PpsVirtualBoundariesPosX[ i ] = pps_virtual_boundariespos_x[ i ] * 8
[0403] The distance between any two vertical virtual boundaries shall be greater than or equal to CtbSizeY luma samples.
[0404] pps_num_hor_virtual_boundaries specifies the number of pps_virtual_boundaries_pos_y[i] syntax elements present in the PPS. When pps_num_hor_virtual_boundaries is absent, it is inferred to be equal to 0.
[0405] pps_virtual_boundaries_pos_y[i] is used to calculate the value of PpsVirtualBoundariesPosY[i], which specifies the position of the i-th horizontal virtual boundary in luminance samples. pps_virtual_boundaries_pos_y[i] shall be in the range of 1 to Ceil(pic_height_in_luma_samples÷8)-1, inclusive.
[0406] The position of the horizontal virtual boundary PpsVirtualBoundariesPosY[i] is derived as follows:
[0407] PpsVirtualBoundariesPosY[i] = pps_virtual boundaries_pos_y[i] * 8
[0408] The distance between any two horizontal virtual boundaries shall be greater than or equal to CtbSizeY luminance samples.
[0409] When slice_header_extension_present_flag is equal to 0, it specifies that the slice header extension syntax element does not exist in the slice header of the coded picture that references the PPS. When slice_header_extension_present_flag is equal to 1, it specifies that the slice header extension syntax element exists in the slice header of the coded picture that references the PPS. In a bitstream compliant with this version of the present specification, slice_header_extension_present_flag shall be equal to 0.
[0410] When pps_extension_flag is equal to 0, it specifies that the pps_extension_data_flag syntax element does not exist in the PPS RBSP syntax structure. When pps_extension_flag is equal to 1, it specifies that the pps_extension_data_flag syntax element exists in the PPS RBSP syntax structure.
[0411] pps_extension_data_flag can have any value. Its presence and value do not affect the decoder's compliance with the profiles specified in this version of the present specification. A decoder compliant with this version of the present specification shall ignore all pps_extension_data_flag syntax elements.
[0412] As described above, pictures and their regions can be classified based on which types of prediction modes (e.g., B type, P type, or I type) can be used to encode their video blocks. As provided in Table 2, a NAL unit can include an access unit delimiter. In JVET-O2001, the access unit delimiter is used to indicate the start of an access unit and the type of slices in the coded picture that are present in the access unit containing the NAL unit with the access unit delimiter. Table 5 shows the syntax of the access unit delimiter provided in JVET-O2001.
[0413]
[0414] Regarding the syntax structure of access_unit_delimiter_rbsp(), JVET-O2001 provides the following semantics:
[0415] The access unit delimiter is used to indicate the start of an access unit and the type of slices in the coded picture that are present in the access unit containing the NAL unit with the access unit delimiter. There is no canonical decoding process associated with the access unit delimiter.
[0416] pic_type indicates the number of slice_type values of all slices of the coded picture in the access unit containing the NAL unit with the access unit delimiter that are in the set listed in Table 6 for the given value of pic_type. In a bitstream compliant with this version of the present specification, the value of pic_type shall be equal to 0, 1, or 2. Other values of pic_type are reserved for future use by ITU-T | ISO / IEC. A decoder compliant with this version of the present specification shall ignore all reserved values of pic_type.
[0417]
[0418] As provided in Table 2, a NAL unit can include coded slices of a picture. The slice syntax structure includes the slice_header() syntax structure and the slice_data() syntax structure. Table 7 shows the syntax of the slice header provided in JVET-O2001.
[0419]
[0420]
[0421]
[0422]
[0423]
[0424] Regarding Table 7, JVET-O2001 provides the following semantics:
[0425] When present, the value of each of the slice header syntax elements slice_pic_parameter_set_id, non_reference_picture_flag, colour_plane_id, slice_pic_order_cnt_lsb, recovery_poc_cnt, no_output_of_prior_pics_flag, pic_output_flag, and slice_temporal_mvp_enabled_flag shall be the same in all slice headers of the coded picture.
[0426] The variable CuQpDeltaVal, which specifies the difference between the luma quantization parameter of the coding unit containing cu_qp_delta_abs and its prediction, is set to be equal to 0. The variables CuQpOffset Cb 、Qp' Cr and Qp' CbCr that specify the values to be used when determining the respective values of the quantization parameters of Qp' Cb 、CuQpOffset Cr and CuQpOffset CbCr are all set to be equal to 0.
[0427] slice_pic_parameter_set_id specifies the value of pps_pic_parameter_set_id of the PPS in use. The value of slice_pic_parameter_set_id shall be in the range of 0 to 63, inclusive of the end values.
[0428] Bitstream conformance requires that the value of Temporaild of the current picture shall be greater than or equal to the value of TemporalId of the PPS that has a pps_pic_parameter_set_id equal to slice_pic_parameter_set_id.
[0429] slice_address specifies the slice address of the slice. When not present, the value of slice_address is inferred to be equal to 0.
[0430] If rect_slice_flag is equal to 0, the following applies:
[0431] - The slice address is the specified tile ID.
[0432] - The length of slice_address is Ceil(Log2 (NumBricksInPic)) bits.
[0433] - The value of slice_address shall be in the range from 0 to NumBricksInPic - 1, inclusive.
[0434] Otherwise (when rect_slice_flag equals 1), the following applies:
[0435] - The slice address is the slice ID of the slice.
[0436] - The length of slice_address is signalled_slice_id_length_minus1 + 1 bits.
[0437] - If signalled_slice_id_flag equals 0, the value of slice_address shall be in the range from 0 to num_slices_in_pic_minus1, inclusive. Otherwise, the value of slice_address shall be in the range from 0 to 2 (signalled _slice_id_length_minus1 +1) - 1, inclusive.
[0438] Bitstream conformance requires the following constraints to apply:
[0439] - The value of slice_address shall not be equal to the value of slice_address of any other coded slice NAL unit of the same coded picture.
[0440] - When rect_slice_flag equals 0, the slices of the picture shall be in the order of increasing slice address value.
[0441] - The shape of the slices of the picture shall be such that each brick, when decoded, shall have its entire left and top boundaries formed by the picture boundary or by the boundaries of previously decoded bricks.
[0442] num_bricksin_slice_minus1, when present, specifies the number of bricks in a slice minus 1. The value of num_bricks in a slice shall be in the range of 0 to NumBrickslnPic - 1 (inclusive). When rect_slice_flag equals 0 and single_brick_per_slice_flag equals 1, it is inferred that the value of num_bricksin_slice_minus1 equals 0. When single_brick_per_slice_flag equals 1, it is inferred that the value of num_bricks_in_slice_minus1 equals 0.
[0443] The variable NumBricksInCurrSlice that specifies the number of bricks in the current slice and the slice brick index SliceBrickIdx[i] that specifies the i-th brick in the current slice are derived as follows:
[0444]
[0445] The variables SubPicIdx, SubPicLeftBoundaryPos, SubPicTopBoundaryPos, SubPicRightBoundaryPos, and SubPicBotBoundaryPos are derived as follows:
[0446]
[0447] non_reference_picture_flag equal to 1 specifies that the picture containing the slice must never be used as a reference picture. non_reference_picture_flag equal to 0 specifies that the picture containing the slice may or may not be used as a reference picture.
[0448] slice_type specifies the coding type of the slice according to Table 8.
[0449]
[0450] When nal_unit_type is a value of nal_unit_type in the range from IDR_W_RADL to CRA_NUT (inclusive) and the current picture is the first picture in an access unit, slice_type shall equal 2.
[0451] When separate_colour_plane_flag equals 1, colour_plane_id specifies the colour plane associated with the current slice RBSP. The value of colour_plane_id shall be in the range of 0 to 2, inclusive. The colour_plane_id values 0, 1, and 2 correspond to the Y, Cb, and Cr planes, respectively.
[0452] Note - There is no correlation between the decoding processes of pictures with different values of colour_plane_id.
[0453] slice_pic_order_cnt_lsb represents the picture order count modulo MaxPicOrderCntLsb of the current picture. The length of the slice_pic_order_cntisb syntax element is log2_max pic_order_cnt_lsb_minus4 + 4 bits. The value of slice_pic_order_cnt_lsh shall be in the range of 0 to MaxPicOrderCntLsb - 1, inclusive.
[0454] recovery_poc_cnt specifies the recovery point of the decoded picture in the output order. If there is a picture picA in the CVS that is after the current GDR picture in the decoding order and has a PicOrderCntVal equal to the value of the PicOrderCntVal of the current GDR picture plus recovery_poc_cnt, then picture picA is called the recovery point picture. Otherwise, the first picture in the output order that has a PicOrderCntVal greater than the value of the PicOrderCntVal of the current picture plus recovery_poc_cnt is called the recovery point picture. The recovery point picture shall not be before the current GDR picture in the decoding order. The value of recovery_poc_cnt shall be in the range of 0 to MaxPicOrderCntLsb - 1, inclusive.
[0455] The variable RpPicOrderCntVal is derived as follows:
[0456] RpPicOrderCntVal - PicOrderCntVal + recovery_poc_cnt
[0457] no_output_of_prior_pics_flag affects the output of previously decoded pictures after the CLVSS picture in the bitstream that is not the first picture in the decoded picture buffer, as specified.
[0458] The pic_output_flag controls the decoding picture output and removal process according to the specified influence. When the pic_output_flag does not exist, it is inferred to be equal to 1.
[0459] When ref_pic_list_sps_flag[ i ] is equal to 1, it specifies that the reference picture list i for the current slice is derived based on a syntax structure in the ref pic_list_struct(listIdx, rplsIdx) in the SPS where listIdx is equal to i. When ref_pic_list_sps_flag[ i ] is equal to 0, it specifies that the reference picture list i for the current slice is derived based on the syntax structure ref_pic_list_struct(listIdx, rplsIdx) (where listIdx is equal to i) that is directly included in the slice header of the current picture.
[0460] When ref_pic_list_sps_flag[ i ] does not exist, the following applies:
[0461] - If num_ref_pic_lists_in_sps[i] is equal to 0, it is inferred that the value of ref_pic_list_sps_flag[i] is equal to 0.
[0462] - Otherwise (num_ref_pic_lists_insps[ i ] is greater than 0), if rpl1_idx_present_flag is equal to 0, it is inferred that the value of ref_pic_list_sps_flag[ 1 ] is equal to ref_pic_list_sps_flag[ 0 ].
[0463] - Otherwise, it is inferred that the value of ref_pic_list_sps_flag[ i ] is equal to pps_ref_pic_list_sps_idc[ i ] - 1.
[0464] The ref_pic_list_idx[ i ] specifies the index of the list of ret_pic_list_struct(listIdx, rplsIdx) syntax structures included in the SPS that are of the ref_pic_list_struct(listIdx, rplsIdx) syntax structure (where listIdx is equal to i) for the reference picture list i used for exporting the current picture. The syntax element ref_pic_list_idx[ i ] is represented by Ceil(Log2(num_ref_pic_lists_in_sps[ i ])) bits. When it is not present, the value of ref_pic_list_idx[ i ] is inferred to be equal to 0. The value of ref_pic_list_idx[ i ] shall be in the range of 0 to num_ref_pic_lists_in_sps [ i ] - 1 (inclusive of the end values). When ref_pic_list_sps_flag[ i ] is equal to 1 and num_ref_pic_lists_in_sps[ i ] is equal to 1, the value of ref_pic_list_idx[ i ] is inferred to be equal to 0. When ref_piclist_sps_flag[ i ] is equal to 1 and rpl1_idx_present_flag is equal to 0, the value of ref_pic_list_idx[ 1 ] is inferred to be equal to ref_pic_list_idx[ 0 ].
[0465] The variable RplsIdx[ i ] is derived as follows:
[0466] RplsIdx[ i ] = ref_pic_list_sps_flag[ i ]? ref_pic_list_idx [i]: num_ref_pic_lists_in_sps[ i]
[0467] The slice_poc_lsb_lt[ i ][ j ] specifies the picture order count modulo MaxPicOrderCntLsb of the j-th LTRP entry in the i-th reference picture list. The slice_poc_lsb_lt[ i ] [ j ] syntax element has a length of log2_max_pic_order_cnt_lsb_minus4 + 4 bits.
[0468] The variable PocLsbLt[ i ][ j ] is derived as follows:
[0469] PocLsbLt[ i ][ j ] = ltrp_in_slice_header_flag[ i ][ RplsIdx[ i ] ]?
[0470] slice_poc_lsb_lt[i ][j ] : rpls_poc_lsb_lt[ listIdx ][ RplsIdx[ i ]][ j ]
[0471] The delta_poc_msb_present_flag[ i ] [ j ] being equal to 1 indicates the existence of delta_poc_msb_cycle_lt[ i ][ j ]. The delta poc_msb_present_flag[ i ][ j ] being equal to 0 indicates the non - existence of delta_poc_msb_cycle_lt[ i ][ j ].
[0472] Let prevTid0Pic be the previous picture in the decoding order that has the same nuh_layer_id as the current picture, has a TemporalId equal to 0 and is not a RASL or RADL picture. Let setOfPrevPocVals be the set consisting of:
[0473] - the PicOrderCntVal of prevTid0Pic,
[0474] - the PicOrderCntVal of each picture that is referenced by an entry in RefPicList[ 0 ] or RefPicList[ 1 ] of prevTid0Pic and has the same nuh_layer_id as the current picture,
[0475] - the PicOrderCntVal of each picture that is after prevTid0Pic in the decoding order, has the same nuh_layer_id as the current picture and is before the current picture in the decoding order.
[0476] When there are more than one value in setOfPrevPocVals for which the value modulo MaxPicOrderCntLsb is equal to PocLsbLt[ i ][ j ], the delta_poc_msb present_flag[ i ] [ j ] value shall be equal to 1.
[0477] delta_poc_msb_cycle_ld[i ][ j ] specifies the value of FullPocLt[ i ][ j ] as follows:
[0478]
[0479] The value of delta_poc_msb_cycle_lt[ i ][ j ] shall be in the range of 0 to 2 (32 - log2_max_pic_order_cntisb_minus4 - 4) (inclusive of the end values). When it does not exist, it is inferred that the value of delta_poc_msb_cycle_lt[ i ][ j ] is equal to 0.
[0480] num_ref_idx_active_override_flag being equal to 1 specifies the existence of the syntax element num_ref_idx_active_minus1[0] for P slices and B slices, and the existence of the syntax element num_ref_idx_active_minus1[1] for B slices. num_ref_idx_active_override_flag being equal to 0 specifies the non-existence of the syntax elements num_ref_idx_active_minus1[0] and num_ref_idx_active_minus1[1]. When it does not exist, it is inferred that the value of num_ref_idx_active_override_flag is equal to 1.
[0481] num_ref_idx_active_minus1[ i ] is used to derive the variable NumRefIdxActive[ i ], as specified in Formulas 7 to 98. The value of num_ref_idx_active_minus1[ i ] shall be in the range of 0 to 14 (inclusive of the end values).
[0482] For i equal to 0 or 1, when the current slice is a B slice, num_ref_idx_active_override_flag is equal to 1 and num_ref_idx_active_minus1[ i ] does not exist, it is inferred that num_ref_idx_active_minus1[ i ] is equal to 0.
[0483] When the current slice is a P slice, num_ref_idx_active_override_flag is equal to 1 and num_ref_idx_active_minus1[0] does not exist, it is inferred that num_ref_idx_active_minus1[0] is equal to 0.
[0484] The variable NumRefIdxActive[i] is derived as follows:
[0485]
[0486] The value of NumRefIdxActive[i] - 1 specifies the maximum reference index of reference picture list i that can be used for decoding a slice. When the value of NumRefIdxActive[i] is equal to 0, the reference indices of reference picture list i are not available for decoding the slice.
[0487] When the current slice is a P slice, the value of NumRefIdxActive[0] shall be greater than 0.
[0488] When the current slice is a B slice, both NumRefIdxActive[0] and NumRefIdxActive[1] shall be greater than 0.
[0489] partition_constraints_override_flag being equal to 1 specifies that partition constraint parameters are present in the slice header. partition_constraints_override_flag being equal to 0 specifies that partition constraint parameters are not present in the slice header. When not present, the value of partition_constraints_override_flag is inferred to be equal to 0.
[0490] slicelog2_diff_min_qt_min_cb_luma specifies the difference between the base-2 logarithm of the minimum size in the luma samples of the luma leaf blocks resulting from the quadtree partitioning of a CTU and the base-2 logarithm of the minimum coded block size in the luma samples of the luma CU in the current slice. The value of slice_log2_diff_min_qt_min_cb_luma shall be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive. When not present, the value of slice_log2_diff_min_qt_min_cb_luma is inferred as follows:
[0491] - If slice_type is equal to 2 (I), then the value of slice_log2_diff_min_qt_min_cbluma is inferred to be equal to sps_log2_diff_min_qt_min_cb_intra_slice_luma.
[0492] - Otherwise (slice_type is equal to 0 (B) or 1 (P)), the value of slice_log2_diff_min_qt_min_eb_luma is inferred to be equal to sps_log2_diff_min_qt_min_cb_inter_slice.
[0493] slice_max_mtt_hierarchy_depth_luma specifies the maximum hierarchical depth of the coding units generated by the multi-type tree segmentation of the quadtree leaves in the current slice. The value of slice_max_mtt_hierarchy_depth_luma shall be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY (including the end values). When it is absent, the value of slice_max_mtt_hierarchy_depth_luma is inferred as follows:
[0494] - If slice_type is equal to 2 (I), then the value of slice_max_mtt_hierarchy_depth_luma is inferred to be equal to sps_max_mtt_hierarchy_depth_intra_slice_luma.
[0495] - Otherwise (slice_type is equal to 0 (B) or 1 (P)), the value of slice_max_mtt_hierarchy_depth_luma is inferred to be equal to sps_max_mtt_hierarchy_depth_inter_slice.
[0496] slice_log2_diff_max_bt_min_qt_luma specifies the difference between the base-2 logarithm of the maximum size (width or height) of the luminance samples in the luminance coding blocks that can use binary splitting and the base-2 logarithm of the minimum size (width or height) of the luminance samples in the luminance leaf blocks generated by the quadtree segmentation of the CTUs in the current slice. The value of slicelog2_diff_max_bt_min_qt_luma shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeY (including the end values). When it is absent, the value of slice_log2_diff_max_bt_min_qt_luma is inferred as follows:
[0497] - If slice_type is equal to 2 (I), then the value of slice_log2_diff_max_bt_min_qt_luma is inferred to be equal to sps_log2_diff_max_bt_min_qt_intra_slice_luma.
[0498] - Otherwise (slice_type equals 0 (B) or 1 (P)), infer that the value of slice_log2_diff_max_bt_min_qt_luma is equal to sps_log2_diff_max_bt_min_qt_inter_slice.
[0499] slice_log2_diff_max_tt_min_qt_luma specifies the difference between the base-2 logarithm of the maximum size (width or height) of the luma samples in a luma coded block that can be ternary split and the base-2 logarithm of the minimum size (width or height) of the luma samples in a luma leaf block resulting from a quadtree split of the CTU in the current slice. The value of slice_log2_diff_max_bt_min_qt_luma shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeY, inclusive. When not present, the value of slice_log2_diff_max_bt_min_qt_luma is inferred as follows:
[0500] - If slice_type equals 2 (I), then infer that the value of slice_log2_diff_max_tt_min_qt_luma is equal to sps_log2_diff_max_tt_min_qt_intra_slice_luma.
[0501] - Otherwise (slice_type equals 0 (B) or 1 (P)), infer that the value of slice_log2_diff_max_tt_min_qt_luma is equal to sps_log2_diff_max_tt_min_qt_inter_slice.
[0502] slice_log2_diff_min_qt_min_cb_chroma specifies the difference between the logarithm to the base 2 of the minimum size in the luma samples of the chroma leaf blocks resulting from the quadtree splitting of the chroma CTUs with treeType equal to DUAL_TREE_CHROMA and the logarithm to the base 2 of the minimum coded block size in the luma samples of the chroma CUs with treeType equal to DUAL_TREE_CHROMA in the current slice. The value of slice_log2_diff_min_qt_min_cb_chroma shall be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY (inclusive of the end values). When not present, the value of slice_log2_diff_min_qt_min_cb_chroma is inferred to be equal to sps_log2_diff_min_qt_min_cb_intra_slice_chroma.
[0503] slice_max_mtt_hierarchy_depth_chroma specifies the maximum hierarchical depth of the coding units resulting from the multi-type tree splitting of the quadtree leaves with treeType equal to DUAL_TREE_CHROMA in the current slice. slice_max_mtt_hierarchy_depth_chroma shall be within CtbLog2SizeY - MinCbLog2SizeY (inclusive of the end values). When not present, slice_max_mtt_hierarchy_depth_chroma is inferred to be equal to sps_max_mtt_hierarchy_depth_intra_slices_chroma.
[0504] slice_log2_diff_max_bt_min_qt_chroma specifies the difference between the base-2 logarithm of the maximum size (width or height) of the luma samples in a chroma coded block that can use binary splitting and the base-2 logarithm of the minimum size (width or height) of the luma samples in a chroma leaf block resulting from quadtree splitting of a chroma CTU with treeType equal to DUAL_TREE_CHROMA in the current slice. The value of slice_log2_diff_max_bt_min_qt_chroma shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeC, inclusive. When not present, it is inferred that the value of slice_log2_diff_max_bt_min_qt_chroma is equal to sps_log2_diff_max_bt_min_qt_intra_slice_chroma.
[0505] slice_log2_diff_max_tt_min_qt_chroma specifies the difference between the base-2 logarithm of the maximum size (width or height) of the luma samples in a chroma coded block that can use ternary splitting and the base-2 logarithm of the minimum size (width or height) of the luma samples in a chroma leaf block resulting from quadtree splitting of a chroma CTU with treeType equal to DUAL_TREE_CHROMA in the current slice. The value of slice_log2_diff_max_tt_min_qt_chroma shall be in the range of 0 to CtbLog2SizeY - MinQtLog2SizeC, inclusive. When not present, it is inferred that the value of slice_log2_diff_max_tt_min_qt_chroma is equal to sps_log2_diff_max_tt_min_qt_intra_slice_chroma.
[0506] The variables MinQtLog2SizeY, MmQtLog2Size C, MinQtSizeY, MinQtSizeC, MaxBtSizeY, MaxBtSizeC, MinBtSizeY, MaxTtSizeY, MaxTtSizeC, MinTtSizeY, MaxMttDepthY and MaxMttDepthC are derived as follows:
[0507] MinQtLog2SizeY = MinCbLog2SizeY + slice_log2_diff_min_qt_min_cb_luma
[0508] MinQtLog2SizeC = MinCbLog2SizeY + slice_log2_diff_min_qt_min_cb_chroma
[0509] MinQtSizeY = 1 << MinQtLog2SizeY
[0510] MinQtSizeC = 1 << MinQtLog2SizeC
[0511] MaxBtSizeY = 1 << (MinQtLog2SizeY + slice_log2_diff_max_bt_min_qt_luma)
[0512] MaxBtSizeC = 1 << (MinQtLog2SizeC + slice_log2_diff_max_bt_min_qt_chroma)
[0513] MinBtSizeY = 1 << MinCbLog2SizeY
[0514] MaxTtSizeY = 1 << (MinQtLog2SizeY + slice_log2_diff_max_tt_min_qt_luma)
[0515] MaxTtSizeC = 1 << (MinQtLog2SizeC + slice_log2_diff_max_tt_min_qt_chroma)
[0516] MinTtSizeY = 1 << MinCbLog2SizeY
[0517] MaxMttDepthY = slice_max_mtt_hierarchy_depth_luma
[0518] MaxMttDepthC = slice_max_mtt_hierarchy_depth_chroma
[0519] The slice_temporal_mvp_enabled_flag specifies whether the temporal motion vector predictor can be used for inter prediction. If the slice_temporal_mvp_enabled_flag is equal to 0, the syntax elements of the current picture shall be constrained such that the temporal motion vector predictor is not used in the decoding of the current picture. Otherwise (slice_temporal_mvp_enabled_flag equal to 1), the temporal motion vector predictor can be used in the decoding of the current picture.
[0520] When the slice_temporal_mvp_enabled_flag is not present, the following applies:
[0521] - When the sps_temporal_mvp_enabled_flag is equal to 0, infer that the value of the slice_temporal_mvp_enabled_flag is equal to 0.
[0522] - Otherwise (sps_temporal_mvp_enabled_flag equal to 1), infer that the value of the slice_temporal_mvp_enabled_flag is equal to pps_temporal_mvp_enabled_idc - 1.
[0523] The mvd_l1_zero_flag being equal to 1 indicates that the mvd_coding(x0, y0, 1) syntax structure is not parsed, and MvdLl[x0 ][ y0][ compIdx ] and MvdL1[x0 ] [ y0 ] [ cpIdx ] [ compIdx ] are set to be equal to 0, where compldx = 0..1 and cpldx = 0..2. The mvd_l1_zero_flag being equal to 0 indicates that the mvd_coding(x0, y0, 1) syntax structure is parsed. When it is not present, infer that the value of the mvd_l1_zero_flag is equal to pps_mvd_l1_zero_idc - 1.
[0524] The cabac_init_flag specifies the method for determining the initialization table used in the initialization process for context variables. When the cabac_init_flag is not present, infer that it is equal to 0.
[0525] A collocated_from_l0_flag equal to 1 specifies that the collocated picture used for temporal motion vector prediction is derived from reference picture list 0. A collocated_from_l0_flag equal to 0 specifies that the collocated picture used for temporal motion vector prediction is derived from reference picture list 1. When the collocated_from_l0_flag is not present, the following applies:
[0526] - If the slice_type is not equal to B, then the value of collocated_from_l0_flag is inferred to be equal to 1.
[0527] - Otherwise (slice_type equal to B), the value of collocated_from_l0_flag is inferred to be equal to pps_collocated_from_l0_idc - 1.
[0528] collocated_ref_idx specifies the reference index of the collocated picture used for temporal motion vector prediction.
[0529] When slice_type is equal to P or when slice_type is equal to B and collocated_from_l0_flag is equal to 1, collocated_ref_idx refers to a picture in list 0, and the value of collocated_ref_idx should be in the range of 0 to NumRefldxActive[0] - 1 (inclusive of the end values).
[0530] When slice_type is equal to B and collocated_from_10_flag is equal to 0, collocated_ref_idx refers to a picture in list 1, and the value of collocated_ref_idx should be in the range of 0 to NumRefldxActive[1] - 1 (inclusive of the end values).
[0531] When collocated_ref_idx is not present, the value of collocated_ref_idx is inferred to be equal to 0.
[0532] Bitstream conformance requires that for all slices of an encoded picture, the picture referred to by collocated_ref_idx should be the same. Bitstream conformance requires that the resolution of the reference picture referred to by collocated_ref_idx and the current picture should be the same.
[0533] six_minus_max_num_merge_cand specifies subtracting from 6 the maximum number of merge motion vector prediction (MVP) candidates supported in a slice. The maximum number of merge MVP candidates, MaxNumMergeCand, is derived as follows:
[0534] MaxNumMergeCand = 6 - six_minus_max_num_merge_cand
[0535] The value of MaxNumMergeCand shall be in the range of 1 to 6, inclusive. When not present, it is inferred that the value of six_minus_max_num_merge_cand is equal to pps_six_minus_max_num_merge_cand_plus1 - 1.
[0536] five_minus_max_num_subblock_merge_cand specifies subtracting from 5 the maximum number of block-based merge motion vector prediction (MVP) candidates supported in a slice.
[0537] When five_minus„max_num_subblock_merge_cand is not present, the following applies:
[0538] - If sps_affine_enabled_flagg is equal to 0, it is inferred that the value of five_minus_max_num subblockmerge_cand is equal to 5 - (sps_sbtmvp_enabled_flag && slice_temporal_mvp_enabled_flag).
[0539] - Otherwise ((sps_affine_enabled_flag is equal to 1), it is inferred that the value of five_minus_max__num_subblock_merge_cand is equal to pps_five_minus _max_nurn_subbIock_mcrge_cand plus1 - 1.
[0540] The maximum number of block-based merge MVP candidates, MaxNumSubblockMergeCand, is derived as follows:
[0541] MaxNumSubblockMergeCand = 5 - five_minus_max_num_subblock_merge_cand
[0542] The value of MaxNumSubblockMergeCand shall be in the range of 0 to 5 (including the end values).
[0543] slice_fpel_mmvd_enabled_flag being equal to 1 specifies that the merge mode with motion vector difference in the current slice uses integer sample precision. slice_fpel_mmvd_enabled_flag being equal to 0 specifies that the merge mode with motion vector difference in the current slice can use fractional sample precision. When it does not exist, it is inferred that the value of slice_fpel_mmvd_enabled_flag is equal to 0.
[0544] slice_disable_bdof_dmvr_flag being equal to 1 specifies that either bidirectional optical flow inter prediction or inter dual prediction based on decoder motion vector correction is enabled in the current slice. slice_disableJbdof_dmvr_flag being equal to 0 specifies that bidirectional optical flow inter prediction and inter dual prediction based on decoder motion vector correction may or may not be enabled in the current slice. When slice_disable_bdof_dmvr_flag does not exist, it is inferred that the value of slice_disable_bdof_dmvr_flag is equal to 0.
[0545] max_num_merge_cand_minus_max_num_triangle_cand specifies subtracting the maximum number of triangle merge mode candidates supported in the slice from MaxNumMergeCand.
[0546] When max_num_merge_cand_rninus_max_num_triangle_cand does not exist, sps_triangle_enabled_flag is equal to 1, and MaxNumMergeCand is greater than or equal to 2, it is inferred that max_num_merge_cand_minus_max_num_triangle_cand is equal to pps_max_num_merge_cand_minus_max_num_triangle_cand_minus1 + 1. The maximum number of triangle merge mode candidates MaxNumTriangleMergeCand is derived as follows:
[0547] MaxNumTriangleMergeCand =
[0548] MaxNumMergeCand - max_num_merge_cand_minus_max_num_triangle_cand
[0549] When max_num_merge_cand_minus_max_num_triangle_cand exists, the value of MaxNumTriangleMergeCand shall be in the range from 2 to MaxNumMergeCand (including the end values).
[0550] When max_num_merge_cand_minus_max_num_triangle_cand does not exist (sps_triangle_enabled_flag equals 0 or MaxNumMergeCand is less than 2), MaxNumTriangleMergeCand is set to be equal to 0.
[0551] When MaxNumTriangleMergeCand equals 0, the current slice is not allowed to use the triangular merge mode.
[0552] slice_six_minus_max_num_ibc_merge_cand specifies subtracting from 6 the maximum number of Inter Block Copy (IBC) merge Block Vector Prediction (BVP) candidates supported in the slice. The maximum number of IBC merge BVP candidates, MaxNumlbcMergeCand, is derived as follows:
[0553] MaxNumlbcMergeCand = 6 - slice_six_minus_max_num_ibc_merge_cand
[0554] The value of MaxNumlbcMergeCand shall be in the range from 1 to 6 (including the end values).
[0555] The slice_joint_cbcr_sign_flag specifies whether the co-located residual samples of the two chrominance components have opposite signs in a transform unit with tu_joint_cbcr_residual_flag[x0][y0] equal to 1. When tu_joint_cbcr_residual_flag[x0][y0] of the transform unit is equal to 1, slice_joint_cbcr_sign_flag equal to 0 specifies that the sign of each residual sample of the Cr (or Cb) component is the same as the sign of the co-located Cb (or Cr) residual sample, and slice_joint_cbcr_sign_flag equal to 1 specifies that the sign of each residual sample of the Cr (or Cb) component is the opposite sign of the co-located Cb (or Cr) residual sample.
[0556] The slice_qp_delta specifies the initial value of the Qp Y to be used for the coded blocks in the slice until modified by the value of CuQpDeltaVal in the coding unit layer. The initial value of the slice's Q py quantization parameter SliceQ PY is derived as follows:
[0557] SliceQ PY = 26 + init_qp_minus26 + slice_qp_delta
[0558] SliceQ PY shall be in the range of –QpBdOffset Y to +63 (including the end values).
[0559] When determining the value of Qp' Cb the slice_cb_qp_offset specifies the difference value to be added to the value of pps_cb_qp_offset. The value of slice_cb_qp_offset shall be in the range of -12 to +12 (including the end values). When slice_cb_qp_offset does not exist, it is inferred to be equal to 0. The value of pps_cb_qp_offset + slice_cb_qp_offset shall be in the range of -12 to +12 (including the end values).
[0560] When determining the value of Qp' CrWhen quantizing the value of the parameter, slice_cr_qp_offset specifies the difference value to be added to the value of pps_cb_qp_offset. The value of slice_cr_qp_offset shall be in the range of -12 to +12 (including the end values). When slice_cr_qp_offset does not exist, it is inferred to be equal to 0. The value of pps_cr_qp_offset + slice_cr_qp_offset shall be in the range of -12 to +12 (including the end values).
[0561] When determining the value of Qp' CbCr slice_joint_cbcr_qp_offset specifies the difference value to be added to the value of ppsjoint_cbcr_qp_offset. The value of slice_joint_cbcr_qp_offset shall be in the range of -12 to +12 (including the end values). When slice_joint_cbcr_qp_offset does not exist, it is inferred to be equal to 0. The value of ppsjoint_cbcr_qp_offset + slice_joint_cbcr_qp_offset shall be in the range of -12 to +12 (including the end values).
[0562] slice_sao_luma_flag being equal to 1 specifies that SAO is enabled for the luma component in the current slice, and slice_sao_luma_flag being equal to 0 specifies that SAO is disabled for the luma component in the current slice. When slice_sao_luma_flag does not exist, it is inferred to be equal to 0.
[0563] slice_sao_chroma_flag being equal to 1 specifies that SAO is enabled for the chroma component in the current slice, and slice_sao_chroma_flag being equal to 0 specifies that SAO is disabled for the chroma component in the current slice. When slice_sao_chroma_flag does not exist, it is inferred to be equal to 0.
[0564] slice_alf_enabled_flag being equal to 1 specifies that the adaptive loop filter is enabled in the slice and can be applied to the Y, Cb, or Cr color components. slice_alf_enabled_flag being equal to 0 specifies that the adaptive loop filter is disabled for all color components in the slice.
[0565] slice_num_alf_aps_ids_luma specifies the number of ALF APSs referenced by the slice. The value of slice_num_alf_aps_ids_luma shall be in the range of 0 to 7 (including the end values).
[0566] slice_alf_aps_id_luma[ i ] specifies the adaptation_parameter_set_id of the i-th ALF APS referenced by the luma component of the slice. The TemporalId of the APS NAL unit having aps_params_type equal to ALF_APS and adaptation_parameter_set_id equal to slice_alf_aps_id_luma[ i ] shall be less than or equal to the TemporalId of the coded slice NAL unit.
[0567] For intra slices and slices in IRAP pictures, slice_alf_aps_id_luma[ i ] shall not reference an ALF APS associated with a picture other than the picture containing the intra slice or the IRAP picture.
[0568] slice_alf_chroma_idc equal to 0 specifies that the adaptive loop filter is not applicable to the Cb and Cr color components. slice_alf_chroma_idc equal to 1 indicates that the adaptive loop filter is applicable to the Cb color component. slice_alf_chroma_idc equal to 2 indicates that the adaptive loop filter is applicable to the Cr color component. slice_alf_chroma_idc equal to 3 indicates that the adaptive loop filter is applicable to both the Cb and Cr color components. When slice_alf_chroma_idc is absent, it is inferred to be equal to 0.
[0569] slice_alf_aps_id_chroma specifies the adaptation_parameter_set_id of the ALF APS referenced by the chroma component of the slice. The TemporalId of the APS NAL unit having aps_params_type equal to ALF_APS and adaptation_parameter_set_id equal to slice_alf_aps_id_chroma shall be less than or equal to the TemporalId of the coded slice NAL unit.
[0570] For intra slices and slices in IRAP pictures, slice_alf_aps_id_chroma shall not reference an ALF APS associated with a picture other than the picture containing the intra slice or the IRAP picture.
[0571] The dep_quant_enabled_flag being equal to 0 specifies that dependency quantization is disabled. The dep_quant_enabled_flag being equal to 1 specifies that dependency quantization is enabled. When it does not exist, it is inferred that the value of the dep_quant_enabled_flag is equal to pps_dep_quant_enable_idc - 1.
[0572] The sign_data_hiding_enabled_flag being equal to 0 specifies that sign bit hiding is disabled. The sign_data_hiding_enabled_flag being equal to 1 specifies that sign bit hiding is enabled. When the sign_data_hiding_enabled_flag does not exist, it is inferred that it is equal to 0.
[0573] The deblocking_filter_override_flag being equal to 1 specifies that deblocking parameters exist in the slice header. The deblocking_filter_pverride_flag being equal to 0 specifies that deblocking parameters do not exist in the slice header. When it does not exist, it is inferred that the value of the deblocking_filter_override_flag is equal to 0.
[0574] The slice_deblocking_filter_disabled_flag being equal to 1 specifies that the operation of the deblocking filter should not be applied to the current slice. The slice_deblocking_filter_disabled_flag being equal to 0 specifies that the operation of the deblocking filter should be applied to the current slice. When the slice_deblocking_filter_disabled_flag does not exist, it is inferred that it is equal to pps_deblocking_filter_disabled_flag.
[0575] slice_beta_offset_div2 and slice_tc_offset_div2 specify the deblocking parameter offsets (divided by 2) of 6 and tC for the current slice. The values of slice_beta_offset_div2 and slice_tc_offset_div2 should both be in the range of -6 to 6 (including the end values). When they do not exist, it is inferred that the values of slice_beta_offset_div2 and slice_tc_offset_div2 are equal to pps_beta_offset_div2 and pps_tc_offset_div2 respectively.
[0576] A slice_lmcs_enabled_flag equal to 1 specifies that luminance mapping with chroma scaling is enabled for the current slice. A slice_lmcs_enabled_flag equal to 0 specifies that luminance mapping with chroma scaling is disabled for the current slice. When the slice_lmcs_enabled_flag is absent, it is inferred to be equal to 0.
[0577] slice_lmcs_aps_id specifies the adaptation_parameter_set_id of the LMCS APS referenced by the slice. The TemporalId of the APS NAL unit with an aps_params_type equal to LMCS_APS and an adaptation_parameter_set_id equal to slice_lmcs_aps_id shall be less than or equal to the TemporalId of the coded slice NAL unit. When present, the value of slice_lmcs_aps_id shall be the same for all slices of a picture.
[0578] A slice_chroma_residual_scale_flag equal to 1 specifies that chroma residual scaling is enabled for the current slice. A slice_chroma_residual_scale_flag equal to 0 specifies that chroma residual scaling is not enabled for the current slice. When the slice_chroma_residual_scale_flag is absent, it is inferred to be equal to 0.
[0579] A slice_scaling_list_present_flag equal to 1 specifies that the scaling list data for the current slice is derived based on the scaling list data included in the reference scaling list APS. A slice_scaling_list_present_flag equal to 0 specifies that the scaling list data for the current picture is the default scaling list data derived as specified in Clause 7.4.3.16. When absent, it is inferred that the value of slice_scaling_list_present_flag is equal to 0.
[0580] The slice_scaling_list_aps_id specifies the adaptation_parameter_set_id of the scaling list APS. The TemporalId of an APS NAL unit with an aps_params_type equal to SCALING_APS and an adaptation_parameter_set_id equal to slice_scaling_list_aps_id shall be less than or equal to the TemporalId of the coded slice NAL unit.
[0581] When present, the value of slice_scaling_list_aps_id shall be the same for all slices of a picture.
[0582] When the entry_point_offsets_present_flag is equal to 1, the variable NumEntryPoints that specifies the number of entry points in the current slice is derived as follows:
[0583]
[0584] offset_len_minus1 plus 1 specifies the length, in bits, of the entry point_offset_minusl[ i ] syntax element. The value of offset_len_minusl shall be in the range of 0 to 31, inclusive.
[0585] entry_point_offset__minus1[ i ] plus 1 specifies the i-th entry point offset, in bytes, and is represented by offset_len_minus1 plus 1 bits. The slice data following the slice header consists of NumEntryPoints + 1 subsets, where the subset index values range from 0 to NumEntryPoints, inclusive. The first byte of the slice data is considered byte 0. When present, for the purpose of subset identification, the anti-collision bytes that appear in the slice data section of the coded slice NAL unit are counted as part of the slice data. Subset 0 consists of bytes 0 to entry_point_offset_minus1[ 0 ] (inclusive) of the coded slice data, and subset k (where k is in the range of 1 to NumEntryPoints - 1, inclusive) consists of bytes firstByte[k] to lastByte[k] (inclusive) of the coded slice data, where firstByte[ k ] and lastByte[ k ] are defined as:
[0586]
[0587] The last subset (where the subset index is equal to NumEntryPoints) consists of the remaining bytes of the coded slice data.
[0588] When entropy_coding_sync_enabled_flag is equal to 0, each subset shall consist of all coded bits of all CTUs within the same tile in the slice, and the number of subsets (i.e., the value of NumEntryPoints + 1) shall be equal to the number of tiles in the slice.
[0589] When entropy_coding_sync_enabled_flag is equal to 1, each subset k (where k ranges from 0 to NumEntryPoints, inclusive) shall consist of all coded bits of all CTUs in the CTU rows within a tile, and the number of subsets (i.e., the value of NumEntryPoints + 1) shall be equal to the total number of tile-specific luma CTU rows in the slice.
[0590] slice_header_extension_length specifies the length (in bytes) of the slice header extension data, which does not include the bits used to signal slice_header_extension_length itself. The value of slice_header_extension_length shall be in the range of 0 to 256, inclusive. When not present, the inferred value of slice_header_extension_length is equal to 0.
[0591] slice_header_extension_data_byte can have any value. A decoder compliant with this version of the specification shall ignore the value of slice_header_extension_data_byte. Its value does not affect the decoder's compliance with the profiles specified in this version of the specification.
[0592] Signaling metadata that describes video coding properties provided in JVET-O2001 is suboptimal. Specifically, as described above, in some cases, precisely including an encoded picture may require access units. Additionally, it may be necessary to signal an access unit delimiter at the start of each access unit. Consequently, it may be necessary to signal an access unit delimiter at the start of each picture. As described above, in WET-O2001, the access unit delimiter is encapsulated in a NAL unit. Therefore, signaling an access unit delimiter for each picture requires signaling at least 3 bytes for each picture (i.e., 2 bytes for the NAL unit header and 1 byte for access_unit_delimiter_rbsp()), which can be inefficient, especially for relatively low bitrate applications. Further, as described above, in JVET-02001, the slice design does not include slice segments (i.e., there are no independent / dependent slice segments). However, slices within the same picture can include redundant information. For example, referring to Table 7, in JVET-O2001, for each slice in a picture, the slice_pic_parameter_set_id is signaled. However, for each slice included in the picture, the slice_pic_parameter_set_id needs to be the same. Thus, the slice headers in WET-O2001 may be inefficient as they may include unnecessary redundant information, such as picture information that is the same for all slices within a picture. This disclosure describes techniques for efficiently signaling picture information.
[0593] Figure 1 is a block diagram illustrating an example of a system that can be configured to encode (e.g., encode and / or decode) video data according to one or more techniques of this disclosure. System 100 represents an example of a video data system that can be encapsulated according to one or more techniques of this disclosure. As Figure 1 shown, system 100 includes a source device 102, a communication medium 110, and a destination device 120. In Figure 1 the example shown, source device 102 can include any device configured to encode video data and transmit the encoded video data to communication medium 110. Destination device 120 can include any device configured to receive the encoded video data via communication medium 110 and decode the encoded video data. Source device 102 and / or destination device 120 can include computing devices equipped for wired and / or wireless communication and can include, for example, a set-top box, a digital video recorder, a television, a desktop computer, a laptop computer or tablet, a game console, a medical imaging device, and a mobile device (including, for example, a smart phone, a cellular phone, a personal gaming device).
[0594] The communication medium 110 may include any combination of wireless and wired communication media and / or storage devices. The communication medium 110 may include coaxial cables, fiber optic cables, twisted pair cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other devices that can be used to facilitate communication between various devices and sites. The communication medium 110 may include one or more networks. For example, the communication medium 110 may include a network configured to allow access to the World Wide Web, such as the Internet. The network may operate according to a combination of one or more telecommunications protocols. The telecommunications protocol may include proprietary aspects and / or may include standardized telecommunications protocols. Examples of standardized telecommunications protocols include Digital Video Broadcasting (DVB) standards, Advanced Television Systems Committee (ATSC) standards, Integrated Services Digital Broadcasting (ISDB) standards, Data Over Cable Service Interface Specification (DOCSIS) standards, Global System for Mobile Communications (GSM) standards, Code Division Multiple Access (CDMA) standards, 3rd Generation Partnership Project (3GPP) standards, European Telecommunications Standards Institute (ETSI) standards, Internet Protocol (IP) standards, Wireless Application Protocol (WAP) standards, and Institute of Electrical and Electronics Engineers (IEEE) standards.
[0595] The storage device may include any type of device or storage medium capable of storing data. The storage medium may include tangible or non-transitory computer-readable media. The computer-readable media may include optical discs, flash memories, magnetic memories, or any other suitable digital storage media. In some examples, the memory device or a portion thereof may be described as non-volatile memory, and in other examples, a portion of the memory device may be described as volatile memory. Examples of volatile memory may include Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), and Static Random Access Memory (SRAM). Examples of non-volatile memory may include magnetic hard disks, optical discs, floppy disks, flash memory, or in the form of Electrically Programmable Read-Only Memory (EPROM) or Electrically Erasable and Programmable (EEPROM) memory. The storage device may include memory cards (e.g., Secure Digital (SD) memory cards), internal / external hard disk drives, and / or internal / external solid-state drives. Data may be stored on the storage device according to a defined file format.
[0596] Figure 4 is a conceptual diagram showing examples of components that may be included in a particular implementation of the system 100. In Figure 4 the exemplary particular implementation shown, the system 100 includes one or more computing devices 402A to 402N, a television service network 404, a television service provider site 406, a wide area network 408, a local area network 410, and one or more content provider sites 412A to 412N. Figure 4The specific implementation shown in [Figure] is an example of a system that can be configured to allow digital media content (such as movies, live sports events, etc.) and the data and applications associated therewith, as well as media presentations, to be distributed to and accessed by a plurality of computing devices (such as computing devices 402A through 402N). In Figure 4 In the example shown, computing devices 402A through 402N can include any device configured to receive data from one or more of a television service network 404, a wide area network 408, and / or a local area network 410. For example, computing devices 402A through 402N can be equipped for wired and / or wireless communication, can be configured to receive services over one or more data channels, and can include televisions, including so-called smart TVs, set-top boxes, and digital video recorders. Additionally, computing devices 402A through 402N can include desktop computers, laptop computers or tablet computers, game consoles, mobile devices (including, for example, "smart" phones, cellular phones, and personal gaming devices).
[0597] Television service network 404 is an example of a network configured to enable the distribution of digital media content that can include television services. For example, television service network 404 can include a public over-the-air television network, a public or subscription-based satellite television service provider network, and a public or subscription-based cable television provider network and / or a cloud or Internet service provider. It should be noted that although in some examples, television service network 404 can be primarily used to allow the provision of television services, television service network 404 can also allow the provision of other types of data and services according to any combination of the telecommunications protocols described herein. Additionally, it should be noted that in some examples, television service network 404 can allow two-way communication between a television service provider site 406 and one or more of computing devices 402A through 402N. Television service network 404 can include any combination of wireless and / or wired communication media. Television service network 404 can include coaxial cables, fiber optic cables, twisted pair cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other devices that can be used to facilitate communication between various devices and sites. Television service network 404 can operate according to a combination of one or more telecommunications protocols. Telecommunications protocols can include proprietary aspects and / or can include standardized telecommunications protocols. Examples of standardized telecommunications protocols include DVB standards, ATSC standards, ISDB standards, DTMB standards, DMB standards, Data Over Cable Service Interface Specification (DOCSIS) standards, HbbTV standards, W3C standards, and UPnP standards.
[0598] Referring again to Figure 4, the television service provider site 406 can be configured to distribute television services via the television service network 404. For example, the television service provider site 406 can include one or more broadcast stations, cable television providers, or satellite television providers, or Internet-based television providers. For example, the television service provider site 406 can be configured to receive transmissions (including television programs) via a satellite uplink / downlink. Additionally, as Figure 4 shown, the television service provider site 406 can communicate with the wide area network 408 and can be configured to receive data from the content provider sites 412A to 412N. It should be noted that in some examples, the television service provider site 406 can include a television studio, and the content can originate from the television studio.
[0599] The wide area network 408 can include a packet-based network and operate according to a combination of one or more telecommunications protocols. The telecommunications protocols can include proprietary aspects and / or can include standardized telecommunications protocols. Examples of standardized telecommunications protocols include Global System for Mobile Communications (GSM) standards, Code Division Multiple Access (CDMA) standards, 3rd Generation Partnership Project (3GPP) standards, European Telecommunications Standards Institute (ETSI) standards, European Standards (EN), IP standards, Wireless Application Protocol (WAP) standards, and Institute of Electrical and Electronics Engineers (IEEE) standards, such as one or more IEEE 802 standards (e.g., Wi-Fi). The wide area network 408 can include any combination of wireless and / or wired communication media. The wide area network 408 can include coaxial cables, fiber optic cables, twisted pair cables, Ethernet cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other devices that can be used to facilitate communication between various devices and sites. In one example, the wide area network 408 can include the Internet. The local area network 410 can include a packet-based network and operate according to a combination of one or more telecommunications protocols. The local area network 410 can be distinguished from the wide area network 408 based on access levels and / or physical infrastructure. For example, the local area network 410 can include a secure home network.
[0600] Referring again to Figure 4, content provider sites 412A to 412N represent examples of sites that can provide multimedia content to the television service provider site 406 and / or computing devices 402A to 402N. For example, a content provider site can include a studio having one or more studio content servers configured to provide multimedia files and / or streams to the television service provider site 406. In one example, content provider sites 412A to 412N can be configured to provide multimedia content using an IP suite. For example, a content provider site can be configured to provide multimedia content to a receiver device according to the Real-Time Streaming Protocol (RTSP), HTTP, etc. Additionally, content provider sites 412A to 412N can be configured to provide data including hypertext-based content, etc. to one or more of the receiver devices 402A to 402N and / or the television service provider site 406 via the wide area network 408. Content provider sites 412A to 412N can include one or more web servers. The data provided by data provider sites 412A to 412N can be defined according to a data format.
[0601] Referring again to Figure 1 , source device 102 includes video source 104, video encoder 106, data encapsulator 107, and interface 108. Video source 104 can include any device configured to capture and / or store video data. For example, video source 104 can include a camera and a storage device operably coupled thereto. Video encoder 106 can include any device configured to receive video data and generate a compliant bitstream representing the video data. A compliant bitstream can refer to a bitstream from which a video decoder can receive and reproduce video data. Aspects of the compliant bitstream can be defined according to a video coding standard. When generating a compliant bitstream, video encoder 106 can compress the video data. The compression can be lossy (perceivable or imperceivable to an observer) or lossless. Figure 5 is a block diagram showing an example of a video encoder 500 that can implement the techniques described herein for encoding video data. It should be noted that although the exemplary video encoder 500 is shown as having different functional blocks, such illustrations are for descriptive purposes and do not limit the video encoder 500 and / or its sub-components to a particular hardware or software architecture. The functions of video encoder 500 can be implemented using any combination of hardware, firmware, and / or software in a specific implementation.
[0602] Video encoder 500 can perform intra prediction coding and inter prediction coding of picture regions and can thus be referred to as a hybrid video encoder. In Figure 5In the example shown, video encoder 500 receives a source video block. In some examples, the source video block may include a picture region that has been partitioned according to an encoding structure. For example, the source video data may include macroblocks, CTUs, CBs, their sub-partitions, and / or other equivalent coding units. In some examples, video encoder 500 may be configured to perform additional subdivision of the source video block. It should be noted that the techniques described herein are generally applicable to video coding regardless of how the source video data is partitioned before and / or during encoding. In Figure 5 the example shown, video encoder 500 includes adder 502, transform coefficient generator 504, coefficient quantization unit 506, inverse quantization and transform coefficient processing unit 508, adder 510, intra prediction processing unit 512, inter prediction processing unit 514, filter unit 516, and entropy coding unit 518. As Figure 5 shown, video encoder 500 receives a source video block and outputs a bitstream.
[0603] In Figure 5 the example shown, video encoder 500 may generate residual data by subtracting a predicted video block from the source video block. The selection of the predicted video block is described in detail below. Summing unit 502 represents the component configured to perform this subtraction operation. In one example, the subtraction of video blocks occurs in the pixel domain. Transform coefficient generator 504 applies a transform such as a discrete cosine transform (DCT), discrete sine transform (DST), or conceptually similar transform (e.g., four 8×8 transforms may be applied to a 16×16 array of residual values) to the residual block or its sub-partitions to produce a set of residual transform coefficients. Transform coefficient generator 504 may be configured to perform any and all combinations of the transforms included in the discrete trigonometric transform family, including approximations thereof. Transform coefficient generator 504 may output the transform coefficients to coefficient quantization unit 506. Coefficient quantization unit 506 may be configured to perform quantization of the transform coefficients. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may change the rate distortion (i.e., the relationship between bitrate and video quality) of the encoded video data. The degree of quantization may be modified by adjusting the quantization parameter (QP). The quantization parameter may be determined based on the slice level value and / or the CU level value (e.g., the CU delta QP value). QP data may include any data used to determine the QP for quantizing a particular set of transform coefficients. As Figure 5 shown, the quantized transform coefficients (which may be referred to as level values) are output to inverse quantization and transform coefficient processing unit 508. Inverse quantization and transform coefficient processing unit 508 may be configured to apply inverse quantization and inverse transform to generate reconstructed residual data. As Figure 5As shown, at the adder 510, the reconstructed residual data can be added to the predicted video block. In this way, the encoded video block can be reconstructed, and the resulting reconstructed video block can be used to evaluate the encoding quality of a given prediction, transform, and / or quantization. The video encoder 500 can be configured to perform multiple encoding passes (e.g., perform encoding while changing one or more of the prediction, transform parameters, and quantization parameters). The rate distortion or other system parameters of the bitstream can be optimized based on the evaluation of the reconstructed video block. In addition, the reconstructed video block can be stored and used as a reference for predicting subsequent blocks.
[0604] Referring again to Figure 5 , the intra prediction processing unit 512 can be configured to select an intra prediction mode for the video block to be encoded. The intra prediction processing unit 512 can be configured to evaluate the frame and determine the intra prediction mode for encoding the current block. As described above, the possible intra prediction modes can include a planar prediction mode, a DC prediction mode, and an angular prediction mode. In addition, it should be noted that in some examples, the prediction mode of the chrominance component can be inferred based on the prediction mode of the luminance prediction mode. The intra prediction processing unit 512 can select the intra prediction mode after performing one or more encoding passes. In addition, in one example, the intra prediction processing unit 512 can select the prediction mode based on rate distortion analysis. As Figure 5 shown, the intra prediction processing unit 512 outputs intra prediction data (e.g., syntax elements) to the entropy coding unit 518 and the transform coefficient generator 504. As described above, the transform performed on the residual data can be mode-dependent (e.g., a quadratic transform matrix can be determined based on the prediction mode).
[0605] Referring again to Figure 5 , the inter prediction processing unit 514 can be configured to perform inter prediction coding for the current video block. The inter prediction processing unit 514 can be configured to receive a source video block and calculate the motion vector of the PU of the video block. The motion vector can indicate the displacement of the PU of the video block within the current video frame relative to the predicted block within the reference frame. The inter prediction coding can use one or more reference pictures. In addition, the motion prediction can be uni-directional prediction (using one motion vector) or bi-directional prediction (using two motion vectors). The inter prediction processing unit 514 can be configured to select the predicted block by calculating the pixel difference determined by, for example, the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. As described above, the motion vector can be determined and specified based on motion vector prediction. As described above, the inter prediction processing unit 514 can be configured to perform motion vector prediction. The inter prediction processing unit 514 can be configured to generate a predicted block using the motion prediction data. For example, the inter prediction processing unit 514 can locate the predicted video block within the frame buffer ( Figure 5(not shown in the figure). It should be noted that the inter-frame prediction processing unit 514 may be further configured to apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values for motion estimation. The inter-frame prediction processing unit 514 may output the motion prediction data of the calculated motion vector to the entropy encoding unit 518.
[0606] As Figure 5 shown, the filter unit 516 receives the reconstructed video block and the coding parameters, and outputs the modified reconstructed video data. The filter unit 516 may be configured to perform deblocking and / or sample adaptive offset (SAO) filtering. SAO filtering is a kind of filtering that can be used to improve the reconstructed non-linear amplitude mapping by adding an offset to the reconstructed video data. It should be noted that, as Figure 5 shown, the intra-frame prediction processing unit 512 and the inter-frame prediction processing unit 514 may receive the modified reconstructed video block via the filter unit 516. The entropy encoding unit 518 receives the quantized transform coefficients and the prediction syntax data (i.e., intra-frame prediction data and motion prediction data). It should be noted that, in some examples, the coefficient quantization unit 506 may perform a scan of the matrix including the quantized transform coefficients before outputting the coefficients to the entropy encoding unit 518. In other examples, the entropy encoding unit 518 may perform the scan. The entropy encoding unit 518 may be configured to perform entropy encoding according to one or more of the techniques described herein. In this way, the video encoder 500 represents an example of a device configured to generate encoded video data according to one or more of the techniques of the present disclosure.
[0607] Referring again to Figure 1 , the data encapsulator 107 may receive the encoded video data and generate a compliant bitstream according to a defined data structure, for example, a sequence of NAL units. A device receiving the compliant bitstream may reproduce video data from it. In addition, as described above, sub-bitstream extraction may refer to the process by which a device receiving a bitstream compliant with ITU-T H.265 forms a new bitstream compliant with ITU-T H.265 by discarding and / or modifying data in the received bitstream. It should be noted that the term compliant bitstream may be used instead of the term compliance bitstream. In one example, the data encapsulator 107 may be configured to generate syntax according to one or more of the techniques described herein. It should be noted that the data encapsulator 107 does not necessarily have to be located in the same physical device as the video encoder 106. For example, the functions described as being performed by the video encoder 106 and the data encapsulator 107 may be distributed in the Figure 4 devices shown.
[0608] As described above, it is less than ideal to signal metadata that describes video coding properties provided in JVET-O2001. In one example, according to the techniques herein, one or more syntax elements that are currently signaled in a slice header can be signaled in an access unit delimiter. This can provide bit savings when a picture consists of multiple slices, which improves coding efficiency. Table 9 shows an exemplary syntax for an access unit delimiter, and Table 10 shows an exemplary syntax for a corresponding slice header according to one example of the techniques herein.
[0609]
[0610]
[0611]
[0612] Regarding Table 9, in one example, the semantics can be based on the following:
[0613] The access unit delimiter is used to indicate the start of an access unit and the type of slices in the coded picture that are present in the access unit containing the access unit delimiter NAL unit.
[0614] pic_type indicates the number of slice_type values of all slices of the coded picture in the access unit containing the access unit delimiter NAL unit that are in the set listed for the given value of pic_type in Table 6. In a bitstream conforming to this version of the present specification, the value of pic_type shall be equal to 0, 1, or 2. Other values of pic_type are reserved for future use by ITU-T | ISO / IEC. A decoder conforming to this version of the present specification shall ignore all reserved values of pic_type.
[0615] pic_parameter_set_id specifies the value of pps_pic_parameter_set_id of the PPS in use. The value of pic_paranaeter_set_id shall be in the range of 0 to 63 (inclusive of the end values). Bitstream conformance requires that the value of Temporand of the current picture shall be greater than or equal to the value of Temporand of the PPS that has a pps_pic_parameter_set_id equal to pic_parameter_set_id.
[0616] non_reference_picture_flag equal to 1 specifies that the picture must never be used as a reference picture. non_reference_picture_flag equal to 0 specifies that the picture may or may not be used as a reference picture.
[0617] When separate_colour_plane_flag equals 1, colour_plane_id specifies the colour plane associated with the current picture. The value of colour_plane_id shall be in the range 0 to 2, inclusive. The values 0, 1, and 2 of colour_plane_id correspond to the Y, Cb, and Cr planes, respectively.
[0618] Note - There is no correlation between the decoding processes of pictures with different values of colour_plane_id.
[0619] pic_order_cnt_lsb specifies the picture order count modulo MaxPicOrderCntLsb of the current picture. The length of the slice_pic_order_cnt_lsb syntax element is log2 max pic_order_cnt_lsb_minus4 + 4 bits. The value of slice_pic_order_cnt_lsb shall be in the range 0 to MaxPicOrderCntLsb - 1, inclusive.
[0620] pic_output_flag affects the decoding picture output and removal process as specified. When pic_output_flag is not present, it is inferred to be equal to 1.
[0621] In another example, with respect to the access unit delimiter shown in Table 9, in one example, only some of the syntax elements shown in Table 9 can be signalled in the access unit delimiter. For example, in one example, only the syntax elements non_reference_picture_flag and pic_output_flag can be present in the access unit delimiter.
[0622] In another example, more syntax elements than those shown in Table 9 can be signalled in the access unit delimiter. For example, in one example, any syntax element in the slice header that is the same for all slices of a picture can be signalled in the access unit delimiter.
[0623] Additionally, in one example, the relative order of the syntax elements in the access unit delimiter can be different compared to that shown in Table 9. For example, in one example, pic_parameter_set_id can be the first syntax element to be signalled in the access unit delimiter.
[0624] Regarding Table 9, in one example, the semantics can be based on the semantics provided above with respect to Table 7.
[0625] Referring to Tables 7 and 9, in one example, according to the techniques herein, the syntax element slice_pic_order_cnt_lsb can be the first element in the slice header, i.e., slice_pic_order_crut_lsb can be the syntax element immediately following "slice_header() {.". It should be noted that in IV-ET-02001, the syntax element slice_pic_order_cnt_lsb is signaled for all picture types including TDR pictures. In one example, according to the techniques herein, an additional requirement can be that in a CVS, the PicOrderCntVal values of any two coded pictures having the same nuh_layer_id value should not be the same. Additionally, in one example, in this case, an additional requirement can be that when present, the value of slice_pic_order_cnt_lsb should be the same in all slice headers of a coded picture. It should be noted that moving slice_pic_order_cnt_lsb to the first element in the slice header and requiring slice_pic_order_cnt_lsb to be the same in all slice headers of a coded picture allows the video decoder to check and determine that a change in the value compared to the previous slice indicates that the slice belongs to a new picture.
[0626] In another example, the syntax element slice_pic_order_ent_lsb can be signaled such that it appears in the slice header as an element after the slice_pic_parameter_set_id syntax element and before other syntax elements. In one example, this can be as shown in Table 10A below:
[0627]
[0628] In one example, the steps for detecting a new picture based on the syntax shown in Table 10A, for example, can be as follows:
[0629] - Parse the value of slice_pic_parameter_set_id. This can be ignored or can be used as follows: If it is a value different from the previous slice, then this is the first slice of the picture because all slices of a picture must use the same PPS.
[0630] - Parse the value of slice_pic_order_cnt_lsb. If it is a value different from the previous slice, then this is the first slice of the picture because all slices of a picture must have the same slice_pic_order_cnt_lsb value.
[0631] The following steps can also be employed during new picture detection:
[0632] - At the start of a coded video sequence (CVS), the SPS will be parsed and log2_max_pic_order_cnt_lsb_minus4 will be parsed. Then from this point on, since log2_max_pic_order_cnt_lsb_minus4 is the same for the entire CVS, parsing of slice_pic_order_cnt_lsb can continue without further checking and / or without parsing slice_pic_parameter_set_id in each slice of the CVS.
[0633] In one example, according to the techniques herein, a flag can be added to the SPS that specifies that each picture that references the SPS precisely includes one slice. In one example, the flag can be encoded as a u(1) syntax element. In one example, the semantics of the flag can be based on the following:
[0634] single_slice_in_pic_flag being equal to 1 specifies that each picture that references the SPS precisely includes one tile. single_brick_per_slice_flag being equal to 0 specifies that pictures that reference the SPS can include more than one brick.
[0635] It should be noted that pictures can reference the SPS by referring to the PPS, and the PPS can reference the SPS.
[0636] Regarding Table 3, in one example, according to the techniques herein, single_slice_in_pic_flag can be immediately before the syntax element pic_width_max_in_luma_samples. In one example, single_slice_in_pic_flag can be immediately after the syntax element pic_height_max_in_luma_samples. In another example, single_slice_in_pic_flag can be signaled at different locations in the SPS.
[0637] It should be pointed out that when single_slice_in_pic_flag or another flag or syntax element specifies that each picture that references the SPS precisely includes one slice, the following requirements can be specified:
[0638] If, in the SPS of a picture reference in CVS, single_slice_in_pic_flag is equal to 1, then all access units in CVS consist of zero or one access unit delimiter NAL unit and one or more layer access units (in increasing order of nuh_layer_id).
[0639] Otherwise (single_slice_in_pic_flag equal to 0), all access units in CVS consist of one access unit delimiter NAL unit and one or more layer access units (in increasing order of nuh_layer_id).
[0640] In addition, in one example, single_slice_in_pic_flag is included in the SPS and is before the syntax element subpics_present_flag, and then the presence of subpics_present_flag can depend on the value of single_slice_in_pic_flag such that when single_slice_in_pic_flag is equal to 1, it does not need to be signaled. This can be achieved, for example, using the following syntax, as follows:
[0641]
[0642] In the semantics of subpics_present_flag, the following can be added: When absent, it is inferred that the value of subpics_present_flag is equal to 0.
[0643] In addition, when single_slice_in_pic_flag etc. are included in the SPS, in the slice header, when it is ensured that there is only one slice, it is not necessary to signal the slice address. That is, in one example, the presence of the syntax element slice_address can be adjusted as follows:
[0644]
[0645] It should be noted that adjusting the presence of the syntax element slice_address on single_slice_in_pic_flag may result in bit savings and ensure that single_slice_in_pic_flag is not a non - canonical syntax element.
[0646] In addition, in one example, when single_slice_in_pic_flag etc. are included in the SPS, in the slice header, the presence of the syntax element num_bricks_in_slice_minus1 can be adjusted, and the value of this syntax element is inferred based on single_slice_in_pic_flag. For example, num_bricks_in_slice_minus1 can be signaled conditionally as follows:
[0647]
[0648] The semantics are as follows:
[0649] num_bricks_in_slice_minus1, when present, specifies the number of bricks in the slice minus 1. The value of num_bricks_in_slice_minus1 should be in the range of 0 to NumBricksInPic - 1 (including the end values). When rect_slice_flag is equal to 0 and single_brick_per_slice_flag is equal to 1, the value of num_bricks_in_slice_minus1 is inferred to be equal to 0. When single_brick_per_slice_flag is equal to 1, the value of num_bricks_in_slice_minus1 is inferred to be equal to 0. When single_slice_in_pic_flag is equal to 1, the value of num_bricks_in_slice_minus1 is inferred to be equal to NumBrickslnPic - 1.
[0650] It should be noted that based on the value of single_slice_in_pic_flag being equal to 1, there can be other ways to describe and / or specify that the inference of the value of num_bricks_in_slice_minus1 is equal to NumBricksInPic - 1.
[0651] In addition, in one example, when single_slice_in_pic_flag etc. are included in the PPS and in the SPS, the presence of the syntax elements num_slices_in_pic_minus1, bottom_right_brick_idx_length_nainus1, bottom_right_brick_idx_deltan[ i ], brick_idx_delta_sign_flag[ i ] can be adjusted, and the values are inferred based on single_slice_in_ouc_flag. For example, bottom_right_brick_idx_length_minus1, bottom_right_brick_iclx_deltafil, brick_idx_delta_sign_flag[ i ] can be signaled conditionally as follows:
[0652]
[0653] The following modifications are made to the semantics with respect to the semantics provided above with respect to Table 4:
[0654] num_slices_in_pic_minus1 plus 1 specifies the number of slices in each picture that reference the PPS. The value of num_slices_in_pic_minus1 should be in the range of 0 to NumBrickslnPic - 1 (including the end values). When it is absent and single_brick_per_slice_flag is equal to 1, the value of num_slices_in_pic_minus1 is inferred to be equal to NumBrickslnPic - 1.
[0655] When it is absent and single_slice_in_pic_flag is equal to 1, the value of num_slices_in_pic_minus1 is inferred to be equal to 0.
[0656] It should be noted that there may be other syntax elements in the PPS, where the presence or derived values of these other syntax elements depend on the value of single_slice_in_pic_flag.
[0657] In one example, a flag can be added to the PPS that specifies that each picture referencing the PPS exactly includes one slice. In one example, the flag can be encoded as a u(1) syntax element. In one example, the semantics of the flag can be based on the following:
[0658] A single_slice_inpic_flag equal to 1 specifies that each picture that references this PPS precisely includes one tile. A single_slice_in_pic_flag equal to 0 specifies that the pictures that reference this PPS may include more than one tile.
[0659] Regarding Table 4, in one example, according to the techniques herein, single_slice_inpic_flag may immediately precede the syntax element single_tide_in_pic_flag. In one example, single_slice_in_pic_flag may immediately follow the syntax element single_tile_in_pic_flag. It should be noted that when single_slice_in_pic_flag is included in the PPS, compliance constraints may be defined such that the same single_slice_inpic_flag value is provided for all PPSs in the CVS. Since it is not possible to code multiple sub - pictures and also not possible to code multiple slices, in one example, it is also proposed to signal subpic_present_flag only when single_slice_in_pic_flag is equal to 0.
[0660]
[0661] Referring to Table 9, in one example, according to the techniques herein, access_unit_delimiter_rbsp() may include a syntax element that specifies the value of vps_video_parameter_set_id of the VPS referenced by the access unit delimiter. Table 11 shows an example of access_unit_delimiter_rbsp() that includes a syntax element that specifies the value of vps_video_parameter_set_id of the VPS referenced by the access unit delimiter.
[0662]
[0663] Regarding Table 11, the semantics may be based on the semantics provided above with respect to Table 9, where the semantics of the syntax element aud_video_parameter_set_id are based on the following:
[0664] When aud_video_parameter_set_id is greater than 0, it specifies the value of vps_video_parameter_set_id of the VPS referenced by the access unit delimiter. When aud_video_parameter_set_id is equal to 0, the access unit delimiter does not reference a VPS, and when decoding a slice, no VPS is referenced. The value of aud_video_parameter_set_id shall be equal to the value of sps_video_parameter_set_id in the SPS where sps_decoding_parameter_set_id is equal to pps_seq_parameter_set_id of each slice of the VCL NAL unit in the access unit.
[0665] Alternatively, in one example, syntax elements that are the same for all VCL NAL units of an access unit can be moved to the access unit delimiter. This can include, for example, the syntax elements: pic_order_cnt_lsb, colour_plane_id, non_reference_picture_flag. In this case, the check if (vps_max_layers_minus1 == 0 || aud_video_parameter_set_id == 0) is not performed.
[0666] In one example, according to the techniques herein, the order of NATs, units, and coded pictures and their association with layer access units and access units can be based on the following:
[0667] A layer access unit consists of one coded picture and zero or more non-VCL NAL units. An access unit consists of an access unit delimiter NAL unit and one or more layer access units (in ascending order of nuh_layer_id).
[0668] The access unit delimiter NAL unit shall have an nuh_layer_id value equal to vps_layer_id[0].
[0669] In another example: The access unit delimiter NAL unit shall have an nuh_layer_id value equal to 0.
[0670] The access unit delimiter NAL unit shall have a TemporalId equal to the TemporalId of the layer access unit that has an nuh_layer_id equal to vps_layer_id[0].
[0671] In another example, an access unit delimiter NAL unit shall have a TemporalId equal to the TemporalId of the access unit that contains the NAL unit.
[0672] Additionally, in one example, the TemporalId of an access unit may be defined as one of the following options:
[0673] 1. For all VCL NAL units of an access unit, bitstream conformance may require that the TemporalId of the access unit be the same. The value of the TemporalId of an encoded picture or access unit is the value of the TemporalId of the VCL NAL units of the encoded picture or access unit.
[0674] 2. Different TemporalIds are allowed for different layer access unit VCL NAL units. The minimum or maximum TemporalId value of the VCL NAL units of a layer access unit is defined as the value of the TemporalId of the access unit.
[0675] The first access unit in the bitstream starts from the first NAL unit of the bitstream. Each access unit starts from an access unit delimiter NAL unit. There shall be at most one access unit delimiter NAL unit in any access unit.
[0676] The order of coded pictures and non-VCL NAL units within a layer access unit or an access unit shall follow the following constraints:
[0677] - When any DPS NAL unit, VPS NAL unit, SPS NAL unit, PPS NAL unit, APS NAL unit, prefix SET NAL unit, NAL unit with nal_unit_type in the range of RSV_NVCL_25..RSV_NVCL_26, or NAL unit with nal_unit_type in the range of UNSPEC28..UNSPEC29 exists in a layer access unit, they shall not be after the last VCL NAL unit of the layer access unit.
[0678] - In a layer access unit, a NAL unit with nal_unit_type equal to SUFFIX_SET_NUT or RSV_NVCL_27 or in the range of UNSPEC30..UNSPEC31 shall not be before the first VCL NAL unit of the layer access unit.
[0679] - When the end of the sequence NAL unit exists in an access unit, this unit shall be the last NAL unit among all NAL units in the access unit except the end of the bitstream NAL units (when present).
[0680] - When the end of the bitstream NAL unit is present in the access unit, the unit shall be the last NAL unit in the access unit.
[0681] Referring to Table 10, in one example, according to the techniques herein, slice_header() may conditionally signal syntax elements based on vps_max_layers_minus1. Table 12 shows an example of the slice_header() syntax structure that conditionally signals syntax elements based on vps_max_layers_minus1.
[0682]
[0683]
[0684] In one example, the condition if(vps_max_layers_minus1>0) in Table 12 above may be replaced with:
[0685] if(vps_max_layers_minus1>0 && aud_video_parameter_set_id !=0)
[0686] In one example, a syntax element may be included in the access unit delimiter syntax structure and may be introduced into the access unit delimiter, which specifies whether certain syntax elements are present in the access unit delimiter syntax structure or in the slice header syntax structure. Tables 13 and 14 show examples of the access unit delimiter syntax structure and the corresponding slice header syntax structure, where the syntax element au_lpicture_header_flag specifies whether certain syntax elements are present in the access unit delimiter syntax structure or in the slice header syntax structure.
[0687]
[0688]
[0689]
[0690] Regarding Tables 13 and 14, the semantics may be based on the semantics provided above, where the semantics of the syntax element aud_picture_header_flag are based on the following:
[0691] When aud_picture_header_flag is equal to 1, it specifies that pic_paramter_set_id, non_reference_picture_flag, colour_plane_id, pic_order_count_lsb, and pic_output_flag are present in the access unit delimiter (AUD). When aud_picture_header_flag_equal is equal to 0, it specifies that these syntax elements are present in the slice header of the picture.
[0692] In one example, it may be required that aud_picture_header_flag be equal to 0 when vps_max_layers_minus1 is greater than 0.
[0693] In another example:
[0694] When aud_picture_heacler_flag is equal to 1, it specifies that pic_paramter_set_id, non_reference_picture_flag, colour_plane_id (when separate_colour_plane_flag is equal to 1), pic_order_count_lsb, and pic_output_flag (when output_flag_present_flag is equal to 1) are present in the AUD. When aud_picture_header_flag is equal to 0, it specifies that these syntax elements are present in the slice header of the picture.
[0695] In another example, a syntax structure called pic_header0 can be signaled. The syntax structure can include only syntax elements that need to have the same value in all slices of the picture. In one example, pic_header() can have the following syntax in Table 15 and semantics.
[0696]
[0697] When the pic_header0 syntax structure exists in the access unit delimiter, the term current picture refers to the picture in the access unit with a nuh_layer_id equal to vps_layer_id[0]. When the pic_header() syntax structure exists in the slice header, the term current picture refers to the picture of which the current slice is a part. The ph_pic_parameter_set_id specifies the value of pps_pic_parameter_set_id of the PPS for the slice of the current picture. The value of slice_ph_pic_parameter_set_id shall be in the range of 0 to 63, inclusive.
[0698] Bitstream conformance requires that the value of TemporalId of the current picture shall be greater than or equal to the value of TemporalId of the PPS that has a pps_pic_parameter_set_id equal to slice_pic_parameter_set_id. A non_reference_picture_flag equal to 1 specifies that the current picture shall never be used as a reference picture. A non_reference_picture_flag equal to 0 specifies that the current picture may or may not be used as a reference picture.
[0699] When separate_colour_plane_flag is equal to 1, colour_plane_id specifies the colour plane associated with the current picture. The value of colour_plane_id shall be in the range of 0 to 2, inclusive. The colour_plane_id values 0, 1, and 2 correspond to the Y, Cb, and Cr planes, respectively.
[0700] Note - There is no correlation between the decoding processes of pictures with different colour_plane_id values.
[0701] ph_pic_order_cnt_lsb specifies the picture order count modulo MaxPicOrderCntLsb of the current picture. The length of the slice_pic_order_cnt_lsb syntax element is log2_max_pic_order_cnt_lsb_minus4 + 4 bits. The value of ph_pic_order_cnt_lsb shall be in the range of 0 to MaxPicOrderCntLsb - 1, inclusive.
[0702] The pic_output_flag affects the decoding picture output and removal process as specified. When pic_output_flag does not exist, it is inferred to be equal to 1.
[0703] In another example, the syntax element aud_pic_header_flag can be added to the AUD NAL unit syntax structure as provided in Table 16 below. When aud_pic_header_flag is equal to 1, the pic_header() syntax structure is present in the AUD NAL unit (and not present in the slice headers of that picture).
[0704]
[0705] In one example, the following semantics can be used for the syntax element aud_pic_header_flag.
[0706] auci_pic_header_flag being equal to 1 specifies that the pic_header() syntax structure is present in the AUD. aud_pic_header_flag being equal to 0 specifies that the pic_header() syntax structure is present in all slice headers of the access unit. When the AUD NAL unit is not present in the access unit, it is inferred that the value of aud_pic_header_flag is equal to 0.
[0707] The slice headers can be modified as provided in Table 17:
[0708]
[0709] In one example, it may be desired that aud_picture_header_flag is equal to 0 when vps_max_layers_minus1 is greater than 0.
[0710] In this way, the source device 102 represents an example of a device that is configured to: signal a flag in the sequence parameter set that has a value indicating whether each picture that references the sequence parameter set exactly includes one slice; and conditionally signal one or more syntax elements in the slice header based on the value of that flag.
[0711] Refer again to Figure 1, Interface 108 can include any device configured to receive data generated by data encapsulator 107 and transmit and / or store the data to a communication medium. Interface 108 can include a network interface card such as an Ethernet card, and can include an optical transceiver, a radio frequency transceiver, or any other type of device that can send and / or receive information. Additionally, interface 108 can include a computer system interface that can enable files to be stored on a storage device. For example, interface 108 can include support for Peripheral Component Interconnect (PCI) and Peripheral Component Interconnect Express (PCIe) bus protocols, proprietary bus protocols, Universal Serial Bus (USB) protocols, I 2 C chipset, or any other logical and physical structure that can be used to interconnect peer devices.
[0712] Referring again to Figure 1 , target device 120 includes interface 122, data de-encapsulator 123, video decoder 124, and display 126. Interface 122 can include any device configured to receive data from a communication medium. Interface 122 can include a network interface card such as an Ethernet card, and can include an optical transceiver, a radio frequency transceiver, or any other type of device that can receive and / or send information. Additionally, interface 122 can include a computer system interface that allows retrieval of a compliant video bitstream from a storage device. For example, interface 122 can include support for PCI and PCIe bus protocols, proprietary bus protocols, USB protocols, I 2 C chipset, or any other logical and physical structure that can be used to interconnect peer devices. Data de-encapsulator 123 can be configured to receive and parse any of the example syntax structures described herein.
[0713] Video decoder 124 can include any device configured to receive a bitstream (e.g., sub-bitstream extraction) and / or an acceptable variant thereof and reproduce video data therefrom. Display 126 can include any device configured to display video data. Display 126 can include various display devices such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or one of another type of display. Display 126 can include a high definition display or an ultra high definition display. It should be noted that although in Figure 1 the example shown, video decoder 124 is described as outputting data to display 126, video decoder 124 can be configured to output video data to various types of devices and / or its sub-components. For example, video decoder 124 can be configured to output video data to any communication medium, as described herein.
[0714] Figure 6FIG. is a block diagram illustrating an example of a video decoder that may be configured to decode video data according to one or more techniques of the present disclosure (e.g., for the decoding process of the reference picture list construction described above). In one example, video decoder 600 may be configured to decode transform data and reconstruct residual data from transform coefficients based on the decoded transform data. Video decoder 600 may be configured to perform intra prediction decoding and inter prediction decoding, and may thus be referred to as a hybrid decoder. Video decoder 600 may be configured to parse any combination of the syntax elements described above in Tables 1 to 17. Video decoder 600 may decode pictures based on or according to the above processes and also based on the parsed values in Tables 1 to 17.
[0715] In Figure 6 the example shown, video decoder 600 includes an entropy decoding unit 602, an inverse quantization unit 604, an inverse transform coefficient processing unit 606, an intra prediction processing unit 608, an inter prediction processing unit 610, a summer 612, a post-filter unit 614, and a reference buffer 616. Video decoder 600 may be configured to decode video data in a manner consistent with a video coding system. It should be noted that although the illustrated exemplary video decoder 600 has different functional blocks, such illustration is for descriptive purposes and does not limit video decoder 600 and / or its sub-components to a particular hardware or software architecture. The functions of video decoder 600 may be implemented using any combination of hardware, firmware, and / or software implementations.
[0716] As Figure 6 shown, entropy decoding unit 602 receives an entropy-coded bitstream. Entropy decoding unit 602 may be configured to decode syntax elements and quantization coefficients from the bitstream according to a process inverse to the entropy coding process. Entropy decoding unit 602 may be configured to perform entropy decoding according to any of the entropy coding techniques described above. Entropy decoding unit 602 may determine the values of the syntax elements in the encoded bitstream in a manner consistent with a video coding standard. As Figure 6 shown, entropy decoding unit 602 may determine quantization parameters, quantization coefficient values, transform data, and prediction data from the bitstream. In this example, as Figure 6 shown, inverse quantization unit 604 and inverse transform coefficient processing unit 606 receive quantization parameters, quantization coefficient values, transform data, and prediction data from entropy decoding unit 602 and output reconstructed residual data.
[0717] Referring again to Figure 6, the reconstructed residual data can be provided to an adder 612. The adder 612 can add the reconstructed residual data to a predicted video block and generate reconstructed video data. The predicted video block can be determined according to prediction video techniques (i.e., intra prediction and inter prediction). The intra prediction processing unit 608 can be configured to receive intra prediction syntax elements and retrieve a predicted video block from a reference buffer 616. The reference buffer 616 can include a memory device configured to store one or more video data frames. The intra prediction syntax elements can identify intra prediction modes, such as the intra prediction modes described above. The inter prediction processing unit 610 can receive inter prediction syntax elements and generate motion vectors to identify a predicted block in one or more reference frames stored in the reference buffer 616. The inter prediction processing unit 610 can generate a motion compensated block, and may perform interpolation based on an interpolation filter. An identifier of the interpolation filter for motion estimation with sub-pixel accuracy can be included in the syntax elements. The inter prediction processing unit 610 can use the interpolation filter to calculate an interpolated value of sub-integer pixels of a reference block. The post-filter unit 614 can be configured to perform filtering on the reconstructed video data. For example, the post-filter unit 614 can be configured to perform deblocking and / or sample adaptive offset (SAO) filtering, e.g., based on parameters specified in a bitstream. In addition, it should be noted that in some examples, the post-filter unit 614 can be configured to perform any dedicated filtering (e.g., visual enhancement such as mosquito noise cancellation). As Figure 6 shown, the video decoder 600 can output a reconstructed video block. In this way, the video decoder 600 represents an example of a device configured to: parse a flag in a sequence parameter set that has a value indicating whether each picture that references the sequence parameter set exactly includes one slice; and conditionally parse one or more syntax elements in a slice header based on the value of the flag.
[0718] In one or more examples, the functions may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium can include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a propagation medium including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. Thus, the computer-readable medium generally can correspond to: (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or a carrier wave. The data storage medium can be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product can include a computer-readable medium.
[0719] By way of example, and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that is accessible by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave from a website, server, or other remote source, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but are directed to non-transitory, tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0720] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. In addition, in some aspects, the functions described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Moreover, these techniques can be implemented entirely within one or more circuits or logic elements.
[0721] The techniques of the present disclosure can be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or a group of ICs (e.g., a chipset). Various components, modules, or units are described in the present disclosure to emphasize functional aspects of devices configured to perform the disclosed techniques, but need not necessarily be implemented by different hardware units. Instead, as described above, the various units can be combined in a codec hardware unit, or provided by a collection of one or more processors as described above in interoperating hardware units, in conjunction with appropriate software and / or firmware.
[0722] In addition, each functional block or various features of the base station device and the terminal device used in each of the above embodiments can be implemented or executed by a circuit (usually one integrated circuit or a plurality of integrated circuits). The circuit designed to execute the functions described in this specification may include a general-purpose processor, a digital signal processor (DSP), an application-specific or general-purpose integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic, or discrete hardware components, or a combination thereof. The general-purpose processor may be a microprocessor, or alternatively, the processor may be a conventional processor, a controller, a microcontroller, or a state machine. The general-purpose processor or each of the above circuits may be configured by a digital circuit or may be configured by an analog circuit. In addition, when a technology for manufacturing an integrated circuit that replaces the current integrated circuit emerges due to the progress of semiconductor technology, it is also possible to use the integrated circuit produced by this technology.
[0723] Various examples have been described. These examples and other examples are within the scope of the following claims.
[0724] <Overview>
[0725] In one example, a method of signaling picture information is provided, the method comprising: signaling a flag in a sequence parameter set, the flag having a value indicating whether each picture that refers to the sequence parameter set exactly includes one slice; and conditionally signaling one or more syntax elements in a slice header based on the value of the flag.
[0726] In one example, a method of decoding picture information for decoding video data is provided, the method comprising: parsing a flag in a sequence parameter set, the flag having a value indicating whether each picture that refers to the sequence parameter set exactly includes one slice; and conditionally parsing one or more syntax elements in a slice header based on the value of the flag.
[0727] In one example, the method is provided, wherein the one or more syntax elements in the slice header include a syntax element indicating a slice address.
[0728] In one example, the method is provided, wherein the one or more syntax elements in the slice header include a syntax element indicating the number of tiles in a slice.
[0729] In one example, the method is provided, wherein a syntax element having a value specifying a picture order count is the first syntax element in the slice header.
[0730] In one example, a device is provided, the device comprising one or more processors configured to perform any and all combinations of the steps.
[0731] In one example, the apparatus is provided, wherein the apparatus includes a video encoder.
[0732] In one example, the apparatus is provided, wherein the apparatus includes a video decoder.
[0733] In one example, a system is provided, the system including: an apparatus including a video encoder; and the apparatus including a video decoder.
[0734] In one example, an apparatus is provided, the apparatus including: means for performing any and all combinations of steps.
[0735] In one example, a non-transitory computer-readable storage medium is provided, the non-transitory computer-readable storage medium including instructions stored thereon that, when executed, cause one or more processors of the apparatus to perform any and all combinations of steps.
[0736] In one example, a method for decoding picture information for decoding video data is provided, the method including: receiving a slice header syntax structure; determining whether the slice header syntax structure includes a picture header syntax structure; and, when the slice header syntax structure includes the picture header syntax structure, parsing, from the picture header syntax structure, a first syntax element specifying a picture parameter set identifier and a second syntax element specifying a picture order count.
[0737] In one example, the method is provided, and when the slice header syntax structure includes the picture header syntax structure, further parsing, from the picture header syntax structure, a third syntax element specifying whether the current picture is a reference picture.
[0738] In one example, the method is provided, and when the slice header syntax structure includes the picture header syntax structure, further determining whether the picture header syntax structure includes a picture output flag syntax element, and when the picture header syntax structure includes the picture output flag syntax element, parsing the picture output flag syntax element from the picture header syntax structure.
[0739] In one example, the method is provided, wherein it is determined whether the picture header syntax structure includes the picture output flag syntax element based on the value of a current flag.
[0740] In one example, the method is provided, wherein it is determined whether the slice header syntax structure includes the picture header syntax structure based on the value of a flag.
[0741] In one example, a device is provided that includes one or more processors configured to: receive a slice header syntax structure; determine whether the slice header syntax structure includes a picture header syntax structure; and, in the case where the slice header syntax structure includes the picture header syntax structure, parse a first syntax element specifying a picture parameter set identifier from the picture header syntax structure and parse a second syntax element specifying a picture order count from the picture header syntax structure.
[0742] In one example, the device is provided, and in the case where the slice header syntax structure includes the picture header syntax structure, the one or more processors are further configured to parse a third syntax element specifying whether the current picture is a reference picture from the picture header syntax structure.
[0743] In one example, the device is provided, where the one or more processors are further configured to: in the case where the slice header syntax structure includes the picture header syntax structure, determine whether the picture header syntax structure includes a picture output flag syntax element; and, in the case where the picture header syntax structure includes the picture output flag syntax element, parse the picture output flag syntax element from the picture header syntax structure.
[0744] In one example, the device is provided, where the one or more processors determine whether the picture header syntax structure includes the picture output flag syntax element based on the value of a current flag.
[0745] In one example, the device is provided, where the one or more processors determine whether the slice header syntax structure includes the picture header syntax structure based on the value of a flag.
[0746] In one example, the device is provided, where the device is a video decoder.
[0747] <Cross - reference>
[0748] This non - provisional patent application claims priority under 35 U.S.C. § 119 to Provisional Application No. 62 / 890,523, filed on August 22, 2019, and Provisional Application No. 62 / 905,307, filed on September 24, 2019, the entire contents of both applications are hereby incorporated by reference.
Claims
1. A device comprising one or more processors, the one or more processors being configured to: Receive a picture header syntax structure; Wherein: (i) The picture header syntax structure is conditionally included in the slice header syntax structure based on the value of a flag, (ii) The picture header syntax structure is different from the slice header syntax structure such that when the slice header syntax structure does not include the picture header syntax structure, the picture header syntax structure is included in a non-video coding layer network abstraction layer unit, (iii) The syntax elements included in the picture header syntax structure are different from the syntax elements included in the picture parameter set syntax structure, and (iv) The picture header syntax structure includes: A first syntax element specifying the value of the picture parameter set identifier of the picture parameter set in use, A second syntax element specifying whether the current picture has never been used as a reference picture, and A third syntax element specifying the least significant bit information of the picture order count of the current picture.
2. A non-transitory computer-readable medium storing a program that causes a processor to perform the following operations: Receive a picture header syntax structure; Wherein: (i) The picture header syntax structure is conditionally included in the slice header syntax structure based on the value of a flag, (ii) The picture header syntax structure is different from the slice header syntax structure such that when the slice header syntax structure does not include the picture header syntax structure, the picture header syntax structure is included in a non-video coding layer network abstraction layer unit, (iii) The syntax elements included in the picture header syntax structure are different from the syntax elements included in the picture parameter set syntax structure, and (iv) The picture header syntax structure includes: A first syntax element specifying the value of the picture parameter set identifier of the picture parameter set in use, A second syntax element specifying whether the current picture has never been used as a reference picture, and A third syntax element specifying the least significant bit information of the picture order count of the current picture.
3. The device according to claim 1, wherein the device is a video decoder.