System and method for signaling a tile structure for pictures of a coded video
The proposed method addresses the limitations in signaling tile structures within current video coding standards by using a flag and syntax elements to specify tile columns and rows, resulting in improved video delivery system performance through reduced bandwidth and enhanced parallel processing capabilities.
Patent Information
- Application Number
- JP2024110897
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-02-15
- Filing Date
- 2024-07-10
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2039-11-26
AI Technical Summary
Current video coding standards, such as ITU-T H.265, have limitations in signaling tile structures for encoded video pictures, which can lead to inefficiencies in video delivery systems, including increased transmission bandwidth and reduced parallelization capabilities in encoding and decoding processes.
The proposed solution involves signaling a flag to indicate the activation of a tile set, along with syntax elements specifying the number of tile columns and rows partitioning the picture. This allows for flexible tile structures, enabling improved video delivery system performance by reducing transmission bandwidth and facilitating parallel processing.
By enabling flexible tile structures, the proposed solution enhances video delivery system performance by reducing transmission bandwidth and facilitating parallel processing in video encoders and decoders, thereby improving overall efficiency and scalability.
Smart Images

Figure 0007699699000107 
Figure 0007699699000108 
Figure 0007699699000109
Abstract
Description
Technical Field
[0001] The present disclosure relates to video coding, and more particularly, to techniques for signaling a tile structure for pictures of encoded video.
[0002] Background Art Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, laptop or desktop computers, tablet computers, digital recording devices, digital media players, video gaming devices, cellular phones including so-called smart phones, medical imaging devices, and the like. Digital video can be encoded according to video coding standards. Video coding standards can incorporate video compression techniques. Examples of video coding standards include ISO / IEC MPEG-4 Visual and ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) and High-Efficiency Video Coding (HEVC). HEVC is High Efficiency Video Coding (HEVC), Rec. ITU-T H. It is described in ITU-T H.265 (December 2016), which is incorporated herein by reference and is referred to as ITU-T H.265 herein. Extensions and improvements to ITU-T H.265 are currently under consideration for the development of the next-generation video coding standard. For example, the ITU-T Video Coding Experts Group (VCEG) and ISO / IEC (Moving Picture Experts Group (MPEG), collectively called the Joint Video Exploration Team (JVET)) are considering the potential need for standardizing future video coding technologies with compression capabilities significantly exceeding those of the current HEVC standard. The Joint Exploration Model 7 (JEM 7), Algorithm Description of Joint Exploration Test Model 7 (JEM 7), ISO / IEC JTC1 / SC29 / WG11 Document: namely, JVET-G1001, July 2017, Torino, IT, which is incorporated herein by reference, describes the coding features that are in the midst of collaborative test model research by the JVET as having the potential to improve video coding technology beyond the capabilities of ITU-T H.265. Note that the coding features of JEM 7 are implemented in the JEM reference software. As used herein, the term JEM may collectively represent the algorithms included in JEM 7 and the implementation of the JEM reference software. Furthermore, in response to the "Joint Call for Proposals on Video Compression with Capabilities beyond HEVC" jointly issued by VCEG and MPEG, multiple descriptions related to video coding have been proposed by various groups at the 10th Meeting of ISO / IEC JTC1 / SC29 / WG11, 16 - 20 April 2018, San Diego, CA.As a result of multiple descriptions of video coding, draft text of the video coding specification is described in "Versatile Video Coding (Draft 1)," 10th Meeting of ISO / IEC JTC1 / SC29 / WG11 16-20 April 2018, San Diego, CA, document JVET-J1001-v2, which is incorporated herein by reference and referred to as JVET-J1001. "Versatile Video Coding (Draft 2)," 11th Meeting of ISO / IEC JTC1 / SC29 / WG11 10-18 July 2018, Ljubljana, SI, document JVET-K1001-v7, which is incorporated herein by reference and referred to as JVET-K1001, is an updated version of JVET-J1001. Further, "Versatile Video Coding (Draft 3)," 12th Meeting of ISO / IEC JTC1 / SC29 / WG11 3-12 October 2018, Macao, CN, document JVET-L1001-v2, which is incorporated herein by reference and referred to as JVET-L1001, is an updated version of JVET-K1001.
[0003] Video compression technology reduces the data requirements for storing and transmitting video data by exploiting the inherent redundancy in a video sequence. The video compression technology can continuously subdivide the video sequence into smaller and smaller parts (i.e., groups of frames within the video sequence, frames within a group of frames, slices within a frame, coded tree units (e.g., macroblocks) within a slice, coded blocks within a coded tree unit, etc.). Intra prediction coding techniques (e.g., within a picture (spatial)) and inter prediction techniques (i.e., between pictures (temporal)) can be used to generate a difference value between a unit of video data to be coded and a reference unit of video data. The difference value may be referred to as residual data. The residual data can be coded as quantized transform coefficients. Syntax elements can associate the residual data with a reference coded unit (e.g., intra prediction mode index, motion vector, and block vector). The residual data and the syntax elements can be entropy coded. The entropy-coded residual data and syntax elements can be included in a compliant bitstream. The compliant bitstream and associated metadata may have a format according to a data structure.
[0004] Summary of the Invention In one embodiment, a method for decoding video data includes decoding a first flag syntax in a picture parameter set, the first flag syntax specifying whether tiles in each of at least one slice are in raster scan order or whether the tiles in each of the at least one slice cover a rectangular region of the picture; determining whether a slice address syntax is present in a slice header based on a value of the first flag syntax; and decoding the slice address syntax when the slice address syntax is present in the slice header.
[0005] In one embodiment, a method for encoding video data includes encoding a first flag syntax within a picture parameter set, where the first flag syntax specifies whether tiles in each of at least one slice are in raster scan order or whether the tiles in each of the at least one slice cover a rectangular region of the picture, and determining whether to encode a slice address syntax in a slice header based on the value of the first flag syntax.
Brief Description of the Drawings
[0006]
Figure 1
Figure 2A
Figure 2B
Figure 2C
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
[0007] Embodiments for Carrying Out the Invention Generally, the present disclosure describes various techniques for encoding video data. Specifically, the present disclosure describes techniques for signaling the tile structure of pictures of encoded video. As used herein, the term tile structure may refer to a particular partitioning of a picture into tiles. According to the techniques described herein, as will be described in more detail below, a picture may be partitioned into variable-size tiles and a tile structure. Signaling of the tile structure according to the techniques described herein can be particularly useful for improving video delivery system performance by reducing the transmission bandwidth and / or facilitating the parallelization of video encoders and / or decoders. The techniques of the present disclosure are described with respect to ITU-T H.264 and ITU-T H.265, but it should be noted that the techniques of the present disclosure are generally applicable to video encoding. For example, the encoding techniques described herein are ITU-T H. It can be incorporated into a video coding system that includes block structures, intra prediction techniques, inter prediction techniques, transform techniques, filtering techniques, and / or entropy coding techniques other than those included in 265 (including video coding systems based on future video coding standards). Accordingly, references to ITU-T H.264 and ITU-T H.265 are for illustrative purposes only and should not be construed as limiting the scope of the techniques described herein. Further, it should be noted that incorporation by reference of documents herein should not be construed as limiting or creating ambiguity with respect to the terms used herein. For example, if the incorporated reference gives a definition of a term that is different from that of another incorporated reference and / or if such term is used herein, then the term should be construed to broadly include each corresponding definition and / or to include each specific definition instead.
[0008] In one embodiment, a method for signaling a tile set structure includes signaling a flag indicating that the tile set is active in the bitstream, signaling a syntax element indicating a number of tile set columns partitioning a picture, and signaling a syntax element indicating a number of tile set rows partitioning a picture.
[0009] In one embodiment, a device comprises one or more processors configured to signal a flag indicating that the tile set is active in the bitstream, signal a syntax element indicating a number of tile set columns partitioning a picture, and signal a syntax element indicating a number of tile set rows partitioning a picture.
[0010] In one embodiment, a non-transitory computer-readable storage medium includes instructions stored thereon that, when executed, cause one or more processors of a device to signal a flag indicating that a tile set is enabled in a bitstream, signal a syntax element indicating a number of tile set columns partitioning a picture, and signal a syntax element indicating a number of tile set rows partitioning a picture.
[0011] In one embodiment, an apparatus comprises means for signaling a flag indicating that a tile set is enabled in a bitstream, means for signaling a syntax element indicating a number of tile set columns partitioning a picture, and means for signaling a syntax element indicating a number of tile set rows partitioning a picture.
[0012] In one embodiment, a method of decoding video data includes parsing a flag indicating that a tile set is enabled in a bitstream, parsing a syntax element indicating a number of tile set columns partitioning a picture, parsing a syntax element indicating a number of tile set rows partitioning a picture, and generating video data based on values of the parsed syntax elements.
[0013] In one embodiment, a device comprises one or more processors configured to parse a flag indicating that a tile set is enabled in a bitstream, parse a syntax element indicating a number of tile set columns partitioning a picture, parse a syntax element indicating a number of tile set rows partitioning a picture, and generate video data based on values of the parsed syntax elements.
[0014] In one embodiment, the non-transitory computer-readable storage medium includes instructions stored thereon that, when executed, cause one or more processors of the device to parse a flag indicating that a tile set is enabled in a bitstream, parse a syntax element indicating a number of tile set columns partitioning a picture, parse a syntax element indicating a number of tile set rows partitioning a picture, and generate video data based on the values of the parsed syntax elements.
[0015] In one embodiment, the apparatus comprises means for parsing a flag indicating that a tile set is enabled in a bitstream, means for parsing a syntax element indicating a number of tile set columns partitioning a picture, means for parsing a syntax element indicating a number of tile set rows partitioning a picture, and means for generating video data based on the values of the parsed syntax elements.
[0016] The details of one or more embodiments are described in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims.
[0017] Video content typically includes a video sequence consisting of a series of frames. A series of frames may also be referred to as a group of pictures (GOP). Each video frame or picture can contain one or more slices, and a slice contains a plurality of video blocks. A video block contains an array of pixel values (sometimes called samples) that can be predicted and encoded. Video blocks can be ordered according to a scanning pattern (e.g., raster scan). A video encoder performs predictive encoding on video blocks and their subdivisions. ITU-T H.264 defines macroblocks that contain 16×16 luma samples. ITU-T H.265 defines a similar Coding Tree Unit (CTU) structure (sometimes called the Largest Coding Unit (LCU)), where a picture can be divided into CTUs of equal size, and each CTU can contain a Coding Tree Block (CTB) with 16×16, 32×32, or 64×64 luma samples. As used herein, the term video block may generally refer to a region of a picture or, more specifically, to the largest array of pixel values that can be predictively encoded, its subdivisions, and / or the corresponding structure. Further, according to ITU-T H.265, each video frame or picture may be partitioned to include one or more tiles, where a tile is a sequence of coding tree units corresponding to a rectangular region of the picture.
[0018] In ITU-T H.265, a CTU is composed of respective CTBs (Coding Tree Blocks) for each component of the video data (e.g., luma (Y) and chroma (Cb and Cr)). Further, ITU-T In H.265, a CTU can be divided according to a quadtree (QT) partitioning structure, and as a result, the CTB of the CTU is divided into coding blocks (CBs). That is, in ITU-T H.265, a CTU can be divided into quadtree leaf nodes. According to ITU-T H.265, one luma CB along with two corresponding chroma CBs and related syntax elements is called a coding unit (CU). In ITU-T H.265, the minimum allowable size of a CB can be signaled. In ITU-T H.265, the smallest minimum allowable size of a luma CB is 8×8 luma samples. In ITU-T H.265, the decision to encode a picture area using intra prediction or inter prediction is made at the CU level.
[0019] In ITU-T H.265, a CU is associated with a prediction unit (PU) structure having its root in the CU. In ITU-T H.265, the PU structure enables splitting of the luma CB and chroma CB for the purpose of generating corresponding reference samples. That is, in ITU-T H.265, the luma CB and chroma CB can be split into respective luma and chroma prediction blocks (PBs), where a PB contains a block of sample values to which the same prediction is applied. In ITU-T H.265, a CB can be split into 1, 2, or 4 PBs. ITU-T H.265 supports PB sizes from 64×64 samples down to 4×4 samples. In ITU-T H.265, square PBs are supported for intra prediction, where the CB can form a PB or be split into 4 square PBs (i.e., the intra prediction PB size types include M×M or M / 2×M / 2, where M is the height and width of the square CB). In ITU-T H.265, in addition to square PBs, rectangular PBs are supported for inter prediction, where the CB can be bisected vertically or horizontally to form a PB (i.e., the inter prediction PB types include M×M, M / 2×M / 2, M / 2×M, or M×M / 2). Further, in ITU-T H.265, 4 non-symmetric PB splits are supported for inter prediction, where the CB is split into 2 PBs at one quarter of its height (upper or lower) or width (left or right) (i.e., the non-symmetric partitions include M / 4×M left, M / 4×M right, M×M / 4 upper, and M×M / 4 lower). Reference sample values and / or predicted sample values for the PB are generated using intra prediction data (e.g., intra prediction mode syntax elements) or inter prediction data (e.g., motion data syntax elements) corresponding to the PB.
[0020] JEM defines a CTU that has a maximum size of 256×256 luma samples. JEM defines a quad-tree + binary-tree (QTBT) block structure. In JEM, the QTBT structure enables further splitting of the quadtree leaf nodes by a binary-tree structure (BT). That is, in JEM, the binary-tree structure enables recursive splitting of the quadtree leaf nodes vertically or horizontally. Therefore, the binary-tree structure in JEM enables square leaf nodes and rectangular leaf nodes, and each leaf node contains one CB. As shown in FIG. 2A, a picture included in a GOP can contain a plurality of slices, each slice contains a series of CTUs, and each CTU can be split according to the QTBT structure. In JEM, the CB is used for prediction without further splitting. That is, in JEM, the CB can be a block of sample values to which the same prediction is applied. Therefore, the JEM QTBT leaf node may be similar to the PB in ITU-T H.265.
[0021] Intra prediction data (e.g., intra prediction mode syntax elements) or inter prediction data (e.g., motion data syntax elements) can be associated with a corresponding reference sample of a PU. Residual data can include respective arrays of difference values corresponding to each component of video data (e.g., luma (Y) and chroma (Cb and Cr)). The residual data can be within a pixel region. Transformations such as discrete cosine transform (DCT), discrete sine transform (DST), integer transform, wavelet transform, or conceptually similar transforms can be applied to the pixel difference values to generate transform coefficients. Note that in ITU-T H.265, a CU can be further split into transform units (TUs). That is, an array of pixel difference values can be split (e.g., four 8×8 transforms can be applied to a 16×16 array of residual values corresponding to a 16×16 luma CB) to generate transform coefficients, and such a split is sometimes called a transform block (TB). The transform coefficients can be quantized according to a quantization parameter (QP). The quantized transform coefficients (which may be called level values) can be entropy coded according to entropy coding techniques (e.g., content adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), probability interval partitioning entropy coding (PIPE), etc.). Further, syntax elements, such as syntax elements indicating a prediction mode, etc., can also be entropy coded. The entropy coded and quantized transform coefficients and the corresponding entropy coded syntax elements can form a compliant bitstream that can be used to regenerate the video data.The binarization process can be performed on the syntax elements as part of the entropy coding process. Binarization refers to the process of converting a syntax value into a sequence of one or more bits. These bits may be referred to as "bins".
[0022] As described above, the intra prediction data or inter prediction data is used to generate reference sample values for a block of sample values. The difference between the sample values included in the current PB or another type of picture area structure and the associated reference samples (e.g., reference samples generated using prediction) may be referred to as residual data. As described above, the intra prediction data or inter prediction data can associate an area of a picture (e.g., a PB or a CB) with the corresponding reference samples. For intra prediction coding, the intra prediction mode can specify the positions of the reference samples within the picture. In ITU-T H.265, the defined possible intra prediction modes include the planar (i.e., surface-conforming) prediction mode (predMode: 0), the DC (i.e., flat overall averaging) prediction mode (predMode: 1), and the 33-angle prediction mode (predMode: 2-34). In JEM, the defined possible intra prediction modes include the planar prediction mode (predMode: 0), the DC prediction mode (predMode: 1), and the 65-angle prediction mode (predMode: 2-66). Note that the planar and DC prediction modes may be referred to as non-directional prediction modes, and the angle prediction modes may be referred to as directional prediction modes. Note that the techniques described herein may generally be applicable regardless of the number of defined possible prediction modes.
[0023] For inter prediction coding, a motion vector (MV) identifies reference samples within a picture other than the picture of the video block to be coded, thereby exploiting the temporal redundancy of the video. For example, a current video block can be predicted from reference block(s) located within a previously coded frame(s), and the motion vector can be used to indicate the position of the reference block. The motion vector and related data can describe, for example, the horizontal component of the motion vector, the vertical component of the motion vector, the resolution for the motion vector (e.g., quarter-pel accuracy, half-pel accuracy, one-pel accuracy, two-pel accuracy, four-pel accuracy), the prediction direction, and / or the index value of the reference picture. Further, for example, coding standards such as ITU-T H.265 can support motion vector prediction. Motion vector prediction enables specifying a motion vector using the motion vectors of adjacent blocks. Examples of motion vector prediction include advanced motion vector prediction (AMVP), temporal motion vector prediction (TMVP), the so-called "combined" mode, and "skip" and "direct" motion estimation. Further, JEM supports advanced temporal motion vector prediction (ATMVP), spatial-temporal motion vector prediction (STMVP), pattern match motion vector derivation (PMMVD) mode which is a special merge mode based on frame rate up-conversion (FRUC) technology, and affine transform motion compensation prediction.
[0024] Residual data can include respective arrays of difference values corresponding to respective components of video data. The residual data can be within a pixel region. Transformations such as discrete cosine transform (DCT), discrete sine transform (DST), integer transform, wavelet transform, or conceptually similar transforms can be applied to the array of difference values to generate transform coefficients. In ITU-T H.265, a coding unit (CU) is associated with a transform unit (TU) structure having its root at the CU level. That is, in ITU-T H.265, as described above, the array of difference values can be re-partitioned for the purpose of generating transform coefficients (e.g., four 8×8 transforms can be applied to a 16×16 array of residual values). It should be noted that in ITU-T H.265, it is not always necessary to align the transform blocks (TBs) with the prediction blocks (PBs).
[0025] Note that in JEM, transform coefficients are generated without further partitioning, using the residual values corresponding to the coding blocks (CBs). That is, in JEM, a quad-tree block transform (QTBT) leaf node can be similar to both the prediction block (PB) and the transform block (TB) in ITU-T H.265. In JEM, the core transform and subsequent secondary transforms can be applied (in a video encoder) to generate transform coefficients. For a video decoder, the order of the transforms is reversed. Further, in JEM, whether to apply the secondary transform to generate transform coefficients can depend on the prediction mode.
[0026] The quantization process can be performed on the transform coefficients. Quantization approximates the transform coefficients by amplitudes limited to a specific set of values. Quantization may be used to vary the amount of data required to represent the group of transform coefficients. Quantization can be achieved by division of the transform coefficients by a scaling factor and any associated rounding function (e.g., rounding to the nearest integer). Quantized transform coefficients may be called coefficient level values. Inverse quantization Quantization (or "dequantization") can include multiplication of coefficient level values by a magnification factor. It should be noted that when used in this specification, the term quantization process may, in some cases, refer to division by a magnification factor to generate a level value, or, in some cases, multiplication by a magnification factor to recover a conversion coefficient. That is, the quantization process may, in some cases, refer to quantization and, in some cases, inverse quantization.
[0027] Regarding the equations used in this specification, the following arithmetic operators may be used.
[0028] + Addition - Subtraction * Multiplication including matrix multiplication x y Exponentiation. Specifies the y-th power of x. In other contexts, such notation is used for superscripts not intended to be interpreted as exponentiation.
[0029] / Integer division with truncation of the result to zero. For example, 7 / 4 and -7 / -4 are truncated to 1, and -7 / 4 and 7 / -4 are truncated to -1.
[0030] ÷ Used to represent division in an equation where truncation or rounding is not intended.
[0031] x / y Used to represent division in an equation where truncation or rounding is not intended.
[0032] Furthermore, the following mathematical functions can be used.
[0033] Log2(x), the logarithm of x with base 2;
[0034]
Number
[0035] Ceil(x) is the smallest integer greater than or equal to x.
[0036] Regarding examples of the syntax used in this specification, the following definitions of logical operators may apply.
[0037] x&&y The Boolean logical "product" of x and y x||y The Boolean logical "sum" of x and y ! Boolean logical "negation" x?y:z If x is true or not equal to 0, evaluate the value of y; otherwise, evaluate the value of z.
[0038] Furthermore, the following relational operators may apply.
[0039] > Greater than >= Greater than or equal to < Less than <= Less than or equal to == Equal to != Not equal to Furthermore, it should be noted that in the syntax descriptors used in this specification, the following descriptors may apply.
[0040] -u(n): Unsigned integer using n bits.
[0041] -ue(v): Unsigned integer Golomb coding syntax element of order 0 with the left bit first.
[0042] A virtual reality (VR) application can include video content that can be rendered on a head-mounted display, and only the area of the omnidirectional video corresponding to the orientation of the user's head is rendered. The VR application can be enabled by an omnidirectional video, also called a 360° spherical video. The omnidirectional video is typically captured by a plurality of cameras covering a scene of up to 360°. A distinct feature of the omnidirectional video compared to a normal video is that typically only a subset of the entire captured video area is displayed, i.e., the area corresponding to the current user's field of view (FOV). The FOV is also sometimes called a viewport. In other cases, the viewport may be a part of the spherical video that is currently being displayed and viewed by the user. Note that the size of the viewport may be below the field of view.
[0043] The most interesting area in the omnidirectional video picture can refer to a subset of the entire video area that is statistically most likely to be rendered to the user when presenting the photo (i.e., most likely to be in the FOV). Note that the most interesting area of the omnidirectional video may be determined by the intention of the director or producer, or derived from user statistics by a service or content provider (e.g., through the statistics of the areas most frequently requested / viewed by users when the omnidirectional video content is provided through a streaming service). The most interesting area can be used for data prefetching in omnidirectional video adaptive streaming by an edge server or a client, and / or for transcoding optimization when the omnidirectional video is transcoded, for example, into different codecs or projection mappings. Therefore, signaling the most interesting area in the omnidirectional video picture can improve system performance by reducing the transmission bandwidth and reducing the decoding complexity. Note that the base area generally refers to the entire area of the encoded video data, e.g., the entire video area.
[0044] As described above, according to ITU-T H.265, each video frame or picture may be partitioned to include one or more slices, and may be further partitioned to include one or more tiles. FIGS. 2A to 2C are conceptual diagrams showing an example of a group of pictures that include slices and further partition the pictures into tiles. In the example shown in FIG. 2A, Pic4 is shown as including two slices (i.e., Slice1 and Slice2), and each slice includes a sequence of CTUs (e.g., in raster scan order). A slice is a sequence of one or more slice segments starting from an independent slice segment, including all subsequent dependent slice segments (if any), and then followed by the next independent slice segment (if any) in the same access unit. Note that a slice segment such as a slice is a sequence of coding tree units. In the embodiments described herein, in some cases, the terms slice and slice segment may be used interchangeably to indicate the result of a coding tree unit. In the example shown in FIG. 2B, Pic4 is shown as including six tiles (i.e., Tile1 to Tile6), and each tile is rectangular and includes a sequence of CTUs. Note that in ITU-T H.265, a tile may consist of coding tree units included in two or more slices, and a slice may consist of coding tree units included in two or more tiles. However, ITU-T H.265 stipulates that one or both of the following conditions must be met. (1) All coding tree units in a slice belong to the same tile, and (2) all coding tree units in a tile belong to the same slice. Thus, for example, with respect to FIG. 2B, all of the tiles may belong to a single slice, or the tiles may belong to multiple slices (e.g., Tile1 to Tile3 may belong to Slice1, and Tile4 to Tile6 may belong to Slice2). Regarding JVET-L1001, it has been proposed that not only does it need to consist of an integer number of complete CTUs, but also a slice needs to be composed of an integer number of complete tiles.Thus, a slice containing a set of CTUs that do not form a rectangular region of a picture may or may not be supported by some video coding techniques. Further, a slice that needs to consist of an integer number of complete tiles is called a tile group. The techniques described herein may be applicable to slices, tiles, and / or tile groups.
[0045] Further, as shown in FIG. 2B, tiles may form a tile set (i.e., Tile2 and Tile3 form a tile set). A tile set can be used to define boundaries for encoding dependencies (e.g., intra prediction dependencies, entropy coding dependencies, etc.), and thus can enable parallel processing in encoding and region of interest encoding. For example, if the video sequence in the example shown in FIG. 2B corresponds to a night news program, the tile set formed by Tile2 and Tile3 may correspond to the region of visual interest including the news anchor reading the news. ITU-T H.265 defines signaling to enable a motion constrained tile set (MCTS). A motion constrained tile set may include a tile set in which inter-picture prediction dependencies are limited to collocated tile sets within a reference picture. As a result, it is possible to perform motion compensation for a given MCTS independently of the decoding of other tile sets outside the MCTS. For example, referring to FIG. 2B, if the tile set formed by Tile2 and Tile3 is an MCTS and each of Pic1 - Pic3 includes a collocated tile set, the motion compensation on Tile2 and Tile3 may be performed independently of the encoding of Tile1, Tile4, Tile5, and Tile6 of Pic4, and the tiles collocated with Tile1, Tile4, Tile5, and Tile6 in each of Pic1 - Pic3. Encoding video data according to MCTS may be useful for video applications including omnidirectional video presentations.
[0046] As shown in FIG. 2C, Tile1 to Tile6 can form the most interesting area of the omnidirectional video. Further, the tile set formed by Tile2 and Tile3 may be the MCTS included in the most interesting area. Viewport-dependent video coding, which may also be called viewport-dependent partial video coding, can be used to enable the multiplexing of only a part of the entire video area. That is, for example, viewport-dependent video coding can be used to provide sufficient information for rendering the current FOV. For example, an omnidirectional video can be encoded using MCTS such that each area that may cover the viewport can be decoded independently of other areas over time. In this case, for example, for a specific current viewport, a minimum set of tiles covering the viewport may be sent to, decoded by, and / or rendered by the client. This process may be called simple tile-based partial decoding (STPD).
[0047] In ITU-T H.265, an encoded video sequence (CVS) may be encapsulated (or structured) as a sequence of access units, and each access unit contains video data structured as network abstraction layer (NAL) units. In ITU-T H.265, a bitstream is described as containing a sequence of NAL units that form one or more CVSs. Note that ITU-T H.265 supports multi-layer extensions including format range extension (RExt), scalability (SHVC), multi-view (MV-HEVC) and 3-D (3D-HEVC). The multi-layer extensions enable the video display to include a base layer and one or more additional enhancement layers. For example, the base layer can enable a video display with a basic level of quality (e.g., high-resolution rendering), and the enhancement layer can enable a video display with an enhanced level of quality (e.g., ultra-high-resolution rendering). In ITU-T H.265, the enhancement layer can be encoded by referring to the base layer. That is, for example, pictures in the enhancement layer may be encoded (e.g., using inter-prediction techniques) by referring to one or more pictures (including scaled-up / down versions thereof) in the base layer. In ITU-T H.265, each NAL unit may include an identifier indicating the layer of the video data with which the NAL unit is associated. Referring to the example shown in Figure 2A, each slice of the video data included in Pic4 (i.e., Slice1 and Slice2) is shown as being encapsulated in an NAL unit. Further, in ITU-T H.265, each of a video sequence, GOP, picture, slice, and CTU may be associated with metadata describing video encoding characteristics. ITU-T H.265 defines parameter sets that can be used to describe video data characteristics and / or video encoding characteristics. In ITU-T H.265, a parameter set may be encapsulated as a special type of NAL unit or signaled as a message.NAL units containing symbolized video data (e.g., slices) may be referred to as VCL (Video Coding Layer) NAL units, and NAL units containing metadata (e.g., parameter sets) may be referred to as non-VCL NAL units. Further, ITU-T H.265 enables the signaling of Supplemental Enhancement Information (SEI) messages. In ITU-T H.265, SEI messages assist processes related to decoding, display, or other purposes, but SEI messages may not be required to create luma or chroma samples by the decoding process. In ITU-T H.265, SEI messages may be signaled in the bitstream using non-VCL NAL units. Further, SEI messages may be conveyed by some means other than by being present within the bitstream (i.e., signaled out-of-band).
[0048] FIG. 3 shows an example of a bitstream including a plurality of CVSs, where each CVS is represented by a NAL unit included in each access unit. In the example illustrated in FIG. 3, non-VCL NAL units include each parameter set unit (i.e., video parameter set (VPS), sequence parameter set (SPS), and picture parameter set (PPS) units), and an access unit delimiter NAL unit. Note that ITU-T H.265 defines NAL unit header semantics that specify the type of raw byte sequence payload (RBSP) data structure included in the NAL unit. As described above, omnidirectional video can be encoded using MCTS. Sub-bitstream extraction may refer to a process in which a device receiving an ITU-T H.265 compliant bitstream forms a new ITU-T H.265 compliant bitstream by discarding and / or modifying data within the received bitstream. For example, as described above, for a particular current viewport, a minimum set of tiles covering the viewport may be transmitted to the client. A new ITU-T H.265 compliant bitstream including the minimum set of tiles can be formed using sub-bitstream extraction. For example, referring to FIG. 2C, if the viewport includes only Tile2 and Tile3, the access units in the bitstream include VCL NAL units of Tile1 to Tile6, Tile1, Tile2, and Tile3 are included in the first slice, and Tile4, Tile5, and Tile6 are included in the second slice, the sub-bitstream extraction process may include generating a new bitstream including only the VCL NAL units of the slice including Tile2 and Tile3 (i.e., the VCL NAL units including the slice including Tile4, Tile5, and Tile6 are removed from the received bitstream).
[0049] As described above, the term tile structure may refer to a specific partitioning of a picture into tiles. Referring to FIG. 2B, the tile structure of Pic4 includes the illustrated tiles Tile1 to Tile6. In some cases, it may be useful to use different tile structures for different pictures. In ITU-T H.265, the tile structure of a picture is signaled using a picture parameter set. Table 1 is part of the syntax of the PPS specified in ITU-T H.265 that includes the relevant syntax elements for signaling the tile structure.
[0050]
Table 1
[0051] ITU-T H.265 defines the following for each of the syntax elements shown in Table 1. pps_pic_parameter_set_id specifies the PPS for reference by other syntax elements. The value of pps_pic_parameter_set_id shall be in the range of 0 to 63, inclusive.
[0052] pps_seq_parameter_set_id specifies the value of sps_seq_parameter_set_id for the active SPS. The value of pps_seq_parameter_set_id shall be in the range of 0 to 15, inclusive.
[0053] tiles_enabled_flag equal to 1 specifies that there are two or more tiles in each picture with reference to the PPS. tiles_enabled_flag equal to 0 specifies that there is only one tile in each picture with reference to the PPS. For all PPSs activated within a CVS, it is a requirement for bitstream conformance that the value of tiles_enabled_flag be the same.
[0054] num_tile_columns_minus1 plus 1 specifies the number of tile columns that partition the picture. num_tile_columns_minus1 shall be in the range of 0 to PicWidthInCtbsY - 1, inclusive of both end values. If the value of num_tile_columns_minus1 does not exist, the value of num_tile_columns_minus1 is presumed to be equal to 0.
[0055] num_tile_rows_minus1 plus 1 specifies the number of tile rows that partition the picture. num_tile_rows minus1 shall be in the range of 0 to PicHeightInCtbsY - 1, inclusive of both end values. If the value of num_tile_rows_minus1 does not exist, the value of num_tile_rows_minus1 is presumed to be equal to 0. When tiles_enabled_flag is equal to 1, both num_tile_columns_minus1 and num_tile_rows_minus1 shall not be equal to 0.
[0056] A uniform_spacing_flag equal to 1 specifies that the tile column boundaries and likewise the tile row boundaries are distributed uniformly across the picture. A uniform_spacing_flag equal to 0 specifies that the tile column boundaries and likewise the tile row boundaries are not distributed uniformly across the picture, but are explicitly signaled using the syntax elements column_width_minus1[i] and row_height_minus1[i]. If the value of uniform_spacing_flag does not exist, the value of uniform_spacing_flag is presumed to be equal to 1.
[0057] column_width_minus1[i] plus 1 specifies the width of the i-th tile column in units of coded tree blocks.
[0058] row_height_minus1[i] plus 1 specifies the height of the i-th tile row in units of the coding tree block.
[0059] Furthermore, in ITU-T H.265, information regarding entry points within the bitstream is signaled using the slice segment header. Table 2 is part of the syntax of the slice segment header specified in ITU-T H.265 that contains relevant syntax elements for signaling entry points.
[0060]
Table 2
[0061] ITU-T H.265 stipulates the following definitions for each syntax element shown in Table 2.
[0062] slice_pic_parameter_set_id specifies the value of pps_pic_parameter_set_id of the PPS in use. The value of slice_pic_parameter_set_id shall be in the range of 0 to 63, including both end values.
[0063] num_entry_point_offsets specifies the number of entry_point_offset_minus1[i] syntax elements within the slice header. If the value of num_entry_point_offsets does not exist, the value of num_entry_point_offsets is presumed to be equal to 0.
[0064] The value of num_entry_point_offset is constrained as follows.
[0065] - When -tiles_enabled_flag is equal to 0 and entropy_coding_sync_enabled_flag is equal to 1, the value of num_entry_point_offsets shall be in the range of 0 to PicHeightInCtbsY - 1, including both end values.
[0066] - Otherwise, when -tiles_enabled_flag is equal to 1 and entropy_coding_sync_enabled_flag is equal to 0, the value of num_entry_point_offsets shall be in the range of 0 to (num_tile_columns_minus1 + 1) * (num_tile_rows_minus1 + 1) - 1, including both end values.
[0067] - Otherwise, when -tiles_enabled_flag is equal to 1 and entropy_coding_sync_enabled_flag is equal to 1, the value of num_entry_point_offsets shall be in the range of 0 to (num_tile_columns_minus1 + 1) * PicHeightInCtbsY - 1, including both end values.
[0068] offset_len_minus1 plus 1 specifies the length (bits) of the entry_point_offset_minus1[i] syntax element. The value of offset_len_minus1 shall be in the range of 0 to 31, including both end values shall be.
[0069] entry_point_offset_minus1[i] plus 1 specifies the i-th entry point offset (in bytes) and is represented by offset_len_minus1 plus 1 bits. The slice segment data follows the slice segment header and consists of num_entry_point_offset + 1 subsets. The subset index values range from 0 to num_entry_point_offset, including both end values. The first byte of the slice segment data is considered byte 0. If present, the emulation prevention bytes that appear in the slice segment data part of the coded slice segment NAL unit are counted as part of the slice segment data for subset identification purposes. Subset 0 consists of bytes 0 to entry_point_offset_minus1[0] (including both end values) of the coded slice segment data. For k having a value in the range of 1 to num_entry_point_offsets - 1 (including both end values), subset k consists of bytes firstByte[k] to lastByte[k] (including both end values) of the coded slice segment data, where firstByte[k] and lastByte[k] are defined as follows.
[0070]
Number
[0071] The last subset (having a subset index equal to num_entry_point_offset) consists of the bytes of the coded slice segment data.
[0072] When tiles_enabled_flag is equal to 1 and entropy_coding_sync_enabled_flag is equal to 0, each subset consists of all the coded bits of all the coded tree units within a slice segment within the same tile, and the number of subsets (i.e., the value of num_entry_point_offsets + 1) is assumed to be equal to the number of tiles containing the coded tree units that are coded slice segments.
[0073] Note 6 - When tiles_enabled_flag is equal to 1 and entropy_coding_sync_enabled_flag is equal to 0, each slice shall either contain a subset of the coded tree units of one tile (in which case the syntax element entry_point_offset_minus1[i] does not exist), or shall contain all the coded tree units of an integer number of complete tiles.
[0074] When tiles_enabled_flag is equal to 0 and entropy_coding_sync_enabled_flag is equal to 1, each subset k having k in the range 0 to num_entry_point_offsets (inclusive) consists of all the coded bits of all the coded tree units within a slice segment containing the luma coded tree blocks within the same luma coded tree block row of the picture, and the number of subsets (i.e., the value of num_entry_point_offsets + 1) is assumed to be equal to the number of luma coded tree block rows of the picture containing the coded tree units within the coded slice segment.
[0075] Note 7 - The last subset (i.e., subset k equal to num_entry_point_offsets) may or may not contain all the coded tree units containing the luma coded tree blocks in the luma coded tree block row of the picture.
[0076] When the tiles_enabled_flag is equal to 1 and the entropy_coding_sync_enabled_flag is equal to 1, for each subset k having k in the range of 1 to num_entry_point_offsets (including both end values), all the coded bits of all the coded tree units in the slice segment that includes the luma coded tree blocks within the same luma coded tree block row of the tile consist of, and the number of subsets (i.e., the value of num_entry_point_offsets + 1) is assumed to be equal to the number of luma coded tree block rows of the tile that includes the coded tree units within the coded slice segment.
[0077] As indicated by the above syntax and semantics, in ITU-T H.265, the tile structure is specified by a number of columns and rows of numbers, and thus is restricted in that each row and column contains the same number of tiles. Restricting the tile structure in this way may not be ideal. According to the techniques described herein, a video encoder can signal the tile structure and tile sets in a manner that provides increased flexibility.
[0078] FIG. 1 is a block diagram illustrating an example of a system that can be configured to encode (encode and / or decode) video data according to one or more techniques of the present disclosure. System 100 represents an example of a system that can encapsulate video data according to one or more techniques of the present disclosure. As shown in FIG. 1, system 100 includes a source device 102, a communication medium 110, and a destination device 120. In the example shown in FIG. 1, source device 102 can include any device configured to encode video data and transmit the encoded video data to communication medium 110. Destination device 120 can include any device configured to receive the encoded video data via communication medium 110 and decode the encoded video data. Source device 102 and / or destination device 120 can include computing devices equipped for wired and / or wireless communication and can include, for example, set-top boxes, digital video recorders, televisions, desktops, laptops, or tablet computers, gaming consoles, medical imaging devices, and mobile devices including, for example, smartphones, cellular phones, personal gaming devices.
[0079] The communication medium 110 can include any combination of wireless and wired communication media and / or storage devices. Examples of the communication medium 110 can include coaxial cables, fiber optic cables, twisted pair cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other device that may be useful for facilitating communication between various devices and sites. The communication medium 110 can include one or more networks. For example, the communication medium 110 can include a network configured to enable access to the World Wide Web, such as the Internet. The network can operate according to a combination of one or more telecommunication protocols. The telecommunication protocol can include proprietary aspects and / or can include standardized telecommunication protocols. Examples of standardized telecommunication protocols include Digital Video Broadcasting (DVB) standards, Advanced Television Systems Committee (ATSC) standards, Integrated Services Digital Broadcasting (ISDB) standards, Data Over Cable Service Interface Specification (DOCSIS) standards, Global System Mobile Communications (GSM) standards, code division multiple access (CDMA) standards, 3rd Generation Partnership Project (3GPP) standards, European Telecommunications Standards Institute (ETSI) standards, Internet Protocol (IP) standards, Wireless Application Protocol (WAP) standards, and Institute of Electrical and Electronics Engineers (IEEE) standards.
[0080] A memory device can include any type of device or storage medium capable of storing data. The storage medium can include a tangible or non-transitory computer-readable medium. Examples of computer-readable media can include optical disks, flash memories, magnetic memories, or any other suitable digital storage media. In some examples, a memory device or a portion thereof may be described as non-volatile memory, and in other examples, a portion of the memory device may be described as volatile memory. Examples of volatile memory can include random access memory (RAM), dynamic random access memory (DRAM), and static random access memory (SRAM). Examples of non-volatile memory can include magnetic hard disks, optical disks, floppy disks, flash memories, or in the form of electrically programmable memory (EPROM) or electrically erasable and programmable memory (EEPROM). Examples of the memory device(s) can include memory cards (e.g., Secure Digital (SD) memory cards), internal / external hard disk drives, and / or internal / external solid state drives. Data can be stored on the memory device according to a defined file format.
[0081] FIG. 4 is a conceptual drawing showing an example of components that may be included in one implementation of system 100. In the exemplary implementation shown in FIG. 4, system 100 includes one or more computing devices 402A - 402N, a television service network 404, a television service provider site 406, a wide area network 408, a local area network 410, and one or more content provider sites 412A - 412N. The implementation shown in FIG. 4 represents an example of a system that may be configured such that digital media content, such as movies, live sports events, as well as data and applications and media presentations associated therewith, are distributed to and accessible by a plurality of computing devices such as computing devices 402A - 402N. In the example shown in FIG. 4, computing devices 402A - 402N can include any device configured to receive data from one or more of television service network 404, wide area network 408, and / or local area network 410. For example, computing devices 402A - 402N may be equipped for wired and / or wireless communication, may be configured to receive services through one or more data channels, and may include televisions including so - called smart TVs, set - top boxes, and digital video recorders. Further, computing devices 402A - 402N may include mobile devices including desktop, laptop or tablet computers, gaming consoles, e.g., "smart" phones, cellular phones, and personal gaming devices.
[0082] The television service network 404 is an example of a network configured to enable the delivery of digital media content, which may include television services. For example, the television service network 404 may include a terrestrial television network, a public or subscription-based satellite television service provider network, and a public or subscription-based cable television provider network and / or an over-the-top service provider or Internet service provider. In some embodiments, the television service network 404 may be primarily used to enable the provision of television services, but it should be noted that the television service network 404 can also enable the provision of other types of data and services based on any combination of the telecommunication protocols described herein. Further, in some embodiments, it should be noted that the television service network 404 can enable two-way communication between the television service provider site 406 and one or more of the computing devices 402A to 402N. The television service network 404 can include any combination of wireless communication media and / or wired communication media. The television service network 404 can include coaxial cables, fiber optic cables, twisted pair cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other device that may be useful for facilitating communication between various devices and sites. The television service network 404 can operate according to a combination of one or more telecommunication protocols. The telecommunication protocols can include dedicated aspects and / or can include standardized telecommunication protocols. Examples of standardized telecommunication protocols include DVB standards, ATSC standards, ISDB standards, DTMB standards, DMB standards, Data Over Cable Service Interface Specification (DOCSIS) standards, HbbTV standards, W3C standards, and UPnP standards.
[0083] Referring again to FIG. 4, the television service provider site 406 can be configured to deliver television services via the television service network 404. For example, the television service provider site 406 can include one or more broadcast stations, cable television providers, or satellite television providers, or Internet-based television providers. For example, the television service provider site 406 can be configured to receive transmissions including television programs via satellite uplink / downlink. Further, as shown in FIG. 4, the television service provider site 406 can communicate with the wide area network 408 and can be configured to receive data from the content provider sites 412A-412N. Note that in some embodiments, the television service provider site 406 can include a television studio and content can be transmitted therefrom.
[0084] The wide area network 408 includes a packet-based network and can operate according to a combination of one or more telecommunication protocols. The telecommunication protocols can include proprietary aspects and / or can include standardized telecommunication protocols. Examples of standardized telecommunication protocols include the Global System Mobile Communications (GSM) standard, the code division multiple access (CDMA) standard, the 3rd Generation Partnership Project (3GPP) standard, the European Telecommunications Standards Institute (ETSI) standard, the European Norm (EN), the IP standard, the Wireless Application Protocol (WAP) standard, and one or more of the Institute of Electrical and Electronics Engineers (IEEE) standards such as, for example, one of the IEEE 802 standards (e.g., Wi-Fi). The wide area network 408 can include any combination of wireless communication media and / or wired communication media. The wide area network 408 can include coaxial cables, fiber optic cables, twisted pair cables, Ethernet cables, wireless transmitters and receivers, routers, switches, repeaters, base stations, or any other device that may be useful for facilitating communication between various devices and sites. In one embodiment, the wide area network 408 may include the Internet. The local area network 410 includes a packet-based network and can operate according to a combination of one or more telecommunication protocols. The local area network 410 can be distinguished from the wide area network 408 based on the level of access and / or physical infrastructure. For example, the local area network 410 may include a secure home network.
[0085] Referring back to FIG. 4, content provider sites 412A - 412N represent examples of sites that can provide multimedia content to television service provider site 406 and / or computing devices 402A - 402N. For example, a content provider site can include a studio having one or more studio content servers configured to provide multimedia files and / or streams to television service provider site 406. In one embodiment, content provider sites 412A - 412N may be configured to provide multimedia content using an IP suite. For example, a content provider site may be configured to provide multimedia content to a receiving device according to the Real - Time Streaming Protocol (RTSP), HTTP, etc. Further, content provider sites 412A - 412N may be configured to provide data, including hypertext - based content, etc., to one or more of a receiving device, computing devices 402A - 402N, and / or television service provider site 406 through wide - area network 408. Content provider sites 412A - 412N may include one or more web servers. The data provided by data provider sites 412A - 412N can be defined according to a data format.
[0086] Referring back to FIG. 1, source device 102 includes video source 104, video encoder 106, data encapsulation device 107, and interface 108. Video source 104 can include any device configured to capture and / or store video data. For example, video source 104 can include a video camera and a storage device operably coupled thereto. Video encoder 106 can include any device configured to receive video data and generate a compliant bitstream representative of the video data. A compliant bitstream may refer to a bitstream that a video decoder can receive and from which it can regenerate the video data. The form of the compliant bitstream can be defined according to a video coding standard. When generating the compliant bitstream, video encoder 106 can compress the video data. The compression can be irreversible (recognizable or unrecognizable to the viewer) or reversible. FIG. 5 is a block diagram showing an example of video encoder 500 that can implement techniques for encoding video data described herein. The example video encoder 500 is shown as having separate functional blocks, but such an illustration is for explanatory purposes and it should be noted that it does not limit video encoder 500 and / or its subcomponents to a particular hardware or software architecture. The functions of video encoder 500 can be realized using any combination of hardware, firmware, and / or software implementations.
[0087] Video encoder 500 may perform intra-prediction coding and inter-prediction coding of a picture area, and may thus be referred to as a hybrid video encoder. In the example shown in FIG. 5, video encoder 500 receives a source video block. In some examples, the source video block may include a portion of a picture that is divided according to an encoding structure. For example, the source video data may include macroblocks, CTUs, CBs, their subdivisions, and / or other equivalent encoding units. In some examples, video encoder 500 may be configured to perform additional subdivision of the source video block. It should be noted that the techniques described herein are generally applicable to video encoding regardless of how the source video data is divided before and / or during encoding. In the example shown in FIG. 5, video encoding apparatus 500 includes an adder 502, a transform coefficient generator 504, a coefficient quantizer 506, an inverse quantization and transform coefficient processor 508, an adder 510, an intra-prediction processor 512, an inter-prediction processor 514, and an entropy encoder 516. As shown in FIG. 5, video encoder 500 receives a source video block and outputs a bitstream.
[0088] In the example shown in FIG. 5, video encoder 500 can generate residual data by subtracting a predicted video block from a source video block. The selection of the predicted video block is described in detail below. Adder 502 represents a component configured to perform this subtraction operation. In one example, the subtraction of video blocks is performed in the pixel domain. Transform coefficient generator 504 applies a transform such as a discrete cosine transform (DCT), discrete sine transform (DST), or conceptually similar transform to the residual block or a subdivision thereof (e.g., four 8×8 transforms can be applied to a 16×16 array of residual values), generating a set of residual transform coefficients. Transform coefficient generator 504 can be configured to perform any and all combinations of transforms included in the family of discrete trigonometric transforms, including approximations of discrete trigonometric transforms. Transform coefficient generator 504 can output the transform coefficients to coefficient quantizer 506. Coefficient quantizer 506 can be configured to perform quantization of the transform coefficients. The quantization process can reduce the bit depth associated with some or all of the coefficients. Depending on the degree of quantization, the rate distortion of the encoded video data (i.e., the video bitrate versus quality) can be changed. The degree of quantization can be adjusted by adjusting the quantization parameter (QP). The quantization parameter can be determined based on a slice level value and / or a CU level value (e.g., a CU delta QP value). The QP data can include any data used to determine the QP for quantizing a particular set of transform coefficients. As shown in FIG. 5, the quantized transform coefficients (which may also be referred to as level values) are output to inverse quantizer and transform coefficient processor 508. Inverse quantizer and transform coefficient processor 508 can be configured to apply inverse quantization and inverse transform to generate the restored residual data. As shown in FIG. 5, in adder 510, the restored residual data can be added to the predicted video block. In this way, the encoded video block can be restored, and the resulting restored video block can be used to evaluate the encoding quality for a given prediction, transform, and / or quantization.The video encoding device 500 can be configured to execute a plurality of encoding paths (for example, execute encoding while changing one or more of prediction, transform parameters, and quantization parameters). The rate distortion of the bitstream or other system parameters can be optimized based on the evaluation of the restored video block. Further, the restored video block can be stored and used as a reference for predicting subsequent blocks.
[0089] Referring to FIG. 5, the intra prediction processing unit 512 can be configured to select an intra prediction mode for the video block to be encoded. The intra prediction processing unit 512 can be configured to evaluate the frame and determine the intra prediction mode to be used for encoding the current block. As described above, the possible intra prediction modes may include a planar prediction mode, a DC prediction mode, and an angular prediction mode. Further, it should be noted that in some examples, the prediction mode for the chroma component can be inferred from the prediction mode for the luma prediction mode. The intra prediction processing unit 512 may select the intra prediction mode after executing one or more encoding paths. Further, in one embodiment, the intra prediction processing unit 512 may select the prediction mode based on rate distortion analysis. As shown in FIG. 5, the intra prediction processing unit 512 outputs intra prediction data (for example, syntax elements) to the entropy encoding unit 516 and the transform coefficient generator 504. As described above, the transform executed on the residual data may be mode-dependent (for example, the secondary transform matrix can be determined based on the prediction mode).
[0090] Referring back to FIG. 5, the inter prediction processing unit 514 can be configured to perform inter prediction encoding on the current video block. The inter prediction processing unit 514 can be configured to receive a source video block and calculate a motion vector for the PU of the video block. The motion vector can indicate the displacement of the PU of the video block in the current video frame relative to the prediction block in the reference frame. Inter prediction encoding can use one or more reference pictures. Further, the motion prediction can be single prediction (using one motion vector) or dual prediction (using two motion vectors). The inter prediction processing unit 514 can be configured to select a prediction block, for example, by calculating the pixel difference determined by the sum of absolute difference (SAD), sum of square difference (SSD), or other difference measurement methods. As described above, the motion vector can be determined and judged according to motion vector prediction. The inter prediction processing unit 514 can be configured to perform motion vector prediction as described above. The inter prediction processing unit 514 can be configured to generate a prediction block using the motion prediction data. For example, the inter prediction processing unit 514 can place the predicted video block in a frame buffer (not shown in FIG. 5). It should be noted that the inter prediction processing unit 514 can be further configured to apply one or more interpolation filters to the restored residual block to calculate pixel values less than an integer for use in motion prediction. The inter prediction processing unit 514 can output the motion prediction data for the calculated motion vector to the entropy encoding unit 516.
[0091] Referring back to FIG. 5, the entropy encoder 518 receives the quantized transform coefficients and prediction syntax data (i.e., intra prediction data, motion prediction data). Note that in some examples, the coefficient quantization unit 506 can perform a scan of the matrix including the quantized transform coefficients before the coefficients are output to the entropy encoder 518. In other examples, the entropy encoder 518 can perform the scan. The entropy encoder 518 can be configured to perform entropy encoding according to one or more of the techniques described herein. Thus, the video encoder device 500 represents an example of a device configured to generate encoded video data according to one or more techniques of the present disclosure. In one embodiment, the video encoder 500 can generate encoded video data including a motion constraint tile set.
[0092] Referring back to FIG. 1, the data encapsulation unit 107 can receive the encoded video data and generate a compliant bitstream, e.g., a compliant bitstream such as a series of NAL units, according to a defined data structure. A device that receives the compliant bitstream can regenerate the video data therefrom. Further, as described above, sub-bitstream extraction can refer to the process by which a device receiving an ITU-T H.265 compliant bitstream forms a new ITU-T H.265 compliant bitstream by discarding and / or modifying data within the received bitstream. Note that the term compliant bitstream can be used instead of the term compliant bitstream.
[0093] As described above, in ITU-T H.265, the tile structure is restricted in that each row and column contains the same number of tiles. In some cases, it may be useful to have different numbers of tiles in a row and / or column. For example, in order to encode a 360° spherical video, it may be useful to have fewer tiles in the polar regions than at the equator of the sphere, in which case this may be useful for varying the number of tile columns from row to row. In one embodiment, the data encapsulation device 107 can be configured to signal a tile structure according to one or more of the techniques described herein. Note that the data encapsulation device 107 need not be located within the same physical device as the video encoder 106. For example, the functions described as being performed by the video encoder 106 and the data encapsulation unit 107 may be distributed among the devices shown in FIG. 4.
[0094] According to the techniques described herein, the data encapsulation device 107 can be configured to signal one or more of the following types of information for a tile set structure.
[0095] A flag indicating whether the tile set is valid. If not valid, the entire picture is assumed to be one tile set.
[0096] If tile set signaling is enabled, the following can be signaled.
[0097] The number of rows of the tile set may be signaled, The number of columns of the tile set may be signaled.
[0098] For each tile set, the following information can be signaled.
[0099] The number of tile rows within the tile set, The number of tile columns within the tile set, An indicator / flag that signals when tiles are evenly spaced along the horizontal direction within a tile set, and / or An indicator / flag that signals when tiles are evenly spaced along the vertical direction within a tile set, and / or When the spacing is uniform, the following can be signaled for each tile set.
[0100] The tile width of the number of CTBs within each tile row in the tile set, The tile height of the number of CTBs within each tile column in the tile set, When the spacing is not uniform, the following can be signaled for each tile within each tile set.
[0101] The number of CTBs in the rows within each tile in the tile set, The number of CTBs in the columns within each tile in the tile set Table 3 shows an example of the syntax of a parameter set that can be used to signal the tile structure according to the technology of this specification. In one embodiment, the exemplary syntax included in Table 3 may be included in the PPS. In other examples, the exemplary syntax included in Table 3 may be included in the VPS or SPS.
[0102] [Table 3]
[0103] Regarding Table 3, it should be noted that the syntax elements tiles_enabled_flag, tilesets_enabled_flag, num_tile_columns_minus1[k][l], num_tile_rows_minus1[k][l], uniform_spacing_flag[k][1], column_width_minus1[k][l][i], row_height_minus1[k][l][i], tile_width_m_ctbsy_minus1[k][l], tile_height_in_ctbsy_minus1[k][l], and loop_filter_across_tiles_enabled_flag[k][l] can be based on the following exemplary definitions.
[0104] A tiles_enabled_flag equal to 1 specifies, with reference to the parameter set, that there are two or more tiles within each picture. A tiles_enabled_flag equal to 0 specifies, with reference to the parameter set, that there is only one tile within each picture.
[0105] A tilesets_enabled_flag equal to 1 specifies, with reference to the parameter set, that there are two or more tilesets within each picture. A tilesets_enabled_flag equal to 0 specifies, with reference to the parameter set, that there is only one tileset within each picture.
[0106] For all parameter sets activated within the CVS, it is a requirement for bitstream compliance that the value of tilesets_enabled_flag be the same.
[0107] When tiles_enabled_flag is equal to 0, it is a requirement for bitstream compliance that tilesets_enabled_flag be equal to 0.
[0108] num_tile_set_columns_minus1 plus 1 specifies the number of tile columns that partition the picture. num_tile_set_columns_minus1 shall be in the range of 0 to PicWidthInCtbsY - 1, inclusive. If the value of num_tile_set_columns_minus1 does not exist, the value of num_tile_set_columns_minus1 is assumed to be equal to 0.
[0109] num_tile_set_rows_minus1 plus 1 specifies the number of tile rows that partition the picture. num_tile_set_rows_minus1 shall be in the range of 0 to PicHeightInCtbsY - 1, inclusive. If the value of num_tile_set_rows_minus1 does not exist, the value of num_tile_set_rows_minus1 is assumed to be equal to 0.
[0110] In one embodiment, when tilesets_enabled_flag is equal to 1, both num_tile_set_columns_minus1 and num_tile_set_rows_minus1 shall not be equal to 0.
[0111] num_tile_columns_minus1[k][1] plus 1 specifies the number of tile columns that partition the tile set associated with the index (k, l). num_tile_columns_minus1[k][1] shall be in the range of 0 to PicWidthInCtbsY - 1, inclusive. num_tile_columns_minus1[k] If the value of [1] does not exist, the value of num_tile_columns_minus1[k][1] is assumed to be equal to 0.
[0112] In another example, num_tile_columns_minus1[k][1] shall be in the range of 0 to PicWidthInCtbsY - num_tile_set_columns_minus1 - 1, inclusive.
[0113] num_tile_rows_minus1[k][1] plus 1 specifies the number of tile rows that partition the tile set associated with the index (k, l). Assume num_tile_rows_minus1[k][1] is in the range of 0 to PicHeightlnCtbsY - 1, inclusive of both end values. If the value of num_tile_rows_minus1[k][1] does not exist, it is assumed that the value of num_tile_rows_minus1[k][1] is equal to 0.
[0114] In another example, assume num_tile_rows_minus1[k][1] is in the range of 0 to PicHeightInCtbsY - num_tile_set_rows_minus1 - 1, inclusive of both end values.
[0115] Equal to 1, uniform_spacing_flag[k][1] specifies that the tile column boundaries and, similarly, the tile row boundaries are uniformly distributed across the tile set associated with the index (k, l). Equal to 0, uniform_spacing_flag[k][1] specifies that the tile column boundaries and, similarly, the tile row boundaries are not uniformly distributed across the tile set associated with the index (k, l), but are explicitly signaled using the syntax elements column_width_minus1[k][1][i] and row_height_minus1[k][1] i]. If the value of uniform_spacing_flag[k][1] does not exist, it is assumed that the value of uniform_spacing_flag[k][1] is equal to 1.
[0116] column_width_minus1[i] plus 1 specifies the width of the i-th tile column in the unit of the coded tree block in the tile set associated with the index (k, l).
[0117] row_height_minus1[k][1][i] plus 1 specifies the height of the i-th tile row in units of the coded tree block within the tile set associated with the index (k, l).
[0118] tile_width_in_ctbsy_minus1[k][1] plus 1 specifies the width of each tile column within the tile set associated with the index (k, l) in units of the coded tree block.
[0119] In one embodiment, for each k in the range of 0 to num_tile_set_rows_minus1, it is a requirement for the bitstream compliance that the sum value of tile_width_in_ctbsy_minus1[k][1] for l in the range of 0 to num_tile_set_columns_minus1 is the same value.
[0120] tile_height_in_ctbsy_minus1[k][1] plus 1 specifies the height of each tile row within the tile set associated with the index (k, l) in units of the coded tree block.
[0121] In one embodiment, for each 1 in the range of 0 to num_tile_ser_columns_minus1, it is a requirement for the bitstream compliance that the sum value of tile_height_in_ctbsy_minus1[k][1] for k in the range of 0 to num_tile_set_rowss_minus1 is the same value.
[0122] A loop_filter_across_tiles_enabled_flag[k][l] equal to 1 specifies that the in-loop filtering operation can be performed across tile boundaries within the tile set associated with index (k, l) with reference to the PPS. A loop_filter_across_tiles_enabled_flag equal to 0 specifies that the in-loop filtering operation cannot be performed across tile boundaries within the tile set associated with index (k, l) with reference to the PPS. The in-loop filtering operation includes deblocking filter and sample adaptive offset filter operations. If the value of loop_filter_across_tiles_enabled_flag[k][1] does not exist, the value of loop_filter_across_tiles_enabled_flag[k][1] is assumed to be equal to 1.
[0123] In another example, the number of tile set columns per tile set row can be different for each tile set row. Table 4 shows an example of the syntax of a parameter set that can be used to signal the tile structure according to the techniques of this specification. In one embodiment, the exemplary syntax included in Table 4 may be included in the PPS. In other examples, the exemplary syntax included in Table 4 may be included in the VPS or SPS.
[0124] [Table 4]
[0125] Regarding Table 4, note that the syntax elements tiles_enabled_flag, tilesets_enabled_flag, num_tile_rows_minus1, uniform_spacing_flag[k][l], column_width_minus1[k][l][i], row_height_minus1[k][l][i], tile_width_in_ctbsy_minus1[k][l], tile_height_in_ctbsy_minus1[k][l], and loop_filter_across_tiles_enabled_flag[k][l] can be based on the definitions provided above regarding Table 3. num_tile_set_columns_minus1[k] can be based on the following exemplary definition.
[0126] num_tile_set_columns_minus1[k] plus 1 specifies the number of tile set columns in tile set row k. num_tile_set_columns_minus1[k] shall be in the range of 0 to PicWidthInCtbsY - 1, inclusive. If the value of num_tile_set_columns_minus1[k] does not exist, the value of num_tile_set_columns_minus1[k] is presumed to be equal to 0 for k in the range of 0 to num_tile_set_rows_minus1, inclusive.
[0127] In another example, the number of tile set rows per tile set column may be different for each tile set row. Table 5 shows an example of the syntax of a parameter set that can be used to signal the tile structure according to the technology of this specification. In one embodiment, the exemplary syntax included in Table 5 may be included in the PPS. In other examples, the exemplary syntax included in Table 5 may be included in the VPS or SPS.
[0128] [Table 5]
[0129] Regarding Table 5, note that the syntax elements tiles_enabled_flag, tilesets_enabled_flag, num_tile_columns_minus1, uniform_spacing_flag[k][l], column_width_minus1[k][l][i], row_height_minus1[k][l][i], tile_width_in_cthsy_minus1[k][l], tile_height_in_ctbsy_minus1[k][l], and loop_filter_across_tiles_enabled_flag[k][l] can be based on the definitions provided above regarding Table 3. num_tile_set_rows_minus1 [k] can be based on the following exemplary definitions.
[0130] num_tile_set_rows_minus1[l] plus 1 specifies the number of tileset rows in tileset column 1. num_tile_set_rows_minus1[l] shall be in the range of 0 to PicHeightlnCtbsY - 1, inclusive. If the value of num_tile_set_rows_minus1[k] does not exist, the value of num_tile_set_rows_minus1[1] is presumed to be equal to 0 for l in the range of 0 to num_tile_set_columns_minus1, inclusive.
[0131] For this example, in another example, the array indices [k][l] for the syntax elements num_tile_columns_mmusl[k][l], num_tile_rows_minus1[k][l], uniform_spacing_flag[k][l], tile_width_in_ctbsy_minus1[k][l], tile_height_in_ctbsy_minus1[k][l], and loop_filter_across_tiles_enabled_flag[k][l] may instead be signaled as indices in the order [l][k], and thus may be signaled as the syntax elements num_tile_columns_minus1[l][k], num_tile_rows_minus1[l][k], uniform_spacing_flag[l][k], tile_width_in_ctbsy_minus1[l][k], tile_height_in_ctbsy_minus1[l][k], loop_filter_across_tiles_enabled_flag[l][k].
[0132] Furthermore, for this example, in another example, the array indices [k][l][i] for the elements column_width_minus1[k][l][i], row_height_minus1[k][l][i] may instead be signaled as indices in the order [l][k][i], and thus may be signaled as the syntax elements column_width_minus1[l][k][i], row_height_minus1[l][k][i].
[0133] In one embodiment, according to the technology of this specification, the raster order of tiles may be row-by-row within a tile set, and the tile set is in raster order within an image. It should be noted that this can help form the encoded data within consecutive tile sets and facilitate the splicing of the bitstreams of the tile sets and the parallel decoding of those bitstream portions. As used herein, the term splicing may refer to the extraction of only portions of the overall bitstream where the extracted portions may correspond to one or more tile sets. In contrast, in ITU-T H.265, the raster order of tiles is row-by-row within an image. FIG. 6 shows a tile raster scan where the raster order of tiles is row-by-row within a tile set, and the tile set is in raster order within an image. In FIG. 6, the subscript numbers shown for each tile provide the order in which the tiles are scanned.
[0134] In one embodiment, according to the technology of this specification, the raster order of encoded tree blocks (CTB / CTU) may be row-by-row in the tile raster scan within a tile set, and the tile set is in raster order within a picture. It should be noted that this can help form the encoded data within consecutive tile sets and facilitate the splicing of the bitstreams of the tile sets and parallel decoding. As used herein, the term splicing may refer to the extraction of only portions of the overall bitstream where the extracted portions may correspond to one or more tile sets. In contrast, in ITU-T H.265, the raster order of encoded tree blocks (CTB / CTU) is row-by-row in the tile raster scan within a picture. FIG. 7 shows that the raster order of encoded tree blocks (CTB / CTU) may be row-by-row in the tile raster scan within a tile set, and the tile set is in raster order within a picture. In FIG. 7, the numbers shown for each CTU provide the order in which the CTUs are scanned. Note that the example of FIG. 7 includes the same tile structure as that shown in FIG. 6.
[0135] The raster order of encoded tree blocks (CTB / CTU), tiles, and tile sets may be based on the following description.
[0136] This sub - item specifies the order of VCL NAL units and their relationship to the coded picture. Each VCL NAL unit is part of a coded picture.
[0137] The order of VCL NAL units within a coded picture is constrained as follows.
[0138] - The first VCL NAL unit of a coded picture shall have a first_slice_segment_in_pic_flag equal to 1.
[0139] - sliceSegAddrA and sliceSegAddrB are the slice_segment_address values of any two coded slice segment NAL units A and B within the same coded picture. If any of the following conditions is true, coded slice segment NAL unit A shall precede coded slice segment NAL unit B.
[0140] - TileId[CtbAddrRsToTs[sliceSegAddrA]] is less than TileId[CtbAddrRsToTs[sliceSegAddrB]].
[0141] - TileId[CtbAddrRsToTs[sliceSegAddrA]] is equal to TileId[CtbAddrRsToTs[sliceSegAddrB]] and CtbAddrRsToTs[sliceSegAddrA] is less than CtbAddrRsToTs[sliceSegAddrB].
[0142] The list colWidtht[k][l][i] for k in the range of 0 to num_tile_set_rows_minus1 inclusive, l in the range of 0 to num_tile_set_columns_minus1 inclusive, and i in the range of 0 to num_tile_columns_minus1 inclusive, which specifies the width of the i-th tile column of the tile set associated with the index (k, l) in units of CTB, is derived as follows.
[0143]
Table 6
[0144] The list rowHeight[k][l][j] for k in the range of 0 to num_tile_set_rows_minus1 inclusive, l in the range of 0 to num_tile_set_columns_minus1 inclusive, and j in the range of 0 to num_tile_rows_minus1 inclusive, which specifies the height of the j-th tile row of the tile set associated with the index (k, l) in units of CTB, is derived as follows.
[0145]
Table 7
[0146] The variable NumTileSets indicating the number of tile sets is derived as follows.
[0147]
Number
[0148] Arrays pwctbsy[k][l], phctbsy[k][l], psizectbsy[k][l] that specify the picture width in the luma CTB, the picture height in the luma CTB, and the picture size in the luma CTB of the tile set associated with the index (k, l), respectively, for k in the range of 0 to num_tile_set_rows_minus1 including both end values, and for l in the range of 0 to num_tile_set_columns_minus1 including both end values, and array ctbAddrRSOffset[k][l] that specifies the cumulative count of CTBs up to the tile set associated with the index (k, l) are derived as follows.
[0149]
Table 8
[0150] The j-th tile set is associated with the index (k, l) as follows. The j-th tile set may be called the tile set with index j.
[0151] Considering the tile set index j, the numbers k and l of the tile set columns of the picture are derived as follows.
[0152]
Equation
[0153] Considering the indices k and l and the number of tile set columns of the picture, the tile set index j is derived as follows.
[0154]
Equation
[0155] For k in the range of 0 to num_tile_set_rows_minus1 including both end values, l in the range of 0 to num_tile_set_columns_minus1 including both end values, and i in the range of 0 to num_tile_columns_minus1 + 1 including both end values, which specify the location of the i-th tile row boundary of the tile set associated with the index (k, l) in units of the symbolized tree block, the list colBd[k][l][i] is derived as follows.
[0156]
Table 9
[0157] For k in the range of 0 to num_tile_set_rows_minus1 including both end values, l in the range of 0 to num_tile_set_columns_minus1 including both end values, and j in the range of 0 to num_tile_rows_minus1 + 1 including both end values, which specify the location of the j-th tile row boundary of the tile set associated with the index (k, l) in units of the symbolized tree block, the list rowBd[j] is derived as follows.
[0158]
Table 10
[0159] For ctbAddrRs in the range of 0 to PicSizelnCtbsY - 1 including both end values, which specifies the conversion from the CTB address in the CTB raster scan of the picture to the CTB address in the tile set and tile scan, the list CtbAddrRsToTs[ctbAddrRs] is derived as follows.
[0160]
Table 11
[0161] The list CtbAddrTsToRs[ctbAddrTs] which specifies the conversion from the CTB address in tile scanning to the CTB address in the CTB raster scan of the picture, and is in the range of 0 to PicSizeInCtbsY-1 including both end values of ctbAddrTs, is derived as follows.
[0162]
Table 12
[0163] The list TileId[ctbAddrTs] which specifies the conversion from the CTB address in tile scanning to the tile ID, and is in the range of 0 to PicSizeInCtbsY-1 including both end values of ctbAddrTs, is derived as follows.
[0164]
Table 13
[0165] In one embodiment according to the technology of this specification, the data encapsulation device 107 may be configured to signal information so that each tileset and tile within the tileset can be processed independently. In one embodiment, the byte range information of each tileset is signaled. In one embodiment, this may be signaled as a list of tileset entry point offsets. FIG. 8 shows an embodiment where T[j] and T[j + 1] respectively indicate the sizes of the bytes belonging to the j-th and (j + 1)-th tilesets. Additionally, for each tile within the tileset, the byte range information of each tileset is signaled. In one embodiment, this may be signaled as a list of tile entry point offsets. In FIG. 8, b[j][k] and b[j + 1][k] respectively indicate the sizes of the bytes belonging to the k-th tile subset of the j-th tileset and the k-th tile subset of the (j + 1)-th tileset. Table 6 shows an exemplary slice segment header that can be used to signal information so that each tileset and tile within the tileset can be processed independently according to the technology of this specification.
[0166]
Table 14
[0167] Regarding Table 6, note that the syntax elements tileset_ofiset_len_minus1, tileset_entry_point_offeet_minus1, num_entry_point_offsets, offset_len_minus1, and entry_point_offset_minus1 can be based on the following definitions.
[0168] tileset_offset_len_minus1 plus 1 specifies the length (in bits) of the tileset_entry_point_offset_minus1[i] syntax element. The value of tileset_offset_len_minus1 shall be in the range of 0 to 31, inclusive of both end values.
[0169] tileset_entry_point_offset_minus1[i] plus 1 specifies the i-th entry point offset (in bytes) for the i-th tileset and is represented by tileset_offset_len_minus1 plus 1 bits. The slice segment data follows the slice segment header and consists of NumTileSets subsets. The tileset index value is in the range of 0 to NumTileSets - 1, inclusive of both end values. The first byte of the slice segment data is considered byte 0. If present, the emulation prevention bytes that appear in the slice segment data part of the coded slice segment NAL unit are counted as part of the slice segment data for the purpose of subset identification. Tileset 0 consists of bytes 0 to tileset_entry_point_offset_minus1[0] (inclusive) of the coded slice segment data. For k having a value in the range of 1 to NumTileSets - 1 (inclusive), tileset k consists of bytes firstTilesetByte[k] to lastTileSetByte[k] (inclusive) of the coded slice segment data, where firstTileSetByte[k] and lastTileSetByte[k] are defined as follows.
[0170]
Table 15
[0171] The last subset (where the subset index is equal to NumTileSets) consists of the remaining bytes of the coded slice segment data, i.e., the bytes from lastTileSetByte[NumTileSets - 1]+1 to the end of the slice segment data.
[0172] Each subset shall consist of all the coded bits of all the coded tree units within the slice segment that are within the same tile set. num_entry_point_offsets[j] specifies the number of the entry_point_offset_minus1[j][i] syntax elements for the j-th tile set in the slice header. If the value of num_entry_point_offsets[j] does not exist, the value of num_entry_point_offsets[j] is assumed to be equal to 0.
[0173] The variables k and l are derived as follows.
[0174]
Number
[0175] The value of num_entry_point_offset[j] is constrained as follows.
[0176] - When - tiles_enabled_flag is equal to 0 and entropy_coding_sync_enabled_flag is equal to 1, the value of num_entry point_offsets[j] shall be in the range of 0 to PicHeightInCtbsY - 1, inclusive of both end values.
[0177] - Otherwise, if tiles_enabled_flag is equal to 1 and entropy_coding_sync_enabled_flag is equal to 0, the value of num_entry_point_offsets[j] shall be in the range of 0 to (num_tile_columns_minus1[k][l] + 1) including both end values. * It is assumed to be in the range of (num_tile_rows_minus1[k][l] + 1) - 1.
[0178] - Otherwise, if tiles_enabled_flag is equal to 1 and entropy_coding_sync_enabled_flag is equal to 1, the value of num_entry_point_offsets shall be in the range of 0 to (num_tile_colunms_minus1[k][l] + 1) including both end values. * It is assumed to be in the range of phctbsy[k][l] - 1.
[0179] offset_len_minus1[j] plus 1 specifies the length (bits) of the entry_point_offset_minus1[j][i] syntax element. The value of offset_len_minus1[j] shall be in the range of 0 to 31 including both end values.
[0180] entry_point_offset_minus1[i] plus 1 specifies the i-th entry point offset (in bytes) in the j-th tile set and is represented by offset_len_minus1[j] + 1 bits. The data in the slice segment data corresponding to the j-th tile set consists of a subset of num_entry_point_offsets[j] + 1 according to the position of tileset_entry_point_offset_minus1[j], and the subset index value of the j-th tile set ranges from 0 to num_entry_point_offsets[j] including both end values. The first byte of the slice segment data is regarded as byte 0. If present, the emulation prevention bytes that appear in the slice segment data part of the coded slice segment NAL unit are counted as part of the slice segment data for the purpose of subset identification. The subset 0 of the j-th tile set consists of bytes tileset_entry_point_offset_minus1[j] + 0 to tileset_entry_point_offset_minus1[j] + entry_point_offset_minus1[0] (including both end values) of the coded slice segment data, and the subset k having k in the range of 1 to num_entry_point_offsets[j] - 1 (including both end values) consists of bytes firstByte[j][k] to lastByte[j][k] (including both end values) of the coded slice segment data, where firstByte[j][k] and lastByte[j][k] are defined as follows.
[0181]
Number
[0182] The last subset (where the subset index is equal to [j][num_entry_point_offsets[j]]) consists of the bytes from lastByte[j][num_entry_point_offsets[j]-1]+l to tileset_entry_point_offset_minus1[j+1]-1, and includes the end values of the j-th coded slice segment data in the range of 0 to NumTileSets-1.
[0183] The last subset of the last tile set (where the subset index is equal to [NumTileSets-1][num_entry_point_offsets[NumTileSets]]) consists of the bytes of the coded slice segment data, i.e., the bytes from lastByte[NumTileSets-1][num_entry_point_offsets[j]-1]+l to the end of the slice segment data.
[0184] When tiles_enabled_flag is equal to 1 and entropy_coding_sync_enabled_flag is equal to 0, each subset consists of all the coded bits of all the coded tree units within a slice segment within the same tile, and the number of subsets (i.e., the value of num_entry_point_offsets[j]+1) is assumed to be equal to the number of tiles containing the coded tree units in the j-th tile set within the coded slice segment.
[0185] Note: When tiles_enabled_flag is equal to 1 and entropy_coding_sync_enabled_flag is equal to 0, each slice must either contain a subset of the coded tree units of one tile (in which case the syntax element entry_point_offset_minus1[j][i] does not exist) or must contain all the coded tree units of an integer number of complete tiles.
[0186] When tiles_enabled_flag is equal to 0 and entropy_coding_sync_enabled_flag is equal to 1, each subset k having k in the range of 1 to num_entry_point_offsets[j] (including both end values) consists of all the coded bits of all the coded tree units in the slice segment that includes the luma coded tree blocks within the same luma coded tree block row of the picture, and the number of subsets (i.e., the value of num_entry_point_offsets[j] + 1) is assumed to be equal to the number of luma coded tree block rows of the picture that include the coded tree units in the coded slice segment.
[0187] Note: The last subset (i.e., subset k equal to num_entry_point_offsets[j]) may or may not include all the coded tree units that include the luma coded tree blocks in the luma coded tree block row of the picture.
[0188] When tiles_enabled_flag is equal to 1 and entropy_coding_sync_enabled_flag is equal to 1, each subset k of the j-th tile set having k in the range of 0 to num_entry_point_offsets[j] (including both end values) consists of all the coded bits of all the coded tree units in the slice segment that includes the luma coded tree blocks within the same luma coded tree block row of the tiles in the j-th tile set, and the number of subsets (i.e., the value of num_entry_point_offsets[j] + 1) is assumed to be equal to the number of luma coded tree block rows of the tiles in the j-th tile set that include the coded tree units in the coded slice segment.
[0189] In one embodiment, the offset length information used for the fixed-length coding of the byte range signaling (tile offset signaling) of the tiles within each tile set can be signaled only once and is applied to all tile sets. The syntax of this example is shown in Table 7.
[0190]
Table 16
[0191] Regarding Table 7, the syntax elements tileset_offeet_len_minus1, tileset_entry_point_offset_minus1, num_entry_point_offsets, and offset_len_minus1 can be based on the definitions provided above regarding Table 6. Note that all_tile_offset_len_minus1 and entry_point_offset_minus1 can be based on the following definitions.
[0192] all_tile_offset_len_minus1 plus 1 specifies the length (in bits) of the entry_point_offset_minus1[j][i] syntax element for each value of j in the range 0 to NumTileSets - 1, inclusive of both end values. The value of all_tile_offset_len_minus1 shall be in the range 0 to 31, inclusive of both end values.
[0193] entry_point_offset_minus1[j][i] plus 1 specifies the i-th entry point offset (in bytes) for the j-th tile set and is represented by all_tile_offset_len_minus1 plus 1 bits.
[0194] In one embodiment, the tile byte range information may be signaled singly for the loop over all tiles in the picture. In this case, a single syntax element may be signaled for the number of byte ranges of the tiles being signaled. Next, other signaled syntax elements can be used to determine whether the number of these tile byte range elements belongs to each tile set.
[0195] In one embodiment, the syntax elements for the number of tile set columns in a picture (unm_tile_set_columns_minus1) and the number of tile set rows (num_tile_set_rows_minus1) may be of fixed length encoded using u(v) coding instead of ue(v) coding. In one embodiment, in this case, additional syntax elements may be signaled to indicate the length in bits used for the fixed length coding of these elements. In another example, the length in bits used for the coding of these syntax elements may not be signaled and instead may be inferred to be equal to the following.
[0196] Ceil(Log2(PicSizeInCtbsY)) bits, where PicSizeInCtbsY indicates the number of CTBs in the picture.
[0197] In one embodiment, the syntax elements for the tile width (tile_width_in_ctbsy_minus1[k][l]) of the number of CTBs of the tile set associated with the index (k, l) in a picture and the tile height (tile_height_in_ctbsy_minus1[k][l]) of the number of CTBs of the tile set associated with the index (k, l) may be of fixed length encoded using u(v) coding instead of ue(v) coding. In one embodiment, in this case, additional syntax elements may be signaled to indicate the length in bits used for the fixed length coding of these elements. In another example, the length in bits used for the coding of these syntax elements may not be signaled and instead may be inferred to be equal to the following.
[0198] Ceil(Log2(PicSizeInCtbsY)) bits, where PicSizeInCtbsY indicates the number of CTBs in the picture.
[0199] In one embodiment, the syntax elements for the column width in the CTB and / or the row height in the CTB may not be signaled for the last tile set column (num_tile_columns_minus1[num_tile_set_rows_minus1][num_tile_set_columns_minus1]) and / or the last tile set row (num_tile_rows_minus1[num_tile_set_rows_minus1][num_tile_set_columns_minus1]) within the picture of the last tile set. In this case, those values can be inferred from the picture height of the CTB and / or the picture width of the CTB.
[0200] Table 8 shows an example of the syntax of a parameter set that can be used to signal the tile structure according to the technology of this specification. In one embodiment, the exemplary syntax included in Table 8 may be included in the PPS. In other examples, the exemplary syntax included in Table 8 may be included in the VPS or SPS.
[0201] [Table 17]
[0202] Regarding Table 8, each syntax element can be based on the following definitions.
[0203] num_tile_columns_minus1 plus 1 specifies the number of tile columns partitioning the picture. num_tile_columns_minus1 is assumed to be in the range of 0 to PicWidthInCtbsY - 1, including both end values. If the value of num_tile_columns_minus1 does not exist, the value of num_tile_columns_minus1 is assumed to be equal to 0.
[0204] num_tile_rows_minus1 plus 1 specifies the number of tile rows that partition the picture. num_tile_rows_minus1 shall be in the range of 0 to PicHeightInCtbsY - 1, inclusive of both end values. If the value of num_tile_rows_minus1 does not exist, the value of num_tile_rows_minus1 is presumed to be equal to 0.
[0205] When tiles_enabled_flag is equal to 1, both num_tile_columns_minus1 and num_tile_rows_minus1 shall not be equal to 0.
[0206] A uniform_spacing_flag equal to 1 specifies that the tile column boundaries and likewise the tile row boundaries are distributed uniformly across the picture. A uniform_spacing_flag equal to 0 specifies that the tile column boundaries and likewise the tile row boundaries are not distributed uniformly across the picture, but are explicitly signaled using the syntax elements column_width_minus1[i] and row_height_minus1[i]. If the value of uniform_spacing_flag does not exist, the value of uniform_spacing_flag is presumed to be equal to 1.
[0207] column_width_minus1[i] plus 1 specifies the width of the i-th tile column in units of coded tree blocks.
[0208] row_height_minus1[i] plus 1 specifies the height of the i-th tile row in units of coded tree blocks.
[0209] The loop_filter_across_tiles_enabled_flag equal to 1 specifies that the in-loop filtering operation can be performed across tile boundaries within a picture with reference to the PPS. The loop_filter_across_tiles_enabled_flag equal to 0 specifies that the in-loop filtering operation cannot be performed across tile boundaries within a picture with reference to the PPS. The in-loop filtering operation includes deblocking filter and sample adaptive offset filter operations. If the loop_filter_across_tiles_enabled_flag does not exist, the value of the loop_filter_across_tiles_enabled_flag is presumed to be equal to 1.
[0210] The tilesets_enabled_flag equal to 1 specifies that there are two or more tile sets within each picture with reference to the PPS. The tilesets_enabled_flag equal to 0 specifies that there is only one tile set within each picture with reference to the PPS.
[0211] In another example, the tilesets_enabled_flag equal to 1 indicates the presence of the syntax elements num_tile_set_rows_minus1, num_tile_sets_columns_minus1, num_tile_rows_in_tileset_minus1[k], num_tile_columns_in_tileset_minus1[l]. The tilesets_enabled_flag equal to 0 indicates the absence of the syntax elements num_tile_set_rows_minus1, num_tile_sets_columns_minus1, num_tile_rows_in_tileset_minus1[k], num_tile_columns_in_tileset_minus1[l].
[0212] In one embodiment, when tilesets_enabled_flag is equal to 0, num_tile_set_rows_minus1 is presumed to be equal to 0 and num_tile_sets_columns_minus1 is presumed to be equal to 0 (i.e., the entire image is a single tile set).
[0213] In one embodiment, for all PPSs activated within the CVS, it is a requirement for bitstream compliance that the value of tilesets_enabled_flag be the same. When tiles_enabled_flag is equal to 0, tilesets_enabled_flag is presumed to be equal to 0.
[0214] In another example, when tiles_enabled_flag is equal to 0, it is a requirement for bitstream compliance that tilesets_enabled_flag be equal to 0.
[0215] num_tile_set_rows_minus1 plus 1 specifies the number of tile rows partitioning the picture. Assume num_tile_set_rows_minus1 is in the range of 0 to num_tile_rows_minus1, inclusive. In another example, assume num_tile_set_rows_minus1 is in the range of 0 to PicHeightInCtbsY - 1, inclusive. If the value of num_tile_set_rows_minus1 does not exist, the value of num_tile_set_rows_minus1 is presumed to be equal to 0.
[0216] num_tile_set_columns_minus1 plus 1 specifies the number of tile columns that partition the picture. num_tile_set_columns_minus1 shall be in the range of 0 to num_tile_columns_minus1, inclusive of both end values. In another example, num_tile_set_columns_minus1 shall be in the range of 0 to PicWidthInCtbsY - 1, inclusive of both end values. If the value of num_tile_set_columns_minus1 does not exist, the value of num_tile_set_columns_minus1 is assumed to be equal to 0.
[0217] In one embodiment, when tilesets_enabled_flag is equal to 1, both num_tile_set_columns_minus1 and num_tile_set_rows_minus1 shall not be equal to 0.
[0218] num_tile_rows_in_tileset_minus1[k] plus 1 specifies the number of tile rows in the tile set associated with the index (k, l) for each k in the range of 0 to num_tile_set_rows_minus1, inclusive of both end values. num_tile_rows_in_tileset_minus1[k] shall be in the range of 0 to num_tile_rows_minus1, inclusive of both end values. In another example, num_tile_rows_in_tileset_minus1[k] shall be in the range of 0 to PicHeightInCtbsY - 1, inclusive of both end values. If the value of num_tile_rows_in_tileset_minus1[k] does not exist, the value of num_tile_rows_in_tileset_minus1[k] is assumed to be equal to num_tile_rows_minus1. In one embodiment, if the value of num_tile_rows_in_tileset_minus1[k] does not exist, the value of num_tile_rows_in_tileset_minus1[k] is assumed to be equal to 0.
[0219] In one embodiment, for k in the range of 0 to num_tile_set_rows_minus1, including the end values, the sum of all (num_tile_rows_in_tileset_minus1[k]+1) is equal to (num_tile_rows_minus1+1), which is a requirement for the compliance of the bitstream.
[0220] In another example, num_tile_rows_minus1[k] is in the range of 0 to PicHeightInCtbsY - num_tile_set_rows_minus1 - 1, including the end values.
[0221] num_tile_columns_in_tileset_minus1[l] plus 1 specifies the number of tile rows in the tile set associated with the index (k, l) for each l in the range of 0 to num_tile_set_columns_minus1, including the end values. num_tile_columns_in_tileset_minus1[l] is in the range of 0 to num_tile_columns_minus1, including the end values. In another example, num_tile_columns_in_tileset_minus1[l] is in the range of 0 to PicWidthInCtbsY - 1, including the end values. If the value of num_tile_columns_in_tileset_minus1[l] does not exist, the value of num_tile_columns_in_tileset_minus1[l] is presumed to be equal to num_tile_columns_minus1. In one embodiment, if the value of num_tile_columns_in_tileset_minus1[l] does not exist, the value of num_tile_columns_in_tileset_minus1[l] is presumed to be equal to 0.
[0222] In one embodiment, for l in the range of 0 to num_tile_set_columns_minus1, inclusive of the end values, the sum of all (num_tile_columns_in_tileset_minus1[l]+1) is equal to (num_tile_columns_minus1+1), which is a requirement for the compliance of the bitstream.
[0223] In another example, num_tile_columns_minus1[l] is assumed to be in the range of 0 to PicWidthInCtbsY-num_tile_set_columns_minus1-1, inclusive of the end values.
[0224] In one embodiment, the tile structure syntax (i.e., the syntax elements from ITU-T H.265) may be signaled in the PPS, and the newly proposed tile set related syntax may be signaled in the SPS.
[0225] In one embodiment, when the syntax is signaled in the PPS as described above, the range of the tile set can be defined as follows: The set of pictures PPSassociatedPicSet is the set of all pictures that are consecutive in the decoding order and for which the associated PPS (by including slice_pic_parameter_set_id) is activated within the slice header. Then, the range of the tile set signaled within the PPS is the set of pictures PPSassociatedPicSet.
[0226] In another embodiment, In the PPS having a pps_pic_parameter_set_id value equal to PPSVaIA, the range of the tile sets signaled is consecutive in decoding order, and its slice header is a set of pictures having a slice_pic_parameter_set_id value equal to PPSVaIA. The pictures before and after this set of pictures in decoding order have slice_pic_parameter_set_id values not equal to PPSVaIA. There may be multiple such sets in the encoded video sequence.
[0227] In one embodiment, the encoded tree block raster and tile scan conversion process may be as follows.
[0228] For k in the range of 0 to num_tile_set_rows_minus1 inclusive, l in the range of 0 to num_tile_set_columns_minus1 inclusive, and i in the range of 0 to num_tile_columns_in_tileset_minus1[l] inclusive, which specify the width of the i-th tile column of the tile set associated with the index (k, l) in units of CTB, the list colWidth[k][l][i] is derived as follows.
[0229] [Table 18]
[0230] For k in the range of 0 to num_tile_set_rows_minus1 inclusive, l in the range of 0 to num_tile_set_columns_minus1 inclusive, and j in the range of 0 to num_tile_rows_in_tileset_minus][k] inclusive, which specify the height of the j-th tile row of the tile set associated with the index (k, l) in units of CTB, the list rowHeight[k][l][j] is derived as follows.
[0231]
Table 19
[0232] In another example, the above derivation can be performed as follows.
[0233]
Number
[0234] For k in the range of 0 to num_tile_set_rows_minus1 including both end values, l in the range of 0 to num_tile_set_columns_minus1 including both end values, and i in the range of 0 to num_tile_columns_in_tileset_minus1[l] including both end values, which specify the width of the i-th tile column of the tile set associated with the index (k, l) in units of CTB, the list colWidth[k][l][i] is derived as follows.
[0235]
Table 20
[0236] In another example, the above derivation can be performed as follows.
[0237]
Number
[0238] For k in the range of 0 to num_tile_set_rows_minus1 including both end values, l in the range of 0 to num_tile_set_columns_minus1 including both end values, and i in the range of 0 to num_tile_columns_in_tileset_minus1[l] including both end values, which specify the width of the i-th tile column of the tile set associated with the index (k, l) in units of CTB, the list colWidth[k][l][i] is derived as follows.
[0239]
Table 21
[0240] For k in the range of 0 to num_tile_set_rows_minus1, inclusive, for l in the range of 0 to num_tile_set_columns_minus1, inclusive, and for j in the range of 0 to num_tile_rows_in_tileset_minus1[k], inclusive, the list rowHeight[k][l][j] that specifies the height of the j-th tile row of the tile set associated with the index (k, l) in units of CTB is derived as follows.
[0241]
Table 22
[0242] In yet another example, the above derivation can be performed as follows.
[0243] For k in the range of 0 to num_tile_set_rows_minus1, inclusive, for l in the range of 0 to num_tile_set_columns_minus1, inclusive, and for i in the range of 0 to num_tile_columns_in_tileset_minus1[l], inclusive, the list colWidth[k][l][i] that specifies the width of the i-th tile column of the tile set associated with the index (k, l) in units of CTB is derived as follows.
[0244]
Table 23
[0245] The list rowHeight[k][l][j] that specifies the height of the j-th tile row of the tile set associated with the index (k, l) in units of CTB, for k in the range of 0 to num_tile_set_rows_minus1 including both end values, l in the range of 0 to num_tile_set_columns_minus1 including both end values, and j in the range of 0 to num_tile_rows_in_tileset_minus1[k] including both end values, is derived as follows.
[0246]
Table 24
[0247] The variable NumTileSets indicating the number of tile sets is derived as follows.
[0248]
Number
[0249] The arrays pwctbsy[k][l], phctbsy[k][l], psizectbsy[k][l] that respectively specify the picture width in the luma CTB, the picture height in the luma CTB, and the picture size in the luma CTB of the tile set associated with the index (k, l), for k in the range of 0 to num_tile_set_rows_minus1 including both end values and l in the range of 0 to num_tile_set_columns_minus1 including both end values, and the array ctbAddrRSOffset [k][l] that specifies the cumulative count of CTBs up to the tile set associated with the index (k, l), for k in the range of 0 to num_tile_set_rows_minus1 including both end values and l in the range of 0 to num_tile_set_columns_minus1 including both end values, is derived as follows.
[0250]
Table 25
[0251] In one embodiment, as described above,
[0252]
Table 26
[0253] The j-th tile set is associated with the index (k, l) as follows. The j-th tile set may be referred to as the tile set having the index j.
[0254] Considering the tile set index or tile set identifier j, the numbers k and l of the tile set columns of the picture are derived as follows.
[0255]
Equation
[0256] Considering the indices k and l and the number of tile set columns of the picture, the tile set index or tile set identifier j is derived as follows.
[0257]
Equation
[0258] For each tile set for k in the range of 0 to num_tile_set_rows_minus1 including both end values, and for l in the range of 0 to num_tile_set_columns_minus1 including both end values, the number of tiles is derived as follows.
[0259]
Table 27
[0260] A list colBd[k][l] for k in the range of 0 to num_tile_set_rows_minus1 including both end values, l in the range of 0 to num_tile_set_columns_minus1 including both end values, and i in the range of 0 to num_tile_columns_in_tileset_minus1[k]+1 including both end values, which specifies the location of the i-th tile column boundary of the tile set associated with the index (k, D) in units of symbol tree blocks. [i] is derived as follows.
[0261]
Table 28
[0262] A list rowBd[j] for k in the range of 0 to num_tile_set_rows_minus1 including both end values, l in the range of 0 to num_tile_set_columns_minus1 including both end values, and j in the range of 0 to num_tile_rows_in_tileset_minus1[k]+1 including both end values, which specifies the location of the j-th tile row boundary of the tile set associated with the index (k, l) in units of symbol tree blocks, is derived as follows.
[0263]
Table 29
[0264] A list CtbAddrRsToTs[ctbAddrRs] for ctbAddrRs in the range of 0 to PicSizeInCtbsY-1 including both end values, which specifies the conversion from the CTB address in the CTB raster scan of the picture to the CTB address in the tile set and tile scan, is derived as follows.
[0265]
Table 30
[0266] The list CtbAddrTsToRs[ctbAddrTs] which specifies the conversion from the CTB address in tile scanning to the CTB address in the CTB raster scan of the picture, with both end values included, and ctbAddrTs in the range of 0 to PicSizeInCtbsY - 1, is derived as follows.
[0267]
Table 31
[0268] The list TileId[ctbAddrTs] which specifies the conversion from the CTB address in tile scanning to the tile ID, with both end values included, and ctbAddrTs in the range of 0 to PicSizeInCtbsY - 1, is derived as follows.
[0269]
Table 32
[0270] In one embodiment, additional calculations may be required for sub-bitstream extraction or other processes. In one embodiment, the further calculations may be as follows.
[0271] For k in the range of 0 to num_tile_set_rows_minus1 with both end values included, and for l in the range of 0 to num_tile_set_columns_minus1 with both end values included, the number of tile columns minus 1 and the number of tile rows minus 1 are derived as follows.
[0272]
Table 33
[0273] Index in units of tiles ( *, l) The list tilecolumnPos[l] of 1 in the range of 0 to num_tile_columns_minus1 + 1, including both end values, which specifies the location of the tile set associated with the boundary, is derived as follows.
[0274]
Table 34
[0275] Index in terms of tile units (k, * ) The list tilerowPos[k] of k in the range of 0 to num_tile_rows_minus1 + 1, including both end values, which specifies the location of the tile set associated with the boundary, is derived as follows.
[0276]
Table 35
[0277] The list colBd [k][l], and The list rowBd [k][l] that specifies the location of the tile column boundary of the tile set associated with the index (k, l) in terms of the coded tree block unit, for k in the range of 0 to num_tile_set_rows_minus1, including both end values, and for l in the range of 0 to num_tile_set_columns_minus1, including both end values, is derived as follows.
[0278]
Table 36
[0279] Table 9 shows an exemplary slice segment header that can be used to signal information so that each of the tile sets and tiles within a tile set can be processed independently according to the techniques of this specification. In this case, a slice always contains a single complete tile set. In this case, the following signaling is done in the slice header (which may alternatively be called a tile set header or segment header, or some other similar name).
[0280] [Table 37]
[0281] Note that with respect to Table 9, the syntax elements tile_set_id, num_entry_point_offsets, offset_len_minus1, and entry_point_offset_minus1 can be based on the following definitions.
[0282] tile_set_id specifies the tile set identifier of this tile set. Considering the tile set indices k and l and the number of tile sets in the picture's tile set column, the tile set index or tile set identifier j is derived as follows.
[0283] [Equation]
[0284] The tile set identifier may alternatively be called a tile set index. The length of the tile_set_id syntax element is Ceil(Log2(NumTileSets)) bits.
[0285] num_entry_point_offsets, offset_len_minus1, and entry_offset_minus1[i] may have semantics similar to those provided in ITU-T H.265.
[0286] In one embodiment, when tiles_enabled_flag is equal to 1, the following is derived.
[0287] If tilesets_enabled_flag is equal to 1:
[0288]
Number
[0289] Otherwise (i.e., if tilesets_enabled_flag is equal to 0)
[0290]
Table 38
[0291] If tiles_enabled_flag is equal to 0: OffsetInfoPresent = 0, NumOffsets = 0 In another example, Table 10 shows an exemplary slice segment header that can be used to signal information so that each of the tilesets and tiles within a tileset can be processed independently according to the techniques of this specification. In this case, a slice may include an integer number of complete tilesets. In this case, the following signaling is done in the slice header (alternatively called a tileset header or segment header, or some other similar name).
[0292]
Table 39
[0293] Regarding Table 10, note that the syntax elements tile_set_id, num_entry_point_offsets, offset_len_minus1, and entry_point_offset_minus1 can be based on the following definitions.
[0294] tile_set_id specifies the tile set identifier for this tile set. Considering the tile set indices k and l and the number of tile set columns of the picture, the tile set index or tile set identifier j is derived as follows.
[0295]
Number
[0296] The tile set identifier may alternatively be called the tile set index.
[0297] The length of the tile_set_id syntax element is Ceil(Log2(NumTileSets)) bits.
[0298] num_tile_set_ids_minus1 plus 1 specifies the number of tile sets present in the slice (in raster scan order of the tile sets). The length of the num_tile_set_ids_minus1 syntax element is Ceil(Log2(NumTileSets - 1)) bits.
[0299] When tiles_enabled_flag is equal to 1, the following is derived.
[0300] When tilesets_enabled_flag is equal to 1:
[0301]
Table 40
[0302] Otherwise (i.e., if tilesets_enabled_flag is equal to 0)
[0303]
Table 41
[0304] When tiles_enabled_flag is equal to 0: OffsetInfoPresent = 0, NumOffsets = 0 It should be noted that "slice segment" may alternatively be referred to as "slice" or "tileset" or "tileset" or "segment" or "multi-ctu group" or "tile group" or "tile list" or "tile collection", etc. Therefore, in some cases, these words may be used interchangeably. Also, data structure names with similar names are interchangeable. "Slice segment header" may alternatively be referred to as "slice header" or "tileset header" or "tileset header" or "segment header" or "multi-ctu group header" or "tile group header" or "tile list header" or "tile collection header", etc. Therefore, these words are used interchangeably. Also, data structure names with similar names may be used interchangeably in some cases.
[0305] Table 11 shows an example of the syntax of a parameter set that can be used to signal the tile structure according to the technology of this specification. In one embodiment, the exemplary syntax included in Table 11 may be included in the PPS. In other examples, the exemplary syntax included in Table 11 may be included in the VPS or SPS or other parameter sets. In other examples, the exemplary syntax included in Table 11 can be included in the tile group header or slice header.
[0306]
Table 42
[0307] Regarding Table 11, each syntax element can be based on the following definitions.
[0308] A tilesets_enabled_flag equal to 1 specifies, with reference to the parameter set, that there are two or more tile sets within each picture. A tilesets_enabled_flag equal to 0 specifies, with reference to the parameter set, that there is only one tile set within each picture. In a variant, a tilesets_enabled_flag equal to 0 may specify that each tile is a tile set within each picture with reference to the PPS.
[0309] For all PPSs activated within the CVS, it is a requirement for bitstream compliance that the value of the tilesets_enabled_flag be the same. When the tiles_enabled_flag is equal to 0, the tilesets_enabled_flag is assumed to be equal to 0.
[0310] num_tile_sets_in_pic_minus1 plus 1 specifies the number of tile sets within the picture.
[0311] top_left_tile_id specifies the tile ID of the tile placed at the upper left corner of the i-th tile set. The length of top_left_tile_id[i] is Ceil(Log2(num_tilesets_in_pic_minus1 + 1)) bits. For any i not equal to j, the value of top_left_tile_id[i] shall not be equal to the value of top_left_tile_id[j].
[0312] num_tile_rows_in_tileset_minus1[i] plus 1 specifies the number of tile rows in the i-th tile set for each i in the range of 0 to (num_tile_sets in pic minus 1) including both end values. num_tile_rows_in_tileset_minus1[i] shall be in the range of 0 to num_tile_rows_minus1 including both end values. The length of num_tile_rows_in_tileset_minus1[i] is Ceil(Log2 (num_tile_rows_minus1 + 1)) bits.
[0313] If the value of num_tile_rows_in_tileset_minus1[i] does not exist, the value of num_tile_rows_in_tileset_minus1[i] is presumed to be equal to 0.
[0314] In a variant, if the value of num_tile_rows_in_tileset_minus1[i] does not exist, the value of num_tile_rows_in_tileset_minus1[i] is presumed to be equal to num_tile_rows_minus1.
[0315] num_tile_columns_in_tileset_minus1[l] plus 1 specifies the number of tile columns in the i-th tile set for each i in the range of 0 to (num_tile_sets_in_pic_minus1) including both end values. num_tile_columns_in_tileset_minus1[l] shall be in the range of 0 to num_tile_columns_minus1 including both end values. The length of num_tile_columnss_in_tileset_minus1[i] is Ceil(Log2(num_tile_columns_minus1 + 1)) bits.
[0316] If the value of num_tile_columns_in_tileset_minus1[l] does not exist, the value of num_tile_columns_in_tileset_minus1[l] is assumed to be equal to 0.
[0317] In a modification, if the value of num_tile_columns_in_tileset_minus1[i] does not exist, the value of num_tile_columns_in_tileset_minus1[l] is assumed to be equal to num_tile_columns_minus1.
[0318] In a modification, num_tile_rows_m_tileset_minus1[i] and num_tile_columns_in_tileset_minus1[l] are ue(v) coded.
[0319] In a modification, one or two separate additional syntax elements signal the number of bits used for num_tile_rows_in_tileset_minus1[i] and / or num_tile_columns_in_tileset_minus1[l].
[0320] Table 12 shows an example of the syntax of a parameter set that can be used to signal the tile structure according to the technology of this specification. In one embodiment, the exemplary syntax included in Table 12 may be included in the PPS. In other examples, the exemplary syntax included in Table 12 may be included in the VPS or SPS or other parameter sets. In other examples, the exemplary syntax included in Table 12 can be included in the tile group header or slice header.
[0321] [Table 43]
[0322] Regarding Table 12, each syntax element can be based on the definitions provided above and the following definitions.
[0323] The remaining_tiles_tileset_flag equal to 1 specifies that all the remaining tiles in the picture other than those explicitly specified within the (num_tilesets_in_pic_minus1 - 1) tile sets signaled by the syntax elements top_left_tile_id[i], num_tile_rows_in_tileset_minus1[i], num_tile_columns_in_tileset_minus1[i] form the last tile set. The remaining_tiles_tileset_flag equal to 0 specifies that all the num_tilesets_in_pic_minus1 tile sets are explicitly specified by signaling the syntax elements top_left_tile_id[i], num_tile_rows_in_tileset_minus1[i], num_tile_columns_in_tileset_minus1[i].
[0324] For each i in the range from 0 to (num_tile_sets_in_pic_minus1 +!remaining_tiles_tileset_flag - 1) inclusive, num_tile_rows_in_tileset_minus1[i] plus 1 specifies the number of tile rows in the i-th tile set. num_tile_rows_in_tileset_minus1[i] shall be in the range from 0 to num_tile_rows_minus1 inclusive. The length of num_tile_rows_in_tileset_minus1[i] is Ceil(Log2(num_tile_rows_minus1 + 1)) bits.
[0325] If the value of num_tile_rows_in_tileset_minus1[i] does not exist, the value of num_tile_rows_in_tileset_minus1[i] is assumed to be equal to 0.
[0326] In the modification example, when the value of num_tile_rows_in_tileset_minus1[i] does not exist, the value of num_tile_rows_in_tileset_minus1[i] is presumed to be equal to num_tile_rows_minus1.
[0327] num_tile_columns_in_tileset_minus1[l] plus 1 specifies the number of tile columns in the i-th tile set for each i in the range of 0 to (num_tile_sets_in_pic_minus1 +!remainmg_tiles_tileset_flag - 1) including both end values. num_tile_columns_in_tileset_minus1[l] is assumed to be in the range of 0 to num_tile_columns_minus1 including both end values. The length of num_tile_columnss_in_tileset_minus1[i] is Ceil(Log2(num_tile_columns_minus1 + 1)) bits.
[0328] When the value of num_tile_columns_in_tileset_minus1[l] does not exist, the value of num_tile_columns_in_tileset_minus1[l] is presumed to be equal to 0.
[0329] In the modification example, when the value of num_tile_columns_in_tileset_minus1[i] does not exist, the value of num_tile_columns_in_tileset_minus1[l] is presumed to be equal to num_tile_columns_minus1.
[0330] In the modification example, it is a requirement for the conformity of the bit stream that each tile in the picture belongs to and only belongs to one of the tile sets having tile sets in the range of 0 to num_tile_set_minus1 (including both end values).
[0331] Table 12A shows an example of the syntax of a parameter set that can be used to signal the tile structure according to the technology of this specification. In one embodiment, the exemplary syntax included in Table 12A may be included in the PPS. In other examples, the exemplary syntax included in Table 12A may be included in the VPS or SPS or other parameter sets. In other examples, the exemplary syntax included in Table 12A can be included in the tile group header or slice header.
[0332] [Table 44]
[0333] Regarding Table 12A, each syntax element can be based on the definitions provided above.
[0334] Table 13 shows an example of the syntax of a parameter set that can be used to signal the tile structure according to the technology of this specification. In one embodiment, the exemplary syntax included in Table 13 may be included in the PPS. In other examples, the exemplary syntax included in Table 13 may be included in the VPS or SPS or other parameter sets. In other examples, the exemplary syntax included in Table 13 can be included in the tile group header or slice header.
[0335] [Table 45]
[0336] Regarding Table 13, each syntax element can be based on the definitions provided above and the following definitions.
[0337] bottom_right_tile_id[i] specifies the tile ID of the tile placed at the bottom - right corner of the i - th tile set. The length of bottom_right_tile_id[i] is Ceil(Log2(num_tilesets_in_pic_minus1 + 1)) bits. For any i not equal to j, the value of bottom_right_tile_id[i] shall not be equal to the value of bottom_right_tile_id[j].
[0338] Table 14 shows an example of the syntax of a parameter set that can be used to signal the tile structure according to the technology of this specification. In one embodiment, the exemplary syntax included in Table 14 may be included in the PPS. In other examples, the exemplary syntax included in Table 14 may be included in the VPS or SPS or other parameter sets. In other examples, the exemplary syntax included in Table 14 can be included in the tile group header or slice header.
[0339] [Table 46]
[0340] Regarding Table 14, each syntax element can be based on the definitions provided above.
[0341] Regarding Tables 11 - 14, Table 14A shows an exemplary syntax of the tile group header.
[0342] [Table 47]
[0343] Regarding Table 14A, each syntax element can be based on the definitions provided above and the following definitions.
[0344] The tile_set_idx specifies the tile set index of this tile set. The length of the tile_set_idx syntax element is Ceil(Log2(NumTileSets)) bits.
[0345] Regarding Table 11, Table 15 shows an exemplary syntax of tile group data.
[0346]
Table 48
[0347] Regarding Table 15, in one embodiment, the conversion from the CTB address to the tile ID in tile scanning can be as follows.
[0348]
Table 49
[0349] Regarding Tables 12 to 14, Table 16 shows an exemplary syntax of tile group data.
[0350]
Table 50
[0351] Regarding Tables 12 to 14, Table 16A shows another exemplary syntax of the tile group header. The main difference between Table 16 and Table 16A is that some of the syntax elements are replaced by derived variables.
[0352]
Table 51
[0353] Regarding Tables 12, 16, and 16A, in one embodiment, the conversion from the CTB address to the tile ID in tile scanning can be as follows.
[0354] [Table 52]
[0355] [Table 53]
[0356] With reference to Tables 12A and 16 and 16A, in one embodiment, the conversion from CTB address to tile ID in a tile scan may be as follows:
[0357] [Table 54]
[0358] [Table 55]
[0359] With reference to Tables 13, 16 and 16A, in one embodiment, the conversion from CTB address to tile ID in a tile scan may be as follows:
[0360] The list TileId[ctbAddrTs], which is a ctbAddrTs in the range of 0 to PicSizeInCtbsY-1, inclusive, that specifies the conversion from CTB addresses to tile IDs in tile scanning, and the list NumCtusInTile[tileIdx], which is a tileIdx in the range of 0 to PicSizeInCtbsY-1, inclusive, that specifies the conversion from tile indexes in a tile to the number of CTUs, are derived as follows:
[0361] [Table 56]
[0362] [Table 57]
[0363] Regarding Tables 14, 16, and 16A, in one embodiment, the conversion from CTB address to tile ID in tile scanning can be as follows.
[0364] A list TileId[ctbAddrTs] of ctbAddrTs in the range of 0 to PicSizeInCtbsY - 1, including both end values, specifying the conversion from CTB address to tile ID in tile scanning, and a list NumCtusInTile[tileIdx] of tileIdx in the range of 0 to PicSizeInCtbsY - 1, including both end values, specifying the conversion from tile index in a tile to the number of CTUs, are derived as follows.
[0365] [Table 58]
[0366] [Table 59]
[0367] In one embodiment, a flag indicating that each tile group consists of only one tile can be signaled. The signaling of syntax elements for the number of tiles within a tile group in the tile group header may be adjusted based on this flag. Thereby, the source of bits is provided. Table 17 shows an example of the syntax of a picture parameter set that can be used to signal a tile structure including a flag indicating that each tile group consists of only one tile.
[0368] [Table 60]
[0369] Regarding Table 17, each syntax element can be based on the definitions provided above and the following definitions.
[0370] A transform_skip_enabled_flag equal to 1 specifies that the transform_skip_flag may be present in the residual coding syntax. A transform_skip_enabled_flag equal to 0 specifies that the transform_skip_flag is not present in the residual coding syntax.
[0371] A single_tile_in_pic_flag equal to 1 specifies that, referring to the PPS, only one tile exists within each picture. A single_tile_in_pic_flag equal to 0 specifies that, referring to the PPS, two or more tiles exist within each picture.
[0372] For all PPSs activated within the CVS, it is a requirement for bitstream compliance that the values of the single_tile_in_pic_flag be the same.
[0373] A one_tile_per_tile_group equal to 1 specifies that each tile group contains one tile. A one_tile_per_tile_group equal to 0 specifies that the tile group may contain two or more tiles.
[0374] In a variant, A one_tile_per_tile_group equal to 1 specifies that each tile group contains one tile. A one_tile_per_tile_group equal to 0 specifies that at least one tile group contains two or more tiles.
[0375] tile_column_width_minus1 plus 1 specifies the width of the i-th tile column in units of the coding tree block.
[0376] tile_row_height_minus1[i] plus 1 specifies the height of the i-th tile row in units of the coded tree block.
[0377] Regarding Table 17, Table 18 shows an exemplary syntax of the tile group header.
[0378]
Table 61
[0379] Regarding Table 18, each syntax element can be based on the definitions provided above and the following definitions.
[0380] If it exists, the value of the tile group header syntax element tile_group_pic_parameter_set_id shall be the same for all tile group headers of the coded picture.
[0381] tile_group_pic_parameter_set_id specifies the value of pps_pic_parameter_set_id of the PPS in use. The value of tile_group_pic_parameter_set_id shall be in the range of 0 to 63, including both end values.
[0382] tile_group_address specifies the tile address of the first tile in the tile group. The length of tile_group_address is Ceil(Log2(NumTilesInPic)) bits. The value of tile_group_address shall be in the range of 0 to NumTilesInPic - 1, including both end values, and the value of tile_group_address shall not be equal to the value of tile_group_address of any other coded tile group NAL unit of the same coded picture. If tile_group_address does not exist, it is assumed to be equal to 0.
[0383] num_tiles_in_tile_group_minus1 plus 1 specifies the number of tiles in the tile group. The value of the number of tiles in tile group_minus1 shall be in the range of 0 to NumTilesInPic - 1, including both end values. If the value of num_tiles_in_tile_group_minus1 does not exist, the value of num_tiles_in_tile_group_minus1 is presumed to be equal to 0.
[0384] In a variant, the syntax element num_tiles_in_tile_group_minus1 for the number of tiles in tile group (minus1) is signaled using fixed-length coding (i.e., u(v)) instead of variable-length coding (i.e., ue(v)). This may enable easier parsing at the system level. Such num_tiles_in_tile_group_minus1 can be based on the following definition.
[0385] num_tiles_in_tile_group_minus1 plus 1 specifies the number of tiles in the tile group. The length of num_tiles_in_tile_group_minus1 is Ceil(Log2(NumTilesInPic)) bits. If the value of num_tiles_in_tile_group_minus1 does not exist, the value of num_tiles_in_tile_group_minus1 is presumed to be equal to 0.
[0386] In one embodiment, the signaling of the number of tiles in the tile group may be adjusted based on the single_tile_in_pic_flag syntax element instead of a variable derived from NumTileInPic. Using a syntax element for signaling simplifies the parsing of the tile group header by not requiring the derivation and use of additional variables to determine whether the syntax element is included. Table 19 shows an exemplary syntax of the tile group header for this embodiment.
[0387]
Table 62
[0388] Regarding Table 19, each syntax element can be based on the definitions provided above.
[0389] Table 20 shows another exemplary syntax of the tile group header.
[0390]
Table 63
[0391] Regarding Table 20, each syntax element can be based on the definitions provided above and the following definitions.
[0392] tile_set_idx specifies the tile set index of this tile set. The length of the tile_set_id syntax element is Ceil(Log2(NumTileSets)) bits.
[0393] As described above, ITU-T H.265 defines signaling that enables motion-constrained tile sets, and the motion-constrained tile sets may include tile sets where the inter-picture prediction dependency is limited to collocated tile sets within the reference picture. In one embodiment, a flag indicating whether the tile set is an MCTS can be signaled. Table 21 indicates that an example of the syntax of the picture parameter set that can be used to signal a tile structure including a flag indicating whether the tile set is an MCTS can be signaled.
[0394]
Table 64
[0395] Regarding Table 21, each syntax element can be based on the following definitions.
[0396] pps_pic_parameter_set_id specifies the PPS for reference by other syntax elements. The value of pps_pic_parameter_set_id shall be in the range of 0 to 63, inclusive of both end values.
[0397] pps_seq_parameter_set_id specifies the value of sps_seq_parameter_set_id for the active SPS. The value of pps_seq_parameter_set_id shall be in the range of 0 to 15, inclusive of both end values.
[0398] transform_skip_enabled_flag equal to 1 specifies that transform_skip_flag may exist in the residual coding syntax. transform_skip_enabled_flag equal to 0 specifies that transform_skip_flag does not exist in the residual coding syntax.
[0399] single_tile_in_pic_flag equal to 1 specifies that, referring to the PPS, only one tile exists within each picture. single_tile_in_pic_flag equal to 0 specifies that, referring to the PPS, two or more tiles exist within each picture.
[0400] For all PPSs activated within the CVS, it is a requirement for bitstream compliance that the value of single_tile_in_pic_flag be the same.
[0401] num_tile_columns_minus1 plus 1 specifies the number of tile columns that partition the picture. num_tile_columns_minus1 shall be in the range of 0 to PicWidthInCtbsY - 1, inclusive of both end values. If the value of num_tile_columns_minus1 does not exist, the value of num_tile_columns_minus1 is presumed to be equal to 0.
[0402] num_tile_rows_minus1 plus 1 specifies the number of tile rows that partition the picture. num_tile_rows_minus1 shall be in the range of 0 to PicHeightInCtbsY - 1, inclusive of both end values. If the value of num_tile_rows_minus1 does not exist, the value of num_tile_rows_minus1 is presumed to be equal to 0.
[0403] The variable NumTilesInPic is set equal to (num_tile_columns_minus1 + 1) * (num_tile_rows_minus1 + 1).
[0404] If single_tile_in_pic_flag is 0, NumTilesInPic shall be greater than 0.
[0405] tile_id_len_minus1 plus 1, if it exists, specifies the number of bits used to represent the syntax elements tile_id_val[i][j], top_left_tile_id[i], and bottom_right_tile_id[i] in the PPS. The value of tile_id_len_minus1 shall be in the range of Ceil(Log2(NumTilesInPic)) to 15, inclusive of both end values.
[0406] An explicit_tile_id_flag equal to 1 specifies that the tile ID of each tile is explicitly signaled. An explicit_tile_id_flag equal to 0 specifies that the tile ID is not explicitly signaled.
[0407] tile_id_val[i][j] specifies the tile ID of the tile at the i-th tile row and the j-th tile column. The length of tile_id_val[i][j] is tile_id_len_minus1 + 1 bits.
[0408] For any integer m in the range from 0 to num_tile_column_minus1 (including both end values), and any integer n in the range from 0 to num_titile_row_minus1 (including both end values), if i is not equal to m or j is not equal to n, tile_id_val[i][j] is not equal to tile_id_val[m][n]. * (num_tile_columns_minus1 + 1) + i is n * If (num_tile_columns_minus1 + 1) + m is less than (num_tile_columns_minus1 + 1) + i, then tile_id_val[i][j] is less than tile_id_val[m][n].
[0409] A uniform_tile_spacing_flag equal to 1 specifies that the tile column boundaries and, similarly, the tile row boundaries are uniformly distributed across the picture. A uniform_tile_spacing_flag equal to 0 specifies that the tile column boundaries and, similarly, the tile row boundaries are not uniformly distributed across the picture, but are explicitly signaled using the syntax elements tile_column_width_minus1[i] and tile_row_height_minus1[i]. If the uniform_tile_spacing_flag does not exist, its value is assumed to be equal to 1.
[0410] tile_column_width_minus1[i] plus 1 specifies the width of the i-th tile column in units of CTB.
[0411] tile_row_height_minus1[i] plus 1 specifies the height of the i-th tile row in units of CTB. The following variables are derived by calling the CTB raster and tile scan conversion processes: - A list ColWidth[i] for i in the range 0 to num_tile_columns_minus1, inclusive, specifying the width of the i-th tile column in units of CTB, - A list RowHeight[j] for j in the range 0 to num_tile_rows_minus1, inclusive, specifying the height of the j-th tile row in units of CTB, - A list ColBd[i] for i in the range 0 to num_tile_columns_minus1 + 1, inclusive, specifying the location of the i-th tile column boundary in units of CTB, - A list RowBd[j] for j in the range 0 to num_tile_rows_minus1 + 1, inclusive, specifying the location of the j-th tile row boundary in units of CTB, A list CtbAddrRsToTs[ctbAddrRs] for ctbAddrRs in the range 0 to PicSizeInCtbsY - 1, inclusive, specifying the conversion from the CTB address in the CTB raster scan of the picture to the CTB address in the tile scan, - A list CtbAddrTsToRs[ctbAddrTs] for ctbAddrTs in the range 0 to PicSizeInCtbsY - 1, inclusive, specifying the conversion from the CTB address in the tile scan to the CTB address in the CTB raster scan of the picture, - A list TileId[ctbAddrTs] for ctbAddrTs in the range 0 to PicSizeInCtbsY - 1, inclusive, specifying the conversion from the CTB address in the tile scan to the tile ID, - A list NumCtusInTile[tileIdx] of tileIdx in the range of 0 to PicSizeInCtbsY-1, including both end values, which specifies the conversion from the tile index in the tile to the number of CTUs. - A set [TileIdToIdx[tileId] for a set of NumTilesInPic tileId values that specifies the conversion from the tile ID to the tile index, and a list FirstCtbAddrTs[tileIdx] for tileIdx in the range of 0 to NumTilesInPic-1, including both end values, which specifies the conversion from the tile ID to the CTB address in the tile scan of the first CTB in the tile. - A list ColumnWidthInLumaSamples[i] for i in the range of 0 to num_tile_columns_minus1, including both end values, which specifies the width of the i-th tile column in units of luma samples. - A list RowHeightInLumaSamples[j] for j in the range of 0 to num_tile_rows_minus1, including both end values, which specifies the height of the j-th tile row in units of luma samples.
[0412] The values of ColumnWidthInLumaSamplesf i ] for i in the range of 0 to num_tile_columns_minus1, including both end values, and the values of RowHeightInLumaSamples[j] for j in the range of 0 to num_tile_rows_minus1, including both end values, shall all be greater than 0.
[0413] A tile_set_flag equal to 0 specifies that each tile is a tile set. A tile_set_flag equal to 1 specifies that the tile set is explicitly specified by the syntax elements num_tile_sets_in_pic_minus1, top_left_tile_id[i], and bottom_right_tile_id[i].
[0414] In another example, a tile_set_flag equal to 0 specifies that each picture is 1 tile.
[0415] num_tilesets_in_pic_minus1 plus 1 specifies the number of tile sets in the picture. The value of num_tilesets_in_pic_minus1 is assumed to be in the range of 0 to (NumTilesInPic - 1) inclusive of both ends. If it does not exist, num_tilesets_in_pic_minus1 is assumed to be (NumTilesInPic - 1). In another embodiment, if it does not exist, num_tilesets_in_pic_minus1 is assumed to be equal to 0.
[0416] A signaled_tile_set_index_flag equal to 1 specifies that the tile set index of each tile set is signaled. A signaled_tile_set_index_flag equal to 0 specifies that the tile set index is not signaled.
[0417] A remaining_tiles_tileset_flag equal to 1 specifies all the remaining tiles in the tile set other than those explicitly specified in the (num_tilesets_in_pic_minus1 - 1) tile sets signaled by the syntax elements top_left_tile_id[i], num_tile_rows_in_tileset_minus1[i], num_tile_columns_in_tileset_minus1[i].
[0418] Form the final tile set. A remaining_tiles_tileset_flag equal to 0 specifies that all num_tilesets_in_pic_minus1 tile sets are explicitly specified by signaling the syntax elements top_left_tile_id[i], num_tile_rows_in_tileset_minus1[i], and num_tile_columns_in_tileset_minus1[i].
[0419] top_left_tile_id[i] specifies the tile ID of the tile placed at the upper left corner of the i-th tile set. The length of top_left_tile_id[i] is tile_id_len_minus1 + 1 bits. For any i not equal to j, the value of top_left_tile_id[i] shall not be equal to the value of top_left_tile_id[j].
[0420] bottom_right_tile_id[i] specifies the tile ID of the tile placed at the lower right corner of the i-th tile set. The length of bottom_right_tile_id[i] is tile_id_len_minus1 + 1 bits.
[0421] The variables NumTileRowsInSlice[top_left_tile_id[i]], NumTileColumnsInSlice[top_left_tile_id[i]], and NumTilesInSlice[top_left_tile_id[i]] are derived as follows:
[0422]
Table 65
[0423] An is_mcts_flag equal to 1 specifies that the i-th tile set is a motion-constrained tile set. An is_mcts_flag equal to 0 specifies that the i-th tile set is not a motion-constrained tile set.
[0424] In one embodiment, an is_mcts_flag equal to 0 specifies that the i-th tile set may or may not be a motion-constrained tile set.
[0425] If the i-th tile set is a motion-constrained tile set, it may have one or more of the constraints such as those described in ITU-T H.265, Clause D.3.30 (i.e., the Temporal Motion Constraint Tile Set SEI message). In one embodiment, the motion-constrained tile set may be referred to as the temporal motion constraint tile set.
[0426] tile_set_index[i] specifies the tile set index of the i-th tile set. The length of the tile_set_index[i] syntax element is Ceil(Log2(num_tile_sets_in_pic_minus1 + 1)) bits. If the value of tile_set_index[i] does not exist, the value of tile_set_index[i] is assumed to be equal to i for each i in the range of 0 to num_tile_sets_in_pic_minus1, inclusive of both end values.
[0427] The loop_filter_across_tiles_enabled_flag equal to 1 specifies that the in-loop filtering operation can be performed across tile boundaries within a picture with reference to the PPS. The loop_filter_across_tiles_enabled_flag equal to 0 specifies that the in-loop filtering operation cannot be performed across tile boundaries within a picture with reference to the PPS. The in-loop filtering operation includes a deblocking filter, a sample adaptive offset filter, and an adaptive loop filter operation. If the loop_filter_across_tiles_enabled_flag does not exist, the value of the loop_filter_across_tiles_enabled_flag is assumed to be equal to 1.
[0428] In one embodiment, the corresponding part of Table 21 may be changed as shown in Table 21A below.
[0429]
Table 66
[0430] Each syntax element may have the definitions provided above. Further, with respect to Table 21, the syntax of the tile group header may be as shown in Table 21B below.
[0431]
Table 67
[0432] With respect to Table 21B, each syntax element can be based on the following semantics and definitions.
[0433] If it exists, the value of the tile_group_pic_parameter_set_id of the tile group header syntax element shall be the same in all tile group headers of the coded picture.
[0434] The tile_group_pic_parameter_set_id specifies the value of pps_pic_parameter_set_id of the PPS in use. The value of tile_group_pic_parameter_set_id shall be in the range of 0 to 63, including both end values.
[0435] The tile_set_idx specifies the tile set index of this tile set. The length of the tile_set_idx syntax element is Ceil(Log2(num_tile_sets_in_pic_minus1+1)) bits.
[0436] In another example, tile_set_idx may be called tile_group_idx and has the following semantics: The tile_group_idx specifies the tile set index of this tile set. The length of the tile_group_idx syntax element is Ceil(Log2(num_tile_sets_in_pic_minus1+1)) bits.
[0437] The tile_group_type specifies the encoding type of the tile group according to Table 21C.
[0438]
Table 68
[0439] When nal_unit_type is equal to IRAP_NUT, i.e., the picture is an IRAP picture and the tile_group type is equal to 2.
[0440] log2_diff_ctu_max_bt_size specifies the difference between the luma CTB size and the maximum luma size (width or height) of the coded block that can be split using binary splitting. The value of log2_diff_ctu_max_bt_size shall be in the range of 0 to CtbLog2SizeY - MinCbLog2SizeY, inclusive of both endpoints.
[0441] If the value of log2_diff_ctu_max_bt_size does not exist, the value of log2_diff_ctu_max_bt_size is assumed to be equal to 2.
[0442] The variables MinQtLog2SizeY, MaxBtLog2SizeY, MinBtLog2SizeY, MaxTtLog2SizeY, MinTtLog2SizeY, MaxBtSizeY, MinBtSizeY, MaxTtSizeY, MinTtSizeY, and MaxMttDepth are derived as follows:
[0443] [Table 69]
[0444] An sbtmvp_size_override_flag equal to 1 specifies that the syntax element log2_sbtmvp_active_size_minus2 exists for the current tile group. An sbtmvp_size_override_flag equal to 0 specifies that the syntax element log2_atmvp_active_size_minus2 does not exist and it is assumed that log2_sbtmvp_size_active_minus2 is equal to log2_sbtmvp_default_size_minus2.
[0445] log2_sbtmvp_active_size_minus2 plus 2 specifies the value of the sub-block size used to derive the motion parameters of the sub-block based TMVP of the current tile group. If log2_sbtmvp_size_active_minus2 does not exist, it is assumed to be equal to log2_sbtmvp_default_size_minus2. The variable is derived as follows.
[0446]
Number
[0447] tile_group_temporal_mvp_enabled_flag specifies whether the temporal motion vector predictor can be used for inter prediction. If tile_group_temporal_mvp_enabled_flag is equal to 0, the syntax elements of the current picture are assumed to be constrained such that the temporal motion vector predictor is not used for decoding the current picture. Otherwise (tile_group_temporal_mvp_enabled_flag is equal to 1), the temporal motion vector predictor may be used for decoding the current picture. If the value of log2_sbtmvp_default_size_minus2 does not exist, the value of log2_sbtmvp_default_size_minus2 is assumed to be equal to 0.
[0448] mvd_l1_zero_flag equal to 1 indicates that the mvd_coding(x0,y0,1) syntax structure is not parsed and MvdL1[x0][y0][compIdx] is set to 0 for compIdx = 0~1. mvd_l1_zero_flag equal to 0 indicates that the mvd_coding(x0,y0,1) syntax structure has been parsed.
[0449] A collocated_from_10_flag equal to 1 specifies that the collocated picture used for temporal motion vector prediction is derived from reference picture list 0. A collocated_from_10_flag equal to 0 specifies that the collocated picture used for temporal motion vector prediction is derived from reference picture list 1. If the collocated_from_10_flag does not exist, it is assumed to be equal to 1.
[0450] six_minus_max_num_merge_cand specifies the maximum number of merge motion vector prediction (MVP) candidates supported in tile groups subtracted from 6. The maximum number of MVP candidates to merge, MaxNumMergeCand, is derived as follows.
[0451] [Number]
[0452] The value of MaxNumMergeCand shall be in the range of 1 to 6, including both end values.
[0453] A dep_quant_enabled_flag equal to 0 specifies that dependent quantization is disabled. A dep_quant_enabled_flag equal to 1 specifies that dependent quantization is enabled.
[0454] A sign_data_hiding_enabled_flag equal to 0 specifies that sign bit hiding is disabled. A sign_data_hiding_enabled_flag equal to 1 specifies that sign bit hiding is enabled. If the sign_data_hiding_enabled_flag does not exist, it is assumed to be equal to 0.
[0455] offset_len_minus1 plus 1 specifies the length (bits) of the entry_point_offset_minus1[i] syntax element. The value of offset_len_minus1 is in the range of 0 to 31, inclusive of both end values. shall be.
[0456] entry_point_offset_minus1[i] plus 1 specifies the i-th entry point offset (bytes) and is represented by offset_len_minus1 plus 1 bits. The slice data follows the slice header and consists of a subset from NumTilesInSlice[tile_set_idx]. The subset index value is in the range of 0 to NumTilesInSlice[tile_set_idx] - 1, inclusive of both end values. The first byte of the slice data is considered byte 0. If present, the emulation prevention bytes that appear in the slice data part of the coded slice NAL unit are counted as part of the slice data for the purpose of subset identification. Subset 0 consists of bytes 0 to entry_point_offset_minus1[0] (inclusive) of the coded slice segment data. For k having a value in the range of 1 to NumTilesInSlice[tile_set_idx] - 2 (inclusive), subset k consists of bytes firstByte[k] to lastByte[k] (inclusive) of the coded slice data, where firstByte[k] and lastByte[k] are defined as follows.
[0457]
Number
[0458] The last subset (subset index equal to NumTilesInSlice[tile_set_idx] - 1) consists of the remaining bytes of the coded slice data.
[0459] Each subset shall consist of all the encoded bits of all the CTUs within a slice that are within the same tile.
[0460] Furthermore, with respect to Table 21, the syntax of tile_group_data() may be as shown in Table 21D below.
[0461]
Table 70
[0462] With respect to Table 21D, each syntax element may be based on the following definitions.
[0463] Assume end_of_tile_one_bit is equal to 1.
[0464] Note that although the terms slice and slice header are used in the description herein, these terms may be replaced by the terms tile group and tile group header respectively. Additionally, the syntax element slice_type may be replaced by the syntax element tile_group_type. In this case, the conditions for using slice_type and other syntax elements can be changed to tile_group_type. Additionally, the variable NumTilesInSlice may be replaced by the syntax element NumTilesInTileGroup. Furthermore, note that although the description of the embodiments described herein uses the term tile set, one or more occurrences of the tile set may be replaced by the term tile group. Additionally, one or more of the syntax elements having names that include the word tile_set may be replaced by names that include the word tile_group instead. Therefore, one or more of the following changes may be made.
[0465] Change "tile set" to "tile group".
[0466] Change 「tile_set」 to 「tile_group」.
[0467] Change NumTilesInSlice to NumTilesInTileGroup.
[0468] Change NumTileRowsInTileSetMinus1 to NumTileRowsInTileGroupMinus1.
[0469] Change NumTileColumnsInTileSetMinus1 to NumTileColumnsInTileGroupMinus1.
[0470] Change tile_set_idx to tile_group_idx.
[0471] Change num_tile_sets_in_pic_minus1 to num_tile_groups_in_pic_minus1.
[0472] Change signaled_tile_set_index_flag to signaled_tile_group_index_flag or explicit_tile_group_index_flag.
[0473] Change remaining_tiles_tileset_flag to remaining_tiles_tilegroup_flag.
[0474] Change tile_set_index to tile_group_index.
[0475] Tables 22A and 22B show examples of the syntax of picture parameter sets that can be used to signal the tile structure according to the techniques of this specification.
[0476]
Table 71
[0477]
Table 72
[0478]
Table 73
[0479] For Tables 22A, 22B, and 22C, each syntax element can be based on the semantics and definitions provided above and the following semantics and definitions.
[0480] one_tile_per_tile_group equal to 1 specifies that each tile group referring to this PPS contains one tile. one_tile_per_tile_group equal to 0 specifies that the tile groups referring to this PPS can contain two or more tiles.
[0481] In another example, one_tile_per_tile_group_flag equal to 1 specifies that each tile group referring to the PPS contains exactly one tile. one_tile_per_tile_group_flag equal to 0 specifies that each tile group referring to the PPS contains one or more tiles.
[0482] rect_tile_group_flag equal to 0 specifies that the tiles within the tile group are in raster scan order and that the tile group information is not signaled in the PPS. rect_tile_group_flag equal to 1 specifies that the rectangular tile group information is explicitly specified by the syntax elements num_tile_groups_in_pic_minus1, top_left_tile_id[i] when bottom_right_tile_id[i] exists.
[0483] A tile_group_flag equal to 0 specifies that each tile is a tile group. A tile_group_flag equal to 1 specifies that the tile group is explicitly specified by the syntax elements num_tile_groups_in_pic_minus1, top_left_tile_id[i], and bottom_right_tile_id[i].
[0484] In another example, A tile_group_flag equal to 0 specifies that each tile is a tile group. A tile_group_flag equal to 1 specifies that the tile group is explicitly specified by the syntax elements num_tile_groups_in_pic_minus1, top_left_tile_id[i], and bottom_right_tile_id[i].
[0485] signalled_tile_group_index_length_minus1 plus 1, if present, specifies the number of bits used to represent the syntax elements tile_group_index[i] and tile_group_id. The value of signalled_tile_group_index_length_minus1 shall be in the range of 0 to 15, inclusive of both end values.
[0486] A signalled_tile_group_index_flag equal to 1 specifies that the tile set group of each tile group is signalled. A signalled_tile_group_index_flag equal to 0 specifies that the tile group index is not signalled.
[0487] num_tile_groups_in_pic_minus1 plus 1 specifies the number of tile groups in the picture. The value of num_tile_groups_in_pic_minus1 shall be in the range of 0 to (NumTilesInPic - 1) including both end values. If it does not exist, num_tile_groups_in_pic_minus1 is assumed to be (NumTilesInPic - 1).
[0488] remaining_tiles_tile_group_flag equal to 1 specifies that all the remaining tiles in the tile set other than those explicitly specified within the (num_tile_groups_in_pic_minus1 · 1) tile sets signaled by the syntax elements top_left_tile_id[i], num_tile_rows_in_tile_group_minus1[i], num_tile_columns_in_ttile_group_minus1 [i] form the last tile set. remaining_tiles_tile_group_flag equal to 0 specifies that all num_tile_groups_in_pic_minus1 tile sets are explicitly specified by signaling the syntax elements top_left_tile_id[i], num_tile_rows_in_tile_group_minus1[i], num_tile_columns_in_tile_group_minus1[i].
[0489] top_left_tile_id[i] specifies the tile index of the tile placed at the upper left corner of the i-th tile set. For any i not equal to j, the value of top_left_tile_id[i] shall not be equal to the value of top_left_tile_id[j].
[0490] bottom_right_tile_id[i] specifies the tile index of the tile placed at the bottom - right corner of the i - th tile set. When one_tile_per_tile_group_flag is equal to 1, bottom_right_tile_id[i] is assumed to be equal to top_left_tile_id[i].
[0491] The variable NumTilesInTileGroup[i] that specifies the number of tiles in a tile group, and related variables, are derived as follows.
[0492] [Table 74]
[0493] tile_group_index[i] specifies the tile - group index of the i - th tile group. The length of the tile_group_index[i] syntax element is signalled_tile_set_index_length_minus1 + 1 bits. If the value of tile_group_index[i] does not exist, the value of tile_group_index[i] is assumed to be equal to i for each i in the range of 0 to num_tile_groups_in_pic_minus1, inclusive of both end values.
[0494] Regarding Tables 22A, 22B, 22C, 23A, and 23B, an exemplary syntax of the tile - group header is shown, and Table 24 shows an exemplary syntax of the tile - group data.
[0495] [Table 75]
[0496] [Table 76]
[0497]
Table 77
[0498] Regarding Table 23A, Table 23B, and Table 24, each syntax element can be based on the semantics and definitions provided above and the following semantics and definitions.
[0499] tile_group_id specifies the tile group ID of this tile group. When Signalled_tile_group_index_flag is equal to 1, the length of the tile_group_idx syntax element is signalled_tile_set_index_length_minus1 + 1 bits. Otherwise, the length of tile_group_idx is equal to Ceil(Log2(num_tile_groups_in_pic_minus1 + 1)) bits.
[0500] In another example, tile_group_id specifies the tile group ID of this tile group. The length of the tile_group_idx syntax element is signalled_tile_set_index_length_minus1 + 1 bits.
[0501] offset_len_minus1 plus 1 specifies the length (in bits) of the entry_point_offset_minus1[i] syntax element. The value of offset_len_minus1 is in the range of 0 to 31, inclusive. shall be
[0502] entry_point_offset_minus1[i] plus 1 specifies the i-th entry point offset (in bytes) and is represented by offset_len_minus1 plus 1 bits. The slice data following the slice header consists of rect_tile_group_flag? (NumTilesInTileGroup[tile_group_id] - 1) : num_tiles_in_tile_group_minus1 + 1 subsets, and the subset index values range from 0 to rect_tile_group_flag? (NumTilesInTileGroup[tile_group_id] - 1) : num_tiles_in_tile_group_minus1 (including both end values). The first byte of the slice data is considered byte 0. If present, the emulation prevention bytes that appear in the slice data part of the coded slice NAL unit are counted as part of the slice data for the purpose of subset identification. Subset 0 consists of bytes 0 to entry_point_offset_minus1[0] (including both end values) of the coded slice segment data, and subset k, where k ranges from 1 to rect_tile_group_flag? (NumTilesInTileGroup[tile_group_id] - 2) : (num_tiles_in_tile_group_minus1 - 1) (including both end values), consists of bytes firstByte[k] to lastByte[k] (including both end values) of the coded slice data, and firstByte[k] and lastByte[k] are defined as follows.
[0503]
Number
[0504] The last subset (where the subset index is equal to rect_tile_group_flag? (NumTilesInSlice NumTilesInTileGrpup[tile_groupset_idx] - 1): num_tiles_in_tile_group_minus1) consists of the remaining bytes of the coded slice data.
[0505] Regarding Tables 22A to 24, in one example, the CTB raster and tile scanning process can be as follows.
[0506] A list ColWidth[i] for i in the range 0 to num_tile_columns_minus1, including both end values, specifying the width of the i-th tile column in units of CTB, is derived as follows.
[0507] [Table 78]
[0508] A list RowHeight[j] for j in the range 0 to num_tile_rows_minus1, including both end values, specifying the height of the j-th tile row in units of CTB, is derived as follows.
[0509] [Table 79]
[0510] A list ColBd[i] for i in the range 0 to num_tile_columns_minus1 + 1, including both end values, specifying the location of the i-th tile column boundary in units of CTB, is derived as follows.
[0511] [Table 80]
[0512] The list RowBd[j] for j in the range 0 to num_tile_rows_minus1+1, including both end values, which specifies the location of the j-th tile row boundary in units of CTB, is derived as follows.
[0513] [Table 81]
[0514] The list CtbAddrRsToTs[ctbAddrRs] for ctbAddrRs in the range 0 to PicSizeInCtbsY-1, including both end values, which specifies the conversion from the CTB address in the CTB raster scan of the picture to the CTB address in the tile scan, is derived as follows.
[0515] [Table 82]
[0516] The list CtbAddrTsToRs[CtbAddrTs] for CtbAddrTs in the range 0 to PicSizeInCtbsY-1, including both end values, which specifies the conversion from the CTB address in the tile scan to the CTB address in the CTB raster scan of the picture, is derived as follows.
[0517] [Table 83]
[0518] The list TileId[ctbAddrTs] for CtbAddrTs in the range 0 to PicSizeInCtbsY-1, including both end values, which specifies the conversion from the CTB address in the tile scan to the tile ID, is derived as follows.
[0519] [Table 84]
[0520] The list NumCtusInTile[tileIdx], which specifies the conversion from the tile index in a tile to the number of CTUs, with tileIdx in the range of 0 to PicSizeInCtbsY-1 including both end values, is derived as follows.
[0521] [Table 85]
[0522] For the set [TileIdToIdx[tileId] that specifies the conversion from the tile ID to the tile index, and the list FirstCtbAddrTs[tileIdx] for tileIdx in the range of 0 to NumTilesInPic1 including both end values that specifies the conversion from the tile ID to the CTB address in the tile scan of the first CTB in the tile, are derived as follows.
[0523] [Table 86]
[0524] [Table 87]
[0525] The value of ColumnWidthInLumaSamples[i], which specifies the width of the i-th tile column in luma samples, for i in the range of 0 to num_tile_columns_minus1 including both end values, is ColWidth[i] << Set to be equal to CtbLog2SizeY.
[0526] The value of RowHeightInLumaSamples[j], which specifies the height of the j-th tile row in luma samples, is set equal to RowHeight[j] << CtbLog2SizeY for j in the range 0 to num_tile_rows_minus1, inclusive of both endpoints.
[0527] In this way, source device 102 represents an example of a device configured to signal a flag indicating that the tile set is active in the bitstream, signal a syntax element indicating the number of tile set columns partitioning the picture, and signal a syntax element indicating the number of tile set rows partitioning the picture.
[0528] Referring again to FIG. 1, interface 108 may include any device configured to receive data generated by data encapsulation unit 107 and transmit and / or store that data on a communication medium. Interface 108 can include a network interface card such as an Ethernet card, and can include an optical transceiver, a radio frequency transceiver, or any other type of device capable of transmitting and / or receiving information. Further, interface 108 can include a computer system interface that enables files to be stored on a storage device. For example, interface 108 can include a chipset that supports Peripheral Component Interconnect (PCI) and Peripheral Component Interconnect Express (PCIe) bus protocols, proprietary bus protocols, Universal Serial Bus (USB) protocol, I 2 C, or any other logical and physical structure that can be used to interconnect peer devices.
[0529] Referring again to FIG. 1, the target device 120 includes an interface 122, a data decapsulation unit 123, a video decoder 124, and a display 126. The interface 122 can include any device configured to receive data from a communication medium. The interface 122 can include a network interface card such as an Ethernet card, an optical transceiver, a radio frequency transceiver, or any other type of device capable of receiving and / or transmitting information. Further, the interface 122 can include an interface for a computer system that enables obtaining a compliant video bitstream from a storage device. For example, the interface 122 can support PCI and PCIe bus protocols, proprietary bus protocols, USB protocol, I 2 C, or any other logical and physical structure that can be used to interconnect peer devices, and can include a chipset. The data decapsulation unit 123 may be configured to receive and analyze any of the exemplary parameter sets described herein.
[0530] The video decoder 124 can include any device configured to receive a bitstream (e.g., MCTS sub-bitstream extraction) and / or an acceptable variation thereof and regenerate video data therefrom. The display 126 can include any device configured to display video data. The display 126 can include one of various display devices such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display. The display 126 can include a high-definition display or an ultra-high-definition display. In the example shown in FIG. 1, the video decoder 124 is described as outputting data to the display 126, but it should be noted that the video decoder 124 can be configured to output video data to various types of devices and / or their sub-components. For example, the video decoder 124 can be configured to output video data to any communication medium as described herein.
[0531] FIG. 9 is a block diagram showing an example of a video decoder that can be configured to decode video data according to one or more techniques of the present disclosure. In one embodiment, video decoder 600 can be configured to decode transform data and restore residual data from the transform coefficients based on the decoded transform data. Video decoder 600 may be configured to perform intra prediction decoding and inter prediction decoding, and for this reason, may be referred to as a hybrid decoder. In the example shown in FIG. 9, video decoder 600 includes an entropy decoding unit 602, an inverse quantization unit and a transform coefficient processing unit 604, an intra prediction processing unit 606, an inter prediction processing unit 608, an addition unit 610, a post filter unit 612, and a reference buffer 614. Video decoder 600 can be configured to decode video data to match a video encoding system. Although the example video decoder 600 is shown as having separate functional blocks, such an illustration is for explanatory purposes, and it should be noted that video decoder 600 and / or its sub-components are not limited to a specific hardware or software architecture. The functions of video decoder 600 can be implemented using any combination of hardware, firmware, and / or software implementations.
[0532] As shown in FIG. 9, the entropy decoding unit 602 receives an entropy-coded bit stream. The entropy decoding unit 602 can be configured to decode syntax elements and quantized coefficients from the bit stream according to a reciprocal process opposite to the entropy encoding process. The entropy decoding unit 602 can be configured to perform entropy decoding according to any of the entropy encoding techniques described above. The entropy decoding unit 602 can determine the values of the syntax elements in the encoded bit stream so as to conform to the video encoding standard. As shown in FIG. 9, the entropy decoding unit 602 can determine quantization parameters, values of quantized coefficients, transform data, and prediction data from the bit stream. In the embodiment shown in FIG. 9, the inverse quantization unit and transform coefficient processing unit 604 receive quantization parameters, values of quantized coefficients, transform data, and prediction data from the entropy decoding unit 602 and output the restored residual data.
[0533] Referring again to FIG. 9, the restored residual data may be provided to an adder 610, which can add the restored residual data to a predicted video block to generate restored video data. The predicted video block can be determined according to prediction video techniques (i.e., intra prediction and inter prediction). The intra prediction processing unit 606 can be configured to receive intra prediction syntax elements and obtain a predicted video block from a reference buffer 614. The reference buffer 614 can include a memory device configured to store one or more frames of video data. The intra prediction syntax elements can identify an intra prediction mode such as the intra prediction mode described above. The inter prediction processing unit 608 can be configured to receive inter prediction syntax elements, generate motion vectors, and identify a predicted block within one or more reference frames stored in the reference buffer 814. The inter prediction processing unit 608 can, in some cases, perform interpolation based on an interpolation filter to generate a motion-compensated block. The syntax elements can include an identifier of an interpolation filter to be used for motion prediction with sub-pixel accuracy. The inter prediction processing unit 808 can use the interpolation filter to calculate an interpolated value for sub-integer pixels of a reference block. The post-filter unit 612 can be configured to perform filtering on the restored video data. For example, the post-filter unit 612 can be configured to perform deblocking and / or sample adaptive offset (SAO) filtering based on parameters defined in the bitstream, for example. Further, note that in some examples, the post-filter unit 612 can be configured to perform its own arbitrary filtering (e.g., visual enhancement such as mosquito noise reduction). As shown in FIG. 9, the restored video block can be output by the video decoder 600.In this way, the video decoder 600 can be configured to parse a flag indicating that a tile set is enabled in the bitstream, parse a syntax element indicating a number of tile set columns partitioning a picture, parse a syntax element indicating a number of tile set rows partitioning a picture, and generate video data based on the values of the parsed syntax elements.
[0534] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored or transmitted as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium or a communication medium that facilitates transfer of a computer program from one place to another according to, for example, a communication protocol. In this way, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0535] By way of example and without limitation, such a computer-readable storage medium can include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage devices, magnetic disk storage devices, other magnetic storage devices, flash memory, or any other medium, i.e., any other medium that can be used to store the desired program code in the form of instructions or data structures and that is accessible by a computer. Also, any connection can be appropriately referred to as a computer-readable medium. For example, when instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that the computer-readable medium and data storage medium do not include connections, carrier waves, signals, or other transient media, but instead are directed to non-transitory tangible storage media. As used in the present invention, disk and disc include Compact Disc (CD), laser disc, optical disc, Digital Versatile Disc (DVD), floppy disk, and Blu-ray (registered trademark) disc, where disk typically magnetically reproduces data and disc optically reproduces data using a laser. The above combinations must also be included within the scope of the computer-readable medium.
[0536] The commands can be executed by one or more processors such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, as used herein, the term "processor" can refer to either the foregoing structures, or any other structure suitable for implementation of the techniques described herein. Additionally, in some aspects, the functions described herein can be provided within dedicated hardware modules and / or software modules configured to encode and decode, or incorporated into a composite codec. Further, the technology can be implemented entirely within one or more circuits or logic elements.
[0537] The techniques of the present disclosure can be implemented in a variety of devices or apparatuses including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chip sets). While various components, modules, or units are shown in the present disclosure to emphasize the functional aspects of a device configured to execute the disclosed techniques, it is not necessarily required that they be realized by different hardware units. Rather, as described above, the various units can be combined in a codec hardware unit, or provided by a collection of interoperating hardware units including one or more of the foregoing processors, together with suitable software and / or firmware.
[0538] Furthermore, each functional block and various functions of the base station apparatus and terminal apparatus used in each of the above-described implementation forms can generally be realized or executed by an integrated circuit or an electric circuit that is a plurality of integrated circuits. A circuit designed to execute the functions described in this specification may include a general-purpose processor, a digital signal processor (DSP), an application-specific or general-purpose application integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gates or transistor logic, or individual hardware components, or a combination thereof. The general-purpose processor may be a microprocessor, or alternatively, the processor may be a conventional processor, controller, microcontroller, or state machine. The general-purpose processor or each circuit described above may be composed of a digital circuit or an analog circuit. Furthermore, if an integrated circuit technology that replaces the current integrated circuit appears due to the progress of semiconductor technology, an integrated circuit using this technology can also be used.
[0539] Various embodiments have been described. These and other embodiments are within the scope of the following claims.
[0540] <Cross-reference> This application, which is a national application based on International Application PCT / JP2019 / 046062 (filed on November 26, 2019) published as WO2020 / 111023, claims priority based on U.S. Provisional Patent Application Nos. 62 / 774,050 (filed on November 30, 2018), 62 / 784,296 (filed on December 21, 2018), 62 / 791,227 (filed on January 11, 2019), 62 / 806,502 (filed on February 15, 2019), and the entire content thereof is incorporated herein by reference.
Claims
1. 1. A method for decoding video data, comprising the steps of: decoding a first flag syntax in a parameter set that, when a value is equal to a first predetermined value, specifies that the tiles in each of at least one slice are in raster scan order, and when a value is different from the first predetermined value, specifies that the tiles in each of the at least one slice cover a rectangular region of the picture; if the value of the first flag syntax is different from the first predetermined value, decoding a first number syntax present in the parameter set, the first number syntax specifying a number of slices in the picture; if the value of the first flag syntax is equal to the first predetermined value, decoding a slice address syntax and a second number syntax present in a slice header, the second number syntax specifying a number of tiles included in the slice; A method comprising:
2. the slice address syntax is present in the slice header if a value of the first flag syntax specifies that the tiles in each of the at least one slice are in raster scan order and the number of tiles in the picture corresponding to the slice header is greater than one. The method of claim 1.
3. a value of the first number syntax plus one specifies a number of slices in each of at least one picture corresponding to the parameter set; The method of claim 1.
4. a value of the second number syntax plus one specifies the number of tiles in each of the at least one slice corresponding to the slice header; The method of claim 1.
5. 1. A device for decoding video data, comprising: At least one processor; a storage device coupled to said at least one processor, said storage device storing a program that, when executed by said at least one processor, causes said at least one processor to perform the method according to any one of claims 1 to 4; 16. A device comprising:
6. 1. A method for encoding video data, comprising the steps of: encoding a first flag syntax in a parameter set that, when a value is equal to a first predetermined value, specifies that the tiles in each of at least one slice are in raster scan order, and, when a value is different from the first predetermined value, specifies that the tiles in each of the at least one slice cover a rectangular region of the picture; if the value of the first flag syntax is different from the first predetermined value, encoding a first number syntax present in the parameter set, the first number syntax specifying a number of slices in the picture; if the value of the first flag syntax is equal to the first predetermined value, encoding a slice address syntax and a second number syntax present in a slice header, the second number syntax specifying a number of tiles included in the slice; The method includes:
7. the slice address syntax is present in the slice header if a value of the first flag syntax specifies that the tiles in each of the at least one slice are in raster scan order and the number of tiles in the picture corresponding to the slice header is greater than one. The method according to claim 6.
8. a value of the first number syntax plus one specifies a number of slices in each of at least one picture corresponding to the parameter set; The method according to claim 6.
9. a value of the second number syntax plus one specifies the number of tiles in each of the at least one slice corresponding to the slice header; The method according to claim 6.
10. 1. A device for encoding video data, comprising: At least one processor; a storage device coupled to said at least one processor, said storage device storing a program that, when executed by said at least one processor, causes said at least one processor to perform a method according to any one of claims 6 to 9; 16. A device comprising:
Citation Information
Patent Citations
JPP7521077B
JPP7307168B