VIDEO ENCODER, VIDEO DECODER, AND CORRESPONDING METHODS - Patent application

The flexible video tiling scheme addresses the challenge of encoding multiple regions at different resolutions by partitioning pictures into first and second level tiles and assigning them to rectangular tile groups, thereby improving compression efficiency and image quality.

JP7682972B2Active Publication Date: 2025-05-26HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023184083
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-12-27
Filing Date
2023-10-26
Publication Date
2025-05-26
Estimated Expiration
2039-12-27

AI Technical Summary

Technical Problem

Existing video coding technologies struggle to efficiently compress and decompress videos with multiple regions encoded at different resolutions, as current slicing and tiling mechanisms do not support this functionality effectively.

Method used

A flexible video tiling scheme that partitions a picture into first level tiles and second level tiles, allowing for multiple tiles with different resolutions within the same picture. This scheme assigns tiles to rectangular tile groups, enabling separate extraction and processing of different content.

Benefits of technology

The flexible tiling scheme enhances the capabilities of both encoders and decoders by supporting pictures with multiple resolutions, improving compression ratios with minimal sacrifice in image quality, and optimizing bandwidth usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682972000010
    Figure 0007682972000010
  • Figure 0007682972000011
    Figure 0007682972000011
  • Figure 0007682972000012
    Figure 0007682972000012
Patent Text Reader

Abstract

To provide a video coding mechanism employing a flexible video tiling method.SOLUTION: A mechanism of the present invention includes partitioning a picture into a plurality of first level tiles. A subset of the first level tiles is partitioned into a plurality of second level tiles. The first level tiles and the second level tiles are assigned to one or more tile groups such that all tiles in an assigned tile group including the second level tiles are constrained to cover a rectangular area of the picture. The first level tiles and the second level tiles are encoded into a bitstream. The bitstream is stored for communication toward a decoder.SELECTED DRAWING: Figure 8B
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] FIELD This disclosure relates generally to video coding, and more particularly to a flexible video tiling scheme that supports multiple tiles with different resolutions within the same picture. [Background technology]

[0002] The amount of video data required to depict even a relatively short video can be substantial, which can pose difficulties when the data is to be streamed or otherwise communicated across communication networks with limited bandwidth capacity. Thus, video data is typically compressed before being communicated across modern telecommunication networks. The size of the video can also be an issue when the video is stored on a storage device, since memory resources may be limited. Video compression devices often use software and / or hardware at the source to code the video data before transmission or storage, thereby reducing the amount of data required to depict a digital video image. The compressed data is then received at the destination by a video decompression device, which decodes the video data. With limited network resources and an ever-increasing demand for higher video quality, improved compression and decompression techniques are desirable that improve compression ratios with little sacrifice in image quality. Summary of the Invention [Means for solving the problem]

[0003] In one embodiment, the present disclosure includes a method implemented in an encoder, the method including: partitioning, by a processor of the encoder, a picture into a plurality of first level tiles; partitioning, by the processor, a subset of the first level tiles into a plurality of second level tiles; assigning, by the processor, the first level tiles and the second level tiles to one or more tile groups such that all tiles in the assigned tile group that includes the second level tiles are constrained to cover a rectangular portion of the picture; encoding, by the processor, the first level tiles and the second level tiles into a bitstream; and storing the bitstream in a memory of the encoder for communication toward a decoder. Video coding systems may employ slices and tiles to partition a picture. Some streaming applications (e.g., virtual reality (VR) and teleconferencing) may be improved if a single image including multiple regions encoded at different resolutions can be sent. Some slicing and tiling mechanisms may not support such functionality because tiles at different resolutions may be treated differently. For example, a tile at a first resolution may contain a single slice of data, while a tile at a second resolution may carry multiple slices of data due to differences in pixel density. The present embodiment employs a flexible tiling scheme that includes first level tiles and second level tiles. The second level tiles are created by partitioning the first level tiles. The tiling scheme allows the first level tiles to contain one slice of data at the first resolution and the first level tiles that include the second level tiles to contain multiple slices at the second resolution. Tiles may be assigned to tile groups. The present embodiment constrains the tile groups that include the second level tiles to be rectangular as opposed to raster scanning. This approach creates boundaries that support separate extraction and processing of different content.For example, a tile group containing content at a first resolution and a tile group containing content at a second resolution are naturally shaped to support side-by-side display on a screen and / or separate extraction for use on a head-mounted display. Thus, the disclosed flexible tiling scheme enables an encoder / decoder (codec) to support pictures containing multiple resolutions, thus enhancing the capabilities of both the encoder and the decoder.

[0004] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that first level tiles outside the subset include picture data at a first resolution and second level tiles include picture data at a second resolution different from the first resolution.

[0005] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that when any first level tile is partitioned into multiple second level tiles, each of one or more tile groups is constrained to cover a rectangular portion of the picture.

[0006] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the first level tiles and the second level tiles are not assigned to one or more tile groups according to a raster scan order that traverses the picture horizontally from the left boundary to the right boundary.

[0007] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that covering a rectangular portion of the picture includes covering less than a full horizontal portion of the picture.

[0008] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that each second level tile includes a single slice of picture data from the picture.

[0009] Optionally, in any of the aforementioned aspects, another implementation of the aspect further includes encoding, by the processor, second level tile rows and second level tile columns for the partitioned first level tiles, providing that the second level tile rows and second level tile columns are encoded in a picture parameter set associated with the picture.

[0010] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that for first level tiles having a width less than a minimum width threshold and a height less than a minimum height threshold, data that explicitly indicates whether the first level tile is partitioned into second level tiles is excluded from the bitstream, and for partitioned first level tiles having a width less than twice the minimum width threshold and a height less than twice the minimum height threshold, second level tile rows and second level tile columns are excluded from the bitstream.

[0011] In one embodiment, the present disclosure includes a method implemented in a decoder, the method including: receiving, via a receiver, a bitstream including a picture partitioned into a plurality of first level tiles by a processor of the decoder, where a subset of the first level tiles is further partitioned into a plurality of second level tiles, and the first level tiles and the second level tiles are assigned to one or more tile groups such that all tiles in the assigned tile group including the second level tiles are constrained to cover a rectangular portion of the picture; determining, by the processor, a configuration of the first level tiles and a configuration of the second level tiles based on the one or more tile groups; decoding, by the processor, the first level tiles and the second level tiles based on the configuration of the first level tiles and the configuration of the second level tiles; and generating, by the processor, a reconstructed video sequence for display based on the decoded first level tiles and the second level tiles. Video coding systems may employ slices and tiles to partition a picture. Some streaming applications (e.g., VR and teleconferencing) may be improved if a single image including multiple regions encoded at different resolutions can be sent. Some slicing and tiling mechanisms may not support such functionality because tiles at different resolutions may be treated differently. For example, a tile at a first resolution may contain a single slice of data, while a tile at a second resolution may carry multiple slices of data due to differences in pixel density. The present embodiment employs a flexible tiling scheme that includes first level tiles and second level tiles. The second level tiles are created by partitioning the first level tiles. The tiling scheme allows the first level tiles to contain one slice of data at the first resolution and the first level tiles, including the second level tiles, to contain multiple slices at the second resolution. Tiles may be assigned to tile groups.The present embodiment constrains tile groups containing second level tiles to be rectangular as opposed to raster scan. This approach creates boundaries that support separate extraction and processing of different content. For example, tile groups containing content at a first resolution and tile groups containing content at a second resolution are naturally shaped to support separate extraction for side-by-side display on a screen and / or use on a head-mounted display. Thus, the disclosed flexible tiling scheme allows codecs to support pictures containing multiple resolutions, thus enhancing the capabilities of both encoders and decoders.

[0012] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that first level tiles outside the subset include picture data at a first resolution and second level tiles include picture data at a second resolution different from the first resolution.

[0013] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that when any first level tile is partitioned into multiple second level tiles, each of one or more tile groups is constrained to cover a rectangular portion of the picture.

[0014] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the first level tiles and the second level tiles are not assigned to one or more tile groups according to a raster scan order that traverses the picture horizontally from the left boundary to the right boundary.

[0015] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that covering a rectangular portion of the picture includes covering less than a full horizontal portion of the picture.

[0016] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that each second level tile includes a single slice of picture data from the picture.

[0017] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides for further including obtaining, by the processor, second level tile rows and second level tile columns for the partitioned first level tiles from a picture parameter set associated with the picture.

[0018] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that for first level tiles having a width less than a minimum width threshold and a height less than a minimum height threshold, data that explicitly indicates whether the first level tile is partitioned into second level tiles is excluded from the bitstream, and for partitioned first level tiles having a width less than twice the minimum width threshold and a height less than twice the minimum height threshold, second level tile rows and second level tile columns are excluded from the bitstream.

[0019] In one embodiment, the present disclosure includes a video coding device comprising a processor, a receiver connected to the processor, and a transmitter connected to the processor, wherein the processor, the receiver, and the transmitter are configured to perform the method of any of the aforementioned aspects.

[0020] In one embodiment, the present disclosure includes a non-transitory computer-readable medium having stored thereon a computer program for use by a video coding device, the computer program including computer-executable instructions that, when executed by a processor, cause the video coding device to perform a method of any of the aforementioned aspects.

[0021] In one embodiment, the present disclosure includes an encoder comprising partitioning means for partitioning a picture into a plurality of first level tiles and partitioning a subset of the first level tiles into a plurality of second level tiles; allocation means for assigning the first level tiles and the second level tiles to one or more tile groups such that all tiles in an assigned tile group that includes a second level tile are constrained to cover a rectangular portion of the picture; encoding means for encoding the first level tiles and the second level tiles into a bitstream; and storage means for storing the bitstream for communication to a decoder.

[0022] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the encoder is further configured to perform a method of any of the aforementioned aspects.

[0023] In one embodiment, the disclosure includes a decoder comprising: receiving means for receiving a bitstream including a picture partitioned into a plurality of first level tiles, where a subset of the first level tiles is further partitioned into a plurality of second level tiles, where the first level tiles and the second level tiles are assigned to one or more tile groups such that all tiles in an assigned tile group including the second level tile are constrained to cover a rectangular portion of the picture; determining means for determining a configuration of the first level tiles and a configuration of the second level tiles based on the one or more tile groups; decoding means for decoding the first level tiles and the second level tiles based on the configuration of the first level tiles and the configuration of the second level tiles; and generating means for generating a reconstructed video sequence for display based on the decoded first level tiles and the second level tiles.

[0024] Optionally, in any of the aforementioned aspects, another implementation of the aspect provides that the decoder is further configured to perform the method of any of the aforementioned aspects.

[0025] For clarity, any one of the above embodiments may be combined with any one or more of the other embodiments above to create new embodiments within the scope of the present disclosure.

[0026] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims.

[0027] For a more complete understanding of the present disclosure, reference is now made to the accompanying drawings and the following brief description, taken in conjunction with the detailed description, in which like reference numerals represent like parts, in which: [Brief description of the drawings]

[0028] [Figure 1] 4 is a flowchart of an exemplary method for coding a video signal. [Diagram 2] 1 is a schematic diagram of an example coding and decoding (codec) system for video coding. [Diagram 3] FIG. 1 is a schematic diagram illustrating an exemplary video encoder. [Figure 4] FIG. 2 is a schematic diagram illustrating an exemplary video decoder. [Diagram 5] 1 is a schematic diagram illustrating an exemplary bitstream including an encoded video sequence. [Figure 6A] FIG. 1 illustrates an example mechanism for creating an extractor track to combine multiple resolution sub-pictures from different bitstreams into a single picture for use in virtual reality (VR) applications. [Figure 6B]FIG. 1 illustrates an example mechanism for creating an extractor track to combine multiple resolution sub-pictures from different bitstreams into a single picture for use in virtual reality (VR) applications. [Figure 6C] FIG. 1 illustrates an example mechanism for creating an extractor track to combine multiple resolution sub-pictures from different bitstreams into a single picture for use in virtual reality (VR) applications. [Figure 6D] FIG. 1 illustrates an example mechanism for creating an extractor track to combine multiple resolution sub-pictures from different bitstreams into a single picture for use in virtual reality (VR) applications. [Figure 6E] FIG. 1 illustrates an example mechanism for creating an extractor track to combine multiple resolution sub-pictures from different bitstreams into a single picture for use in virtual reality (VR) applications. [Figure 7] FIG. 1 illustrates an example video conferencing application in which multiple resolution pictures from different bitstreams are stitched together into a single picture for display. [Figure 8A] FIG. 2 is a schematic diagram illustrating an example flexible video tiling scheme capable of supporting multiple tiles with different resolutions within the same picture. [Figure 8B] FIG. 2 is a schematic diagram illustrating an example flexible video tiling scheme capable of supporting multiple tiles with different resolutions within the same picture. [Figure 9] 1 is a schematic diagram of an example video coding device. [Figure 10] 1 is a flowchart of an exemplary method for encoding an image by employing a flexible tiling scheme. [Figure 11]1 is a flowchart of an exemplary method for decoding an image by employing a flexible tiling scheme. [Figure 12] 1 is a schematic diagram of an example system for coding a video sequence by employing a flexible tiling scheme. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0029] First, while example implementations of one or more embodiments are provided below, it should be understood that the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or in existence. The present disclosure should in no way be limited to the example implementations, drawings, and techniques set forth below, including the example designs and implementations shown and described herein, but may be modified within the scope of the appended claims together with their full scope of equivalents.

[0030] Various acronyms are employed herein, such as coding tree block (CTB), coding tree unit (CTU), coding unit (CU), coded video sequence (CVS), Joint Video Experts Team (JVET), motion constrained tile set (MCTS), maximum transfer unit (MTU), network abstraction layer (NAL), picture order count (POC), raw byte sequence payload (RBSP), sequence parameter set (SPS), versatile video coding (VVC), and working draft (WD).

[0031] Many video compression techniques may be employed to reduce the size of video files with minimal loss of data. For example, video compression techniques may include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or remove data redundancy in a video sequence. For block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, which may also be referred to as tree blocks, coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are coded using spatial prediction with reference to reference samples in neighboring blocks in the same picture. Video blocks in an inter-coded unidirectionally predicted (P) or bidirectionally predicted (B) slice of a picture may be coded by employing spatial prediction with reference to reference samples in neighboring blocks in the same picture or temporal prediction with reference to reference samples in other reference pictures. A picture may be referred to as a frame and / or an image, and a reference picture may be referred to as a reference frame and / or a reference image. Spatial or temporal prediction results in a predictive block that represents an image block. The residual data represents pixel differences between the original image block and the predictive block. Thus, an inter-coded block is coded according to a motion vector that points to a block of reference samples that form the predictive block, and the residual data that indicates the difference between the coded block and the predictive block. An intra-coded block is coded according to an intra-coding mode and the residual data. For further compression, the residual data may be transformed from the pixel domain to a transform domain. These result in residual transform coefficients, which may be quantized. The quantized transform coefficients may be initially arranged in a two-dimensional array. The quantized transform coefficients may be scanned to produce a one-dimensional vector of transform coefficients.To achieve even further compression, entropy coding may be applied, and such video compression techniques are described in further detail below.

[0032] To ensure that the encoded video can be accurately decoded, the video is encoded and decoded according to a corresponding video coding standard. Video coding standards include Advanced Video Coding (AVC), also known as International Telecommunications Union (ITU) Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group (MPEG)-1 Part 2, ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC), also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC) and Multiview Video Coding Plus Depth (MVC+D), as well as three-dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The Joint Video Experts Team (JVET) of ITU-T and ISO / IEC has begun developing a video coding standard called Versatile Video Coding (VVC). VVC is included in the Working Drafts (WDs) including JVET-L1001-v5.

[0033] To code a video image, the image is first partitioned and the partition is coded into a bitstream. Various picture partitioning schemes are available. For example, an image may be partitioned into normal slices, dependent slices, tiles, and / or according to Wavefront Parallel Processing (WPP). For simplicity, HEVC restricts the encoder to only use normal slices, dependent slices, tiles, WPP, and combinations thereof when partitioning slices into groups of CTBs for video coding. Such partitioning may be applied to support maximum transmission unit (MTU) size alignment, parallel processing, and reduced end-to-end delay. The MTU indicates the maximum amount of data that can be transmitted in a single packet. If a packet payload exceeds the MTU, the payload is split into two packets through a process called fragmentation.

[0034] A normal slice, also simply referred to as a slice, is a partitioned portion of an image that can be reconstructed independently from other normal slices in the same picture, despite some interdependencies due to loop filtering operations. Each normal slice is encapsulated in its own network abstraction layer (NAL) unit for transmission. Furthermore, intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies across slice boundaries can be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, normal slice-based parallelization employs minimal inter-processor or inter-core communication. However, each normal slice is independent, and each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of the slice header per slice and due to the lack of prediction across slice boundaries. Furthermore, normal slices can be employed to support alignment against MTU size requirements. In particular, because regular slices may be encapsulated in separate NAL units and coded independently, each regular slice should be smaller than the MTU in the MTU scheme to avoid breaking the slice into multiple packets. Thus, the goals of parallelization and MTU size matching may impose conflicting demands on slice layout within a picture.

[0035] Dependent slices are similar to normal slices, but have a shortened slice header and allow for partitioning of picture treeblock boundaries without breaking intra-picture prediction. Thus, dependent slices allow normal slices to be fragmented into multiple NAL units, which results in reduced end-to-end delay by allowing parts of a normal slice to be sent out before the encoding of the entire normal slice is completed.

[0036] A tile is a partitioned portion of an image created by horizontal and vertical boundaries that create columns and rows of tiles. Tiles may be coded in raster scan order (right to left and top to bottom). The scan order of the CTBs is local within the tile. Thus, the CTB in the first tile is coded in raster scan order before proceeding to the CTB in the next tile. As with normal slices, tiles break intra-picture prediction dependencies as well as entropy decoding dependencies. However, tiles may not be included in individual NAL units, and therefore tiles may not be used for MTU size alignment. Each tile may be processed by one processor / core, and inter-processor / inter-core communication employed for intra-picture prediction between processing units that decode adjacent tiles may be limited to conveying shared slice headers (when adjacent tiles are in the same slice) and performing loop filter processing-related sharing of reconstructed samples and metadata. When more than one tile is included in a slice, the entry point byte offset for each tile, other than the first entry point offset in the slice, may be signaled in the slice header. For each slice and tile, at least one of the following conditions should be fulfilled: 1) all coded tree blocks in a slice belong to the same tile, and 2) all coded tree blocks in a tile belong to the same slice.

[0037] In WPP, a picture is partitioned into a single row of CTBs. The entropy decoding and prediction mechanism may use data from CTBs in other rows. Parallel processing is enabled through parallel decoding of CTB rows. For example, a current row may be decoded in parallel with a previous row. However, the decoding of the current row is delayed by two CTBs from the decoding process of the previous row. This delay ensures that data related to the CTB above the current CTB and the CTBs above and to the right of the current CTB in the current row are available before the current CTB is coded. This approach looks like a wavefront when represented graphically. This staggered beginning allows parallelization with up to as many processors / cores as the picture contains CTB rows. Since intra-picture prediction between adjacent treeblock rows in a picture is allowed, inter-processor / inter-core communication to enable intra-picture prediction can be substantial. WPP partitioning takes into account the NAL unit size. Therefore, WPP does not support MTU size alignment. However, regular slices may be used with WPP, with some coding overhead, to implement MTU size alignment if necessary.

[0038] A tile may also include a motion constrained tile set. A motion constrained tile set (MCTS) is a tile set designed such that associated motion vectors are restricted to point to full sample locations inside the MCTS and fractional sample locations that require only full sample locations inside the MCTS for interpolation. In addition, the use of motion vector candidates for temporal motion vector prediction derived from blocks outside the MCTS is rejected. In this way, each MCTS can be decoded independently without the presence of tiles not included in the MCTS. A temporal MCTS supplemental enhancement information (SEI) message can be used to indicate the presence of an MCTS in a bitstream and to signal the MCTS. The MCTS SEI message provides additional information that can be used in MCTS sub-bitstream extraction (specified as part of the semantics of the SEI message) to generate a conforming bitstream for the MCTS. The information includes several extraction information sets, each of which defines several MCTSs and includes raw byte sequence payload (RBSP) bytes of a replacement video parameter set (VPS), a sequence parameter set (SPS), and a picture parameter set (PPS) to be used during the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) may be rewritten or replaced and the slice headers may be updated, since one or all of the slice address related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) may adopt different values ​​in the extracted sub-bitstream.

[0039] Various tiling schemes may be employed when partitioning a picture for further encoding. As a particular example, tiles may be assigned to tile groups, which in some examples may replace slices. In some examples, each tile group may be extracted independently of other tile groups. Thus, tile grouping may support parallelization by allowing each tile group to be assigned to a different processor. Tile grouping may also be employed when a decoder may not want to decode the entire image. As a particular example, a video coding scheme may be employed to support virtual reality (VR) video, which may be encoded according to the Omnidirectional Media Application Format (OMAF).

[0040] In a VR video, one or more cameras may record the environment around the camera. The user may then view the VR video as if the user were present in the same location as the camera. In a VR video, the picture encompasses the entire environment around the user. The user then sees a sub-portion of the picture. For example, the user may employ a head-mounted display that changes the sub-portion of the picture displayed based on the user's head movement. The portion of the video being displayed may be referred to as a viewport.

[0041] Thus, a distinct feature of omnidirectional video is that only a viewport is displayed at any particular time. This is in contrast to other video applications that may display the entire video. This feature may be exploited to improve the performance of omnidirectional video systems, for example, through selective delivery according to the user's viewport (or any other criteria, such as recommended viewport timed metadata). Viewport-dependent delivery may be enabled, for example, by employing per-region packing and / or viewport-dependent video coding. The performance improvement may result in smaller transmission bandwidth, lower decoding complexity, or both, when compared to other omnidirectional video systems when employing the same video resolution / quality.

[0042] An exemplary viewport-dependent operation is an MCTS-based approach to achieve an effective equirectangle projection (ERP) resolution of 5000 samples (e.g., 5120×2560 luma samples) resolution (5K) using an HEVC-based viewport-dependent OMAF video profile. This approach is described in more detail below. In general, however, this approach partitions the VR video into tile groups and encodes the video at multiple resolutions. The decoder can indicate the viewport currently used by the user during streaming. The video server providing the VR video data can then transfer the tile groups associated with the viewport at high resolution and transfer the unseen tile groups at a lower resolution. This allows the user to watch the VR video at high resolution without requiring the entire picture to be sent at high resolution. The unseen sub-portions are discarded, and thus the user may not be aware of the lower resolution. However, if the user changes the viewport, the lower resolution tile groups may be displayed to the user. The resolution of the new viewport may then be increased as the video progresses. To implement such a system, a picture should be created that contains both higher resolution and lower resolution tile groups.

[0043] In another example, a video conferencing application may be designed to transfer a picture that includes multiple resolutions. For example, a video conference may include multiple participants. The participant currently speaking may be displayed in a higher resolution, and other participants may be displayed in a lower resolution. To implement such a system, a picture should be created that includes both higher resolution and lower resolution tile groups.

[0044] Various flexible tiling mechanisms are disclosed herein to support creating pictures with sub-pictures coded at multiple resolutions. For example, video may be coded at multiple resolutions. Video may also be coded by employing slices at each resolution. Lower resolution slices are smaller than higher resolution slices. To create a picture with multiple resolutions, a picture may be partitioned into first level tiles. Slices from the highest resolution may be included directly within the first level tiles. Furthermore, the first level tiles may be partitioned into second level tiles that are smaller than the first level tiles. Thus, the smaller second level tiles can directly accept the lower resolution slices. In this way, slices from each resolution may be compressed into a single picture via tile index relationships without requiring tiles of different resolutions to be dynamically re-addressed to use a coherent addressing scheme. The first level tiles and the second level tiles may be implemented as MCTS and thus may accept motion constrained image data at different resolutions. This disclosure includes many aspects. As a specific example, the first level tiles are divided into second level tiles. The second level tiles are then constrained to each contain a single rectangular slice of picture data (e.g., at a lower resolution). As used herein, a tile is a partitioned portion of a picture created by horizontal and vertical boundaries (e.g., by columns and rows). A rectangular slice is a slice that is constrained to maintain a rectangular shape and is therefore coded based on horizontal and vertical picture boundaries. Thus, a rectangular slice is not coded based on a raster scan group (which may contain CTUs in lines from left to right and lines from top to bottom and not maintain a rectangular shape). A slice is a spatially distinct region of a picture / frame that is coded separately from any other region in the same frame / picture.In a further aspect, the first level tiles and the second level tiles are assigned to tile groups. Tile groups employed with the flexible tiling scheme are constrained to be rectangular as opposed to raster scan. For example, first level tiles are included in rectangular tile groups and corresponding second level tiles are constrained to be part of the same tile group as the first level tiles from which such second level tiles are partitioned. This approach works from left to right and top to bottom to create rectangular boundaries rather than raster scan boundaries which are generally not rectangular. By constraining the tile groups to be rectangular, the tile groups result in shapes that support the extraction and display of sub-pictures. Thus, tile groups containing sub-pictures at different resolutions are naturally shaped to support side-by-side display on a screen and / or separate extraction for use on a head-mounted display.

[0045] FIG. 1 is a flow chart of an exemplary operational method 100 of coding a video signal. In particular, a video signal is encoded in an encoder. The encoding process compresses the video signal by employing various mechanisms to reduce the video file size. The smaller file size allows the compressed video file to be transmitted towards a user while reducing the associated bandwidth overhead. A decoder then decodes the compressed video file to reconstruct the original video signal for display to an end user. The decoding process generally mirrors the encoding process to allow the decoder to consistently reconstruct the video signal.

[0046] In step 101, a video signal is input into the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live streaming of the video. The video file may include both audio and video components. The video component includes a series of image frames that, when viewed in sequence, give the visual impression of motion. The frames include pixels that are expressed in terms of light, referred to herein as luma components (or luma samples), and colors, referred to herein as chroma components (or color samples). In some examples, the frames may also include depth values ​​to support three-dimensional viewing.

[0047] In step 103, the video is partitioned into blocks. Partitioning involves subdividing pixels in each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also referred to as H.265 and MPEG-H Part 2), a frame may first be partitioned into coding tree units (CTUs), which are blocks of a predefined size (e.g., 64 pixels by 64 pixels). CTUs include both luma and chroma samples. A coding tree may be employed to partition the CTUs into blocks and then recursively subdivide the blocks until a configuration that supports further encoding is achieved. For example, the luma component of a frame may be subdivided until each block contains relatively homogenous light values. Additionally, the chroma component of a frame may be subdivided until each block contains relatively homogenous color values. Thus, the partitioning mechanism varies depending on the content of the video frame.

[0048] In step 105, various compression mechanisms are employed to compress the image blocks partitioned in step 103. For example, inter-prediction and / or intra-prediction may be employed. Inter-prediction is designed to take advantage of the fact that objects in a common scene tend to appear in successive frames. Thus, a block showing an object in a reference frame does not need to be described repeatedly in adjacent frames. In particular, an object such as a table may remain in a constant position across multiple frames. Thus, the table may be described once and adjacent frames may refer back to the reference frame. A pattern matching mechanism may be employed to match objects across multiple frames. Furthermore, a moving object may be depicted across multiple frames, for example due to object movement or camera movement. As a particular example, a video may show a car moving across the screen across multiple frames. A motion vector may be employed to represent such movement. A motion vector is a two-dimensional vector that provides an offset from the coordinates of the object in the frame to the coordinates of the object in the reference frame. Thus, inter-prediction may encode an image block in a current frame as a set of motion vectors that indicate an offset from a corresponding block in a reference frame.

[0049] Intra prediction encodes blocks within a common frame. Intra prediction exploits the fact that luma and chroma components tend to cluster within a frame. For example, green fragments within a portion of a tree tend to be located adjacent to similar fragments of green. Intra prediction employs multiple directional prediction modes (e.g., 33 in HEVC), planar mode, and direct current (DC) mode. The directional modes indicate that the current block is similar / same as samples of neighboring blocks in the corresponding direction. The planar mode indicates that a series of blocks (e.g., planes) along a row / column may be interpolated based on neighboring blocks at the edge of the row. The planar mode actually indicates a smooth transition of light / color across the row / column by employing a relatively constant gradient in changing values. The DC mode is employed for boundary smoothing, indicating that the block is similar / same as the average value associated with samples of all neighboring blocks associated with the angular direction of the directional prediction mode. Thus, intra prediction blocks can represent image blocks as various prediction mode values ​​that indicate relationships rather than actual values. Additionally, inter prediction blocks can represent image blocks as motion vector values ​​rather than actual values. In either case, the prediction block may not exactly represent the image block in some cases: any differences are stored in the residual block, and a transform may be applied to the residual block to further compress the file.

[0050] In step 107, various filtering techniques may be applied. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction described above may result in the generation of blocky images at the decoder. Furthermore, the block-based prediction scheme may code a block and then reconstruct the coded block for later use as a reference block. The in-loop filtering scheme iteratively applies noise suppression filters, deblocking filters, adaptive loop filters, and sample adaptive offset (SAO) filters to the block / frame. These filters mitigate such blocking artifacts so that the coded file can be accurately reconstructed. Furthermore, these filters mitigate artifacts in the reconstructed reference block, so that the artifacts are less likely to generate additional artifacts in subsequent blocks that are coded based on the reconstructed reference block.

[0051] Once the video signal has been partitioned, compressed, and filtered, the resulting data is coded in a bitstream at step 109. The bitstream includes the data described above, as well as any signaling data desired to support proper video signal reconstruction at the decoder. For example, such data may include partition data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in a memory for transmission to the decoder upon request. The bitstream may also be broadcast and / or multicast to multiple decoders. Creation of the bitstream is an iterative process. Thus, steps 101, 103, 105, 107, and 109 may occur continuously and / or simultaneously across many frames and blocks. The order shown in FIG. 1 is presented for clarity and ease of explanation and is not intended to limit the video coding process to any particular order.

[0052] In step 111, the decoder receives the bitstream and starts the decoding process. In particular, the decoder employs an entropy decoding scheme to convert the bitstream into corresponding syntax and video data. In step 111, the decoder employs syntax data from the bitstream to determine a partition for the frame. The partition should match the result of the block partition in step 103. The entropy encoding / decoding as employed in step 111 is now described. The encoder makes many choices during the compression process, such as selecting a block partitioning scheme from several possible options based on the spatial positioning of values ​​in the input image. Signaling the exact choice may employ a number of bins. A bin, as used herein, is a binary value (e.g., a bit value that may change depending on the context) that is treated as a variable. Entropy coding allows the encoder to discard any options that are clearly not viable for a particular case, leaving a set of acceptable options. Each acceptable option is then assigned a codeword. The length of the codeword is based on the number of allowable options (e.g., one bin for 2 options, two bins for 3-4 options, etc.). The encoder then encodes a codeword for the selected option. This scheme reduces the size of the codeword because the codeword is as large as desired to uniquely indicate a choice from a small subset of allowable options, as opposed to uniquely indicating a choice from a potentially large set of all possible options. The decoder then decodes the choices by determining the set of allowable options in a similar manner as the encoder. By determining the set of allowable options, the decoder can read the codeword and determine the selection made by the encoder.

[0053] In step 113, the decoder performs block decoding. In particular, the decoder employs an inverse transform to generate a residual block. Then, the decoder employs the residual block and a corresponding prediction block to reconstruct an image block according to the partition. The prediction block may include both intra-prediction blocks and inter-prediction blocks, as generated in the encoder in step 105. The reconstructed image block is then placed in a frame of the reconstructed video signal according to the partition data determined in step 111. The syntax for step 113 may also be signaled in the bitstream via entropy coding as described above.

[0054] In step 115, filtering is performed on the frames of the reconstructed video signal in a manner similar to step 107 in the encoder. For example, noise suppression filters, deblocking filters, adaptive loop filters, and SAO filters may be applied to the frames to remove blocking artifacts. Once the frames have been filtered, the video signal may be output to a display in step 117 for viewing by an end user.

[0055] FIG. 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. In particular, the codec system 200 provides functionality to support the implementation of the operational method 100. The codec system 200 is generalized to show components employed in both an encoder and a decoder. The codec system 200 receives and segments a video signal, as described with respect to steps 101 and 103 in the operational method 100, which results in a segmented video signal 201. The codec system 200 then compresses the segmented video signal 201 into a coded bitstream when acting as an encoder, as described with respect to steps 105, 107, and 109 in the method 100. When acting as a decoder, the codec system 200 generates an output video signal from the bitstream, as described with respect to steps 111, 113, 115, and 117 in the operational method 100. The codec system 200 includes a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header formatting and context adaptive binary arithmetic coding (CABAC) component 231. Such components are connected as shown. In FIG. 2, the black lines indicate the movement of data to be encoded / decoded, and the dashed lines indicate the movement of control data that controls the operation of the other components. All of the components of the codec system 200 may be present in an encoder. A decoder may include a subset of the components of the codec system 200. For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223.These components are now described.

[0056] The partitioned video signal 201 is a captured video sequence that has been partitioned into blocks of pixels by a coding tree. The coding tree employs various partitioning modes to subdivide the blocks of pixels into smaller blocks of pixels. These blocks may then be further subdivided into smaller blocks. The blocks may be referred to as nodes on the coding tree. Larger parent nodes are partitioned into smaller child nodes. The number of times a node is subdivided is referred to as the depth of the node / coding tree. The partitioned blocks may possibly be contained within a coding unit (CU). For example, a CU may be a sub-part of a CTU that includes a luma block, a red differential chroma (Cr) block, and a blue differential chroma (Cb) block along with corresponding syntax instructions for the CU. The partitioning modes may include a binary tree (BT), a triple tree (TT), and a quad tree (QT), which are employed to partition a node into two, three, or four child nodes, respectively, of varying shapes depending on the partitioning mode employed. The segmented video signal 201 is forwarded to a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.

[0057] The generic coder control component 211 is configured to make decisions related to the coding of images of a video sequence into a bitstream according to application constraints. For example, the generic coder control component 211 manages the optimization of bitrate / bitstream size versus reconstruction quality. Such decisions may be made based on storage space / bandwidth availability and image resolution requirements. The generic coder control component 211 also manages buffer utilization in light of transmission rate to mitigate buffer under-run and over-run issues. To manage these issues, the generic coder control component 211 manages segmentation, prediction, and filtering by other components. For example, the generic coder control component 211 may dynamically increase compression complexity for higher resolution and higher bandwidth usage, or may decrease compression complexity for lower resolution and bandwidth usage. Thus, the generic coder control component 211 controls other components of the codec system 200 to balance video signal reconstruction quality with bitrate issues. The generic coder control component 211 creates control data that controls the operation of other components. Control data is also forwarded to the header formatting and CABAC component 231 to be encoded in the bitstream for signaling parameters for decoding at the decoder.

[0058] The partitioned video signal 201 is also sent to a motion estimation component 221 and a motion compensation component 219 for inter prediction. A frame or slice of the partitioned video signal 201 may be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter predictive coding of the received video blocks relative to one or more blocks in one or more reference frames to perform temporal prediction. The codec system 200 may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.

[0059] The motion estimation component 221 and the motion compensation component 219 may be highly integrated, but are illustrated separately for conceptual purposes. Motion estimation performed by the motion estimation component 221 is a process of generating motion vectors that estimate motion for a video block. A motion vector may indicate, for example, the displacement of a coded object relative to a predictive block. A predictive block is a block that is found to closely match a block to be coded in terms of pixel difference. A predictive block is sometimes referred to as a reference block. Such pixel difference may be determined by sum of absolute difference (SAD), sum of square difference (SSD), or other difference metric. HEVC employs several coded objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU may be divided into CTBs, which may then be divided into CBs for inclusion in a CU. A CU may be coded as a prediction unit (PU), which includes predictive data, and / or a transform unit (TU), which includes transformed residual data for the CU. The motion estimation component 221 generates the motion vectors, PUs, and TUs by using the rate-distortion analysis as part of a rate-distortion optimization process. For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. for a current block / frame, and may select the reference block, motion vector, etc. with the best rate-distortion characteristics. The best rate-distortion characteristics balance both the quality of the video reconstruction (e.g., the amount of data lost due to compression) and the coding efficiency (e.g., the size of the final encoding).

[0060] In some examples, the codec system 200 may calculate values ​​for sub-integer pixel positions of a reference picture stored in the decoded picture buffer component 223. For example, the video codec system 200 may interpolate values ​​for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of a reference picture. Thus, the motion estimation component 221 may perform motion searches for full pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy. The motion estimation component 221 calculates motion vectors for PUs of video blocks in an inter-coded slice by comparing the positions of the PUs with the positions of the predictive blocks of the reference pictures. The motion estimation component 221 outputs the calculated motion vectors as motion data to the header formatting and CABAC component 231 for encoding, and to the motion compensation component 219.

[0061] The motion compensation performed by the motion compensation component 219 may involve fetching or generating a predictive block based on a motion vector determined by the motion estimation component 221. Again, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated in some examples. Upon receiving a motion vector for a PU of a current video block, the motion compensation component 219 may locate a predictive block to which the motion vector points. A residual video block is then formed by subtracting pixel values ​​of the predictive block from pixel values ​​of the current video block being coded to form pixel difference values. In general, the motion estimation component 221 performs motion estimation on the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma and luma components. The predictive block and the residual block are forwarded to the transform scaling and quantization component 213.

[0062] The partitioned video signal 201 is also sent to an intra picture estimation component 215 and an intra picture prediction component 217. As with the motion estimation component 221 and the motion compensation component 219, the intra picture estimation component 215 and the intra picture prediction component 217 may be highly integrated, but are illustrated separately for conceptual purposes. The intra picture estimation component 215 and the intra picture prediction component 217 intra predict a current block relative to blocks in a current frame, as an alternative to the inter prediction performed by the motion estimation component 221 and the motion compensation component 219 between frames as described above. In particular, the intra picture estimation component 215 determines an intra prediction mode to be used to encode the current block. In some examples, the intra picture estimation component 215 selects an appropriate intra prediction mode for encoding the current block from a plurality of tested intra prediction modes. The selected intra prediction mode is then forwarded to the header formatting and CABAC component 231 for encoding.

[0063] For example, the intra picture estimation component 215 uses a rate-distortion analysis to calculate rate-distortion values ​​for various intra prediction modes tested, and selects an intra prediction mode with the best rate-distortion characteristics among the tested modes. The rate-distortion analysis generally determines the amount of distortion (or error) between a coding block and an original uncoded block that was coded to produce the coding block, as well as the bit rate (e.g., number of bits) used to produce the coding block. The intra picture estimation component 215 calculates a ratio from the distortion and rate for various coding blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block. In addition, the intra picture estimation component 215 may be configured to code a depth block of a depth map using a depth modeling mode (DMM) based on a rate-distortion optimization (RDO).

[0064] The intra-picture prediction component 217 may generate a residual block from the prediction block based on a selected intra-prediction mode determined by the intra-picture estimation component 215 when implemented on an encoder, or may read the residual block from the bitstream when implemented on a decoder. The residual block includes the difference in values ​​between the prediction block and the original block, represented as a matrix. The residual block is then forwarded to the transform scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both luma and chroma components.

[0065] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block to produce a video block comprising residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform may convert the residual information from a pixel value domain to a transform domain, such as a frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling involves applying a scale factor to the residual information, such that different frequency information is quantized with different granularity, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, the transform scaling and quantization component 213 may then perform a scan of a matrix containing the quantized transform coefficients, which are forwarded to the header formatting and CABAC component 231 to be encoded in the bitstream.

[0066] The scaling and inverse transform component 229 applies the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transformation, and / or quantization to reconstruct a residual block in the pixel domain, for example, for later use as a reference block that can become a predictive block for another current block. The motion estimation component 221 and / or the motion compensation component 219 may compute a reference block by adding the residual block back to the corresponding predictive block for use in motion estimation of a later block / frame. A filter is applied to the reconstructed reference block to mitigate artifacts created during the scaling, quantization, and transformation. Such artifacts may potentially cause inaccurate predictions (and may create additional artifacts) when subsequent blocks are predicted.

[0067] The filter control analysis component 227 and the in-loop filter component 225 apply filters to the residual block and / or the reconstructed image block. For example, a transformed residual block from the scaling and inverse transform component 229 may be combined with a corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. A filter may then be applied to the reconstructed image block. In some examples, the filter may be applied to the residual block instead. As with the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 may be highly integrated and implemented together, but are shown separately for conceptual purposes. The filters applied to the reconstructed reference block are applied to a particular spatial region and include multiple parameters to adjust how such filters are applied. The filter control analysis component 227 analyzes the reconstructed reference block and sets the corresponding parameters to determine where such filters should be applied. Such data is forwarded to the header formatting and CABAC component 231 as filter control data for encoding. The in-loop filter component 225 applies such filters based on the filter control data. The filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Such filters may be applied in the spatial / pixel domain (e.g., to reconstructed pixel blocks) or in the frequency domain, depending on the example.

[0068] When operating as an encoder, the filtered reconstructed image blocks, residual blocks, and / or predictive blocks are stored in the decoded picture buffer component 223 for later use in motion estimation as described above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed filtered blocks and forwards them towards the display as part of the output video signal. The decoded picture buffer component 223 may be any memory device capable of storing predictive blocks, residual blocks, and / or reconstructed image blocks.

[0069] The header formatting and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into a coded bitstream for transmission towards the decoder. In particular, the header formatting and CABAC component 231 generates various headers for encoding control data, such as general control data and filter control data. Furthermore, prediction data, including intra prediction and motion data, and residual data in the form of quantized transform coefficient data are all encoded in the bitstream. The final bitstream contains all information desired by the decoder to reconstruct the partitioned original video signal 201. Such information may also include an intra prediction mode index table (also called a codeword mapping table), definitions of coding contexts for various blocks, indications of most probable intra prediction modes, indications of partition information, etc. Such data may be encoded by employing entropy coding. For example, the information may be encoded by employing context adaptive variable length coding (CAVLC), CABAC, syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding technique. Following entropy coding, the coded bitstream may be transmitted to another device (e.g., a video decoder) or archived for later transmission or retrieval.

[0070] 3 is a block diagram illustrating an example video encoder 300. The video encoder 300 may be employed to perform the encoding functions of the codec system 200 and / or to perform steps 101, 103, 105, 107, and / or 109 of the method of operation 100. The encoder 300 segments an input video signal, resulting in a segmented video signal 301 that is substantially similar to the segmented video signal 201. The segmented video signal 301 is then compressed and encoded into a bitstream by components of the encoder 300.

[0071] In particular, the partitioned video signal 301 is forwarded to an intra-picture prediction component 317 for intra prediction. The intra-picture prediction component 317 may be substantially similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. The partitioned video signal 301 is also forwarded to a motion compensation component 321 for inter prediction based on a reference block in a decoded picture buffer component 323. The motion compensation component 321 may be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction block and the residual block from the intra-picture prediction component 317 and the motion compensation component 321 are forwarded to a transform and quantization component 313 for transform and quantization of the residual block. The transform and quantization component 313 may be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual block and the corresponding prediction block (together with associated control data) are forwarded to an entropy coding component 331 for coding into a bitstream. The entropy coding component 331 may be substantially similar to the header formatting and CABAC component 231 .

[0072] The transformed and quantized residual block and / or the corresponding prediction block are also forwarded from the transform and quantization component 313 to the inverse transform and quantization component 329 for reconstruction into a reference block for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially similar to the scaling and inverse transform component 229. An in-loop filter in the in-loop filter component 325 is also applied to the residual block and / or the reconstructed reference block depending on the example. The in-loop filter component 325 may be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters as described with respect to the in-loop filter component 225. The filtered block is then stored in the decoded picture buffer component 323 for use as a reference block by the motion compensation component 321. The decoded picture buffer component 323 may be substantially similar to the decoded picture buffer component 223.

[0073] 4 is a block diagram illustrating an example video decoder 400. The video decoder 400 may be employed to perform the decoding functions of the codec system 200 and / or to perform steps 111, 113, 115, and / or 117 of the method of operation 100. The decoder 400 receives a bitstream, e.g., from the encoder 300, and generates a reconstructed output video signal based on the bitstream for display to an end user.

[0074] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to perform an entropy decoding scheme, such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 may employ header information to provide a context for interpreting additional data encoded as codewords in the bitstream. The decoded information includes any desired information for decoding a video signal, such as general control data, filter control data, partition information, motion data, prediction data, and quantized transform coefficients from residual blocks. The quantized transform coefficients are forwarded to the inverse transform and quantization component 429 for reconstruction into residual blocks. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.

[0075] The reconstructed residual block and / or predictive block are forwarded to the intra picture prediction component 417 for reconstruction into an image block based on an intra prediction operation. The intra picture prediction component 417 may be similar to the intra picture estimation component 215 and the intra picture prediction component 217. In particular, the intra picture prediction component 417 employs a prediction mode to identify the location of a reference block in a frame and applies a residual block to the result to reconstruct an intra predicted image block. The reconstructed intra predicted image block and / or the residual block and corresponding inter prediction data are forwarded via the in-loop filter component 425 to the decoded picture buffer component 423, which may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image block, the residual block, and / or the predictive block, and such information is stored in the decoded picture buffer component 423. The reconstructed image blocks from the decoded picture buffer component 423 are forwarded to the motion compensation component 421 for inter prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. In particular, the motion compensation component 421 employs motion vectors from a reference block to generate a prediction block and applies a residual block to the result to reconstruct an image block. The resulting reconstructed block may also be forwarded to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconstructed image blocks, which may be reconstructed into frames via the partition information. Such frames may also be placed into sequences. The sequences are output to a display as a reconstructed output video signal.

[0076] 5 is a schematic diagram illustrating an exemplary bitstream 500 including an encoded video sequence. For example, the bitstream 500 may be generated by the codec system 200 and / or the encoder 300 for decoding by the codec system 200 and / or the decoder 400. As another example, the bitstream 500 may be generated by the encoder in step 109 of the method 100 for use by the decoder in step 111.

[0077] The bitstream 500 includes a sequence parameter set (SPS) 510, multiple picture parameter sets (PPS) 512, a tile group header 514, and image data 520. The SPS 510 includes sequence data common to all pictures in a video sequence included in the bitstream 500. Such data may include picture size determination, bit depth, coding tool parameters, bit rate constraints, etc. The PPS 512 includes parameters specific to one or more corresponding pictures. Thus, each picture in the video sequence may reference one PPS 512. The PPS 512 may indicate coding tools, quantization parameters, offsets, picture-specific coding tool parameters (e.g., filter control), etc. available for tiles in the corresponding picture. The tile group header 514 includes parameters specific to each tile group in the picture. Thus, there may be one tile group header 514 for each tile group in the video sequence. The tile group header 514 may include tile group information, picture order count (POC), reference picture list, prediction weights, tile entry points, deblocking parameters, etc. Note that some systems refer to the tile group header 514 as a slice header and use such information to support slices rather than tile groups.

[0078] The image data 520 includes video data encoded according to inter-prediction and / or intra-prediction, as well as corresponding transformed and quantized residual data. Such image data 520 is sorted according to a partition used to partition the image before encoding. For example, an image in the image data 520 is divided into tiles 523. The tiles 523 are further divided into coding tree units (CTUs). The CTUs are further divided into coding blocks based on the coding tree. The coding blocks may then be encoded / decoded according to a prediction mechanism. An image / picture may include one or more tiles 523.

[0079] A tile 523 is a partitioned portion of a picture created by horizontal and vertical boundaries. The tile 523 may be rectangular and / or square. In particular, the tile 523 includes four sides connected at right angles. The four sides include two pairs of parallel sides. Furthermore, the sides in a parallel side pair are equal in length. Thus, the tile 523 may be any rectangular shape, where a square is a special case of a rectangle where all four sides are equal in length. A picture may be arranged in rows and columns of tiles 523. A tile row is a set of tiles 523 arranged horizontally adjacent to create a continuous line from the left boundary to the right boundary of the picture (or vice versa). A tile column is a set of tiles 523 arranged vertically adjacent to create a continuous line from the top boundary to the bottom boundary of the picture (or vice versa). A tile 523 may or may not allow prediction based on other tiles 523, depending on the example. Each tile 523 may have a unique tile index within a picture. The tile index is a procedurally selected numeric identifier that may be used to distinguish one tile 523 from another tile 523. For example, the tile index may increase numerically in a raster scan order. The raster scan order is left-to-right and top-to-bottom. Note that in some examples, tiles 523 may also be assigned a tile identifier (ID). The tile ID is an assigned identifier that may be used to distinguish one tile 523 from another tile 523. In some examples, the calculation may employ the tile ID rather than the tile index. Furthermore, in some examples, the tile ID may be assigned to have the same value as the tile index. The tile index and / or tile ID may be signaled to indicate the tile group that includes the tile 523. For example, the tile index and / or tile ID may be employed to map picture data associated with the tile 523 to an appropriate location for display.A tile group is a related set of tiles 523 that may be extracted and coded separately, for example, to support display of regions of interest and / or to support parallel processing. Tiles 523 within a tile group may be coded without reference to tiles 523 outside the tile group. Each tile 523 may be assigned to a corresponding tile group, and thus a picture may contain multiple tile groups.

[0080] 6A-6E illustrate an example mechanism 600 for creating an extractor track 610 for combining multiple resolution sub-pictures from different bitstreams into a single picture for use in virtual reality (VR) applications. The mechanism 600 may be employed to support example use cases of the method 100. For example, the mechanism 600 may be employed to generate a bitstream 500 for transmission from the codec system 200 and / or the encoder 300 towards the codec system 200 and / or the decoder 400. As a particular example, the mechanism 600 may be employed for use with VR, OMAF, 360-degree video, etc.

[0081] In VR, only a portion of the video is displayed to the user. For example, the VR video may be shot to include a sphere surrounding the user. The user may employ a head mounted display (HMD) to view the VR video. The user may point the HMD towards a target area. The target area is displayed to the user, and other video data is discarded. In this way, at any given moment, the user sees only the portion of the VR video that the user selected. This approach mimics the user's perception, and thus allows the user to experience the virtual environment in a way that mimics the real environment. One of the problems with this approach is that although the entire VR video may be transmitted to the user, only the current viewport of the video is actually used, and the rest is discarded. To increase signaling efficiency for streaming applications, the user's current viewport may be transmitted at a higher first resolution, and other viewports may be transmitted at a lower second resolution. In this way, the viewports that may be discarded occupy less bandwidth than the viewports that may be seen by the user. When the user selects a new viewport, the lower resolution content may be shown until the decoder may request that a different current viewport be transmitted at the higher first resolution. To support this functionality, a mechanism 600 may be employed to create an extractor track 610 as shown in Figure 6E. The extractor track 610 is a track of image data that encapsulates pictures at multiple resolutions for use as described above.

[0082] The mechanism 600 encodes the same video content in a first resolution 611 and a second resolution 612, as shown in FIG. 6A and FIG. 6B, respectively. As a specific example, the first resolution 611 may be 5120×2560 luma samples, and the second resolution 612 may be 2560×1280 luma samples. Pictures of the video may be partitioned into tiles 601 in the first resolution 611 and tiles 603 in the second resolution 612, respectively. In the illustrated example, the tiles 601 and 603 are each partitioned into a 4×2 grid. Furthermore, an MCTS may be coded for the location of each tile 601 and 603. The pictures in the first resolution 611 and the second resolution 612 each result in an MCTS sequence that represents the video over time at the corresponding resolution. Each coded MCTS sequence is stored as a sub-picture track or a tile track. The mechanism 600 can then use the pictures to create segments to support viewport-adaptive MCTS selection. For example, each range of viewing orientations that causes a different selection of high-resolution and low-resolution MCTSs is considered. In the illustrated example, four tiles 601 containing MCTSs in a first resolution 611 and four tiles 603 containing MCTSs in a second resolution 613 are obtained.

[0083] The mechanism 600 can then create an extractor track 610 for each possible viewport-adaptive MCTS selection. Figures 6C and 6D show an exemplary viewport-adaptive MCTS selection. In particular, a set of selected tiles 605 and 607 are selected in a first resolution 611 and a second resolution 612, respectively. The selected tiles 605 and 607 are shown in shades of gray. In the illustrated example, the selected tile 605 is a tile 601 in the first resolution 611 that will be shown to the user, and the selected tile 607 is a tile 603 in the second resolution 612 that may be discarded but is retained to support display if the user selects a new viewport. The selected tiles 605 and 607 are then combined into a single picture that includes image data in both the first resolution 611 and the second resolution 612. Such pictures are combined to create the extractor track 610. FIG. 6E shows, for illustrative purposes, a single picture from a corresponding extractor track 610. As shown, the picture in the extractor track 610 includes selected tiles 605 and 607 from a first resolution 611 and a second resolution 612. As mentioned above, FIGS. 6C-6E show a single viewport adaptive MCTS selection. To allow user selection of any viewport, an extractor track 610 should be created for each possible combination of selected tiles 605 and 607.

[0084] In the illustrated example, each selection of tiles 603 encapsulating content from a bitstream of a second resolution 612 includes two slices. A RegionWisePackingBox may be included in the extractor track 610 to create a mapping between the packed picture and the projected picture in the ERP format. In the presented example, the resolved bitstream from the extractor track has a resolution of 3200x2560. Thus, a 4000 sample (4K) capable decoder may decode content in which the viewport is extracted from a coded bitstream having a 5000 sample 5K (5120x2560) resolution.

[0085] As shown, the extractor track 610 includes two rows of high resolution tiles 601 and four rows of low resolution tiles 603. Thus, the extractor track 610 includes two slices of high resolution content and four slices of low resolution content. Uniform tiling does not support such use cases. Uniform tiling is defined by a set of tile columns and a set of tile rows. The tile columns extend from the top of the picture to the bottom of the picture. Similarly, the tile rows extend from the left of the picture to the right of the picture. Although such a structure may be easily defined, this structure cannot effectively support advanced use cases, such as the use case described by the mechanism 600. In the illustrated example, different numbers of rows are employed in different sections of the extractor track 610. If uniform tiling were employed, the tiles on the right side of the extractor track 610 would have to be rewritten to accommodate two slices each. This approach is inefficient and computationally complex.

[0086] The present disclosure includes a flexible tiling scheme as described below that does not require tiles to be rewritten to include a different number of slices. The flexible tiling scheme allows a tile 601 to include content in a first resolution 611. The flexible tiling scheme also allows the tile 601 to be partitioned into smaller tiles that can each be directly mapped to a tile 603 in a second resolution 612. This direct mapping is more efficient because such an approach does not require tiles to be rewritten / readdressed when different resolutions are combined as described above.

[0087] FIG. 7 illustrates an exemplary video conferencing application 700 that stitches together pictures of multiple resolutions from different bitstreams into a single picture for display. The application 700 may be employed to support an exemplary use case of the method 100. For example, the application 700 may be employed in the codec system 200 and / or the decoder 400 to display video content from the bitstream 500 from the codec system 200 and / or the encoder 300. The video conferencing application 700 displays a video sequence to a user. The video sequence includes pictures displaying a speaking participant 701 and other participants 703. The speaking participant 701 is displayed in a higher first resolution and the other participants 703 are displayed in a smaller second resolution. To code such a picture, the picture should include a portion with a single row and a portion with three rows. To support such a scenario with uniform tiling, the picture is partitioned into a left tile and a right tile. The right tile is then rewritten / readdressed to include three rows. Such re-addressing incurs both compression and performance penalties. The flexible tiling scheme described below allows a single tile to be partitioned into smaller tiles and mapped to tiles in the sub-picture bitstreams associated with other participants 703. In this way, the speaking participant 701 can be mapped directly to a first level tile, and the other participants 703 can be mapped to second level tiles split off from the first tile without such rewriting / re-addressing.

[0088] 8A-8B are schematic diagrams illustrating an example flexible video tiling scheme 800 capable of supporting multiple tiles with different resolutions in the same picture. The flexible video tiling scheme 800 may be adopted to support more efficient coding mechanisms 600 and applications 700. Thus, the flexible video tiling scheme 800 may be adopted as part of the method 100. Furthermore, the flexible video tiling scheme 800 may be adopted by the codec system 200, the encoder 300, and / or the decoder 400. The result of the flexible video tiling scheme 800 may be stored in the bitstream 500 for transmission between the encoder and the decoder.

[0089] As shown in FIG. 8A, a picture (e.g., a frame, image, etc.) may be partitioned into first level tiles 801, also called level 1 tiles. As shown in FIG. 8B, the first level tiles 801 may be selectively partitioned to create second level tiles 803, also called level 2 tiles. The first level tiles 801 and the second level tiles 803 may then be employed to create a picture having sub-pictures coded at multiple resolutions. The first level tiles 801 are tiles generated by completely partitioning the picture into a set of columns and a set of rows. The second level tiles 803 are tiles generated by partitioning the first level tiles 801.

[0090] As described above, in various scenarios, for example, in VR and / or teleconferencing, a video may be coded at multiple resolutions. The video may also be coded by employing slices at each resolution. The lower resolution slices are smaller than the higher resolution slices. To create a picture with multiple resolutions, the picture may be partitioned into first level tiles 801. Slices from the highest resolution may be included directly in the first level tiles 801. Furthermore, the first level tiles 801 may be partitioned into second level tiles 803 that are smaller than the first level tiles 801. Thus, the smaller second level tiles 803 may directly accept slices with lower resolutions. In this way, slices from each resolution may be compressed into a single picture, for example, via a tile index relationship, without requiring tiles with different resolutions to be dynamically readdressed to use a coherent addressing scheme. The first level tiles 801 and the second level tiles 803 may be implemented as MCTS and may therefore accept motion constrained image data at different resolutions.

[0091] The first level tiles 801 and the second level tiles 803 may also be assigned to tile groups 811, 812, and / or 813. The tile groups 811, 812, and / or 813 are selections of related tiles that undergo similar processing during encoding and decoding. As an example, the tile groups 811, 812, and / or 813 may be employed to store different sub-pictures at different resolutions. The different tile groups 811, 812, and / or 813 may have different parameters, and thus the tile groups 811, 812, and / or 813 allow the different sub-pictures to be treated differently by the encoder and / or decoder. The tile groups 811, 812, and / or 813 may be constrained to be rectangular rather than in raster scan order to support the features described herein. It should be noted that a square is a special case of a rectangle, and thus a rectangular shape should be understood to include a square shape. As a specific example, a first level tile 801 may be assigned to rectangular tile groups 811, 812, and / or 813. Each second level tile 803 generated by splitting the first level tile 801 may then be assigned to a rectangular tile group 811, 812, and / or 813 associated with the corresponding first level tile 801 from which the second level tile 803 was split. As mentioned above, the tile groups 811, 812, and / or 813 are rectangular and not raster scanned. The raster scan order proceeds in CTU order across the picture from left to right until the right side of the picture is reached. The raster scan then moves down to the next line of the CTU, moving from the left side of the picture towards the right side. For example, a raster scan order tile group cannot include only first level tiles 801 or only second level tiles 803, as shown in FIG. 8B, because the raster scan order will proceed across the picture and then move down.Thus, rectangular tile groups 811, 812, and / or 813 allow separate processing of first level tiles 801 and only second level tiles 803, which may contain sub-pictures of different resolution.

[0092] This disclosure includes many aspects. As a specific example, a first level tile 801 is divided into second level tiles 803. The second level tiles 803 may then each be constrained to include a single rectangular slice of picture data (e.g., at a smaller resolution). A rectangular slice is a slice that is constrained to maintain a rectangular shape and is therefore coded based on horizontal and vertical picture boundaries. Thus, a rectangular slice is not coded based on a raster scan group (which may include CTUs in lines from left to right and lines from top to bottom and not maintain a rectangular shape). A slice is a spatially distinct region of a picture / frame that is coded separately from any other region in the same frame / picture. In another example, a first level tile 801 may be divided into two or more complete second level tiles 803. In such a case, the first level tile 801 may not include a partial second level tile 803. In another example, the configuration of the first level tiles 801 and the second level tiles 803 may be signaled in a parameter set in the bitstream, such as a PPS associated with the picture partitioned to create the tile. In one example, a split indication, such as a flag, may be coded in the parameter set for each first level tile 801. The indication indicates which first level tiles 801 are further split into second level tiles 803. In another example, the configuration of the second level tiles 803 may be signaled as a number of second level tile columns and a number of second level tile rows.

[0093] In another example, the first level tiles 801 and the second level tiles 803 may be assigned to tile groups. Such tile groups may be constrained such that all tiles in the corresponding tile group are constrained to cover a rectangular region of the picture (e.g., as opposed to raster scan). For example, some systems may add tiles to the tile group in raster scan order. This includes adding an initial tile in the current row, proceeding to add each tile in the row until the left picture boundary of the current row is reached, proceeding to the right boundary of the next row, and adding each tile in the next row until the final tile is reached, etc. This approach may result in non-rectangular shapes that extend beyond the picture. Such shapes may not be useful for creating pictures with multiple resolutions as described herein. Instead, the present example may constrain the tile groups such that any first level tile 801 and / or second level tile 803 may be added to the tile group (e.g., in any order), but the resulting tile group must be rectangular or square (e.g., including four sides connected at right angles). This constraint may ensure that second level tiles 803 that are partitioned from a single first level tile 801 are not placed into different tile groups.

[0094] In another example, when the first level tile width is smaller than twice the minimum width threshold and the first level tile height is smaller than twice the minimum height threshold, data explicitly indicating the number of second level tile columns and the number of second level tile rows may be omitted from the bitstream because a first level tile 801 that satisfies such a condition may not be split into two or more columns or one row, respectively, and therefore such information may be inferred by a decoder. In another example, partitioning indications indicating which first level tiles 801 are partitioned into second level tiles 803 may be omitted from the bitstream for some first level tiles 801. For example, when a first level tile 801 has a first level tile width smaller than the minimum width threshold and a first level tile height smaller than the minimum height threshold, such data may be omitted because a first level tile 801 that satisfies such a condition is too small to be partitioned into second level tiles 803, and therefore such information may be inferred by a decoder.

[0095] As described above, the flexible video tiling scheme 800 supports merging sub-pictures from different bitstreams into a picture containing multiple resolutions. The following describes various embodiments that support such functionality. Generally, this disclosure describes methods for signaling and coding tiles in video coding that partition pictures in a more flexible manner than the tiling schemes in HEVC. More specifically, this disclosure describes some tiling schemes in which tile columns may not extend uniformly from the top to the bottom of a coded picture, and similarly tile rows may not extend uniformly from the left to the right of a coded picture.

[0096] For example, based on the HEVC tiling approach, some tiles should be further divided into multiple tile rows to support the functionality described in Figures 6A-6E and 7. Furthermore, depending on how the tiles are arranged, the tiles should be further divided into tile columns. For example, in Figure 7, in some cases participants 2-4 may be arranged below participant 1, which may be supported by dividing the tiles into columns. To satisfy these scenarios, the first level tiles may be divided into tile rows and tile columns of second level tiles as described below.

[0097] For example, the tile structure may be relaxed as follows: Tiles in the same picture are not required to be a specific number of tile rows. Furthermore, tiles in the same picture are not required to be a specific number of tile columns. For signaling of flexible tiles, the following steps may be used: The first level tile structure may be defined by tile columns and tile rows as defined in HEVC. The tile columns and tile rows may be uniform or non-uniform in size. Each of these tiles may be referred to as a first level tile. A flag may be signaled to specify whether each first level tile is further split into one or more tile columns and one or more tile rows. If a first level tile is further split, the tile columns and tile rows may be either uniform or non-uniform in size. The new tiles resulting from the split of the first level tiles are referred to as second level tiles. The flexible tile structure may be limited to only second level tiles, and thus, in some examples, further splitting of any second level tiles is not allowed. In other examples, further division of the second level tiles may be applied to create subsequent level tiles in a manner similar to the creation of the second level tiles from the first level tiles.

[0098] For simplicity, when a first level tile is divided into two or more second level tiles, the division may always use uniform sized tile columns and uniform tile rows. The derivation of the tile locations, sizes, indexes, and scan order of the flexible tiles defined by this approach is described below. For simplicity, when such a flexible tile structure is used, the tile group may be constrained to include one or more complete first level tiles. In this example, when a tile group includes a second level tile, all second level tiles resulting from the division of the same first level tile should be included in the tile group. When such a flexible tile structure is used, the tile group may be further constrained to include one or more tiles, where all tiles together belong to a tile group that covers a rectangular area of ​​the picture. In another aspect, when such a flexible tile structure is used, the tile group may include one or more first level tiles, where all tiles together belong to a tile group that covers a rectangular area of ​​the picture.

[0099] In one example, signaling of flexible tiles may be as follows: A minimum tile width and a minimum tile height are defined values. A first level tile structure may be defined by tile columns and tile rows. The tile columns and tile rows may be uniform or non-uniform in size. Each of these tiles may be referred to as a first level tile. A flag may be signaled to specify whether any of the first level tiles may be further divided. This flag may be absent when the width of each first level tile is less than or equal to twice the minimum tile width and the height of each first level tile is less than or equal to twice the minimum tile height. When absent, the value of the flag is inferred to be equal to 0.

[0100] In one example, the following applies for each first level tile: A flag may be signaled to specify whether the first level tile is further divided into one or more tile columns and one or more tile rows. The presence of the flag may be constrained as follows: The flag is present / signaled if the first level tile width is greater than the minimum tile width or the first level tile height is greater than the minimum tile height. Otherwise, the flag is not present and the value of the flag is inferred to be equal to 0 indicating that the first level tile is not further divided.

[0101] If the first level tile is further divided, the number of tile columns and the number of tile rows for this division may be further signaled. The tile columns and tile rows may be either uniform or non-uniform in size. The tiles resulting from the division of the first level tiles are called second level tiles. The existence of the number of tile columns and the number of tile rows may be constrained as follows: When the first level tile width is smaller than twice the minimum tile width, the number of tile columns may not be signaled and the number of tile column values ​​may be inferred to be equal to 1. The signaling may employ a _minus1 syntax element, such that the signaled syntax element value may be 0 and the number of tile columns is the value of the syntax element + 1. This approach may further compress the signaling data. When the first level tile height is smaller than twice the minimum tile height, the number of tile rows may not be signaled and the value of the number of tile rows may be inferred to be equal to 0. The signaled syntax element value may be 0 and the number of tile rows may be the syntax element value + 1 to further compress the signaling data. The tiles resulting from the division of the first level tiles may be referred to as second level tiles. The flexible tile structure may be limited to only second level tiles such that further division of any second level tiles is not allowed. In other examples, further division of second level tiles may be applied in a manner similar to the division of first level tiles into second level tiles.

[0102] In one example, signaling of a flexible tile structure may be as follows: When a picture includes more than one tile, a signal such as a flag may be employed in a parameter set that is directly or indirectly referenced by the corresponding tile group. The flag may specify whether the corresponding tile structure is a uniform tile structure or a non-uniform tile structure (e.g., a flexible tile structure as described herein). The flag may be referred to as uniform_tile_structure_flag. When uniform_tile_structure_flag is equal to 1, HEVC-style uniform tile structure signaling is employed, e.g., by signaling num_tile_columns_minus1 and num_tile_rows_minus1 to indicate a single level of uniform tiles. When uniform_tile_structure_flag is equal to 0, the following information may also be signaled: The number of tiles in a picture may be signaled by the syntax element num_tiles_minus2, which indicates that the number of tiles in the picture (NumTilesInPic) is equal to num_tiles_minus2+2. This may result in bit savings in signaling, since by default pictures may be considered to be tiles. For each tile, the addresses of the first coding block (e.g., CTU) and the last coding block of the tile, except for the last one, are signaled. The addresses of the coding blocks may be the index of the block in the picture (e.g., the index of the CTU in the picture). The syntax elements for such coding blocks may be tile_first_block_address[i] and tile_last_block_address[i]. These syntax elements may be coded as ue(v) or u(v). When the syntax elements are coded as u(v), the number of bits used to represent each of the syntax elements is ceil(log2(maximum number of coding blocks in a picture)).The addresses of the first and last coding blocks of the last tile do not need to be signaled, but instead may be derived based on the picture size, among the luma samples, and the collection of all other tiles in the picture.

[0103] In one example, rather than signaling the addresses of the first and last coding blocks of the tile, except for the last one, for each tile, the address of the first coding block of the tile, as well as the width and height of the tile, may be signaled. In another example, rather than signaling the addresses of the first and last coding blocks of the tile, except for the last one, for each tile, the offset of the top-left point of the tile (e.g., the top-left of the picture) relative to the picture original, as well as the width and height of the tile, may be signaled. In yet another example, rather than signaling the addresses of the first and last coding blocks of the tile, except for the last one, for each tile, the following information may be signaled: The width and height of the tile may be signaled. Also, the location of each tile may not be signaled. Instead, a flag may be signaled to specify whether the tile should be placed immediately to the right or immediately below the previous tile. If the tile can only be to the right or only below the previous tile, this flag may not be present. The top-left offset of the first tile may be set to always be the origin / top-left of the picture (eg, x=0 and y=0).

[0104] For signaling efficiency, a set of unique tile sizes (e.g., width and height) may be signaled. This list of unique tile sizes may be referenced by an index from a loop that contains the signaling of each tile size. In some examples, the tile locations and tile sizes as derived from the signaled tile structure must constrain partitioning to ensure that no gaps or overlaps occur between any tiles.

[0105] The following constraints may also apply: The tile shape may be required to be rectangular (e.g., not raster scan shaped). The units of tiles in a picture must cover the picture without any gaps between tiles and without any overlap. When decoding is performed with only one core, for coding of a current coding block (e.g., CTU) that is not at the left edge of the picture, the left neighboring coding block must be decoded before the current coding block. When decoding is performed with only one core, for coding of a current coding block (e.g., CTU) that is not at the top edge of the picture, the top neighboring coding block must be decoded before the current coding block. When two tiles have tile indices (e.g., idx3 and idx4) that are adjacent to each other, one of the following is true: When two tiles share a vertical edge and / or a first tile has a top-left location at (Xa,Ya) with size (Wa and Ha representing its width and height) and a second tile has a top-left location at (Xb,Yb), then Yb=Ya+Ha.

[0106] The following constraints may also be applied: When a tile has more than one left neighbor, the tile's height must be equal to the sum of the heights of all its left neighbors. When a tile has more than one right neighbor, the tile's height must be equal to the sum of the heights of all its left neighbors. When a tile has more than one above neighbor, the tile's width must be equal to the sum of the widths of all its above neighbors. When a tile has more than one below neighbor, the tile's width must be equal to the sum of the widths of all its below neighbors.

[0107] The following are specific example embodiments of the above aspects: The CTB raster and tile traversal process may be as follows: A list ColWidth[i], for i ranging from 0 to num_level1_tile_columns_minus1 inclusive, specifying the width of the i-th first level tile column in units of CTB, may be derived as follows: if(uniform_level1_tile_spacing_flag) for(i=0;i<=num_level1_tile_columns_minus1;i++) ColWidth[i]=((i+1)*PicWidthInCtbsY) / (num_level1_tile_columns_minus1+1)-(i*PicWidthInCtbsY) / (num_level1_tile_columns_minus1+1) else { ColWidth[num_level1_tile_columns_minus1]=PicWidthInCtbsY (6-1) for(i=0;i <num_level1_tile_columns_minus1;i++){ ColWidth[i]=tile_level1_column_width_minus1[i]+1 ColWidth[num_tile_level1_columns_minus1] -= ColWidth[i] } }

[0108] The list RowHeight[j], for j ranging from 0 to num_level1_tile_rows_minus1 inclusive, specifying the height of the jth tile row in CTB units, may be derived as follows: if(uniform_level1_tile_spacing_flag) for(j=0;j<=num_level1_tile_rows_minus1;j++) RowHeight[j]=((j+1)*PicHeightInCtbsY) / (num_level1_tile_rows_minus1+1)-(j*PicHeightInCtbsY) / (num_level1_tile_rows_minus1+1) else { RowHeight[num_level1_tile_rows_minus1]=PicHeightInCtbsY (6-2) for(j=0;j <num_level1_tile_rows_minus1;j++){ RowHeight[j]=tile_level1_row_height_minus1[j]+1 RowHeight[num_level1_tile_rows_minus1] -= RowHeight[j] } }

[0109] The list colBd[i], for i ranging from 0 to num_level1_tile_columns_minus1+1 inclusive, specifying the location of the i-th tile column boundary in units of CTB, may be derived as follows: for(colBd[0]=0,i=0;i<=num_level1_tile_columns_minus1;i++) colBd[i+1]=colBd[i]+ColWidth[i] (6-3)

[0110] The list rowBd[j], for j ranging from 0 to num_level1_tile_rows_minus1+1 inclusive, specifying the location of the jth tile row boundary in units of CTB, may be derived as follows: for(rowBd[0]=0,j=0;j<=num_level1_tile_rows_minus1;j++) rowBd[j+1]=rowBd[j]+RowHeight[j] (6-4)

[0111] The variable NumTilesInPic, which specifies the number of tiles in the picture with reference to the PPS, and the lists TileColBd[i], TileRowBd[i], TileWidth[i], and TileHeight[i], for i ranging from 0 to NumTilesInPic-1 inclusive, which specify the location of the i-th tile column boundary in units of CTBs, the location of the i-th tile row boundary in units of CTBs, the width of the i-th tile column in units of CTBs, and the height of the i-th tile column in units of CTBs, may be derived as follows: for(tileIdx=0,i=0;i <NumLevel1Tiles;i++){ tileX=i%(num_level1_tile_columns_minus1+1) tileY=i / (num_level1_tile_columns_minus1+1) if(!level2_tile_split_flag[i]){ (6-5) TileColBd[tileIdx]=colBd[tileX] TileRowBd[tileIdx]=rowBd[tileY] TileWidth[tileIdx]=ColWidth[tileX] TileHeight[tileIdx]=RowHeight[tileY] tileIdx++ }else{ for(k=0;k<=num_level2_tile_columns_minus1[i];k++) colWidth2[k]=((k+1)*ColWidth[tileX]) / (num_level2_tile_columns_minus1[i]+1)-(k*ColWidth[tileX]) / (num_level2_tile_columns_minus1[i]+1) for(k=0;k<=num_level2_tile_rows_minus1[i];k++) rowHeight2[k]=((k+1)*RowHeight[tileY]) / (num_level2_tile_rows_minus1[i]+1)-(k*RowHeight[tileY]) / (num_level2_tile_rows_minus1[i]+1) for(colBd2[0]=0,k=0;k<=num_level2_tile_columns_minus1[i];k++) colBd2[k+1]=colBd2[k]+colWidth2[k] for(rowBd2[0]=0,k=0;k<=num_level2_tile_rows_minus1[i];k++) rowBd2[k+1]=rowBd2[k]+rowHeight2[k] numSplitTiles=(num_level2_tile_columns_minus1[i]+1)*(num_level2_tile_rows_minus1[i]+1) for(k=0;k<numSplitTiles;k++){ tileX2=k%(num_level2_tile_columns_minus1[i]+1) tileY2=k / (num_level2_tile_columns_minus1[i]+1) TileColBd[tileIdx]=colBd[tileX]+colBd2[tileX2] TileRowBd[tileIdx]=rowBd[tileY]+rowBd2[tileY2] TileWidth[tileIdx]=colWidth2[tileX2] TileHeight[tileIdx]=rowHeight2[tileY2] tileIdx++ } } } NumTilesInPic=tileIdx

[0112] A list CtbAddrRsToTs[ctbAddrRs] for ctbAddrRs ranging from 0 to PicSizeInCtbsY-1, inclusive, that specifies the conversion from CTB addresses in the CTB raster scan of a picture to CTB addresses in the tile scan can be derived as follows: for(ctbAddrRs=0;ctbAddrRs <PicSizeInCtbsY;ctbAddrRs++){ tbX=ctbAddrRs%PicWidthInCtbsY tbY=ctbAddrRs / PicWidthInCtbsY tileFound=FALSE for(tileIdx=NumTilesInPic-1,i=0;i <NumTilesInPic-1 && !tileFound;i++){ (6-6) tileFound=tbX<(TileColBd[i]+TileWidth[i]) && tbY<(TileRowBd[i]+TileHeight[i]) if(tileFound) tileIdx=i } CtbAddrRsToTs[ctbAddrRs]=0 for(i=0;i <tileIdx;i++) CtbAddrRsToTs[ctbAddrRs] += TileHeight[i]*TileWidth[i] CtbAddrRsToTs[ctbAddrRs] += (tbY-TileRowBd[tileIdx])*TileWidth[tileIdx]+tbX-TileColBd[tileIdx] }

[0113] The list CtbAddrTsToRs[ctbAddrTs] for ctbAddrTs ranging from 0 to PicSizeInCtbsY-1 inclusive, which specifies the conversion from CTB addresses in the tile scan to CTB addresses in the CTB raster scan of the picture, can be derived as follows: for(ctbAddrRs=0;ctbAddrRs <PicSizeInCtbsY;ctbAddrRs++) (6-7) CtbAddrTsToRs[CtbAddrRsToTs[ctbAddrRs]]=ctbAddrRs

[0114] A list TileId[ctbAddrTs] for ctbAddrTs ranging from 0 to PicSizeInCtbsY-1, inclusive, that specifies the conversion from CTB addresses to tile IDs in a tile traversal can be derived as follows: for(i=0,tileIdx=0;i<=NumTilesInPic;i++,tileIdx++) for(y = TileRowBd[i];y <TileRowBd[i+1];y++) (6-8) for(x = TileColBd[i];x <TileColBd[i+1];x++) TileId[CtbAddrRsToTs[y*PicWidthInCtbsY+x]]=tileIdx

[0115] The list NumCtusInTile[tileIdx], for tileIdx ranging from 0 to NumTilesInPic-1 inclusive, which specifies the conversion from tile index to the number of CTUs in the tile, can be derived as follows: for(i=0,tileIdx=0;i <NumTilesInPic;i++,tileIdx++) (6-9) NumCtusInTile[tileIdx]=TileColWidth[tileIdx]*TileRowHeight[tileIdx]

[0116] An example picture parameter set RBSP syntax is as follows:

[0117] [Table 1]

[0118] Exemplary picture parameter set RBSP semantics are: num_level1_tile_columns_minus1+1 specifies the number of level 1 tile columns that partition the picture. num_level1_tile_columns_minus1 must be in the range of 0 to PicWidthInCtbsY-1, inclusive. When not present, the value of num_level1_tile_columns_minus1 is inferred to be equal to 0. num_level1_tile_rows_minus1+1 specifies the number of level 1 tile rows that partition the picture. num_level1_tile_rows_minus1 must be in the range of 0 to PicHeightInCtbsY-1, inclusive. When not present, the value of num_level1_tile_rows_minus1 is inferred to be equal to 0. The variable NumLevel1Tiles is set equal to (num_level1_tile_columns_minus1+1)*(num_level1_tile_rows_minus1+1). When single_tile_in_pic_flag is equal to 0, NumTilesInPic must be greater than 1. uniform_level1_tile_spacing_flag is set equal to 1 to specify that the level 1 tile column borders, and similarly the level 1 tile row borders, are uniformly distributed across the picture. uniform_level1_tile_spacing_flag is equal to 0 to specify that the level 1 tile column borders, and similarly the level 1 tile row borders, are not uniformly distributed across the picture but are explicitly signaled using the syntax elements level1_tile_column_width_minus1[i] and level1_tile_row_height_minus1[i]. When not present, the value of uniform_level1_tile_spacing_flag is inferred to be equal to 1. level1_tile_column_width_minus1[i]+1 specifies the width of the i-th level 1 tile column in CTB units.level1_tile_row_height_minus1[i]+1 specifies the height of the i-th tile level 1 row in CTB units. level2_tile_present_flag specifies that one or more level 1 tiles are split into more tiles. level2_tile_split_flag[i]+1 specifies that the i-th level 1 tile is split into two or more tiles. num_level2_tile_columns_minus1[i]+1 specifies the number of tile columns that partition the i-th tile. num_level2_tile_columns_minus1[i] must be in the range 0 to ColWidth[i] inclusive. When not present, the value of num_level2_tile_columns_minus1[i] is inferred to be equal to 0. num_level2_tile_rows_minus1[i]+1 specifies the number of tile rows that partition the i-th tile. num_level2_tile_rows_minus1[i] must be in the range 0 to RowHeight[i], inclusive. When not present, the value of num_level2_tile_rows_minus1[i] is inferred to be equal to 0.

[0119] The CTB raster and tile scan conversion process is invoked by passing the following variables: a list ColWidth[i] for i ranging from 0 to num_level1_tile_columns_minus1 inclusive that specifies the width of the ith level 1 tile column in CTB units; a list RowHeight[j] for j ranging from 0 to num_level1_tile_rows_minus1 inclusive that specifies the height of the jth level 1 tile row in CTB units; a variable NumTilesInPic that specifies the number of tiles in the picture by reference to the PPS; a list TileWidth[i] for i ranging from 0 to NumTilesInPic inclusive that specifies the width of the ith tile in CTB units; a list TileHeight[i] for i ranging from 0 to NumTilesInPic inclusive that specifies the height of the ith tile in CTB units; a list TileColBd[i] for j ranging from 0 to NumTilesInPic, inclusive, that specifies the location of the i-th tile row boundary in CTB units; a list TileRowBd[i] for j ranging from 0 to NumTilesInPic, inclusive, that specifies the conversion from CTB addresses in the CTB raster scan of the picture to CTB addresses in the tile scan; a list CtbAddrRsToTs[ctbAddrRs] for ctbAddrRs ranging from 0 to PicSizeInCtbsY-1, inclusive, that specifies the conversion from CTB addresses in the tile scan to CTB addresses in the tile scan A list CtbAddrTsToRs[ctbAddrTs] for ctbAddrTs ranging from 0 to PicSizeInCtbsY-1 inclusive, specifying the conversion to CTB addresses for the CTB raster scan of the texture; a list TileId[ctbAddrTs] for ctbAddrTs ranging from 0 to PicSizeInCtbsY-1 inclusive, specifying the conversion from CTB addresses to tile IDs for the tile scan; a conversion from tile index to the number of CTUs in the tile;A list NumCtusInTile[tileIdx] for tileIdx ranging from 0 to PicSizeInCtbsY-1 inclusive, and a list FirstCtbAddrTs[tileIdx] for tileIdx ranging from 0 to NumTilesInPic-1 inclusive, which specifies the conversion from the tile ID to the CTB address in the tile traversal of the first CTB in the tile, are derived.

[0120] Exemplary tile group header semantics are as follows: tile_group_address specifies the tile address of the first tile in the tile group, where the tile address is equal to TileId[firstCtbAddrTs] as specified by Equation 6-8, where firstCtbAddrTs is the CTB address in the tile traversal of the CTB of the first CTU in the tile group. The length of tile_group_address is Ceil(Log2(NumTilesInPic)) bits. The value of tile_group_address must be in the range of 0 to NumTilesInPic-1, inclusive, and the value of tile_group_address must not be equal to the value of tile_group_address of any other coded tile group NAL unit of the same coded picture. When tile_group_address is not present, it is inferred to be equal to 0.

[0121] The following is a second specific exemplary embodiment of the above aspect. An exemplary CTB raster and tile traversal process is as follows: A variable NumTilesInPic, which specifies the number of tiles in the picture with reference to the PPS, and lists TileColBd[i], TileRowBd[i], TileWidth[i], and TileHeight[i], for i ranging from 0 to NumTilesInPic-1 inclusive, which specify the location of the i-th tile column boundary in units of CTBs, the location of the i-th tile row boundary in units of CTBs, the width of the i-th tile column in units of CTBs, and the height of the i-th tile column in units of CTBs, are derived as follows: for(tileIdx=0,i=0;i <NumLevel1Tiles;i++){ tileX=i%(num_level1_tile_columns_minus1+1) tileY=i / (num_level1_tile_columns_minus1+1) if(!level2_tile_split_flag[i]){ (6-5) TileColBd[tileIdx]=colBd[tileX] TileRowBd[tileIdx]=rowBd[tileY] TileWidth[tileIdx]=ColWidth[tileX] TileHeight[tileIdx]=RowHeight[tileY] tileIdx++ }else{ if(uniform_level2_tile_spacing_flag[i]){ for(k=0;k<=num_level2_tile_columns_minus1[i];k++) colWidth2[k]=((k+1)*ColWidth[tileX]) / (num_level2_tile_columns_minus1[i]+1)-(k*ColWidth[tileX]) / (num_level2_tile_columns_minus1[i]+1) for(k=0;k<=num_level2_tile_rows_minus1[i];k++) rowHeight2[k]=((k+1)*RowHeight[tileY]) / (num_level2_tile_rows_minus1[i]+1)-(k*RowHeight[tileY]) / (num_level2_tile_rows_minus1[i]+1) }else{ colWidth2[num_level2_tile_columns_minus1[i]]=ColWidth[tileX]) for(k=0;k<=num_level2_tile_columns_minus1[i];k++){ colWidth2[k]=tile_level2_column_width_minus1[k]+1 colWidth2[k] -= colWidth2[k] } rowHeight2[num_level2_tile_rows_minus1[i]]=RowHeight[tileY]) for(k=0;k<=num_level2_tile_rows_minus1[i];k++){ rowHeigh2[k]=tile_level2_column_width_minus1[k]+1 rowHeight2[k] -= rowHeight2[k] } } for(colBd2[0]=0,k=0;k<=num_level2_tile_columns_minus1[i];k++) colBd2[k+1]=colBd2[k]+colWidth2[k] for(rowBd2[0]=0,k=0;k<=num_level2_tile_rows_minus1[i];k++) rowBd2[k+1]=rowBd2[k]+rowHeight2[k] numSplitTiles=(num_level2_tile_columns_minus1[i]+1)*(num_level2_tile_rows_minus1[i]+1) for(k=0;k <numSplitTiles;k++){ tileX2=k%(num_level2_tile_columns_minus1[i]+1) tileY2=k / (num_level2_tile_columns_minus1[i]+1) TileColBd[tileIdx]=colBd[tileX]+colBd2[tileX2] TileRowBd[tileIdx]=rowBd[tileY]+rowBd2[tileY2] TileWidth[tileIdx]=colWidth2[tileX2] TileHeight[tileIdx]=rowHeight2[tileY2] tileIdx++ } } } NumTilesInPic=tileIdx

[0122] An example picture parameter set RBSP syntax is as follows:

[0123] [Table 2]

[0124] Exemplary picture parameter set RBSP semantics are as follows: uniform_level2_tile_spacing_flag[i] is set equal to 1 to specify that the level 2 tile column boundary of the ith level 1 tile and similarly the level 2 tile row boundary of the ith level 1 tile are uniformly distributed across the picture. uniform_level2_tile_spacing_flag[i] may be set equal to 0 to specify that the level 2 tile column boundary of the ith level 1 tile and similarly the level 2 tile row boundary of the ith level 1 tile are not uniformly distributed across the picture but are explicitly signaled using the syntax elements level2_tile_column_width_minus1[j] and level2_tile_row_height_minus1[j]. When not present, the value of uniform_level2_tile_spacing_flag[i] is inferred to be equal to 1. level2_tile_column_width_minus1[j]+1 specifies the width of the jth level 2 tile column of the ith level 1 tile in CTB units. level2_tile_row_height_minus1[j]+1 specifies the height of the jth tile level 2 row of the ith level 1 tile in CTB units.

[0125] The following is a third specific exemplary embodiment of the above aspect: An exemplary picture parameter set RBSP syntax is as follows:

[0126] [Table 3]

[0127] An example picture parameter set RBSP semantic is as follows: Bitstream adaptation may require that the following constraints apply: The value MinTileWidth specifies the minimum tile width and shall be equal to 256 luma samples. The value MinTileHeight specifies the minimum tile height and shall be equal to 64 luma samples. The values ​​of minimum tile width and minimum tile height may vary according to the profile and level definition. The variable Level1TilesMayBeFurtherSplit may be derived as follows: Level1TilesMayBeFurtherSplit=0 for(i=0,!Level1TilesMayBeFurtherSplit && i=0;i <NumLevel1Tiles;i++) if((ColWidth[i]*CtbSizeY>=(2*MinTileWidth))||(RowHeight[i]*CtbSizeY>=(2*MinTileHeight))) Level1TilesMayBeFurtherSplit=1

[0128] level2_tile_present_flag specifies that one or more level tiles are split into more tiles. When not present, the value of level2_tile_present_flag is inferred to be equal to 0. level2_tile_split_flag[i]+1 specifies that the i-th level 1 tile is split into two or more tiles. When not present, the value of level2_tile_split_flag[i] is inferred to be equal to 0.

[0129] The following is a fourth specific exemplary embodiment of the above aspect. Each tile location and each tile size may be signaled. A syntax for supporting such tile structure signaling may be as tabulated below: tile_top_left_address[i] and tile_bottom_right_address[i] are CTU indices in the picture, indicating the rectangular area covered by the tile. The number of bits for signaling these syntax elements should be sufficient to represent the maximum number of CTUs in the picture.

[0130] [Table 4]

[0131] Each tile location and each tile size may be signaled. A syntax to support such tile structure signaling may be as tabulated below: tile_top_left_address[i] is the CTU index of the first CTU in the tile in the order of CTU raster scan of the picture. Tile width and tile height specify the size of the tile. By signaling the tile size common unit first, some bits may be saved when signaling these two syntax elements.

[0132] [Table 5]

[0133] Alternatively, the signaling could be as follows:

[0134] [Table 6]

[0135] In another example, each tile size may be signaled as follows: To signal a flexible tile structure, the location of each tile may not be signaled. Instead, a flag may be signaled to specify whether the tile should be placed immediately to the right or immediately below the previous tile. If the tile can only be to the right or below the current tile, this flag may not be present.

[0136] The values ​​of tile_x_offset[i] and tile_y_offset[i] may be derived by the following ordered steps: tile_x_offset[0] and tile_y_offset[0] are set equal to 0. maxWidth is set equal to tile_width[0], and maxHeight is set equal to tile_height[0]. runningWidth is set equal to tile_width[0] and runningHeight is set equal to tile_height[0]. lastNewRowHeight is set equal to 0. TilePositionCannotBeInferred=false. For i>0 the following applies: Let the value isRight be set to: If runningWidth+tile_width[i]<=PictureWidth, then isRight==1, Otherwise, isRight==0. Let the value isBelow be set to: If runningHeight+tile_height[i]<=PictureHeight, then isBelow==1, Otherwise, isBelow==0. If isRight==1 && isBelow==1, then TilePositionCannotBeInferred=true. If isRight==1 && isBelow==0, then: right_tile_flag[i]=1, tile_x_offset[i]=runningWidth, tile_y_offset[i]=(runningWidth==maxWidth) ? 0 : lastNewRowHeight, lastNewRowHeight=(runningWidth==maxWidth) ? 0 : lastNewRowHeight is applied, Otherwise, if isRight==0 && isBelow==1, then: right_tile_flag[i]=0, tile_y_offset[i]=runningHeight, tile_x_offset[i]=(runningHeight==maxHeight) ? 0 : tile_x_offset[i-1], lastNewRowHeight=(runningHeight==maxHeight && runningWidth==maxWidth) ? runningHeight : lastNewRowHeight is applied, Otherwise, if isRight==1 && isBelow==1 && right_tile_flag[i]==1, then: tile_x_offset[i]=runningWidth, tile_y_offset[i]=(runningWidth==maxWidth) ? 0 : lastNewRowHeight, lastNewRowHeight=(runningWidth==maxWidth) ? 0 : lastNewRowHeight is applied, Otherwise (i.e., isRight==1 && isBelow==1 && right_tile_flag[i]==0), tile_y_offset[i]=runningHeight, tile_x_offset[i]=(runningHeight==maxHeight) ? 0 : tile_x_offset[i-1], lastNewRowHeight=(runningHeight==maxHeight && runningWidth==maxWidth) ? runningHeight : lastNewRowHeight is applied, If right_tile_flag[i]==1, then: runningWidth = runningWidth + tile_width[i] is applied, If runningWidth>maxWidth, set maxWidth equal to runningWidth, runningHeight is equal to tile_y_offset[i] + tile_height[i], Otherwise (i.e. right_tile_flag[i]==0), runningHeight = runningHeight + tile_height[i] is applied, If runningHeight>maxHeight, set maxHeight equal to runningHeight, runningWidth is equal to tile_x_offset[i] + tile_width[i].

[0137] The above can be written in pseudocode as follows: tile_x_offset[0]=0 tile_y_offset[0]=0 maxWidth=tile_width[0] maxHeight=tile_height[0] runningWidth=tile_width[0] runningHeight=tile_height[0] lastNewRowHeight=0 isRight=false isBelow=false TilePositionCannotBeInferred=false for(i=1;i <num_tiles_minus2+2;i++){ TilePositionCannotBeInferred=false isRight=(runningWidth+tile_width[i]<=PictureWidth) ? true : false isbelow=(runningHeight+tile_height[i]<=PictureHeight) ? true : false if(!isRight && !isBelow) / / Error, this case should never occur. if(isRight && isBelow) TilePositionCannotBeInferred=true if(isRight && !isBelow) { right_tile_flag[i]=true tile_x_offst[i]=runningWidth tile_y_offset[i]=(runningWidth==maxWidth) ? 0 : lastNewRowHeight lastNewRowHeight=tile_y_offset[i] } else if(!isRight && isBelow){ right_tile_flag[i]=false tile_y_offset[i]=runningHeight tile_x_offset[i]=(runningHeight==maxHeight) ? 0 : tile_x_offset[i-1] lastNewRowHeight=(runningHeight==maxHeight && runningWidth==maxWidth) ? runningHeight : lastNewRowHeight } else if(right_tile_flag[i]){ tile_x_offst[i]=runningWidth tile_y_offset[i]=(runningWidth==maxWidth) ? 0 : lastNewRowHeight lastNewRowHeight=tile_y_offset[i] } else{ tile_y_offset[i]=runningHeight tile_x_offset[i]=(runningHeight==maxHeight) ? 0 : tile_x_offset[i-1] lastNewRowHeight=(runningHeight==maxHeight && runningWidth==maxWidth) ? runningHeight : lastNewRowHeight } } if(right_tile_flag[i]){ runningWidth += tile_width[i] if(runningWidth>maxWidth)maxWidth=runningWidth runningHeight=tile_y_offset[i]+tile_height[i] } else{ runningHeight += tile_height[i] if(runningHeight>maxHeight)maxHeight=runningHeight runningWidth=tile_x_offset[i]+tile_width[i] }

[0138] [Table 7]

[0139] Below is one implementation in pseudocode that derives the size of the final tile. tile_x_offset[0]=0 tile_y_offset[0]=0 maxWidth=tile_width[0] maxHeight=tile_height[0] runningWidth=tile_width[0] runningHeight=tile_height[0] lastNewRowHeight=0 isRight=false isBelow=false TilePositionCannotBeInferred=false for(i=1;i <num_tiles_minus2+2;i++){ currentTileWidth=(i==num_tiles_minus2+1) ? (PictureWidth-runningWidth)%PictureWidth : tile_width[i] currentTileHeight=(i==num_tiles_minus2+1) ? (PictureHeight-runningHeight)%PictureHeight : tile_Height[i] isRight = (runningWidth + currentTileWidth <= PictureWidth)? true : false isbelow = (runningHeight + currentTileHeight <= PictureHeight)? true : false if (!isRight &&!isBelow) / / Error. This case should not occur. if (isRight && isBelow) TilePositionCannotBeInferred = true if (isRight &&!isBelow){ right_tile_flag[i] = true tile_x_offst[i] = runningWidth tile_y_offset[i] = (runningWidth == maxWidth)? 0 : lastNewRowHeight lastNewRowHeight = tile_y_offset[i] } else if (!isRight && isBelow){ right_tile_flag[i] = false tile_y_offset[i] = runningHeight tile_x_offset[i] = (runningHeight == maxHeight)? 0 : tile_x_offset[i - 1] lastNewRowHeight = (runningHeight == maxHeight && runningWidth == maxWidth)? runningHeight : lastNewRowHeight } else if (right_tile_flag[i]){ tile_x_offst[i] = runningWidth tile_y_offset[i]=(runningWidth==maxWidth)? 0 : lastNewRowHeight lastNewRowHeight=tile_y_offset[i] } else{ tile_y_offset[i]=runningHeight tile_x_offset[i]=(runningHeight==maxHeight)? 0 : tile_x_offset[i-1] lastNewRowHeight=(runningHeight==maxHeight && runningWidth==maxWidth)? runningHeight : lastNewRowHeight } } if(right_tile_flag[i]){ runningWidth += currentTileWidth if(runningWidth>maxWidth)maxWidth=runningWidth runningHeight=tile_y_offset[i]+currentTileHeight } else{ runningHeight += currentTileHeight if(runningHeight>maxHeight)maxHeight=runningHeight runningWidth=tile_x_offset[i]+currentTileWidth }

[0140]

Table 8

[0141] For further signaling bit savings, the number of unique tile sizes may be signaled to support tabulation of unit tile sizes. The tile sizes may then be referenced by index only.

[0142] [Table 9]

[0143] FIG. 9 is a schematic diagram of an example video coding device 900. The video coding device 900 is suitable for implementing the disclosed examples / embodiments as described herein. The video coding device 900 comprises a downstream port 920, an upstream port 950, and / or a transceiver unit (Tx / Rx) 910 including a transmitter and / or a receiver for communicating data upstream and / or downstream over a network. The video coding device 900 also includes a processor 930 including a logic unit and / or a central processing unit (CPU) for processing data, and a memory 932 for storing data. The video coding device 900 may also comprise electrical components, optical-to-electrical (OE) components, electrical-to-optical (EO) components, and / or wireless communication components connected to the upstream port 950 and / or the downstream port 920 for communication of data over an electrical communication network, an optical communication network, or a wireless communication network. Video coding device 900 may also include input and / or output (I / O) devices 960 for communicating data to and from a user. I / O devices 960 may include output devices, such as a display for displaying video data, speakers for outputting audio data, etc. I / O devices 960 may also include input devices, such as a keyboard, mouse, trackball, etc., and / or corresponding interfaces for interacting with such output devices.

[0144] The processor 930 is implemented by hardware and software. The processor 930 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 930 is in communication with downstream ports 920, Tx / Rx 910, upstream ports 950, and memory 932. The processor 930 comprises a coding module 914. The coding module 914 implements the disclosed embodiments described herein, such as methods 100, 1000, and 1100, mechanism 600, and / or application 700, which may employ bitstream 500 and / or images partitioned according to flexible video tiling scheme 800. The coding module 914 may also implement any other method / mechanism described herein. Additionally, the coding module 914 may implement codec system 200, encoder 300, and / or decoder 400. For example, the coding module 914 can partition a picture into first level tiles and partition the first level tiles into second level tiles. The coding module 914 can also assign such tiles to rectangular tile groups to support separate sub-picture extraction. The coding module 914 can also signal supporting data to indicate the configuration of the first level tiles and the second level tiles. The coding module 914 further supports employing such mechanisms to combine sub-pictures at different resolutions into a single picture for various use cases as described herein. Thus, the coding module 914 improves the functionality of the video coding device 900 and addresses problems specific to video coding techniques. Additionally, the coding module 914 effects the transformation of the video coding device 900 into different states.Alternatively, the coding module 914 may be implemented as instructions stored in memory 932 and executed by the processor 930 (eg, as a computer program product stored on a non-transitory medium).

[0145] Memory 932 comprises one or more memory types such as a disk, a tape drive, a solid state drive, a read only memory (ROM), a random access memory (RAM), a flash memory, a ternary content addressable memory (TCAM), a static random access memory (SRAM), etc. Memory 932 may be used as an overflow data storage device for storing such programs when such programs are selected for execution and for storing instructions and data read during program execution.

[0146] 10 is a flowchart of an example method 1000 of encoding an image by employing a flexible tiling scheme, such as flexible tiling scheme 800. Method 1000 may be employed by an encoder, such as codec system 200, encoder 300, and / or video coding device 900, when executing method 100, mechanism 600, and / or supporting application 700. Additionally, method 1000 may be employed to generate bitstream 500 for transmission to a decoder, such as decoder 400.

[0147] Method 1000 may begin when an encoder receives a video sequence including multiple images and determines, for example, based on user input, that the video sequence should be encoded into a bitstream. As an example, the video sequence, and therefore the images, may be encoded at multiple resolutions. In step 1001, a picture is partitioned into multiple first level tiles. In step 1003, a subset of the first level tiles is partitioned into multiple second level tiles. Each second level tile may include a single rectangular slice of picture data. In some examples, the first level tiles outside the subset include picture data at the first resolution and the second level tiles include picture data at a second resolution different from the first resolution. In some examples, each first level tile in the subset of the first level tiles includes two or more complete second level tiles.

[0148] In step 1005, the first level tiles and the second level tiles are assigned to one or more tile groups such that all tiles in the assigned tile groups, including the second level tiles, are constrained to cover a rectangular portion of the picture. As a particular example, when any first level tile is partitioned into multiple second level tiles (e.g., whenever flexible tiling is employed), each of the one or more tile groups may be constrained to cover a rectangular portion of the picture. Thus, with a raster scan order that traverses the picture horizontally from the left boundary to the right boundary, the first level tiles and the second level tiles are not assigned to one or more tile groups. Furthermore, covering the rectangular portion of the picture may include covering less than a full horizontal portion of the picture.

[0149] In step 1007, the first level tiles and the second level tiles are encoded in the bitstream. In some examples, data indicating the configuration of the second level tiles may be encoded in the bitstream in a picture parameter set associated with the picture. The configuration of the second level tiles may be signaled (e.g., in the PPS) as a number of second level tile columns and a number of second level tile rows for the partitioned first level tiles. In some examples, when the first level tile width is less than twice the minimum width threshold and the first level tile height is less than twice the minimum height threshold, data explicitly indicating the number of second level tile columns and the number of second level tile rows is omitted from the bitstream. In some examples, a split indication indicating the first level tiles to be partitioned into second level tiles may be encoded in the bitstream (e.g., in the PPS). In some examples, when the first level tile width is less than the minimum width threshold and the first level tile height is less than the minimum height threshold, the split indication may be omitted from the bitstream for the corresponding first level tile.

[0150] The bitstream may be stored in a memory for communication to a decoder in step 1009. The bitstream may be transmitted to the decoder upon request.

[0151] 11 is a flowchart of an example method 1100 of decoding an image by employing a flexible tiling scheme, such as flexible video tiling scheme 800. Method 1100 may be employed by a decoder, such as codec system 200, decoder 400, and / or video coding device 900, when executing method 100, mechanism 600, and / or supporting application 700. Additionally, method 1100 may be employed upon receiving bitstream 500 from an encoder, such as encoder 300.

[0152] Method 1100 may begin when a decoder begins receiving a bitstream of coded data representing a video sequence, e.g., as a result of method 1000. The bitstream may include video data from a video sequence coded at multiple resolutions. In step 1101, the bitstream is received. The bitstream includes a picture partitioned into multiple first level tiles. A subset of the first level tiles is further partitioned into multiple second level tiles. Each second level tile may include a single rectangular slice of picture data. In some examples, the first level tiles outside the subset include picture data at the first resolution and the second level tiles include picture data at a second resolution different from the first resolution. In some examples, each first level tile in the subset of first level tiles includes two or more complete second level tiles. The first level tiles and the second level tiles may be assigned to one or more tile groups such that all tiles in an assigned tile group that includes a second level tile are constrained to cover a rectangular portion of the picture. For example, when any first level tile is partitioned into multiple second level tiles (e.g., flexible tiling is employed), one or more tile groups may each be constrained to cover a rectangular portion of the picture. Furthermore, due to a raster scan order that traverses the picture horizontally from the left boundary to the right boundary, the first level tiles and the second level tiles may not be assigned to one or more tile groups. In some cases, covering a rectangular portion of the picture includes covering less than a full horizontal portion of the picture.

[0153] In step 1103, a configuration of first level tiles and a configuration of second level tiles are determined based on one or more tile groups. For example, data indicating a configuration of the second level tiles may be obtained from a picture parameter set associated with the picture. In some examples, the configuration of the second level tiles is obtained from data indicating a number of second level tile columns and a number of second level tile rows for the partitioned first level tiles. In some examples, when the first level tile width is less than twice the minimum width threshold and the first level tile height is less than twice the minimum height threshold, data explicitly indicating the number of second level tile columns and the number of second level tile rows is omitted from the bitstream for the corresponding tile. In some examples, as part of determining the configuration of the first level tiles and the second level tiles, a partitioning indication may be obtained from the bitstream. The partitioning indication may indicate the first level tiles to be partitioned into second level tiles. In some examples, when a first level tile width is less than a minimum width threshold and a first level tile height is less than a minimum height threshold, splitting instruction data that explicitly indicates whether a first level tile is partitioned into a second level tile is excluded from the bitstream.

[0154] In step 1105, the first level tiles and the second level tiles are decoded based on the first level tile configuration and the second level tile configuration. In step 1107, a reconstructed video sequence is generated for display based on the decoded first level tiles and second level tiles.

[0155] 12 is a schematic diagram of an example system 1200 for coding a video sequence by employing a flexible tiling scheme, such as flexible video tiling scheme 800. System 1200 may be implemented by an encoder and a decoder, such as codec system 200, encoder 300, decoder 400, and / or video coding device 900. Additionally, system 1200 may be employed when implementing methods 100, 1000, 1100, mechanism 600, and / or application 700. System 1200 may also encode data into a bitstream, such as bitstream 500, and decode such bitstream for display to a user.

[0156] The system 1200 includes a video encoder 1202. The video encoder 1202 comprises a partition module 1201 for partitioning a picture into a plurality of first level tiles and partitioning a subset of the first level tiles into a plurality of second level tiles. The video encoder 1202 further comprises an allocation module 1203 for assigning the first level tiles and the second level tiles to one or more tile groups such that all tiles in the assigned tile group, including the second level tiles, are constrained to cover a rectangular portion of the picture. The video encoder 1202 further comprises an encoding module 1205 for encoding the first level tiles and the second level tiles into a bitstream. The video encoder 1202 further comprises a storage module 1207 for storing the bitstream for communication towards a decoder. The video encoder 1202 further comprises a transmission module 1209 for transmitting the bitstream towards the decoder. The video encoder 1202 may be further configured to perform any of the steps of the method 1000.

[0157] The system 1200 also includes a video decoder 1210. The video decoder 1210 comprises a receiving module 1211 for receiving a bitstream including a picture partitioned into a plurality of first level tiles, a subset of the first level tiles being further partitioned into a plurality of second level tiles, and the first level tiles and the second level tiles being assigned to one or more tile groups such that all tiles in the assigned tile group including the second level tiles are constrained to cover a rectangular portion of the picture. The video decoder 1210 further comprises a determining module 1213 for determining a configuration of the first level tiles and a configuration of the second level tiles based on the one or more tile groups. The video decoder 1210 further comprises a decoding module 1215 for decoding the first level tiles and the second level tiles based on the configuration of the first level tiles and the configuration of the second level tiles. The video decoder 1210 further comprises a generating module 1217 for generating a reconstructed video sequence for display based on the decoded first level tiles and the second level tiles. The video decoder 1210 may be further configured to perform any of the steps of the method 1100.

[0158] A first component is directly coupled to a second component when there are no intervening components between the first and second components, other than a line, trace, or another medium. A first component is indirectly coupled to a second component when there are intervening components between the first and second components, other than a line, trace, or another medium. The term "coupled" and variations thereof include both directly coupled and indirectly coupled. Use of the term "about" refers to a range that includes ±10% of the succeeding number unless otherwise specified.

[0159] It should also be understood that the steps of the exemplary methods described herein need not necessarily be performed in the order described, and the order of steps of such methods should be understood to be merely exemplary. Similarly, additional steps may be included in such methods, and certain steps may be excluded or combined in methods consistent with various embodiments of the present disclosure.

[0160] Although several embodiments are provided in this disclosure, it can be understood that the disclosed systems and methods can be embodied in many other specific forms without departing from the spirit or scope of the disclosure. The examples should be considered as illustrative rather than restrictive, and the intention should not be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or certain features may be excluded or not implemented.

[0161] Additionally, the techniques, systems, subsystems, and methods described and illustrated in various embodiments as separate or distinct may be combined or integrated with other systems, components, techniques, or methods without departing from the scope of the present disclosure. Other examples of changes, substitutions, and alterations will be ascertainable by one of ordinary skill in the art and may be made without departing from the spirit and scope of the present disclosure. [Explanation of symbols]

[0162] 100 How it works 200 Coding and Decoding (Codec) Systems 201 Segmented Video Signal 211 General purpose coder control components 213 Transform Scaling and Quantization Components 215 Intra-picture Estimation Components 217 Intra-picture prediction components 219 Motion Compensation Components 221 Motion Estimation Components 223 Decoded Picture Buffer Components 225 In-loop filter components 227 Filter Control Analysis Components 229 Scaling and Inverse Transformation Components 231 Header Formatting and Context-Adaptive Binary Arithmetic Coding (CABAC) Components 300 Video Encoder 301 Segmented Video Signal 313 Transformation and Quantization Components 317 Intra-picture prediction components 321 Motion Compensation Components 323 Decoded Picture Buffer Components 325 In-loop filter components 329 Inverse Transformation and Quantization Components 331 Entropy Coding Components 400 Video Decoder 417 Intra-picture prediction components 421 Motion Compensation Components 423 Decoded Picture Buffer Components 425 In-loop filter components 429 Inverse Transformation and Quantization Components 433 Entropy Decoding Components 500 bitstream 510 Sequence Parameter Set (SPS) 512 Picture Parameter Set (PPS) 514 Tile Group Header 520 Image data 523 tiles 600 Mechanism 601,603,605,607 tiles 610 Extractor Truck 611 First Resolution 612 Second Resolution 700 Videoconferencing Applications 701 Talking Participants 703 other participants 800 Flexible Video Tiling Method 801 1st level tiles 803 Second Level Tiles 811,812,813 Tile Group 900 Video Coding Device 910 Transceiver unit (Tx / Rx) 914 Coding Module 920 downstream ports 930 Processor 932 Memory 950 Upstream Ports 960 Input and / or Output (I / O) Devices 1000 How it works 1100 How it works 1200 System 1201 Division Module 1202 Video Encoder 1203 Allocation Module 1205 Encoding Module 1207 Memory Module 1209 Transmitter module, transmitter 1210 Video Decoder 1211 Receiving module, receiver 1213 Decision Module 1215 Decryption Module 1217 Generation Module

Claims

1. A method for transmitting an encoded bitstream, the bitstream including a picture partitioned into a plurality of first level tiles, a subset of the first level tiles being further partitioned into a plurality of second level tiles, the second level tiles being assigned to one or more tile groups defined by corresponding tile group addresses, and all second level tiles created from a single first level tile being assigned to the same tile group such that all tiles in an assigned tile group are constrained to cover a rectangular area of ​​the picture, transmitting the encoded bitstream.

2. 2. The method of claim 1, wherein the first level tiles excluding the subset contain picture data at a first resolution and the second level tiles contain picture data at a second resolution different from the first resolution.

3. 3. The method of claim 1 or 2, wherein when any first level tile is partitioned into multiple second level tiles, each of the one or more tile groups is constrained to cover a rectangular region of the picture.

4. A method according to any one of claims 1 to 3, wherein the second level tiles are not assigned to the one or more tile groups in a raster scan order that traverses the picture horizontally from the left boundary to the right boundary to avoid the one or more tile groups resulting in a non-rectangular shape.

5. 5. A method according to claim 1, wherein each second level tile comprises a single slice of picture data from the picture.

6. A method according to any one of claims 1 to 5, wherein second level tile rows and second level tile columns for partitioned first level tiles are included in a picture parameter set associated with the picture.

7. 1. An encoder comprising: A memory storing instructions; one or more processors coupled to said memory; Equipped with The one or more processors execute the instructions to cause the encoder to: Partitioning a picture into a plurality of first level tiles; partitioning the subset of first level tiles into a plurality of second level tiles; assigning the second level tiles to one or more tile groups defined by corresponding tile group addresses, where all second level tiles created from a single first level tile are assigned to the same tile group such that all tiles in an assigned tile group are constrained to cover a rectangular area of ​​the picture; encoding the second level tiles into a bitstream according to the one or more tile groups; The encoder is configured to:

8. The one or more processors execute the instructions to cause the encoder to: storing said bitstream for communication to a decoder. The encoder of claim 7 , further configured to:

9. 9. An encoder as claimed in claim 7 or 8, wherein the first level tiles excluding the subset contain picture data at a first resolution and the second level tiles contain picture data at a second resolution different from the first resolution.

10. 10. An encoder as claimed in claim 7, wherein when any first level tile is partitioned into multiple second level tiles, each of the one or more tile groups is constrained to cover a rectangular region of the picture.

11. An encoder as described in any one of claims 7 to 10, wherein the second level tiles are not assigned to the one or more tile groups in a raster scan order that traverses the picture horizontally from the left boundary to the right boundary to avoid the one or more tile groups resulting in a non-rectangular shape.

12. 12. An encoder as claimed in any one of claims 7 to 11, wherein each second level tile comprises a single slice of picture data from the picture.

13. The one or more processors execute the instructions to cause the encoder to: encoding second level tile rows and second level tile columns for the partitioned first level tiles, the second level tile rows and the second level tile columns being encoded in a picture parameter set associated with the picture; 13. The encoder of claim 7, further configured to:

14. A decoder comprising: A memory storing instructions; one or more processors coupled to said memory; Equipped with The one or more processors execute the instructions to cause the decoder to: receiving a bitstream including a picture partitioned into a plurality of first level tiles, a subset of the first level tiles being further partitioned into a plurality of second level tiles, the second level tiles being assigned to one or more tile groups defined by corresponding tile group addresses, and all second level tiles created from a single first level tile being assigned to the same tile group such that all tiles in an assigned tile group are constrained to cover a rectangular area of ​​the picture; determining a configuration of the second level tiles based on parameters in the bitstream; and decoding the second level tiles based on the configuration of the second level tiles and the one or more tile groups; generating a reconstructed video sequence for display based on the decoded second level tiles; and The decoder is configured to:

15. 15. The decoder of claim 14, wherein the first level tiles excluding the subset contain picture data at a first resolution and the second level tiles contain picture data at a second resolution different from the first resolution.

16. 16. A decoder as claimed in claim 14 or 15, wherein when any first level tile is partitioned into multiple second level tiles, each of the one or more tile groups is constrained to cover a rectangular region of the picture.

17. A decoder as described in any one of claims 14 to 16, wherein the second level tiles are not assigned to the one or more tile groups in a raster scan order that traverses the picture horizontally from the left boundary to the right boundary to avoid the one or more tile groups resulting in a non-rectangular shape.

18. 18. A decoder as claimed in any one of claims 14 to 17, wherein each second level tile contains a single slice of picture data from the picture.

19. The one or more processors execute the instructions to cause the decoder to: obtaining second level tile rows and second level tile columns for the partitioned first level tiles from a picture parameter set associated with the picture; 19. A decoder according to any one of claims 14 to 18, further configured to: