VIDEO ENCODER, VIDEO DECODER, AND CORRESPONDING METHODS - Patent application

A flag-based system for encoding tile information in video coding optimizes tile group structure signaling, addressing redundancy and adapting to bandwidth and latency demands, enhancing coding efficiency and compression ratios.

JP7790648B2Active Publication Date: 2025-12-23HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023177002
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-07-05
Filing Date
2023-10-12
Publication Date
2025-12-23
Estimated Expiration
2039-12-12

AI Technical Summary

Technical Problem

Existing video coding techniques face challenges in efficiently managing tile group structure signaling, leading to redundant information and conflicting requirements for parallel processing and MTU size adaptation, especially in applications with varying bandwidth and latency demands.

Method used

Implementing a flag system to specify whether tile information is encoded in a parameter set or tile group header, allowing for optimized signaling of tile group structure, reducing redundancy and improving efficiency in video coding.

Benefits of technology

Enhances video coding efficiency by minimizing redundant information and adapting to different bandwidth and latency requirements, facilitating improved compression ratios with minimal image quality sacrifice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007790648000004
    Figure 0007790648000004
  • Figure 0007790648000005
    Figure 0007790648000005
  • Figure 0007790648000006
    Figure 0007790648000006
Patent Text Reader

Abstract

To provide a method for decoding a video bit stream.SOLUTION: A bitstream includes coding data of at least one picture, and each picture includes at least one tile group. A method includes a step of parsing a flag that specifies whether tile information of a coded picture is present in a parameter set or in a tile group header. The tile information indicates which tiles of a picture are included in a tile group. The method parses the tile information from the parameter set or the tile group header based on the flag. The method obtains decoded data of the coded picture based on the tile information.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD This disclosure relates generally to video coding, and more particularly to tile group signaling in video coding. [Background technology]

[0002] The amount of video data required to render even a relatively short video can be substantial. This can pose challenges when data is streamed or otherwise communicated across communications networks with limited bandwidth capabilities. Therefore, video data is typically compressed before being communicated across today's telecommunications networks. When video is stored on a storage device, video size can also be an issue because memory resources may be limited. Video compression devices often use software and / or hardware to code video data at the source before transmission or storage, thereby reducing the amount of data needed to represent a digital video image. The compressed data is then received at the destination by a video decompressor, which decodes the video data. With limited network resources and ever-increasing demands for higher video quality, improved compression and decompression techniques that increase compression ratios with little or no sacrifice in image quality are desirable. Summary of the Invention

[0003] A first aspect of the present disclosure is a method implemented in an encoder, the method comprising: encoding, by a processor of the encoder, a flag specifying whether tile information for a coding picture is present in a parameter set or in a tile group header, the tile information indicating which tiles of the picture are included in a tile group; encoding, by the processor, the tile information only within the parameter set in response to determining that the flag specifies that the tile information of a coding picture is to be encoded within the parameter set; encoding, by the processor, the tile information only within the tile group header in response to determining that the flag specifies that the tile information of a coding picture is to be encoded within the tile group header; encoding, by the processor, the picture within the video bitstream based on the tile information; transmitting the video bitstream along a network towards a decoder; This aspect provides a mechanism to improve tile group structure signaling and reduce redundant information.

[0004] Optionally, in a first aspect, the step of encoding the tile information only within the parameter set by the processor includes a step of encoding by the processor a tile identifier (ID) of the first tile of each tile group within the picture.

[0005] Optionally, in the first aspect, the step of encoding, by the processor, the tile information only in the parameter set comprises: parsing, by the processor, a second flag specifying whether the current tile group referencing the parameter set includes more than one tile; encoding, by the processor, a tile ID of a last tile of the current tile group in the picture in response to determining that the second flag specifies that the current tile group referencing the parameter set includes more than one tile; Further includes:

[0006] Optionally, in the first aspect, the step of encoding, by the processor, the tile information only in the parameter set comprises: parsing, by the processor, a second flag specifying whether the current tile group referencing the parameter set includes more than one tile; encoding, by the processor, in response to determining that the second flag specifies that the current tile group referencing the parameter set includes more than one tile, the number of tiles in the current tile group within the picture; Further includes:

[0007] Optionally, in the first aspect, the step of encoding, by the processor, the tile information only in the tile group header comprises: encoding, by the processor, a tile ID of a first tile of a tile group in the picture into a tile group header; determining, by the processor, whether the flag specifies that the tile information of the coding picture is coded in the tile group header and whether the second flag specifies that the current tile group referencing the parameter set includes more than one tile; In response to determining that the flag specifies that the tile information of the coding picture is coded in the tile group header and that the second flag specifies that the current tile group referencing the parameter set includes more than one tile, coding a tile ID of a last tile of the tile group in the picture in the tile group header picture; Includes:

[0008] A second aspect of the present disclosure is a method implemented in a decoder for decoding a video bitstream, the bitstream including coding data for at least one picture, each picture including at least one tile group, the method comprising: parsing, by the decoding processor, a flag specifying whether tile information for a coding picture is present in a parameter set or in a tile group header, the tile information indicating which tiles of the picture are included in a tile group; parsing, by the processor, the tile information from the parameter set in response to determining that the flag specifies that the tile information of a coding picture is encoded within the parameter set; parsing, by the processor, the tile information from the tile group header in response to determining that the flag specifies that the tile information of a coding picture is encoded within the tile group header; obtaining the decoded data of the coded picture based on the tile information; The method includes:

[0009] Optionally, in a second aspect, parsing the tile information in the parameter set by the processor includes decoding a tile identifier (ID) of a first tile of each tile group in the picture.

[0010] Optionally, in a second aspect, the step of parsing, by the processor, the tile information in the parameter set comprises: parsing, by the processor, a second flag specifying whether the current tile group referencing the parameter set includes more than one tile; in response to determining that the second flag specifies that the current tile group referencing the parameter set includes more than one tile, decoding a tile ID of a last tile of the current tile group within the picture; Further includes:

[0011] Optionally, in a second aspect, the step of parsing, by the processor, the tile information in the parameter set comprises: parsing, by the processor, a second flag specifying whether the current tile group referencing the parameter set includes more than one tile; decoding a number of tiles in the current tile group within the picture in response to determining that the second flag specifies that the current tile group referencing the parameter set includes more than one tile; Further includes:

[0012] Optionally, in a second aspect, the step of parsing the tile information in the tile group header by the processor comprises: decoding a tile ID of a first tile of a tile group in the picture in a tile group header; determining whether the flag specifies that the tile information of the coding picture is coded in the tile group header, and whether the second flag specifies that the current tile group referencing the parameter set includes more than one tile; in response to determining that the flag specifies that the tile information of the coding picture is coded in the tile group header and that the second flag specifies that the current tile group referencing the parameter set includes more than one tile, decoding, in the tile group header picture, a tile ID of a last tile of the tile group in the picture; Includes:

[0013] Optionally, in a second aspect, the step of parsing by the processor the flag specifying whether tile information for the coding picture is present in a parameter set or in a tile group header includes a step of inferring, in response to determining that the flag is not present in the parameter set, that the flag specifies that the tile information for the coding picture is present only in the tile group header.

[0014] Optionally, in any of the above aspects, the flag is called tile_group_info_in_pps_flag.

[0015] Optionally, in any of the above aspects, the second flag is called single_tile_per_tile_group_flag.

[0016] Optionally, in any of the preceding aspects, the parameter set is a picture parameter set.

[0017] Optionally, in any of the preceding aspects, the parameter set is a sequence parameter set.

[0018] Optionally, in any of the preceding aspects, the parameter set is a video parameter set.

[0019] A third aspect of the present disclosure is a video coding device, comprising: A video coding apparatus including a processor, a receiver connected to the processor, and a transmitter connected to the processor, the processor and the transmitter configured to perform the method of any of the above aspects.

[0020] A fourth aspect of the present disclosure includes a non-transitory computer-readable medium including a computer program product for use by a video coding device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor, cause the video coding device to perform a method of any of the preceding aspects.

[0021] For purposes of clarity, any one of the above-described embodiments may be combined with any one or more of the other above-described embodiments to create new embodiments that are within the scope of the present disclosure.

[0022] The above and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. [Brief explanation of the drawings]

[0023] For a more complete understanding of the present disclosure, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.

[0024] [Figure 1] 1 is a flowchart of an exemplary method for coding a video signal.

[0025] [Figure 2] 1 is a schematic diagram of an example coding and decoding (codec) system for video coding.

[0026] [Figure 3] FIG. 1 is a schematic diagram illustrating an exemplary video encoder.

[0027] [Figure 4] FIG. 1 is a schematic diagram illustrating an exemplary video decoder.

[0028] [Figure 5]FIG. 1 is a schematic diagram illustrating an exemplary bitstream containing an encoded video sequence.

[0029] [Figure 6] FIG. 1 is a schematic diagram illustrating a picture partitioned into exemplary tile groups.

[0030] [Figure 7] 1 is a schematic diagram of an exemplary video coding device;

[0031] [Figure 8] 1 is a flowchart of an exemplary method for encoding an image into a bitstream along with a flag indicating the location of tile group information.

[0032] [Figure 9] 1 is a flowchart of an exemplary method for decoding an image from a bitstream with a flag indicating the location of tile group information.

[0033] [Figure 10] 1 is a schematic diagram of an example system for coding a video sequence of images in a bitstream; DETAILED DESCRIPTION OF THE INVENTION

[0034] It should be understood at the outset that, although illustrative implementations of one or more embodiments apply below, the disclosed systems and / or methods may be implemented using any number of technologies, whether currently known or existing. The present disclosure should in no way be limited to the illustrative implementations, drawings, and technologies described below, including the exemplary designs and implementations shown and described herein, but may be modified within the scope of the appended claims, along with their full range of equivalents.

[0035] Many video compression techniques may be utilized to reduce the size of video files with minimal data loss. For example, video compression techniques may include performing spatial (e.g., intra-picture) prediction and / or temporal (e.g., inter-picture) prediction to reduce or remove data redundancy within a video sequence. In block-based video coding, a video slice (e.g., a video picture or portion of a video picture) may be partitioned into video blocks, which may also be referred to as tree blocks, coding blocks (CBs), coding tree blocks (CTBs), coding tree units (CTUs), coding units (CUs), and / or coding nodes. A coding block (CB) may be an M×N block of samples for some values ​​of M and N. Consequently, the division of a CTB into coding blocks is a partition. Correspondingly, a coding tree block (CTB) may be an N×N block of samples for some value of N. Consequently, the division of a component into CTBs is a partition. A coding tree unit (CTU) is a CTB of luma samples, two corresponding CTBs of chroma samples for a picture with a three-sample arrangement, or a CTB of samples for a monochrome picture or a picture coded using a syntax structure used to code three distinct color planes and samples. A coding unit (CU) is a coding block of luma samples, two corresponding coding blocks of chroma samples for a picture with a three-sample arrangement, or a coding block of samples for a monochrome picture or a picture coded using a syntax structure used to code three distinct color planes and samples.

[0036] Video blocks in intra-coded (I) slices of a picture are coded using spatial prediction with respect to reference samples in neighboring blocks within the same picture. Video blocks in inter-coded unidirectionally predicted (P) or bidirectionally predicted (B) slices of a picture may be coded using spatial prediction with respect to reference samples in neighboring blocks within the same picture or temporal prediction with respect to reference samples in other reference pictures. A picture may be referred to as a frame and / or an image, and a reference picture may be referred to as a reference frame and / or a reference image. Spatial or temporal prediction results in a predictive block that represents an image block. Residual data represents pixel differences between the original image block and the predictive block. Thus, inter-coded blocks are coded according to a motion vector that points to a block of reference samples that form the predictive block, and residual data that indicates the difference between the coding block and the predictive block. Intra-coded blocks are coded according to an intra-coding mode and the residual data. For further compression, the residual data may be transformed from the pixel domain to a transform domain, resulting in residual transform coefficients that may be quantized. The quantized transform coefficients may be initially arranged into a two-dimensional array. The quantized transform coefficients may be scanned to generate a one-dimensional vector of transform coefficients. Entropy coding may be applied to achieve even greater compression. Such video compression techniques are discussed in more detail below.

[0037] To ensure that the encoded video is decoded accurately, the video is encoded and decoded according to a corresponding video coding standard, including International Telecommunication Union (ITU) Standardization Sector (ITU-T) H.261, International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) Motion Picture Experts Group (MPEG)-1 Part 2, Advanced Video Coding (AVC), also known as ITU-T H.262 or ISO / IEC MPEG-2 Part 2, ITU-T H.263, ISO / IEC MPEG-4 Part 2, ITU-T H.264 or ISO / IEC MPEG-4 Part 10, and High Efficiency Video Coding (HEVC), also known as ITU-T H.265 or MPEG-H Part 2. AVC includes extensions such as Scalable Video Coding (SVC), Multiview Video Coding (MVC), and Multiview Video Coding plus Depth (MVC+D), as well as three dimensional (3D) AVC (3D-AVC). HEVC includes extensions such as Scalable HEVC (SHVC), Multiview HEVC (MV-HEVC), and 3D HEVC (3D-HEVC). The ITU-T and ISO / IEC joint video experts team (JVET) have begun developing a video coding standard called Versatile Video Coding (VVC).VVC is included in the Working Draft (WD) including JVET-L1001-v9.

[0038] To code a video image, the image is first partitioned, and the partitions are coded into a bitstream. Various picture partitioning schemes are available. For example, an image can be partitioned into normal slices, dependent slices, tiles, and / or according to Wavefront Parallel Processing (WPP). For simplicity, HEVC constrains the encoder to use only normal slices, dependent slices, tiles, WPP, and combinations thereof when partitioning slices into groups of CTBs for video coding. Such partitioning can be applied to support Maximum Transfer Unit (MTU) size adaptation, parallel processing, and reduced end-to-end delay. The MTU indicates the maximum amount of data that can be transmitted in a single packet. If a packet payload exceeds the MTU, the payload is split into two packets through a process called fragmentation.

[0039] A normal slice, also referred to simply as a slice, is a partitioned portion of an image that can be reconstructed independently of other normal slices within the same picture, despite any interdependencies due to loop filtering operations. A slice contains an integer number of CTUs ordered consecutively in a raster scan. Each slice is encapsulated within its own Network Abstraction Layer (NAL) unit for transmission. Furthermore, intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies across slice boundaries may be disabled to support independent reconstruction. Such independent reconstruction supports parallelization. For example, parallelization based on normal slices utilizes minimal inter-processor or inter-core communication. However, because each normal slice is independent, each slice is associated with a separate slice header. The use of normal slices can incur significant coding overhead due to the bit cost of a slice header per slice and the lack of prediction across slice boundaries. Furthermore, normal slices may be used to support compliance with MTU size requirements. Specifically, because regular slices are encapsulated in separate NAL units and can be coded independently, each regular slice should be smaller than the MTU in the MTU scheme to prevent the slices from being split into multiple packets. Thus, the objectives of parallelization and MTU size adaptation may impose conflicting requirements on the slice layout within a picture.

[0040] Dependent slices are similar to normal slices but have a shortened slice header, allowing partitioning on picture tree block boundaries without breaking intra-picture prediction. Dependent slices therefore allow normal slices to be fragmented into multiple NAL units, which results in reduced end-to-end delay by allowing parts of a normal slice to be sent before the encoding of the entire normal slice is complete.

[0041] A tile is a partitioned portion of an image / picture generated by horizontal and vertical boundaries that create tile columns and rows. A tile contains a rectangular area of ​​CTUs within a particular tile column and a particular tile row within a picture. Tiles may be coded in raster scan order (right to left and top to bottom). The scan order of CTBs is local within a tile. Thus, a CTB in a first tile is coded in raster scan order before proceeding to a CTB in the next tile. Similar to regular slices, tiles break intra-picture prediction dependencies as well as entropy decoding dependencies. However, tiles may not be included in individual NAL units, and therefore tiles may not be used for MTU size adaptation. Each tile can be processed by one processor / core, and inter-processor / inter-core communication utilized for inter-picture prediction between processing units decoding neighboring tiles may be limited to carrying a shared slice header (when adjacent tiles are in the same slice) and performing sharing of reconstructed samples and metadata related to loop filtering. When more than one tile is included in a slice, the entry point byte offset of each tile other than the first entry point offset in the slice may be signaled in the slice header.

[0042] A given coded video sequence cannot contain both tiles and wavefronts for most of the profiles specified in HEVC. For each slice and tile, one or both of the following conditions should be met: 1) all coding tree blocks in a slice belong to the same tile, and 2) all coding tree blocks in a tile belong to the same slice. A wavefront segment contains exactly one CTB row, and when WPP is used, if a slice starts in a CTB row, it must end in the same CTB row.

[0043] In WPP, a picture is partitioned into a single row of CTBs. The entropy decoding and prediction mechanisms may use data from CTBs in other rows. Parallel processing is enabled through parallel decoding of CTB rows. For example, the current row may be decoded in parallel with the previous row. However, the decoding of the current row is delayed from the decoding process of the previous row by two CTBs. This delay ensures that data related to the CTBs above and to the right of the current CTB in the current row are available before the current CTB is coded. This approach can be represented graphically as a wavefront. This staggered start allows parallelization by up to as many processors / cores as the picture contains CTB rows. Because intra-picture prediction between neighboring treeblock rows within a picture is allowed, inter-processor / inter-core communication to enable intra-picture prediction can be important. WPP partitioning takes into account the NAL unit size. Therefore, WPP does not support MTU size adaptation. However, regular slicing can be used in conjunction with WPP to implement the desired MTU size adaptation, with some coding overhead involved.

[0044] A tile may also include a motion constrained tile set (MCTS). A motion constrained tile set (MCTS) is a tile set designed such that associated motion vectors are restricted to point to full sample positions within the MCTS and to fractional sample positions that require only full sample positions within the MCTS for interpolation. Furthermore, the use of motion vector candidates for temporal motion vector prediction derived from blocks outside the MCTS is not permitted. In this way, each MCTS may be decoded independently, without the presence of tiles not included in the MCTS. HEVC specifies three MCTS-related supplemental enhancement information (SEI) messages: the temporal MCTS SEI message, the MCTS extraction information set SEI message, and the MCTS extraction information nest SEI message.

[0045] The temporal MCTS SEI message may be used to indicate the presence of an MCTS in a bitstream and to signal the MCTS. The MCTS SEI message provides supplemental information that can be used in MCTS sub-bitstream extraction (specified as part of the semantics of the SEI message) to generate a confirmation bitstream of the MCTS. The information includes the number of extraction information sets, each of which defines the number of MCTSs and contains raw bytes sequence payload (RBSP) bytes of the replacement video parameter set (VPS), sequence parameter set (SPS), and picture parameter set (PPS) to be used during the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, parameter sets (VPS, SPS, and PPS) may be rewritten or replaced and slice headers may be updated because one or all of the syntax elements related to slice addresses (including first_slice_segment_in_pic_flag and slice_segment_address) may use different values ​​in the extracted sub-bitstream.

[0046] As mentioned above, a tile group contains an integer number of tiles of a picture, either in a tile raster scan of the picture or in a rectangular grouping. The tiles within a tile group are contained exclusively in a single NAL unit. A tile group can replace slices in some examples. There are various tiling schemes (i.e., approaches to tile grouping) that may be utilized when partitioning a picture for further encoding. As a specific example, the tile grouping may be constrained so that tiles grouped together in a tile group form a rectangular region in the picture (referred to herein as a rectangular tile group). The tiles included in a tile group may be signaled by indicating the first and last tiles of the tile group. In such a case, the tile index of the first tile may be a smaller value than the tile index of the last tile.

[0047] There are two possibilities for signaling the tile group structure (e.g., the address / position of the tile group within a picture, if needed, and the number of tiles in the tile group). The first is to signal the tile group structure within a parameter set, for example, within the picture parameter set (PPS). The second is to signal the tile group structure within the header of each tile group. The two signaling possibilities must not be used simultaneously. Each possibility has its own advantages. For example, the first option is advantageous for applications like 360° video, where tile groups are usually coded as MCTS to allow transmitting only parts of each picture. In this case, the allocation of tiles to video coding layer (VCL) NAL units is usually a known encoding of the picture. The second option is advantageous in application scenarios where the allocation of tiles to VCL NAL units may need to depend on the actual size in bits of the tiles, e.g., in ultra-low latency applications like wireless display. When the tile group structure is signaled in a parameter set, some syntax elements in the tile group header may not be necessary and may therefore be removed or their presence may need to be adjusted.

[0048] In some applications, a picture may be encapsulated in several VCL NAL units, with each VCL NAL unit containing one tile group. In such applications, parallel processing may not be a primary goal / concern, since each tile group in this application may contain only one tile. An example may be a 360-degree video application with viewpoint-dependent delivery optimization. In such situations, the current signaling of the tile group structure, whether signaled in the parameter set or in the tile group header, may have some redundancy.

[0049] To solve the above-mentioned problems, various mechanisms for improving tile group structure signaling are disclosed herein. As described further, in embodiments, an encoder can encode a flag (e.g., a flag called single_tile_per_tile_group_flag) in a parameter set directly or indirectly referenced by a tile group to specify whether each of the tile groups that reference the parameter set contains only one tile. For example, if single_tile_per_tile_group_flag is set to one (1) or true, certain syntax elements (e.g., syntax elements specifying the number of tiles in the tile group) are excluded in the tile group header. Other syntax elements can also be excluded from the tile header, as described further herein. Furthermore, in some embodiments, an encoder can encode a flag (e.g., a flag called tile_group_info_in_pps_flag) in a parameter set to specify whether tile group structure information is present in the parameter set. The value of Tile_group_info_in_pps_flag is used to coordinate the presence of syntax elements related to tile group structure in the parameter set and the tile group header. In certain embodiments, the tile_group_info_in_pps_flag may be coded instead of or in addition to the single_tile_per_tile_group_flag.

[0050] FIG. 1 is a flowchart of an exemplary operational method 100 for coding a video signal. Specifically, a video signal is encoded by an encoder. The encoding process compresses the video signal by utilizing various mechanisms to reduce the video file size. The smaller file size allows for the transmission of the compressed video file to a user while reducing the associated bandwidth overhead. A decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process typically mirrors the encoding process, allowing the decoder to consistently reconstruct the video signal.

[0051] In step 101, a video signal is input to an encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support live streaming of the video. The video file may include both an audio component and a video component. The video component includes a series of image frames that, when viewed in sequence, create the visual impression of movement. The frames include pixels represented in terms of light, referred to herein as luma components (or luma samples), and color, referred to herein as chroma components (or chroma samples). In some examples, the frames may also include depth values ​​to support three-dimensional displays.

[0052] In step 103, the video is partitioned into blocks. Partitioning involves subdividing pixels within each frame into square and / or rectangular blocks for compression. For example, in High Efficiency Video Coding (HEVC) (also known as H.265 and MPEG-H Part 2), a frame can first be divided into coding tree units (CTUs), which are blocks of a predetermined size (e.g., 64 pixels by 64 pixels). A CTU contains both luma and chroma samples. The coding tree may be used to divide the CTUs into blocks and then iteratively subdivide the blocks until a structure that supports further encoding is achieved. For example, the luma component of a frame may be subdivided until each block contains relatively homogeneous light values. Furthermore, the chroma component of a frame may be subdivided until each block contains relatively homogeneous color values. Thus, the partitioning mechanism varies depending on the content of the video frame.

[0053] In step 105, various compression mechanisms are utilized to compress the image blocks partitioned in step 103. For example, inter-prediction and / or intra-prediction may be utilized. Inter-prediction is designed to take advantage of the fact that objects in a common scene tend to appear in consecutive frames. Thus, blocks depicting an object in a reference frame need not be repeatedly shown in adjacent frames. Specifically, an object such as a table may remain in a constant position across multiple frames. Thus, the table may be shown once, and adjacent frames can reference back to the reference frame. A pattern matching mechanism may be utilized to match objects across multiple frames. Furthermore, a moving object may be displayed across multiple frames, e.g., due to object motion or camera motion. As a specific example, a video may show a car moving across the screen across multiple frames. To indicate such motion, a motion vector may be utilized. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of the object in a reference frame. Thus, inter-prediction allows an image block in a current frame to be coded as a set of motion vectors indicating its offset from a corresponding block in a reference frame.

[0054] Intra prediction encodes blocks within a common frame. Intra prediction exploits the fact that luma and chroma components tend to be clustered within a frame. For example, a patch of green in a tree tends to be located adjacent to other similar patches of green. Intra prediction utilizes multiple directional prediction modes (e.g., 33 in HEVC), planar mode, and direct current (DC) mode. Directional mode indicates that the current block is similar / identical to samples from neighboring blocks in the corresponding direction. Planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on neighboring blocks at the end of the row. Planar mode effectively indicates a smooth transition of light / color across a row / column by utilizing a relatively constant gradient of changing values. DC mode is used for boundary smoothing and indicates that the block is similar / identical to the average value associated with samples from all neighboring blocks associated with the angular direction of the directional prediction mode. Therefore, intra-predicted blocks can represent image blocks as various related prediction modes instead of actual values. Furthermore, inter-predicted blocks can represent image blocks as motion vector values ​​instead of actual values. In either case, the prediction block may not accurately represent the image in some cases. Any differences are stored in a residual block. To further compress the file, a transform may be applied to the residual block.

[0055] Various filtering techniques may be applied in step 107. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction described above may result in the generation of blocky images at the decoder. Furthermore, block-based prediction schemes may encode blocks and then reconstruct the encoded blocks for later use as reference blocks. In-loop filtering schemes iteratively apply noise suppression filters, deblocking filters, adaptive loop filters, and sample adaptive offset (SAO) filters to blocks / filters. These filters mitigate such blocky artifacts so that the encoded file can be accurately reconstructed. Furthermore, these filters mitigate artifacts in the reconstructed reference blocks so that they are less likely to introduce additional artifacts in subsequent blocks that are coded based on the reconstructed reference blocks.

[0056] Once the video signal has been partitioned, compressed, and filtered, the resulting data is encoded into a bitstream at step 109. The bitstream includes the data described above and any desired signaling data to support proper video signal reconstruction at the decoder. For example, such data may include partition data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may also be broadcast and / or multicast to multiple decoders. Generating the bitstream is an iterative process. Thus, steps 101, 103, 105, 107, and 109 may occur sequentially and / or simultaneously across multiple frames and blocks. The order shown in FIG. 1 is presented for clarity and ease of discussion and is not intended to limit the video coding process to any particular order.

[0057] In step 111, the decoder receives the bitstream and begins the decoding process. Specifically, the decoder uses an entropy decoding scheme to convert the bitstream into corresponding syntax and video data. In step 111, the decoder uses the syntax from the bitstream to determine the frame partitions. The partitions should match the block partition results from step 103. The entropy encoding / decoding used in step 111 is described below. During the compression process, the encoder generates many options, such as selecting a block partitioning scheme from several possible options based on the spatial location of values ​​within the input image. Signaling the exact option may utilize a large number of bins. As used herein, a bin is a binary value treated as a variable (e.g., a bit value that can change depending on the context). Entropy coding allows the encoder to discard any options that are clearly not feasible in a particular case, leaving a set of acceptable options. Each acceptable option is then assigned a codeword. The length of the codeword is based on the number of allowable choices (e.g., one bin for two choices, two bins for three or four choices, etc.). The encoder then encodes a codeword for the selected choice. This scheme reduces the size of the codeword because the codeword is desirably large to uniquely indicate a choice from a small subset of the possible choices, as opposed to uniquely indicating a choice from a potentially large set of all possible choices. The decoder then decodes the choices by determining the set of allowable choices in a similar manner to the encoder. By determining the set of allowable choices, the decoder can read the codeword and determine the choice made by the encoder.

[0058] In step 113, the decoder performs block decoding. Specifically, the decoder generates a residual block using an inverse transform. Then, the decoder uses the residual block and the corresponding prediction block to reconstruct an image block according to the partition. The prediction block may include both the intra-prediction block and the inter-prediction block generated in step 105 in the encoder. The reconstructed image block is then positioned into a frame of the reconstructed video signal according to the partition data determined in step 111. The syntax of step 113 may also be signaled in the bitstream by entropy coding as described above.

[0059] In step 115, filtering is performed on the frames of the reconstructed video signal in a manner similar to step 107 in the encoder. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter may be applied to the frames to remove blocking artifacts. Once the frames have been filtered, the video signal can be output to a display in step 117 for viewing by an end user.

[0060] 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. Specifically, codec system 200 provides functionality to support the implementation of operational method 100. Codec system 200 is generalized to show components utilized in both an encoder and a decoder. Codec system 200 receives and partitions a video signal as described above with respect to steps 101 and 103 in operational method 100, resulting in a partitioned video signal 201. When operating as an encoder as described above with respect to steps 105, 107, and 109 in method 100, codec system 200 then compresses partitioned video signal 201 into a coding bitstream. When operating as a decoder, codec system 200 generates an output video signal from the bitstream as described above with respect to steps 111, 113, 115, and 117 in operational method 100. Codec system 200 includes a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header format and context adaptive binary arithmetic coding (CABAC) component 231. These components are coupled as shown. In FIG. 2, black lines indicate the movement of data to be encoded / decoded, while dashed lines indicate the movement of control data that controls the operation of other components. The components of codec system 200 may all reside within an encoder. A decoder may include some of the components of codec system 200.For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components are described herein.

[0061] The partitioned video signal 201 is a captured video sequence that has been partitioned into blocks of pixels by a coding tree. The coding tree utilizes various partitioning modes to subdivide blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into smaller blocks. Blocks may be referred to as nodes on the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is subdivided is referred to as the depth of the node / coding tree. The partitioned blocks may be included in a coding unit (CU) in some cases. For example, a CU may be a subpart of a CTU that includes a luma block, a red differential chroma (Cr) block, and a blue differential chroma (Cb) block, as well as the corresponding syntax instructions for the CU. Partitioning modes may include a binary tree (BT), a triple tree (TT), and a quad tree (QT), which are used to partition a node into two, three, or four child nodes, each of which has a shape that varies depending on the partitioning mode used. The partitioned video signal 201 is forwarded to a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.

[0062] The generic coder control component 211 is configured to make decisions related to the coding of images of a video sequence into a bitstream according to application constraints. For example, the generic coder control component 211 manages the optimization of bitrate / bitstream size versus reconstruction quality. Such decisions may be based on storage space / bandwidth availability and image resolution requirements. The generic coder control component 211 also manages buffer utilization in terms of conversion speed to mitigate buffer underrun and overrun issues. To address these issues, the generic coder control component 211 manages partitioning, prediction, and filtering by other components. For example, the generic coder control component 211 may dynamically increase compression complexity to increase resolution, or increase bandwidth usage or decrease compression complexity to reduce resolution and bandwidth usage. Thus, the generic coder control component 211 controls other components of the codec system 200 to balance video signal reconstruction quality and bitrate concerns. The generic coder control component 211 generates control data that controls the operation of other components. Control data is also forwarded to the header format and CABAC component 231 to be encoded into the bitstream to signal parameters for decoding at the decoder.

[0063] The partitioned video signal 201 is also sent to a motion estimation component 221 and a motion compensation component 219 for inter-prediction. A frame or slice of the partitioned video signal 201 may be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter-predictive coding of the received video blocks relative to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.

[0064] The motion estimation component 221 and the motion compensation component 219 may be highly integrated but are shown separately for conceptual purposes. Motion estimation, performed by the motion estimation component 221, is the process of generating motion vectors that estimate motion for a video block. A motion vector may indicate the location of a coding object with respect to a predictive block, for example. A predictive block is a block that is found to closely match a block to be coded in terms of pixel differences. A predictive block may also be referred to as a reference block. Such pixel differences may be determined by sum of absolute difference (SAD), sum of square difference (SSD), or other difference metrics. HEVC utilizes several coding objects, including CTUs, coding tree blocks (CTBs), and CUs. For example, a CTU can be divided into CTBs, which can then be divided into CBs for inclusion in CUs. A CU may be coded as a prediction unit (PU), which contains prediction data, and / or a transform unit (TU), which contains transformed residual data of the CU. The motion estimation component 221 uses rate-distortion analysis as part of a rate-distortion optimization process to generate motion vectors, PUs, and TUs. For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame and may select the reference block, motion vector, etc. with the optimal rate-distortion characteristics. The optimal rate-distortion characteristics balance both the quality of the video reconstruction (e.g., the amount of data loss due to compression) and the coding efficiency (e.g., the size of the final encoding).

[0065] In some examples, the codec system 200 may calculate values ​​for sub-integer picture positions of reference pictures stored in the decoded picture buffer component 223. For example, the video codec system 200 may interpolate values ​​for quarter-pixel positions, eighth-pixel positions, or other fractional-pixel positions of the reference pictures. Accordingly, the motion estimation component 221 may perform motion searches relative to full-pixel and fractional-pixel positions and output motion vectors with fractional-pixel accuracy. The motion estimation component 221 calculates motion vectors for PUs of video blocks in inter-coding slices by comparing the positions of the PUs with the positions of predictive blocks in the reference pictures. The motion estimation component 221 outputs the calculated motion vectors as motion data to the header format and CABAC component 231 for encoding, and motion to the motion compensation component 219.

[0066] The motion compensation performed by motion compensation component 219 may include fetching or generating a predictive block based on the motion vector determined by motion estimation component 221. Again, motion estimation component 221 and motion compensation component 219 may be functionally integrated in some examples. Upon receiving the motion vector of the PU of the current video block, motion compensation component 219 may locate the predictive block pointed to by the motion vector. A residual video block is then formed by subtracting pixel values ​​of the predictive block from pixel values ​​of the current video block being coded to form pixel difference values. Generally, motion estimation component 221 performs motion estimation with respect to the luma component, and motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma and luma components. The predictive block and residual block are forwarded to transform scaling and quantization component 213.

[0067] The partitioned video signal 201 is also sent to an intra-picture estimation component 215 and an intra-picture prediction component 217. Like the motion estimation component 221 and the motion compensation component 219, the intra-picture estimation component 215 and the intra-picture prediction component 217 may be highly integrated but are shown separately for conceptual purposes. Instead of inter-prediction performed by the inter-frame motion estimation component 221 and the motion compensation component 219 as described above, the intra-picture estimation component 215 and the intra-picture prediction component 217 intra-predict the current block relative to blocks within the current frame. In particular, the intra-picture estimation component 215 determines an intra-prediction mode to use to encode the current block. In some examples, the intra-picture estimation component 215 selects an appropriate intra-prediction mode for encoding the current block from multiple tested intra-prediction modes. The selected intra-prediction mode is then forwarded to the header format and CABAC component 231 for encoding.

[0068] For example, the intra picture estimation component 215 may calculate rate-distortion values ​​for various tested intra prediction modes using rate-distortion analysis and select an intra prediction mode with optimal rate-distortion characteristics among the tested modes. The rate-distortion analysis generally determines the amount of distortion (or error) between a coded block and the original uncoded block that was coded to generate the coded block, as well as the bit rate (e.g., number of bits) used to generate the coded block. The intra picture estimation component 215 may calculate a ratio from the distortion and rate for various coded blocks to determine which intra prediction mode exhibits the optimal rate-distortion value for the block. Furthermore, the intra picture estimation component 215 may be configured to code depth blocks of the depth map using depth modeling mode (DMM) based on rate-distortion optimization (RDO).

[0069] The intra-picture prediction component 217, when implemented in an encoder, may generate a residual block from the prediction block based on the selected intra-prediction mode determined by the intra-picture estimation component 215, or, when implemented in a decoder, may read the residual block from the bitstream. The residual block contains value differences between the prediction block and the original block, represented as a matrix. The residual block is then forwarded to the transform scaling and quantization component 213. The intra-picture estimation component 215 and the intra-picture prediction component 217 may operate on both luma and chroma components.

[0070] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a discrete cosine transform (DCT), a discrete sine transform (DST), or a conceptually similar transform, to the residual block to generate a video block containing residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform may convert the residual information from the pixel value domain to a transform domain, such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling includes applying a scaling factor to the residual information. As a result, different frequency information is quantized with different granularity, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bitrate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be varied by adjusting a quantization parameter. In some examples, the transform scaling and quantization component 213 may then perform a scan of a matrix containing the quantized transform coefficients. The quantized transform coefficients are forwarded to the header format and CABAC component 231 for encoding into the bitstream.

[0071] The scaling and inverse transform component 229 applies the inverse processing of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transform, and / or quantization to reconstruct a residual block in the pixel domain for later use as a reference block, which may become a prediction block for another current block, for example. The motion estimation component 221 and / or motion compensation component 219 may calculate a reference block by adding the residual block back to the corresponding prediction block for use in motion estimation of a later block / frame. A filter is applied to the reconstructed reference block to reduce artifacts produced during scaling, quantization, and transform. Such artifacts may otherwise cause inaccurate predictions (and generate additional artifacts) when subsequent blocks are predicted.

[0072] The filter control analysis component 227 and the in-loop filter component 225 apply filters to residual blocks and / or to reconstructed image blocks. For example, a transformed residual block from the scaling and inverse transform component 229 may be combined with a corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. The filter may then be applied to the reconstructed image block. In some examples, the filter may instead be applied to the residual block. Like the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 may be highly integrated and implemented together, but are shown separately for conceptual purposes. The filters applied to reconstructed reference blocks are applied to specific spatial regions and include multiple parameters for adjusting how such filters are applied. The filter control analysis component 227 analyzes the reconstructed reference blocks to determine when such filters should be applied and set the corresponding parameters. Such data is forwarded to the header format and CABAC component 231 as filter control data for encoding. The in-loop filter component 225 applies such filters based on the filter control data. The filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Such filters may be applied in the spatial / pixel domain (e.g., on reconstructed pixel blocks) or in the frequency domain, depending on the example.

[0073] When operating as an encoder, the filtered reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation, as described above. When operating as a decoder, the decoded picture buffer component 223 stores and forwards the reconstructed and filtered blocks to a display as part of the output video signal. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.

[0074] The header format and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into a coding bitstream for transmission to a decoder. Specifically, the header format and CABAC component 231 generates various headers for encoding control data, such as general control data and filter control data. Additionally, prediction data, including intra-prediction and motion data, and residual data in the form of quantized transform coefficient data are all encoded within the bitstream. The final bitstream contains all information required by a decoder to reconstruct the original partitioned video signal 201. Such information may also include an intra-prediction mode index table (also called a codeword mapping table), definitions of the coding contexts of various blocks, indications of the most likely intra-prediction mode, indications of partition information, etc. Such data may be encoded using entropy coding. For example, the information may be encoded using context adaptive variable length coding (CAVLC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding technique. Following entropy coding, the coded bitstream may be transmitted to another device (e.g., a video decoder) or stored for later transmission or retrieval.

[0075] 3 is a block diagram illustrating an example video encoder 300. Video encoder 300 may be utilized to implement the encoding functionality of codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of method of operation 100. Encoder 300 partitions an input video signal to produce a partitioned video signal 301 that is substantially similar to partitioned video signal 201. Partitioned video signal 301 is then compressed and encoded into a bitstream by components of encoder 300.

[0076] Specifically, the partitioned video signal 301 is forwarded to an intra-picture prediction component 317 for intra prediction. The intra-picture prediction component 317 may be substantially similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. The partitioned video signal 301 is also forwarded to a motion compensation component 321 for inter prediction based on reference blocks in a decoded picture buffer component 323. The motion compensation component 321 may be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks from the intra-picture prediction component 317 and the motion compensation component 321 are forwarded to a transform and quantization component 313 for transforming and quantizing the residual blocks. The transform and quantization component 313 may be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual blocks and corresponding prediction blocks (together with associated control data) are forwarded to an entropy coding component 313 for coding into a bitstream. The entropy coding component 331 may be substantially similar to the header format and CABAC component 231 .

[0077] The transformed and quantized residual block and / or the corresponding prediction block are also forwarded from the transform and quantization component 313 to the inverse transform and quantization component 329 for reconstructing into a reference block for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially similar to the scaling and inverse transform component 229. An in-loop filter in the in-loop filter component 325 is also applied to the residual block and / or the reconstructed reference block, depending on the example. The in-loop filter component 325 may be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters as discussed with respect to the in-loop filter component 225. The filtered block is then stored in the decoded picture buffer component 323 for use as a reference block by the motion compensation component 321. The decoded picture buffer component 323 may be substantially similar to the decoded picture buffer component 223.

[0078] 4 is a block diagram illustrating an exemplary video decoder 400. Video decoder 400 may be utilized to implement the decoding functionality of codec system 200 and / or to implement steps 111, 113, 115, and / or 117 of method of operation 100. Decoder 400 receives a bitstream, for example, from encoder 300, and generates a reconstructed output video signal based on the bitstream for display to an end user.

[0079] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding scheme such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 may use header information to provide context for interpreting additional data encoded as codewords in the bitstream. The decoded information includes any desired information for decoding the video signal, such as general control data, filter control data, partition information, motion data, prediction data, and quantized transform coefficients from the residual block. The quantized transform coefficients are forwarded to the inverse transform and quantization component 429 for reconstruction into the residual block. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.

[0080] The reconstructed residual block and / or prediction block are forwarded to the intra-picture prediction component 417 for reconstructing into an image block based on an intra-prediction operation. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 uses a prediction mode to identify the location of a reference block within a frame and applies the residual block to the result to reconstruct an intra-predicted image block. The reconstructed intra-predicted image block and / or residual block, and the corresponding inter-prediction data, are forwarded to the decoded picture buffer component 423 via an in-loop filter component 425, which may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image block, residual block, and / or prediction block, and such information is stored in the decoded picture buffer component 423. The reconstructed image block from the decoded picture buffer component 423 is forwarded to the motion compensation component 421 for inter-prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 uses a motion vector from a reference block to generate a prediction block and provides a residual block as the result to reconstruct an image block. The resulting reconstructed block may be transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 may continue to store additional reconstructed image blocks that can be reconstructed into frames according to the partition information. Such frames may be arranged in a sequence. The sequence is output to a display as a reconstructed output video signal.

[0081] 5 is a schematic diagram illustrating an exemplary bitstream 500 including a coded video sequence. For example, bitstream 500 may be generated by codec system 200 and / or encoder 300 for decoding by codec system 200 and / or decoder 400. As another example, bitstream 500 may be generated by an encoder in step 109 of method 100 for use in step 111 by a decoder.

[0082] The bitstream 500 includes a sequence parameter set (SPS) 510, multiple picture parameter sets (PPS) 512, a tile group header 514, and image data 520. The SPS 510 includes sequence data common to all pictures in a video sequence included in the bitstream 500. Such data may include picture sizing, bit depth, coding tool parameters, bit rate constraints, etc. The PPS 512 includes parameters specific to one or more corresponding pictures. Thus, each picture in a video sequence may reference one PPS 512. The PPS 512 may indicate available coding tools, quantization parameters, offsets, picture-specific coding tool parameters (e.g., filter control), etc. for tiles in the corresponding picture. The tile group header 514 includes parameters specific to each tile group in a picture. Thus, there may be one tile group header 514 for each tile group in a video sequence. The tile group header 514 may include tile group information, a picture order count (POC), a reference picture list, prediction weights, tile entry points, deblocking parameters, etc. It should be noted that some systems refer to the tile group header 514 as a slice header and use such information to support slices instead of tile groups.

[0083] The image data 520 includes video data coded according to inter-prediction and / or intra-prediction, as well as corresponding transformed and quantized residual data. Such image data 520 is stored according to the partitions used to partition the image before encoding. For example, an image in the image data 520 is divided into tile groups 523. The tiles 523 are further divided into coding tree units (CTUs). The CTUs are further divided into coding blocks based on the coding tree. The coding blocks can then be coded / decoded according to a prediction mechanism. An image / picture may include one or more tiles 523.

[0084] A tile 523 is a partitioned portion of a picture created by horizontal and vertical boundaries. The tiles 523 may be coded in raster scan order, and depending on the example, partitioning based on other tiles 523 may or may not be possible. Each tile 523 may have a unique tile index 524 within the picture. The tile index 524 is a procedurally selected numeric identifier that can be used to distinguish one tile 523 from another. For example, the tile index 524 may increase numerically in raster scan order. The raster scan order is left-to-right and top-to-bottom. Note that in some examples, the tiles 523 may also be assigned a tile identifier (ID). The tile ID is an assigned identifier that can be used to distinguish one tile 523 from another. In some examples, calculations may utilize the tile ID instead of the tile index 524. In some examples, the tile ID may also be assigned to have the same value as the tile index 524.

[0085] The tile index 524 may be signaled to indicate the tile group that includes the tile 523. The first tile index and the last tile index can be signaled in the tile group header 514. In some examples, the first tile index and the last tile index are signaled by the corresponding tile ID. The decoder can then determine the tile group configuration based on the flags, the first tile index, and the last tile index. The encoder can use a similar process as the decoder during the rate-distortion optimization process to predict the decoding results at the decoder when selecting the optimal coding approach. By signaling only the first tile index and the last tile index instead of the complete membership of the tile group, a significant number of bits can be saved. This improves coding efficiency and therefore reduces memory and network resource usage in both the encoder and decoder.

[0086] 6 illustrates an exemplary picture 601 partitioned into exemplary tile groups, according to an embodiment of this disclosure. For example, picture 601 may be a single picture within a video sequence that is encoded into and then decoded into bitstream 500, e.g., by codec system 200, encoder 300, and / or decoder 400. Furthermore, picture 601 may be partitioned to support encoding and decoding according to method 100.

[0087] Picture 601 may be partitioned into tiles 603. Tiles 603 may be substantially similar to tiles 523. Tiles 603 may be rectangular and / or square. Tiles 603 are each assigned a tile index that increases in raster scan order. In the illustrated embodiment, the tile indexes range from 0 to 23 (0-23). ​​Such tile indexes are by way of example and are provided for clarity of discussion and therefore should not be considered limiting.

[0088] Picture 601 includes a left boundary 601a that includes tiles 0, 6, 12, and 18, a right boundary 601b that includes tiles 5, 11, 17, and 23, a top boundary 601c that includes tiles 0 through 5, and a bottom boundary 601d that includes tiles 18 through 23. The left boundary 601a, the right boundary 601b, the top boundary 601c, and the bottom boundary 601d form the edges of picture 601. Furthermore, tiles 603 may be partitioned into tile rows 605 and tile columns 607. Tile row 605 is a set of tiles 603 positioned horizontally adjacent to each other, creating a continuous line from left boundary 601a to right boundary 601b (or vice versa). Tile column 607 is a set of tiles 603 positioned vertically adjacent to each other, creating a continuous line from top boundary 601c to bottom boundary 601d (or vice versa).

[0089] Tiles 603 can be included in one or more tile groups 609. A tile group 609 is a set of related tiles 603 that can be extracted and coded separately, for example, to support display of a region of interest and / or to support parallel processing. Tiles 603 within a tile group 609 can be coded without reference to tiles 603 outside the tile group 609. Each tile 603 may be assigned to a corresponding tile group 609, and thus, picture 601 can include multiple tile groups 609. However, for clarity of discussion, this disclosure refers to tile groups 609 shown as shaded regions containing tiles 603 with indexes seven through ten (7-10) and thirteen through sixteen (13-16).

[0090] Thus, a tile group 609 of a picture 601 can be signaled by a first tile index of 7 and a last tile index of 16. A decoder may want to determine the configuration of the tile group 609 based on the first tile index and the last tile index. As used herein, the tile group 609 configuration refers to the rows, columns, and tiles 603 in the tile group 609. To determine the tile group 609 configuration, a video coding device can utilize a predetermined algorithm. For example, the video coding device can determine the number of tiles 603 in a tile group 609 partitioned from a picture 601 by setting a delta tile index as the difference between the index of the last tile in the tile group 609 and the index of the first tile in the tile group 609. The number of tile rows 605 in the tile group 609 can be determined by dividing the delta tile index by the number of tile columns 607 in the picture 601 plus one. Additionally, the number of tile columns 607 in a tile group 609 can be determined as the delta tile index modulo the number of tile columns 607 in the picture 601 plus 1. The number of tiles 603 in a tile group 609 can be determined by multiplying the number of tile columns 607 in the tile group 609 by the number of tile rows 605 in the tile group 609.

[0091] As mentioned above, in certain situations, the current signaling of the tile group structure, whether signaled in the parameter set or in the tile group header, contains some redundant information. To solve this problem, the following aspects, taken alone or applied in combination in one or more embodiments, are proposed in the present disclosure to solve the above-mentioned problem.

[0092] In an embodiment, when there is more than one tile per picture, a flag is signaled within a parameter set directly or indirectly referenced by a tile group to specify whether each of the tile groups that reference the parameter set contains only one tile. The parameter set may be a sequence parameter set, a picture parameter set, or any other parameter set directly or indirectly referenced by a tile group. In an embodiment, this flag may be referred to as single_tile_per_tile_group_flag. In an embodiment, when the value of single_tile_per_tile_group_flag is equal to 1, the following syntax elements in the tile group header are not present: (1) a syntax element that specifies the number of tiles in the tile group, (2) a syntax element that specifies the identity of the last tile in the tile group, and (3) a syntax element that specifies the tile identity of any tile other than the first tile in the tile group.

[0093] In another embodiment, tile group structure information may be signaled in a parameter set referenced directly or indirectly by each tile group or directly in the tile group header. When the number of tiles in a picture is greater than one, a flag may be present in the parameter set to specify whether tile group structure information is present in the parameter set. The parameter set may be a sequence parameter set, a picture parameter set, or any other parameter set referenced directly or indirectly by a tile group. In an embodiment, the flag may be referred to as tile_group_info_in_pps_flag. In an embodiment, when the flag is not present in the parameter set (e.g., when the picture contains only one tile), the value of tile_group_info_in_pps_flag is inferred to be equal to 0. The value of Tile_group_info_in_pps_flag is used to coordinate the presence of syntax elements related to tile group structure in the parameter sets and the tile group header. These syntax elements may include: (1) a syntax element that specifies the number of tiles in the tile group; (2) a syntax element that specifies the identification of the last tile in the tile group; and (3) a syntax element that specifies the tile identification of any tile other than the first tile in the tile group.

[0094] The syntax and semantics of the relevant syntax elements in the PPS and tile group headers according to an embodiment are as follows: The description is relative to the base text set forth in the JVET contribution JVET-L0686, titled "Draft text of video coding specification." That is, only delta or additive changes are described, while text in the base text not mentioned below applies as they are described in the base text. Changed text for the pic_parameter_set_rbsp( ) function relative to the base text is highlighted with an asterisk (*). [Table 1]

[0095] In the above-mentioned picture parameter set RBSP, the syntax element pps_pic_parameter_set_id identifies a PPS for reference by other syntax elements. In an embodiment, the value of pps_pic_parameter_set_id shall be in the range of 0 to 63, inclusive. The syntax element pps_seq_parameter_set_id specifies the value of sps_seq_parameter_set_id of the active SPS. The value of pps_seq_parameter_set_id shall be in the range of 0 to 15, inclusive. In an embodiment, if the syntax element transform_skip_enabled_flag is equal to 1, the transform_skip_flag syntax element may be present in the residual coding syntax. If the syntax element transform_skip_enabled_flag is equal to 0, the transform_skip_flag syntax element is not present in the residual coding syntax. When the syntax element single_tile_in_pic_flag is equal to 1, it indicates that there is only one tile in each picture that references the PPS. When the syntax element single_tile_in_pic_flag is equal to 0 (i.e., if(!single_tile_in_pic_flag)), it specifies that there is more than one tile in each picture that references the PPS. In this case, the syntax element single_tile_per_tile_group_flag is used to indicate whether each tile group that references the PPS contains exactly one tile. For example, when the syntax element single_tile_per_tile_group_flag is equal to 1, it specifies that each tile group that references the PPS contains exactly one tile. When the syntax element single_tile_per_tile_group_flag is equal to 0, it indicates that each tile group that references the parameter set contains more than one tile.The syntax element tile_group_info_in_pps_flag is used to indicate whether tile group information is present in the PPS or in a tile group header that references the PPS. In an embodiment, if the syntax element tile_group_info_in_pps_flag is equal to 1, this specifies that tile group information is present in the PPS but not in a tile group header that references the PPS. If the syntax element tile_group_info_in_pps_flag is equal to 0, it indicates that tile group information is not present in the PPS but is present in a tile group header that references the PPS.

[0096] In the illustrated embodiment, if the syntax element tile_group_info_in_pps_flag is equal to 1 (i.e., if(tile_group_info_in_pps_flag)), which specifies that tile group information is present in the PPS and not in the tile group header that references the PPS, then the variable num_tile_groups_in_pic_minus1 is set equal to one less than the number of tile groups in the picture. This variable is used in a for-loop to loop through each of the tile groups in the picture. The syntax element pps_first_tile_id[ i ] specifies the tile ID of the first tile of the i-th tile group in the picture. In an embodiment, the length of pps_first_tile_id[ i ] is Ceil( Log2( NumTilesInPic ) ) bits. The value of pps_first_tile_id[ i ] must not be equal to the value of pps_first_tile_id[ j ] for any i not equal to j. Unless otherwise specified, the tile ID of the first tile of the ith tile group in a picture must not be the same as the tile ID of the first tile of any other tile group in the picture.

[0097] For each i-th tile group in a picture, if the syntax element single_tile_per_tile_group_flag is equal to 0 (i.e., if(!single_tile_per_tile_group_flag)), which indicates that the i-th tile group referring to the parameter set contains one or more tiles, the syntax element pps_num_tiles_in_tile_group_minus1[ i ] plus 1 is used to specify the number of tiles in the i-th tile group. The value of pps_num_tiles_in_tile_group_minus1[ i ] should be in the range 0 to NumTilesInPic - 1, inclusive. In an embodiment, when the syntax element pps_num_tiles_in_tile_group_minus1[ i ] is not present, the value of pps_num_tiles_in_tile_group_minus1[ i ] is inferred to be equal to 0.

[0098] As another example, the syntax and semantics of relevant syntax elements in the PPS and tile group header according to the second embodiment are as follows: As mentioned above, the description is relative to the base text set forth in the JVET contribution JVET-L0686, titled "Draft text of video coding specification." That is, only delta or additive changes are described, while text in the base text not mentioned below applies as they are set forth in the base text. Changed text for the pic_parameter_set_rbsp( ) function relative to the base text is highlighted with an asterisk (*).

[0099] [Table 2]

[0100] The syntax elements pps_pic_parameter_set_id, pps_seq_parameter_set_id, transform_skip_enabled_flag, and single_tile_in_pic_flag are as described above previously in accordance with the base text (JVET contribution JVET-L0686). In this embodiment, if the syntax element single_tile_in_pic_flag is equal to 0 (i.e., if(!single_tile_in_pic_flag)), this indicates that each tile group that references the parameter set contains one or more tiles, and the single_tile_per_tile_group_flag syntax element is used to indicate whether each tile group that references the PPS contains exactly one tile. In an embodiment, if single_tile_per_tile_group_flag is equal to 1, it indicates that each tile group that references the PPS contains exactly one tile. When single_tile_per_tile_group_flag is equal to 0, it specifies that each tile group that references a PPS contains one or more tiles. The tile_group_info_in_pps_flag syntax element is used to indicate whether tile group information is present in the PPS or in the tile group header that references the PPS. In an embodiment, when tile_group_info_in_pps_flag is equal to 1, it specifies that tile group information is present in the PPS but not in the tile group header that references the PPS. When tile_group_info_in_pps_flag is 0, it indicates that tile group information is not present in the PPS but is present in the tile group header that references the PPS.

[0101] In the illustrated embodiment, if tile_group_info_in_pps_flag is equal to 1 (i.e., if(tile_group_info_in_pps_flag)), which indicates that tile group information is present in the PPS and not in the tile group header that references the PPS, then the variable num_tile_groups_in_pic_minus1 is set equal to one less than the number of tile groups in the picture. This variable is used in a for-loop to loop through each of the tile groups in the picture. The syntax element pps_first_tile_id[ i ] specifies the tile ID of the first tile of the ith tile group in the picture. The value of pps_first_tile_id[ i ] must not be equal to the value of pps_first_tile_id[ j ] for any i not equal to j. In this embodiment, if there is more than one tile per tile group (i.e., if(!single_tile_per_tile_group_flag)), the syntax element pps_last_tile_id[ i ] specifies the tile ID of the last tile in the i-th tile group. The length of pps_first_tile_id[ i ] and pps_last_tile_id[ i ] is Ceil( Log2( NumTilesInPic ) ) bits.

[0102] In an embodiment, the tile group header and RBSP syntax and semantics are as follows: Modified text for the tile_group_header( ) function relative to the base text is highlighted with an asterisk (*). [Table 3]

[0103] In the embodiment shown, if the number of tiles in a picture (NumTilesInPic) is greater than 1, the syntax element first_tile_id is used to specify the tile ID of the first tile in a tile group. The length of first_tile_id is Ceil(Log2(NumTilesInPic)) bits. The value of first_tile_id for a tile group must not be equal to the value of first_tile_id for any other tile group in the same picture. If single_tile_per_tile_group_flag specifies that there is more than one tile per tile group (i.e., if(!single_tile_per_tile_group_flag)) and tile_group_info_in_pps_flag indicates that the tile group information is not present in the PPS but is present in the tile group header that references the PPS (i.e., !tile_group_info_in_pps_flag), then the syntax element last_tile_id is used to specify the tile ID of the last tile in a tile group. In an embodiment, the length of last_tile_id is Ceil(Log2(NumTilesInPic)) bits.

[0104] In an embodiment, when NumTilesInPic is equal to 1 or single_tile_per_tile_group_flag is equal to 1, the value of last_tile_id is inferred to be equal to first_tile_id. In an embodiment, when tile_group_info_in_pps_flag is equal to 1, the value of last_tile_id is inferred to be equal to the value of pps_first_tile_id[ i ], where i is a value such that first_tile_id is equal to pps_first_tile_id[ i ]. In this embodiment, each tile group may be further constrained to contain a rectangular region of the picture. In this case, first_tile_id specifies the tile ID of the tile located in the upper left corner of the tile group, and last_tile_id specifies the tile ID of the tile located in the lower right corner of the tile group.

[0105] In an embodiment, a syntax element can be further signaled within the PPS to specify a tile group mode, which allows at least the following two tile group modes: In a first mode, called rectangular tile group mode, each tile group is further constrained to contain a rectangular region of the picture, where first_tile_id specifies the tile ID of the tile located in the upper left corner of the tile group and last_tile_id specifies the tile ID of the tile located in the lower right corner of the tile group. In a second mode, called tile raster scan mode, no additional modifications are made, and the tiles included in each tile group are consecutive tiles in the tile raster scan of the picture.

[0106] 7 is a schematic diagram of an exemplary video coding device 700. The video coding device 700 is suitable for implementing examples / embodiments of the disclosure as described herein. The video coding device 700 includes a downstream port 720, an upstream port 750, and / or a transceiver unit (Tx / Rx) 710 including a transmitter and / or receiver for communicating data upstream and / or downstream over a network. The video coding device 700 also includes a processor 730 including a logic unit and / or central processing unit (CPU) for processing data, and a memory 732 for storing data. The video coding device 700 may also include electrical, optical-to-electrical (OE) components, electrical-to-optical (EO) components, and / or wireless communication components connected to the upstream port 750 and / or the downstream port 720 for communication of data over an electrical, optical, or wireless communication network. Video coding device 700 may also include input and / or output (I / O) devices 760 for communicating data to and from a user. I / O devices 760 may include output devices such as a display for displaying video data, speakers for outputting audio data, etc. I / O devices 760 may also include input devices such as a keyboard, mouse, trackball, etc., and / or corresponding interfaces for interfacing with such output devices.

[0107] The processor 730 is implemented in hardware and software. The processor 730 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), field-programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). The processor 730 communicates with the downstream ports 720, the Tx / Rx 710, the upstream ports 750, and the memory 732. The processor 730 includes a coding module 714. The coding module 714 implements embodiments of the disclosure described herein, such as the methods 100, 800, and 900, which may utilize the bitstream 500 and / or the image partitioned into tile groups 609. The coding module 714 may also implement any other method / mechanism described herein. Additionally, the coding module 714 may implement the codec system 200, the encoder 300, and / or the decoder 400. For example, when operating as an encoder, coding module 714 may encode a video bitstream including coding data for at least one picture that includes at least one tile group. Coding module 714 may further encode a flag in a parameter set that specifies whether tile information for the coding picture is present in the parameter set or in a tile group header. When operating as a decoder, coding module 714 may read the flag that indicates whether tile information for the coding picture is present in the parameter set or in a tile group header. Thus, coding module 714 solves problems specific to the field of video coding and improves the functionality of video coding device 700 to reduce or remove data redundancy in a video sequence (thus improving coding efficiency) without adversely affecting tile group signaling.Additionally, coding module 714 performs transformations to different states of video coding device 700. Alternatively, coding module 714 can be implemented as instructions stored in memory 732 and executed by processor 730 (e.g., as a computer program product stored on a non-transitory medium).

[0108] Memory 732 may include one or more memory types such as a disk, a tape drive, a solid state drive, read only memory (ROM), random access memory (RAM), flash memory, ternary content-addressable memory (TCAM), static random-access memory (SRAM), etc. Memory 732 may be used as overflow data storage for storing programs when they are selected for execution and for storing instructions and data read during the execution of programs.

[0109] 8 is a flowchart of an example method 800 for encoding an image, such as picture 601, into a bitstream, such as bitstream 500. Method 800 may be utilized by an encoder, such as codec system 200, encoder 300, and / or video coding device 700, when performing method 100.

[0110] Method 800 may begin when an encoder receives a video sequence including multiple images and decides to encode the video sequence into a bitstream, for example, based on user input. The video sequence is partitioned into pictures / images / frames for further partitioning before encoding. In step 801, a picture is partitioned into multiple tiles. The tiles may be further partitioned into multiple CTUs. The CTUs may be further partitioned into coding blocks for applying prediction-based compression. Groups of tiles are further assigned to tile groups.

[0111] In step 803, the tile groups are coded into a bitstream containing coding data for at least one picture. Each of the pictures contains at least one tile group. Furthermore, a flag is coded in a parameter set in the bitstream. The flag indicates whether tile information for the coded picture is present in the parameter set or in the tile group header. The tile information indicates which tiles of the picture are included in the tile group. As a specific example, the flag is tile_group_info_in_pps_flag. For example, the flag can be coded in the PPS associated with the picture.

[0112] In step 805, the tile information is coded into either the parameter set or the tile group header based on a flag. In an embodiment, when the syntax element tile_group_info_in_pps_flag is equal to 1, this specifies that the tile group information is present in the parameter set but not in the tile group header that references the parameter set. When the syntax element tile_group_info_in_pps_flag is equal to 0, it indicates that the tile group information is not present in the parameter set but is present in the tile group header that references the parameter set. In an embodiment, when the flag is not present in the parameter set (e.g., when the picture contains only one tile), the value of tile_group_info_in_pps_flag is inferred to be equal to 0. The tile group information may include a syntax element that specifies the number of tiles in the tile group, a syntax element that specifies the identification of the last tile in the tile group, and a syntax element that specifies the tile identification of any tiles other than the first tile in the tile group.

[0113] In step 807, the video bitstream is transmitted or sent along the network towards the decoder. In an embodiment, the video bitstream is transmitted on demand. The video bitstream can also be automatically pushed out to the decoder by the encoder. In an embodiment, the coded video bitstream can be stored temporarily or permanently at the encoder.

[0114] 9 is a flowchart of an example method 900 for decoding an image, such as picture 601, from a bitstream, such as bitstream 500. Method 900 may be utilized by a decoder, such as codec system 200, decoder 400, and / or video coding device 700, when performing method 100.

[0115] Method 900 begins in step 901 when a decoder begins receiving a bitstream of coding data representing a video sequence, for example as a result of method 800. For example, the coding data includes coding data for at least one picture, each picture including at least one tile group.

[0116] At step 903, a flag is parsed from a parameter set in the bitstream. For example, the flag can be obtained from a PPS associated with the picture. The terms parse or parse, as used herein, can include the process of identifying or determining whether a flag or other syntax element is present in a parameter set, obtaining a value corresponding to the flag or other syntax element, and determining a condition associated with the value of the flag or other syntax element. In an embodiment, by parsing the flag, method 900 can determine, based on the flag, whether tile information for the coding picture is present in the parameter set or in a tile group header.

[0117] In step 905, the method 900 obtains tile information in either a parameter set or a tile group header based on the flag. For example, when the flag specifies that the tile information of the coding picture is encoded in the parameter set, the method 900 can parse the tile information from the parameter set. Similarly, when the flag specifies that the tile information of the coding picture is encoded in the tile group header, the method 900 can parse the tile information from the tile group header.

[0118] In step 907, the tile group can be decoded to reconstruct a portion of the picture, which can then be included as part of a reconstructed video sequence. The resulting reconstructed video sequence can be transferred to a display device for display to a user. The resulting reconstructed video sequence can also be temporarily or permanently stored in a memory or data storage unit of the decoder.

[0119] 10 is a schematic diagram of an example system 1000 for coding a video sequence of images, such as picture 601, in a bitstream, such as bitstream 500. System 1000 may be implemented by an encoder and decoder, such as codec system 200, encoder 300, decoder 400, and / or video coding device 700. Furthermore, system 1000 may be utilized when performing methods 100, 800, and / or 900.

[0120] The system 1000 includes a video encoder 1002. The video encoder 1002 includes a partition module 1001 that partitions a first picture into multiple tiles. The video encoder 1002 further includes an allocation module 1003 that allocates groups of tiles to tile groups. The video encoder 1002 further includes an encoding module 1005 that encodes the tile groups into a bitstream and encodes a flag in a parameter set in the bitstream to indicate whether tile information for the encoded picture is present in the parameter set or in a tile group header. The video encoder 1002 further includes a storage module 1007 that stores the bitstream for communication to a decoder. The video encoder 1002 further includes a transmission module 1009 that transmits the bitstream to the decoder. The video encoder 1002 may be further configured to perform any of the steps of the method 800.

[0121] The system 1000 also includes a video decoder 1010. The video decoder 1010 includes a receiving module 1011 that receives a bitstream including tile groups, each including a group of tiles partitioned from a picture. The video decoder 1010 further includes an obtaining module 1013 that obtains a flag from a parameter set in the bitstream, the flag indicating whether tile information for the coded picture is present in the parameter set or in a tile group header. The video decoder 1010 further includes a determining module 1015 that determines whether certain conditions are present when they relate to the location of the tiling information. For example, the determining module 1015 can determine whether there is a single tile per tile group by parsing a single_tile_per_tile_group_flag. The video decoder 1010 further includes a decoding module 1017 that decodes the tile groups to generate a reconstructed video sequence for display. The video decoder 1010 may be further configured to perform any of the steps of the method 900.

[0122] A first component is directly connected to a second component when there are no intervening components other than wires, traces, or another medium between the first and second components. A first component is indirectly connected to a second component when there are intervening components other than wires, traces, or another medium between the first and second components. The term "connected" and variations thereof include both direct and indirect connections. The use of the term "about" means a range that includes ±10% of the subsequent numerical value, unless otherwise specified.

[0123] It should be further understood that the steps of the exemplary methods described herein do not necessarily have to be performed in the order described, and the order of the steps of such methods should be understood to be exemplary only. Similarly, additional steps may be included in such methods, and certain steps may be omitted or combined in methods according to various embodiments of the present disclosure.

[0124] Although several embodiments have been provided in this disclosure, it will be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the disclosure. The examples of the present invention should be considered illustrative and not restrictive, and are not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.

[0125] Additionally, the techniques, systems, subsystems, and methods described and illustrated in various embodiments may be combined or integrated with other systems, components, techniques, or methods without departing from the scope of the present disclosure. Other examples of changes, substitutions, and alterations will be ascertainable by those skilled in the art and may be made without departing from the spirit and scope of the present disclosure. [Explanation of symbols]

[0126] 500 bitstream 510 Sequence Parameter Set (SPS) 512 Picture Parameter Set (PPS) 514 Tile Group Header 520 Image Data 524 tile index 523 tiles

Claims

1. 1. A method, implemented in an encoder, for encoding a video bitstream, the video bitstream comprising encoded data for a plurality of pictures, each of the plurality of pictures comprising at least one slice, one slice of the at least one slice comprising a plurality of tiles, the method comprising: obtaining a number of tiles of a picture among the plurality of pictures; When the number of tiles of the picture is greater than one, encoding a flag indicating whether tile information of the picture exists in a picture parameter set; if the flag indicates that the tile information for the picture is present in the picture parameter set, encoding the tile information in the picture parameter set; if the flag indicates that the tile information for the picture is not present in the picture parameter set, encoding the tile information in a slice header of a slice in the picture; encoding data of the picture into the video bitstream based on the tile information; A method comprising:

2. The method of claim 1 , wherein the tile information indicates the number of tiles in the slice.

3. The method of claim 2 , wherein the tile information further indicates an address or location of the slice.

4. 1. A decoder-implemented method for decoding a video bitstream, the video bitstream comprising coded data for a plurality of pictures, each of the plurality of pictures comprising at least one slice, one slice of the at least one slice comprising a plurality of tiles, the method comprising: obtaining a number of tiles of a picture of the plurality of pictures; When the number of tiles of the picture is greater than one, parsing a flag indicating whether tile information of the picture exists in a picture parameter set; if the flag indicates that the tile information for the picture is present in the picture parameter set, parsing the tile information from the picture parameter set; if the flag indicates that the tile information for the picture is not present in the picture parameter set, parsing the tile information from a slice header of a slice in the picture; obtaining decoded data of the picture based on the tile information; A method comprising:

5. The method of claim 4 , wherein the tile information indicates the number of tiles in the slice.

6. The method of claim 5 , wherein the tile information further indicates an address or location of the slice.

7. 1. An encoder comprising: a memory storing instructions; one or more processors coupled to the memory; Including, The one or more processors execute the instructions to cause the encoder to perform the method of any one of claims 1 to 3.

8. A decoder comprising: a memory storing instructions; one or more processors coupled to the memory; Including, The one or more processors execute the instructions to cause the decoder to perform the method of any one of claims 4 to 6.

9. 1. A method for storing an encoded bitstream, the method comprising: generating the bitstream, the bitstream including coded data of a plurality of pictures, each of the plurality of pictures including at least one slice, the bitstream including a flag indicating whether tile information of the picture exists in a picture parameter set when the number of tiles in the picture is greater than one, the tile information being included in the picture parameter set if the flag indicates that the tile information of the picture exists in the picture parameter set, and the tile information being included in a slice header of the slice if the flag indicates that the tile information of the picture does not exist in the picture parameter set; storing the bitstream on a storage medium; A method comprising:

10. 1. An apparatus for storing an encoded bitstream, comprising: at least one storage medium; and at least one processor; the at least one processor is configured to generate the bitstream; the at least one storage medium is configured to store the bitstream, the bitstream including encoded data of a plurality of pictures, each of the plurality of pictures including at least one slice, the bitstream including a flag indicating whether tile information of the picture exists in a picture parameter set when the number of tiles in the picture is greater than one, and if the flag indicates that the tile information of the picture exists in the picture parameter set, the tile information is included in the picture parameter set, and if the flag indicates that the tile information of the picture does not exist in the picture parameter set, the tile information is included in a slice header of the slice.

Citation Information

Patent Citations

  • Mpeg recording / Reproducing method, mpeg recording / reproducing apparatus and mpeg recording / Reproducing program

    JP2003018548A

  • JPP7368477B